Production hardening: API security, per-question degradation, observability, golden-set eval gate - #3
Open
ashishsharma-reborg wants to merge 1 commit into
Conversation
Production hardening of the /api/evaluate pipeline plus a reproducible measurement layer, so prompt/model/calibration changes are gated by regression checks instead of guesswork. Engine (src/lib): - logprob readout: one call per question, option scores from first-token top-logprobs (logprob.ts), replacing per-option micro-calls by default - per-question degradation: questions whose primary readout fails fall back individually to micro-scorers; a request only errors when every question fails; neutral placeholders recorded in meta.failed_questions - bounded concurrency (OPENJEV_CONCURRENCY), client-disconnect abort propagation, per-question readout honesty in meta (effective readout, fallback_questions, retries, calibration_t) - calibration: temperature scaling with fitTemperature (NLL grid search), Brier/ECE metrics, calibration.ts CLI to fit OPENJEV_CALIBRATION_T - retry: SDK retries disabled; own withRetry with exponential backoff, Retry-After, fail-fast on non-retryable 4xx - shuffle: deterministic option shuffling for position-bias checks Server: - zod request validation (.strict, size bounds) with SSRF guard on client-supplied base_url/decider_url - bearer auth (OPENJEV_API_KEYS, timing-safe), per-IP rate limiting - structured JSON logs with request ids, /api/metrics counters and latency percentiles, graceful shutdown, health/metrics stay open for orchestrator probes Ops/tests: - contract tests against a fake OpenAI-compatible server; unit tests for parsing, calibration, retry, fallback, shuffle, golden harness (44 tests); test runner fixed to tsx --test - CI workflow (typecheck + tests + eval-regression gate), Dockerfile, committed lockfile, .env.example Golden-set harness: - eval/golden/dataset.jsonl (8 rows / 22 labeled questions), baseline comparison with thresholds (accuracy -2pt, Brier +0.02, ECE +0.02), scripts/eval-golden.ts CLI, mock OpenAI server for zero-cost demos - NOTE: seeded labels are placeholders to be validated by the team, and the committed baseline comes from the mock server; recreate it against a real provider before trusting the gate Generated with Codebuff 🤖 Co-Authored-By: Codebuff <noreply@codebuff.com>
|
Hi, I have added some hardening, golden evals and logprob readout stack |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Three layers of work on the evaluate pipeline:
/api/evaluate— closes the security and reliability gaps that made the service unsafe to expose: unauthenticated SSRF-capable endpoint, unvalidated client input, request-level fallback that failed whole batches, unbounded outbound concurrency, and zero observability.top_logprobs(letters A–Z), degrading per question to micro-scorers only when needed. Plus temperature-scaling calibration with a fit CLI and honest per-request telemetry inmeta.Notable changes
Security
.strict(), size bounds) with SSRF guard on client-suppliedbase_url/decider_urlOPENJEV_API_KEYS, timing-safe compare), per-IP rate limitingReliability
meta.failed_questionsOPENJEV_CONCURRENCY, default 6), client-disconnect abort propagation, graceful shutdownRetry-After, fail-fast on non-retryable 4xx; SDK retries disabled)Observability
GET /api/metricswith latency p50/p95, fallback/retry/token countersmetanow reports the effective readout,fallback_questions,retries,calibration_tTesting/ops
tsx --test.env.exampleEval harness
eval/golden/dataset.jsonl(8 rows / 22 labeled questions),npm run eval:golden, thresholds: accuracy −2pt / Brier +0.02 / ECE +0.02eval/golden/mock-server.ts) for zero-cost demos of the gate (verified: sharp-vs-flat baseline fails with exit 1)baseline.jsonwas produced by the mock server — recreate it against a real provider:npm run eval:golden -- --create --model <your-model>Verification
npm run typecheckclean🤖 Generated with Codebuff
Co-Authored-By: Codebuff noreply@codebuff.com