Skip to content

Production hardening: API security, per-question degradation, observability, golden-set eval gate - #3

Open
ashishsharma-reborg wants to merge 1 commit into
SiliconLabAI:mainfrom
ashishsharma-reborg:feat/production-hardening
Open

ashishsharma-reborg wants to merge 1 commit into
SiliconLabAI:mainfrom
ashishsharma-reborg:feat/production-hardening

Conversation

@ashishsharma-reborg

Copy link
Copy Markdown

Summary

Three layers of work on the evaluate pipeline:

  1. Production hardening of /api/evaluate — closes the security and reliability gaps that made the service unsafe to expose: unauthenticated SSRF-capable endpoint, unvalidated client input, request-level fallback that failed whole batches, unbounded outbound concurrency, and zero observability.
  2. Logprob readout stack — default parallel mode now scores a question in one call via first-token top_logprobs (letters A–Z), degrading per question to micro-scorers only when needed. Plus temperature-scaling calibration with a fit CLI and honest per-request telemetry in meta.
  3. Golden-set eval harness — a labeled dataset + Brier/ECE/accuracy regression gate so prompt, model, and calibration changes are measured, not guessed.

Notable changes

Security

  • Zod request validation (.strict(), size bounds) with SSRF guard on client-supplied base_url/decider_url
  • Bearer auth (OPENJEV_API_KEYS, timing-safe compare), per-IP rate limiting
  • Health/metrics routes stay unauthenticated for orchestrator probes

Reliability

  • Per-question degradation: one bad question no longer fails the batch; failures recorded in meta.failed_questions
  • Bounded concurrency (OPENJEV_CONCURRENCY, default 6), client-disconnect abort propagation, graceful shutdown
  • Retry policy owned in-app (backoff + Retry-After, fail-fast on non-retryable 4xx; SDK retries disabled)

Observability

  • Structured JSON logs with request IDs; GET /api/metrics with latency p50/p95, fallback/retry/token counters
  • meta now reports the effective readout, fallback_questions, retries, calibration_t

Testing/ops

  • 44 tests: unit suites + contract tests against a fake OpenAI-compatible server (which caught 2 real bugs pre-merge); test runner fixed to tsx --test
  • CI workflow (typecheck → tests → eval-regression), Dockerfile, committed lockfile, .env.example

Eval harness

  • eval/golden/dataset.jsonl (8 rows / 22 labeled questions), npm run eval:golden, thresholds: accuracy −2pt / Brier +0.02 / ECE +0.02
  • Mock OpenAI server (eval/golden/mock-server.ts) for zero-cost demos of the gate (verified: sharp-vs-flat baseline fails with exit 1)

⚠️ Before trusting the gate

  • Seeded dataset labels are sensible placeholders — validate them with the team
  • The committed baseline.json was produced by the mock server — recreate it against a real provider: npm run eval:golden -- --create --model <your-model>

Verification

  • npm run typecheck clean
  • 44/44 tests pass
  • Live smoke test: 401/400/500 paths, metrics, health, abort propagation all confirmed against a running server
  • Regression gate demoed end-to-end with the mock (baseline create → fail → pass)

🤖 Generated with Codebuff
Co-Authored-By: Codebuff noreply@codebuff.com

Production hardening of the /api/evaluate pipeline plus a reproducible
measurement layer, so prompt/model/calibration changes are gated by
regression checks instead of guesswork.

Engine (src/lib):
- logprob readout: one call per question, option scores from first-token
  top-logprobs (logprob.ts), replacing per-option micro-calls by default
- per-question degradation: questions whose primary readout fails fall
  back individually to micro-scorers; a request only errors when every
  question fails; neutral placeholders recorded in meta.failed_questions
- bounded concurrency (OPENJEV_CONCURRENCY), client-disconnect abort
  propagation, per-question readout honesty in meta (effective readout,
  fallback_questions, retries, calibration_t)
- calibration: temperature scaling with fitTemperature (NLL grid search),
  Brier/ECE metrics, calibration.ts CLI to fit OPENJEV_CALIBRATION_T
- retry: SDK retries disabled; own withRetry with exponential backoff,
  Retry-After, fail-fast on non-retryable 4xx
- shuffle: deterministic option shuffling for position-bias checks

Server:
- zod request validation (.strict, size bounds) with SSRF guard on
  client-supplied base_url/decider_url
- bearer auth (OPENJEV_API_KEYS, timing-safe), per-IP rate limiting
- structured JSON logs with request ids, /api/metrics counters and
  latency percentiles, graceful shutdown, health/metrics stay open for
  orchestrator probes

Ops/tests:
- contract tests against a fake OpenAI-compatible server; unit tests
  for parsing, calibration, retry, fallback, shuffle, golden harness
  (44 tests); test runner fixed to tsx --test
- CI workflow (typecheck + tests + eval-regression gate), Dockerfile,
  committed lockfile, .env.example

Golden-set harness:
- eval/golden/dataset.jsonl (8 rows / 22 labeled questions), baseline
  comparison with thresholds (accuracy -2pt, Brier +0.02, ECE +0.02),
  scripts/eval-golden.ts CLI, mock OpenAI server for zero-cost demos
- NOTE: seeded labels are placeholders to be validated by the team, and
  the committed baseline comes from the mock server; recreate it against
  a real provider before trusting the gate

Generated with Codebuff 🤖
Co-Authored-By: Codebuff <noreply@codebuff.com>
@Ashishsharma-12

Copy link
Copy Markdown

Hi, I have added some hardening, golden evals and logprob readout stack

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants