Skip to content

Add the Score primitive and stop publishing nouls as confidence - #135

Merged
mrmps merged 1 commit into
mainfrom
codex/jev-score-primitive
Sep 25, 2026
Merged

mrmps merged 1 commit into
mainfrom
codex/jev-score-primitive

Conversation

@mrmps

@mrmps mrmps commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

src/jev.ts implemented two of Jev's three primitives. This adds Score and fixes the two noul/choice scale-mixing blemishes found while auditing whether the missing primitive hid a bug (it didn't — escalation was already gated !multi, deliberately).

Score, 1-1 with TypeSafe's API

  • Third Question variant: { type: "score", instructions, criteria: string[] } — ordered levels, lowest first (the API takes 2–10).
  • validAnswer: requires a finite score within [0, levels-1], a confidence, and a probability for every level index, keeping the "never trust a partial upstream response" contract.
  • questionCost: score levels get the same 16-token answer-echo term as choice options.
  • Gateway translation both directions, pinned to the evaluation-model spec v4 (EvaluationModelV4Question/Answer in vercel/ai, and the gateway's Jev declares all three types supported): a score is a score there too, same ordered criteria array; the answer carries score + level-indexed probabilities, no legend, with the calibrated confidence in providerMetadata.typesafe and the top level's probability as the stand-in — the analogue of a choice's winning option.
  • Live-verified: a score question through jevAsk against real TypeSafe (jev-1.13.0) round-tripped the exact documented shape through the packer and validator.

No public API change: the typesafe-compat endpoint was already a transparent pass-through (and openapi.ts already documented score there); this closes the gap in our own protocol layer so internal callers (jevAsk, skills-style questions) can ask graded questions.

Multi-label published a noul as confidence

jevClassify and runLaya returned confidence: scores[best] for multi-label — the max per-label noul. Per CONTEXT.md's own vocabulary, confidence "is unavailable when no comparable estimate exists," and a noul is not comparable to choice confidence. JevResult.confidence is now number | null and null for multi, matching what the public Result already said (index.ts was overriding it to null at the edge; now no layer manufactures one).

The uncertain digest counter mixed scales

record()'s uncertain counted every confidence === null row — so all multi-label traffic counted as uncertain, including answers whose nouls are all near zero, which are a confident "none apply". Extracted as countUncertain(): low choice confidence (< ESCALATE_BELOW, the scale the threshold was tuned on) plus escalated rows, which were measurably uncertain before the reasoning model re-asked them. Absence of a comparable estimate no longer counts. Telemetry-only; nothing branches on it.

Tests

  • Score: TypeSafe pass-through with raw validated answer, malformed rejection (missing level probability, out-of-range score), gateway vocabulary + metadata confidence, stand-in confidence without metadata.
  • Multi confidence null pinned on both the TypeSafe and gateway doors.
  • countUncertain semantics pinned.
  • Full suite: 822 pass / 0 fail, tsc --noEmit clean, CLI tests clean.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added support for score-based questions, including validation of scores and level probabilities.
    • Improved gateway handling of choice and score questions, with confidence fallback for score answers.
  • Updates

    • Multi-label results now report confidence as unavailable when no comparable confidence estimate exists.
    • Analytics now counts escalated results and results with confidence below 0.7 as uncertain.
    • Irrelevant screening results are skipped only when confidence and relevance scores meet the required threshold.

Jev's protocol layer implemented two of the model's three primitives.
Score — ordered levels, a fractional position, a probability per level —
now matches TypeSafe's API 1-1: a third Question variant, answer
validation (score within the levels, a probability for every level),
per-level packing cost, and both directions of the gateway translation.
The evaluation-model spec v4 names it score and takes the same ordered
criteria array; the gateway's answer carries no legend and its calibrated
confidence rides in provider metadata, with the top level's probability
as the stand-in, the analogue of a choice's winning option. Verified
against live TypeSafe through jevAsk.

Two scale-mixing blemishes found while auditing the noul/choice boundary:

- Multi-label published the max per-label noul under `confidence` in
  jevClassify and runLaya. Nouls are absolute yes/no probabilities, not
  distribution concentration, so by CONTEXT.md's own vocabulary no
  comparable estimate exists: it is now null, as the public Result
  already said.
- The digest's `uncertain` counter treated a null confidence as
  uncertain, so every multi-label answer counted — including one whose
  nouls are all near zero, a confident "none apply". It now counts only
  what ESCALATE_BELOW was tuned for: low choice confidence, plus
  escalated answers, which were measurably uncertain before the
  reasoning model re-asked them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 25, 2026

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 6e37c528-5aac-4e5d-b4ad-b149eb6fd29c

📥 Commits

Reviewing files that changed from the base of the PR and between e0413f8 and 6cd7e6e.

📒 Files selected for processing (7)
  • src/index.ts
  • src/jev.ts
  • src/laya.ts
  • src/long-context-job.ts
  • src/long-context.ts
  • test/jev.test.ts
  • test/record.test.ts
 __________________
< Zero bugs given. >
 ------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@mrmps
mrmps merged commit 43818f8 into main Sep 25, 2026
2 of 3 checks passed
@mrmps
mrmps deleted the codex/jev-score-primitive branch September 25, 2026 08:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant