Add the Score primitive and stop publishing nouls as confidence - #135
Merged
Merged
Conversation
Jev's protocol layer implemented two of the model's three primitives. Score — ordered levels, a fractional position, a probability per level — now matches TypeSafe's API 1-1: a third Question variant, answer validation (score within the levels, a probability for every level), per-level packing cost, and both directions of the gateway translation. The evaluation-model spec v4 names it score and takes the same ordered criteria array; the gateway's answer carries no legend and its calibrated confidence rides in provider metadata, with the top level's probability as the stand-in, the analogue of a choice's winning option. Verified against live TypeSafe through jevAsk. Two scale-mixing blemishes found while auditing the noul/choice boundary: - Multi-label published the max per-label noul under `confidence` in jevClassify and runLaya. Nouls are absolute yes/no probabilities, not distribution concentration, so by CONTEXT.md's own vocabulary no comparable estimate exists: it is now null, as the public Result already said. - The digest's `uncertain` counter treated a null confidence as uncertain, so every multi-label answer counted — including one whose nouls are all near zero, a confident "none apply". It now counts only what ESCALATE_BELOW was tuned for: low choice confidence, plus escalated answers, which were measurably uncertain before the reasoning model re-asked them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
src/jev.tsimplemented two of Jev's three primitives. This adds Score and fixes the two noul/choice scale-mixing blemishes found while auditing whether the missing primitive hid a bug (it didn't — escalation was already gated!multi, deliberately).Score, 1-1 with TypeSafe's API
Questionvariant:{ type: "score", instructions, criteria: string[] }— ordered levels, lowest first (the API takes 2–10).validAnswer: requires a finitescorewithin[0, levels-1], a confidence, and a probability for every level index, keeping the "never trust a partial upstream response" contract.questionCost: score levels get the same 16-token answer-echo term as choice options.EvaluationModelV4Question/Answerin vercel/ai, and the gateway's Jev declares all three types supported): a score is a score there too, same ordered criteria array; the answer carriesscore+ level-indexedprobabilities, no legend, with the calibrated confidence inproviderMetadata.typesafeand the top level's probability as the stand-in — the analogue of a choice's winning option.jevAskagainst real TypeSafe (jev-1.13.0) round-tripped the exact documented shape through the packer and validator.No public API change: the typesafe-compat endpoint was already a transparent pass-through (and
openapi.tsalready documented score there); this closes the gap in our own protocol layer so internal callers (jevAsk, skills-style questions) can ask graded questions.Multi-label published a noul as
confidencejevClassifyandrunLayareturnedconfidence: scores[best]for multi-label — the max per-label noul. Per CONTEXT.md's own vocabulary, confidence "is unavailable when no comparable estimate exists," and a noul is not comparable to choice confidence.JevResult.confidenceis nownumber | nulland null for multi, matching what the publicResultalready said (index.ts was overriding it to null at the edge; now no layer manufactures one).The
uncertaindigest counter mixed scalesrecord()'suncertaincounted everyconfidence === nullrow — so all multi-label traffic counted as uncertain, including answers whose nouls are all near zero, which are a confident "none apply". Extracted ascountUncertain(): low choice confidence (< ESCALATE_BELOW, the scale the threshold was tuned on) plus escalated rows, which were measurably uncertain before the reasoning model re-asked them. Absence of a comparable estimate no longer counts. Telemetry-only; nothing branches on it.Tests
countUncertainsemantics pinned.tsc --noEmitclean, CLI tests clean.🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Updates