Skip to content

calibration: surface reliability curves in the advisor + trend, and evaluate ladder replacement #8227

Description

@JSONbored

Problem

Epic #8211 track E capstone. Curves are only useful where operators look, and the candidate-ladder question ('should suggestions replace hand-picked candidates?') is an authority-adjacent decision needing evidence side-by-side first.

Requirements

⚠️ Required pattern. Surfacing first, authority later: the advisor's knob recs and the internal trend/knobs endpoints show each live knob's current curve summary + derived suggestion NEXT TO the ladder-based proposal; a soak period comparing the two on real corpora precedes any evaluator change, with the comparison recorded on this issue before the decision.

  • Aggregates only on every surface; curve buckets are numbers, never examples.
  • If the decision lands on replacement: ladder stays as the hard bound set (suggestion clamps to candidates' range) — evidence proposes, bounds constrain.

Deliverables

  • Advisor + endpoint surfacing with tests
  • Soak comparison posted here + the go/no-go recorded

Links & Resources

#8211 (epic), the curves sub-issue (blocked-by), src/review/loosening-recs.ts, src/services/rule-calibration-trend.ts

Boundaries

No evaluator behavior change without the recorded soak decision.

maintainer-only — evaluator-authority decision.

Metadata

Metadata

Assignees

Labels

maintainer-onlyOwner-only work — yields no Gittensor points.roadmapOn the Wave-2 agent-layer roadmap board (project 9)

Projects

No projects

Relationships

None yet

Development

No branches or pull requests

Issue actions