Skip to content

feat(core): add auto reasoning effort coordinated with RedRouter - #373

Merged
filipeforattini merged 5 commits into
mainfrom
reasoning-auto
Sep 23, 2026
Merged

filipeforattini merged 5 commits into
mainfrom
reasoning-auto

Conversation

@filipeforattini

@filipeforattini filipeforattini commented Sep 23, 2026 •

Copy link
Copy Markdown

Summary

Adds an opt-in auto reasoning variant. With auto, every request carries one of the model's own effort variants, chosen per turn. The string auto is never sent to a provider. Redcode and RedRouter now agree on which one of them decides.

Choosing auto

  • TUI: the variant picker lists Auto first, and ctrl+t cycles through it, when the model has two or more effort levels (ReasoningAuto.options).
  • CLI: --variant auto. Saved and session values are kept when the model offers auto.
  • Agent config: variant: "auto".
  • The session stores auto. Subagents inherit the variant the person picked, not the level a turn applied.

decideEffort (packages/core/src/session/reasoning-auto.ts, pure)

  • Base level comes from deliberation = max(complexity, consequence), using the red-router bands (<.15 minimal, <.35 low, <.6 medium, <.8 high, else xhigh). Impact ≥ 2/3 raises the base to at least medium. Without an assessment, the level the session already has is kept (fail open); the first turn defaults to medium.
  • Trouble at the boundary, one step up that skips the dwell: feedback corrects or rejects, then frustration ≥ 2/3.
  • Heavy context, one step up, never on top of trouble: estimated input tokens above 50% of the window. Tool definitions that the count leaves out are added at 250 tokens each.
  • Plan agent or an explicit request to think ("think hard" or "pense bem", PT+EN, Unicode boundaries, code removed first): at least high.
  • One step down when the user agrees with a healthy turn, or when must_clarify ≥ 0.8.
  • Hysteresis: the level changes only at a user-message boundary and holds for 2 user turns. Inside one user turn it holds, except for one step up per turn when the loop guard fires, 2 or more tool failures come in a row, the S1 tool review reports failed_result, or a response-quality repair runs.
  • Bounds: the level is clamped to the model's own variants within reasoning.auto.floor/ceiling (default minimal..xhigh; none and max only when configured). A missing level snaps to the nearest variant; a tie goes to the one with more thought.
  • Per-session state (level, dwell, loop step, decider) lives in memory, keyed by session and bounded.

S1

  • New prompt question user_feedback: agrees | corrects | rejects | neutral, about how the latest user message judges the agent's previous turn.
  • promptContext adds a feedback line (optional, so older evaluations still render).
  • routerHint adds stall, feedback and frustration.
  • effortAssessment maps the answers to decideEffort inputs.

RedRouter coordination (ReasoningAuto.coordinate, following red-router#53)

setup decider wire
dual (S1) + RedRouter redcode concrete variant in the body; x-red-router-reasoning: off on direct models; <level> plus hint on auto/smart combos
single + RedRouter accepting auto router x-red-router-reasoning: auto plus hint (stall, and feedback when known), also on direct models; X-RedRouter-Reasoning is parsed and adopted
single + older RedRouter whose autopilot applies router no header
redcode alone redcode dual: S1; single: deterministic signals
manual or default variant nobody x-red-router-reasoning: off whenever a router with the reasoning capability is detected
  • Capabilities now map to new detection features: reasoning, reasoning-auto (accepts includes auto), reasoning-applies, and hint-signals (hint_keys lists stall, feedback and frustration).
  • validHint accepts effort, stall, feedback and frustration. Signal keys are dropped for routers that do not list them.
  • reported() parses X-RedRouter-Reasoning: <from|->-><level>; cause=…[; shadow]. Shadow reports are never adopted.
  • Legacy runtime: plain OpenAI-compatible servers are still never probed every turn. Only auto sessions, RedRouter's own connection, or models that a router described trigger a probe. The V2 runner reads the cached detection only.

Display

  • The prompt footer shows auto → high · frustration or router: medium, read from the session's metadata.reasoning. The legacy runtime writes it only when it changes.
  • Each assistant message footer shows auto → <level>: the assistant message's variant now records the concrete level its turn applied.

Runtimes

  • Legacy runtime: decides in prompt.ts and applies the level through the user passed to the processor, so request.ts merges a real variant.
  • V2: SessionRunnerModel.resolve(session, choose) resolves auto through the chooser (the default applies when there is no chooser). The runner decides once per turn attempt, and the published assistant model carries the concrete variant.

Notes on the red-router#53 contract

Tests

  • packages/core/test/session-reasoning-auto.test.ts: decideEffort tables (bands, feedback, frustration, impact, clarify, context, plan and explicit think, dwell, the in-turn loop step, variant clamping), explicit-think detection in PT and EN, progress extraction, options, the coordinate table, and session-state adopt and note.
  • provider-router.test.ts: capabilities to features, the reasoning header and hint filtering, the hint grammar, and X-RedRouter-Reasoning parsing.
  • intelligence.test.ts: the user_feedback question, hint signals, effortAssessment, and the feedback line.
  • session-runner-model.test.ts: auto resolves through the chooser, and falls back to the default without one.
  • redcode/test/session/llm.test.ts: a local Bun.serve RedRouter fixture that counts decision-model decisions, checking one decider per turn across the manual, dual-direct, dual-combo and single rows, and that the reported level is adopted.
  • CLI variant tests, and the TUI label tests (packages/tui/test/util/reasoning.test.ts).

I did not run tests or typecheck locally; CI validates.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

@filipeforattini
filipeforattini enabled auto-merge (squash) September 23, 2026 05:34
@filipeforattini
filipeforattini merged commit ac557c8 into main Sep 23, 2026
22 checks passed
@filipeforattini
filipeforattini deleted the reasoning-auto branch September 23, 2026 06:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant