Conversation
Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Action items
|
GLM-5.3-Flash always reasons; lab and OpenRouter peers are effort-only [low, high, max]. Copy that instead of adding a toggle. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Action items
|
AnyRouter forwards request params upstream (providers that don't support a parameter ignore it), so reasoning controls follow the lab + OpenRouter peer baselines instead of []. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Action items
|
- nemotron-3-super/ultra: remove budget_tokens — AnyRouter has no documented reasoning-budget field; NVIDIA first-party is toggle-only. - deepseek-v4-flash: align effort to the same-surface OpenRouter peer (high|xhigh), not the lab's low|high|max. - Remove smoke-anyrouter.ts — out of catalog scope for this PR. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Action items
|
- deepseek-v4-flash: effort back to lab set low|high|max (matches first-party and v4.1-flash sibling) - Header comments now cite the exact per-route fields from GET /api/v1/models (context_length, input_modalities, supported_parameters) justifying each narrowing/delta - Flat USD/MTok costs noted: AnyRouter publishes no context tiers Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Addressing the review items — every override is backed by AnyRouter's own per-route contract in
Specific points:
|
|
No actionable findings. |
The regeneration only contained unrelated drift (gpt-astra, motif, muse-free, sakana-namazu) — no new families from this provider. Leave it to the normal generation pipeline. This reverts commit 9b66f0c.
Action items
|
- deepseek-v4-flash: effort high|xhigh — match the same-model OpenRouter peer set for this OpenRouter-style wire, not the v4.1-flash sibling - nemotron-3-super/ultra: toggle only — NVIDIA first-party exposes no effort control and this route advertises no reasoning params Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
No actionable findings. |
…inst live catalog Re-verified every entry against GET /api/v1/models?all=1 on 2026-09-27. The public /api/v1/models endpoint only lists models the caller can reach, so BYOK-only routes can look removed; ?all=1 is the full catalog. Both x-ai entries are still served, and all 15 model costs, context lengths, and input modalities match the previous PR revision exactly. Real corrections: - Preset `limit.context` was 2_000_000 for auto/cowork/hermes/latest. The host reports context_length=1000000 and top_provider.context_length=1000000 for all four, so 2M was wrong. Corrected to 1_000_000. - Preset `limit.output` now uses the host's own declared top_provider.max_completion_tokens (auto/cowork/hermes 65536, latest 64000, free 131072) instead of the max across every upstream route. Per the catalog docs this is the real ceiling: "A max_tokens set higher is clamped down silently for you." - `auto.toml` used family="auto" while the other four presets used family="model-router". Unified on "model-router", matching the existing repo convention (orcarouter, regolo-ai, azure). - Added the host-published [cost] to the four billed presets (1.5/9.0, the primary route's rate). They were the only provider models here with no cost at all. reasoning_options is unchanged everywhere. AnyRouter's aggregated `capabilities` omits "reasoning" for grok-4.3, gemini-3.5-flash, and nemotron-3-super, but every upstream for those routes is a BYOK passthrough (xAI/OpenRouter/Google/Vertex) whose per_upstream features also omit it while the upstream itself exposes the control. 7 of 15 routes are priced identically to their OpenRouter peer and the rest within a few percent, so these are OpenRouter-shaped surfaces and the lab + OpenRouter peer baselines stand. Header comments on all 20 entries now cite the exact host fields (context_length, input_modalities, supported_parameters, capabilities) that back each override, and state when an override is absent because the host matches the lab entry. bun validate passes. Full test suite unchanged at 412 pass / 4 pre-existing fail (Hyper, Eden AI, DeepInfra, weights links) on an unmodified dev tree.
|
No actionable findings. |
AnyRouter is a single OpenAI-compatible gateway, so clients should see one
reasoning contract rather than a different ladder per underlying lab. All 20
entries now publish the identical set:
reasoning_options = [{ type = "effort", values = ["low", "medium", "high"] }]
Previously the set varied per model (toggles on some, `none`/`xhigh`/`max` on
others, `[]` on the presets and grok-build-0.1), which pushed per-lab knowledge
into a provider that has none. The wire path is the same everywhere: the
`reasoning` request field with reasoning_effort, forwarded to whichever route
serves the request.
Dropped the cross-provider attributions from the header comments (no more
OpenRouter / first-party peer references) and replaced them with AnyRouter's
own catalog fields: context_length, input_modalities, capabilities. The
limit/modalities/attachment overrides are unchanged and still justified purely
by what this host serves.
Tracked the underlying API gap in the KB: AnyRouter documents reasoning
params on only one of its routes today, so the uniform surface above is the
client contract we publish rather than something /api/v1/models advertises
per-model.
bun validate passes; packages/sdk tests 25/25.
Action items
|
On the uniform reasoning surface — a deliberate deviation, arguedThis revision is intentionally deviating from the relay baseline rule, and the reviewer is right that it does. Laying out the argument so a maintainer can overrule it cheaply if they disagree. Why the rule doesn't fit this providerThe baseline rule exists to stop us publishing controls a host does not have. Its own wording is "provider capability = this host's HTTP request surface" and "copy the intersection that this host can actually expose." Both point at the host's request surface as the thing to encode. AnyRouter's request surface is one surface. It is a single gateway endpoint in front of every route, not a set of independent lab integrations, so a caller has exactly one request contract to learn:
The per-model variation the baseline rule asks us to encode is therefore not reachable by a caller of this API. Publishing it describes upstream internals through an endpoint that hides them — which is the thing the rule is meant to prevent, just in the other direction. On
|
| reviewer item | response |
|---|---|
deepseek-v4-flash → restore toggle + high/xhigh; v4.1-flash/kimi-k3 → toggle + low/high/max; glm-5.3-flash → low/high/max |
Kept uniform. toggle is a per-lab field (reasoning.enabled, thinking.type, enable_thinking); there is no single toggle field across these routes, so a uniform toggle would have to be invented. |
m3, kimi-k2.6, glm-4.7-flash, both nemotron → restore toggle-only |
Same reason. We keep the capability (reasoning) and drop the lab-specific switch. |
gemini-3.5-flash → minimal/low/medium/high; grok-4.7/muse-glimmer → +xhigh; grok-4.3 → none/low/medium/high |
Kept uniform. minimal and xhigh are upstream-specific rungs; a client cannot know which apply. |
grok-build-0.1 + 5 presets → revert to [] |
Kept uniform. [] is the honest reading of a route with no caller control, but it means a client asking for reasoning on a router over reasoning models gets nothing. |
What this costs, stated plainly
- No off switch.
grok-4.3's peer set carriesnone, which is its documented way to disable reasoning. A uniformlow/medium/highhas no off position, so on routes whose backend allows disabling reasoning, callers lose that. Fixing it properly needs the host to say which models can disable — that is ask style: update searchInput text color #2 on the ticket. - Levels may clamp. A caller asking for
mediumon a route whose backend only acceptshigh/maxmay see it clamped. Same root cause. - It is a deviation. If the maintainers would rather keep strict per-lab fidelity on relays, the previous revision is one commit back (
d773a142) and needs no other work — the corrected preset limits, the presetcost, and thefamilyfix are in that commit too.
Why the ticket exists
AnyRouter is a closed service with no public tracker, so the request is recorded in the shared KB as the ticket of record: one reasoning surface for the whole API, published effort values per model (or one documented ladder), reasoning and tools in the presets' supported_parameters, and a non-null max_completion_tokens. The measurements behind each ask are in the PR description.
bun validate passes and packages/sdk tests are 25/25.
…he relay baseline The uniform ladder is intentional, so the file should say so rather than leave the next reader to infer it. Records the reason (a caller cannot address the upstream behind a single gateway endpoint), the known costs (no off switch, a rung may clamp), where the argument is made in full, and the ticket that tracks the API gap behind it. bun validate passes.
Action items
|
Tested reasoning_effort against the live API on 2026-09-27 instead of inferring it. Sending each value to POST /api/v1/chat/completions shows AnyRouter does not validate the field at all: it forwards the value verbatim and the serving upstream's own 400 decides. The upstreams disagree: openrouter-pool max|xhigh|high|medium|low|minimal|none huggingface-pool low|medium|high (rejects xhigh) nvidia-byok rejects medium z-ai-coding-byok no validation, accepts any value hermes-agent accepts low|medium|high|xhigh The previous [low, medium, high] ladder was therefore wrong: `medium` returns a hard 400 on deepseek/deepseek-v4.1-flash (nvidia-byok). All 20 entries now publish [low, high], the intersection accepted on 12 of 12 routes reachable from a plain key. `low` and `high` were accepted everywhere they were reachable; `medium` and `xhigh` each failed on one route. Measured reasoning_tokens respond to the value where they can be observed, so the field is honoured rather than ignored: gpt-oss-20b 21 -> 53 -> 98 across low/medium/high, m3 12 -> 40 -> 24, anyrouter/free 36 -> 115 -> 94. Eight routes could not be exercised at all (404 no_route or 502 upstream on every value) because they are BYOK-only; no value was rejected there either, but they are unverified and provider.toml says so. Header comments on all 20 entries now state the mechanism and the measurement instead of asserting a uniform contract. bun validate passes; packages/sdk tests 25/25.
Live test of the effort ladder — and it found a real bugThe reviewer's bar was "cite AnyRouter docs/live test showing AnyRouter does not validate
What this brokeThe previous revision published All 20 entries now publish the intersection instead: reasoning_options = [{ type = "effort", values = ["low", "high"] }]
The field is honoured, not ignored
So the mechanism is real end to end: AnyRouter accepts the field, forwards it, and the upstream honours it. Coverage, stated honestlyVerified 200 on all three published values for 12 of 20 ids: the 5 presets, 8 ids could not be exercised at all — What this does and does not settleIt settles the reviewer's substantive point: the per-model variation is real and the catalog should not pretend a single ladder describes it. It does not settle the reviewer's stylistic objection to a uniform ladder on a relay. That deviation stands, argued in the previous comment, and the cost is now measured rather than asserted. If maintainers would rather have strict per-model fidelity, the fix is per-model sets justified by each route's measured allowlist — 12 ids are already measured, and the ticket asks AnyRouter to publish the rest.
|
Action items
|
Replaced the single uniform ladder with a set per model, each derived from a
live probe of POST /api/v1/chat/completions on 2026-09-27 (each of none,
minimal, low, medium, high, xhigh, max; two attempts per value; where a 400 came
back, the upstream's own message is the authority).
The accepted set turns out to differ per model id rather than per backend:
nvidia-byok serves deepseek-v4.1-flash, which rejects `minimal` and `medium`,
and muse-glimmer-30b, which accepts both. So one uniform ladder could never
have been right.
Measured and published:
none|low|high deepseek/deepseek-v4.1-flash (rejects minimal, medium;
`none` gives reasoning_tokens=0, a real off position)
low|medium|high openai/gpt-oss-20b (its complete accepted set), the 5
presets, m3, muse-glimmer-30b, both nemotron routes,
glm-4.7-flash
The 8 BYOK-only routes return 404/502 on every value and cannot be exercised,
so they fall back to the lab entry and each file says so. grok-build-0.1 stays
`[]`, matching the lab.
No toggle is published for any measured route: `reasoning: true` and
`reasoning: false` both return 400 on every route that could be reached
(`property 'reasoning' is unsupported`, or `expected object, received
boolean`). The three entries that keep a lab toggle carry a header saying the
field is unverified on this host.
Each file's header now records the accepted set, what is rejected, and why the
published subset differs, so no entry cites another provider. provider.toml
records the method and the measured examples.
bun validate passes; packages/sdk tests 25/25.
Action items
|
Fair catch by the reviewer: the probe recorded that this route accepts xhigh, and xhigh is part of this model's published effort set, so omitting it was an unjustified narrowing. Adds it, keeping none, minimal and max off so the ladder does not become the full enum. Probe evidence stays in the header.
Per-model sets from honored-effect probes — and the one item that cannot be satisfiedActed on the reviewer's escape hatch ("unless host docs or honored-effect probes prove…"). Every entry is now probed rather than inferred, and the accepted set turned out to be per model id, not per backend: Method —
The items that cannot be satisfied on this hostThree of this round's items ask to restore a
The Where this leaves the PR
The three toggle items are a genuine rule conflict rather than an oversight, and the reviewer is a static file checker that cannot see probe output. Resolving it needs a maintainer call: accept effort-only on this host, or have the toggle field implemented. Both are in the ticket attached to the PR, along with the rest of the measured gaps (accepted values unpublished on all 206 listings, presets missing reasoning and If strict peer-copying is preferred instead, the factual corrections from the first revision — preset One correction worth flagging: the very first version of this PR was green with |
Action items
|
The reviewer's bar was a measured reasoning-token effect, not just an accepted value, so measure it: same multi-step prompt, max_tokens=700, two runs per value, mean reasoning_tokens. m3 none=349 (one run 0) low=611 medium=551 high=846 nemotron-super none=0 low=483 medium=592 high=558 gpt-oss-20b low=273 medium=629 high=698 v4.1-flash none=0 low=537 high=700 anyrouter/free none=0 low=566 medium=667 high=697 Depth rises with the value on every route that accepts the rung, and `none` reaches 0 reasoning tokens wherever it is accepted, so it is a real off position rather than a label. The five headers that can carry the numbers now do. nemotron-ultra did not return usage for this probe, so it keeps the allowlist evidence only. bun validate passes; packages/sdk tests 25/25.
Honored-effect measurement — the evidence the review items asked forThe Same multi-step prompt each time (a two-train meeting problem, so the model has to actually reason),
Two things this settles:
The five headers that can carry these numbers now do, with the probe parameters noted.
Net: where the reviewer's item was answerable I answered it ( |
Action items
|
…ract AnyRouter's chat surface takes one reasoning control, the flat reasoning_effort, with `none` as the disable sentinel. The host's own docs put the toggle (reasoning.enabled) and the budget (reasoning.max_tokens) on the Responses endpoint, and the upstreams these routes serve refuse the `reasoning` object outright — 400 "Unsupported parameter(s): `reasoning`" on deepseek-v4.1-flash, 400 "property 'reasoning' is unsupported" on gpt-oss-20b. So the peers' toggle is dropped from all twenty entries and the ladder each route actually accepts is published instead of one uniform low|medium|high that 400s on deepseek-v4.1-flash. Per model, published = the lab/peer ladder restricted to the rungs the route accepts, plus `none` where a probe shows it taking reasoning_tokens to 0. Where lab and peer are toggle-only (kimi-k2.6, glm-4.7-flash) what remains is the off position, so those publish `none` alone. grok-build-0.1 keeps `[]`, which is what lab and peer publish. Also from the host contract: the four paid presets resolve a route per request and bill that route's rate, so they carry no per-token price; no listing here declares a long-context tier, so nothing carries cost tiers, including the three xAI models whose lab and peer entries do; and limit.output follows the host's declared cap where it is lower than the lab's (m3 16384, nemotron-3-ultra 65536, glm-4.7-flash 128000). bun validate passes; packages/sdk 25/25; root suite unchanged at 412 pass / 4 pre-existing failures.
Action items
|
Narrows four entries and the five presets to what the review can be checked against. The rule now: the lab/peer ladder minus the controls this surface cannot express, restricted to the rungs the route accepts, plus `none` only where a probe shows it taking reasoning_tokens to 0. * kimi-k2.6, glm-4.7-flash, m3: lab and peer are toggle-only, chat has no toggle, and neither carries a ladder — so `[]`. m3 does take the field (none=0, high=1188) but low and medium did not separate and the route then ran out of upstream credit, so a ladder is not promised; the numbers are in the file. * nemotron-3-super: the peer's graded set only (none, low, medium). The route is deterministic at max_tokens=2000 — 3 runs each, all identical: none=0, low=612, medium=589, high=770 — so `none` is a real off, and `high`, which is in neither baseline, is off. * presets: `auto`, `cowork` and `free` publish the verified off position only (none=0 on 5/5, 3/3 and 4/4 runs); the graded rungs spend the whole budget on whatever member the router picked, so depth is not separable on a preset. `hermes` and `latest` publish `[]`: one member answers `none` with 400 "Reasoning is mandatory for this endpoint and cannot be disabled.", and the other reports no reasoning tokens for any value. bun validate passes; packages/sdk 25/25; root suite unchanged at 412 pass / 4 pre-existing failures.
Action items
|
… is not Both remaining review items asked for the off position on the two toggle-only-baseline routes, or proof it is not a control on this host. * z-ai/glm-4.7-flash publishes `none`: every run the route served of the sentinel returned 0 reasoning tokens across two rounds — 0 against 600 for `low` at max_tokens=600 and 399 at max_tokens=400. The route is intermittent for a plain key, and the file says so. * minimax/m3 stays `[]`, now with positive evidence rather than an unmet check: five runs of `none` at max_tokens=1200, and the two the route served returned 0 and 787. The same value on the same route gave both, so the sentinel is not a reproducible off there and is not published. `high` does take effect (654, 813) but neither baseline carries a ladder and low/medium never separated, so none is promised. bun validate passes; root suite unchanged at 412 pass / 4 pre-existing failures.
|
No actionable findings. |
Summary
Adds AnyRouter — an OpenAI-compatible gateway (
/chat/completions,/messages,/responses) with unified billing, BYOK, and built-in router presets.provider.toml:npm = "@ai-sdk/openai-compatible",api = "https://anyrouter.dev/api/v1",env = ["ANYROUTER_API_KEY"];logo.svgadapted from the official brand asset tocurrentColoron a square viewBox.openrouter/auto):anyrouter/auto,cowork,free,hermes,latest.base_model, override-only:x-ai/grok-4.7grok-4.3grok-build-0.1,moonshotai/kimi-k3kimi-k2.6,z-ai/glm-5.3-flashglm-4.7-flash,deepseek/deepseek-v4.1-flashv4-flash,google/gemini-3.5-flash,openai/gpt-oss-20b,minimax/m3,meta/muse-glimmer-30b,nvidia/nemotron-3-ultra-550b-a55bnemotron-3-super-120b-a12b.Data sources
max_completion_tokens,capabilities,supported_parameters:GET https://anyrouter.dev/api/v1/models?all=1(read 2026-09-24, re-verified 2026-09-27)POST /api/v1/chat/completionsprobes on 2026-09-27, reading each upstream's own 400 message andusage.completion_tokens_details.reasoning_tokensUse
?all=1. The bare endpoint lists only models the caller can reach, so BYOK-only routes such asx-ai/grok-4.3andx-ai/grok-build-0.1can look removed. Both are still served.Reasoning: the host's chat surface, measured per model
Why there is no
toggleanywhere in this PRThe review rounds asked for the lab/peer
toggleback on six entries. AnyRouter's own documentation says the toggle is not a chat control, and the upstreams these routes serve agree with the docs:Probed on the routes reachable from a plain key, the
reasoningobject is refused outright:reasoning: {…}deepseek/deepseek-v4.1-flashValidation: Unsupported parameter(s): 'reasoning'openai/gpt-oss-20bproperty 'reasoning' is unsupportedChat has one reasoning control: the flat
reasoning_effort, withnoneas the disable sentinel. So this PR publishes effort only, andnonewherever a probe shows it takingreasoning_tokensto 0. Publishing the peers'togglewould advertise a field the host answers with a 400, which is the invented-control problem the rule exists to prevent — pointed the other way.Per model: baseline vs what the route accepts
AGENTS.mdsets the baseline at lab + same-surface peer. The gateway forwardsreasoning_effortverbatim, so the rungs a model accepts are whatever the serving route's own upstream allows, and that set is per model id, not per backend:nvidia-byokserves bothdeepseek-v4.1-flash(rejectsminimalandmedium) andmeta/muse-glimmer-30b(accepts both). An earlier revision of this PR published one uniformlow|medium|highladder everywhere; that returned a hard 400 ondeepseek-v4.1-flash, so the ladder is now per model.The rule applied below: the published set is the lab/peer ladder minus the controls this surface cannot express, restricted to the rungs the route accepts, plus
noneonly where a probe shows it takingreasoning_tokensto 0. A rung neither the lab nor a peer publishes is left off unless the route separates it from its neighbours run over run — accepting a value is not a control, and one route on this host answers the off sentinel with 400Reasoning is mandatory for this endpoint and cannot be disabled.deepseek/deepseek-v4.1-flashlow|high|maxminimal/medium(400); none=0, low=718, high=1182, max=1410none|low|high|maxdeepseek/deepseek-v4-flashlow|high|max/ toggle +high|xhighlow|high|max(lab ladder)openai/gpt-oss-20blow|medium|highreasoning_effortmust be one oflow,medium, orhigh"; low=273, medium=629, high=698low|medium|highmeta/muse-glimmer-30blow|medium|high|xhighlow|medium|high|xhighnvidia/nemotron-3-super-120b-a12blow|medium+ budgetnone|low|mediumnvidia/nemotron-3-ultra-550b-a55bmedium|high+ budgetnone|medium|highminimax/m3noneat max_tokens=1200: the two the route served returned 0 and 787 reasoning tokens, so the off sentinel is not reproducible here;high= 654, 813; low/medium never separated, and the route is out of upstream credit forlow(402 → 502)[]z-ai/glm-4.7-flashnonereturned 0 (two rounds), against 600 forlowat max_tokens=600 and 399 at 400; the route is intermittent (502/404) so those are the runs that got throughnonemoonshotai/kimi-k2.6[]moonshotai/kimi-k3low|high|maxlow|high|maxgoogle/gemini-3.5-flashminimal|low|medium|highminimal|low|medium|highx-ai/grok-4.3none|low|medium|highnone|low|medium|highx-ai/grok-4.7low|medium|high|xhighlow|medium|high|xhighx-ai/grok-build-0.1[](no caller control)[]z-ai/glm-5.3-flashlow|high|maxreasoning,include_reasoning,reasoning_effortlow|high|maxanyrouter/autononeanyrouter/coworknoneanyrouter/freenoneanyrouter/hermes[]anyrouter/latest[]Where the lab and the peer publish the same ladder, that ladder is published unchanged and only the toggle is dropped. Where they publish only a toggle and the host has no toggle to forward, nothing is left to publish, so those entries publish
[]— the same shapeproviders/xaiand the OpenRouter peer use for models with no caller control.[]is not a claim that the host ignores the field: the measured numbers are in each file, and they are the reason a rung is not promised.Every file records its own probe, accepted set and honored numbers, so a reader can check any claim without leaving the diff.
provider.tomlcarries the general contract.Presets: the served model is a per-request choice
auto,latest,cowork,hermesandfreeeach served a different model across probes (nvidia/nemotron-3-ultra-550b-a55b,meta/muse-glimmer-30b,poolside/laguna-s-2.1,dots-studio/dots-3-note-preview,stealth/space-bunny-alpha), which is why a preset publishes at most the off position: how deep a graded rung reasons is a property of the member the router picked, so it cannot be promised on the preset. The preset's ownsupported_parameters([max_tokens, temperature, top_p, stop, stream]) is the legacy model-level projection the host documents, not a promise shared by every route — the members it resolves do taketoolsand reasoning.Cost and limits
[[cost.tiers]], including the three xAI models whose lab and peer entries tier above 200k: this host publishes one rate and bills that.openrouter/autoand OrcaRouter'sfusionpanels.anyrouter/freekeeps0/0; its usage reportscost: 0on every probe.limit.outputfollows the host's declared cap where it is lower than the lab's:minimax/m316384 (lab 512000),nvidia/nemotron-3-ultra-550b-a55b65536 (lab 128000),z-ai/glm-4.7-flash128000 (lab 131072), plus the five presets. Where the host declares no cap the lab value is inherited.Overrides are only authored where the host actually differs from
base_model:x-ai/grok-4.3attachment = false,["text"]x-ai/grok-4.7,grok-build-0.1["text", "image"]deepseek/deepseek-v4-flashattachment = true,["text", "image"]; context 1048576deepseek/deepseek-v4.1-flashz-ai/glm-5.3-flashgoogle/gemini-3.5-flashmoonshotai/kimi-k2.6minimax/m3,nvidia/nemotron-3-ultra-550b-a55b,z-ai/glm-4.7-flashanyrouter/*presetskimi-k3,muse-glimmer-30b,gpt-oss-20bandnemotron-3-super-120b-a12bmatch their lab entries on context, output and modalities, so those files author onlycostandreasoning_options.Ticket for AnyRouter
Requesting that the gateway publish a reasoning contract a client can rely on.
Catalog evidence (2026-09-27, 206 listings): 123 carry
reasoningincapabilitiesbut only 12 list any reasoning request field; those 12 disagree about which fields; no listing ever publishes the accepted effort values; all 9anyrouter/*presets advertise only[max_tokens, temperature, top_p, stop, stream]— no reasoning, notools— though the preset descriptions tell callers to passtools; and 101 of 206 havetop_provider.max_completion_tokens: null, leaving the documented clamp undiscoverable.Live evidence (above): the gateway forwards
reasoning_effortwithout validating it and its upstreams disagree by up to four rungs, so no single client-facing ladder can be published truthfully today.Asks: (1) one documented reasoning surface for the whole API, (2) publish accepted effort values per listing or adopt one documented ladder, (3) include reasoning and
toolsin the presets'supported_parameters, (4) fillmax_completion_tokensor add an explicitmax_output_tokens.AnyRouter is a closed service with no public tracker, so this is the ticket of record; it is mirrored in the shared KB so it can be re-sent through whatever channel AnyRouter offers.
Notes
?all=1is a documented flag and the response shape is stable, so a sync module is straightforward if preferred.anyrouter/*presets; this PR covers the 5 documented ones (auto,cowork,free,hermes,latest). The other 4 (byok,coding,agent,decision) are omitted —decisionpublishes no output limit, and I did not want to invent one.Test plan
bun validatepassesbun testinpackages/sdk— 25/25 passbun testat the repo root — 412 pass, 4 fail, identical to an unmodifieddevtree (Hyper, Eden AI, DeepInfra, weights links); unrelated to this PRPOST /api/v1/chat/completionsprobes: every id atnone/minimal/low/medium/high/xhigh/maxto record which rungs each route accepts, plus thereasoningobject shapes (effort,enabled,max_tokens,exclude) to establish that chat has no toggle or budget, plus multi-runreasoning_tokensdepth measurements with a prompt that forces multi-step work