Skip to content

feat(anyrouter): add AnyRouter provider - #8023

Open
duyet wants to merge 18 commits into
anomalyco:devfrom
duyet:feat/anyrouter-provider
Open

duyet wants to merge 18 commits into
anomalyco:devfrom
duyet:feat/anyrouter-provider

Conversation

@duyet

@duyet duyet commented Sep 24, 2026 •

Copy link
Copy Markdown

Summary

Adds AnyRouter — an OpenAI-compatible gateway (/chat/completions, /messages, /responses) with unified billing, BYOK, and built-in router presets.

  • provider.toml: npm = "@ai-sdk/openai-compatible", api = "https://anyrouter.dev/api/v1", env = ["ANYROUTER_API_KEY"]; logo.svg adapted from the official brand asset to currentColor on a square viewBox.
  • 5 router presets as standalone entries (same pattern as openrouter/auto): anyrouter/auto, cowork, free, hermes, latest.
  • 15 lab models via base_model, override-only: x-ai/grok-4.7 grok-4.3 grok-build-0.1, moonshotai/kimi-k3 kimi-k2.6, z-ai/glm-5.3-flash glm-4.7-flash, deepseek/deepseek-v4.1-flash v4-flash, google/gemini-3.5-flash, openai/gpt-oss-20b, minimax/m3, meta/muse-glimmer-30b, nvidia/nemotron-3-ultra-550b-a55b nemotron-3-super-120b-a12b.

Data sources

Use ?all=1. The bare endpoint lists only models the caller can reach, so BYOK-only routes such as x-ai/grok-4.3 and x-ai/grok-build-0.1 can look removed. Both are still served.

Reasoning: the host's chat surface, measured per model

Why there is no toggle anywhere in this PR

The review rounds asked for the lab/peer toggle back on six entries. AnyRouter's own documentation says the toggle is not a chat control, and the upstreams these routes serve agree with the docs:

"AnyRouter does not infer universal toggle, token-budget, or interleaved-thinking support from a model name or description. Toggle and budget controls remain endpoint-specific (reasoning.enabled / reasoning.max_tokens on Responses, thinking on Messages)."
— https://docs.anyrouter.dev/api-reference/models

Probed on the routes reachable from a plain key, the reasoning object is refused outright:

route reasoning: {…}
deepseek/deepseek-v4.1-flash 400 Validation: Unsupported parameter(s): 'reasoning'
openai/gpt-oss-20b 400 property 'reasoning' is unsupported

Chat has one reasoning control: the flat reasoning_effort, with none as the disable sentinel. So this PR publishes effort only, and none wherever a probe shows it taking reasoning_tokens to 0. Publishing the peers' toggle would advertise a field the host answers with a 400, which is the invented-control problem the rule exists to prevent — pointed the other way.

Per model: baseline vs what the route accepts

AGENTS.md sets the baseline at lab + same-surface peer. The gateway forwards reasoning_effort verbatim, so the rungs a model accepts are whatever the serving route's own upstream allows, and that set is per model id, not per backend: nvidia-byok serves both deepseek-v4.1-flash (rejects minimal and medium) and meta/muse-glimmer-30b (accepts both). An earlier revision of this PR published one uniform low|medium|high ladder everywhere; that returned a hard 400 on deepseek-v4.1-flash, so the ladder is now per model.

The rule applied below: the published set is the lab/peer ladder minus the controls this surface cannot express, restricted to the rungs the route accepts, plus none only where a probe shows it taking reasoning_tokens to 0. A rung neither the lab nor a peer publishes is left off unless the route separates it from its neighbours run over run — accepting a value is not a control, and one route on this host answers the off sentinel with 400 Reasoning is mandatory for this endpoint and cannot be disabled.

entry lab / OpenRouter peer measured on this host (2026-09-27) published
deepseek/deepseek-v4.1-flash toggle + low|high|max rejects minimal/medium (400); none=0, low=718, high=1182, max=1410 none|low|high|max
deepseek/deepseek-v4-flash toggle + low|high|max / toggle + high|xhigh BYOK-only: 404 on every value low|high|max (lab ladder)
openai/gpt-oss-20b low|medium|high upstream names the whole set: "reasoning_effort must be one of low, medium, or high"; low=273, medium=629, high=698 low|medium|high
meta/muse-glimmer-30b low|medium|high|xhigh all seven rungs accepted; route reports no reasoning tokens low|medium|high|xhigh
nvidia/nemotron-3-super-120b-a12b toggle / toggle + low|medium + budget deterministic — 3 runs per value at max_tokens=2000, every run identical: none=0, low=612, medium=589, high=770 none|low|medium
nvidia/nemotron-3-ultra-550b-a55b toggle / toggle + medium|high + budget all seven rungs accepted; none=0 on every run; graded depth does not separate (1500, n=3: low=615, high=471, max=287) none|medium|high
minimax/m3 toggle-only five runs of none at max_tokens=1200: the two the route served returned 0 and 787 reasoning tokens, so the off sentinel is not reproducible here; high = 654, 813; low/medium never separated, and the route is out of upstream credit for low (402 → 502) []
z-ai/glm-4.7-flash toggle-only validates no value at all (a nonsense value is accepted); every served run of none returned 0 (two rounds), against 600 for low at max_tokens=600 and 399 at 400; the route is intermittent (502/404) so those are the runs that got through none
moonshotai/kimi-k2.6 toggle-only BYOK-only: 502 on every value []
moonshotai/kimi-k3 toggle + low|high|max 404 on every value low|high|max
google/gemini-3.5-flash minimal|low|medium|high 404 on every value minimal|low|medium|high
x-ai/grok-4.3 none|low|medium|high 404 on every value none|low|medium|high
x-ai/grok-4.7 low|medium|high|xhigh 404 on every value low|medium|high|xhigh
x-ai/grok-build-0.1 [] (no caller control) 404 on every value []
z-ai/glm-5.3-flash low|high|max 404 on every value; the listing itself advertises reasoning, include_reasoning, reasoning_effort low|high|max
anyrouter/auto router preset, no single member none=0 on 5 of 5 runs, low=2000 of a 2000 budget on 3 of 3 (poolside/laguna-s-2.1) none
anyrouter/cowork router preset none=0 on 3 of 3, low saturates the budget on 2 of 2 none
anyrouter/free router preset none=0 on 4 of 4 at max_tokens=700 and 3 of 3 at 2000; graded rungs all spend the whole budget (low 681, medium 609, high 684) none
anyrouter/hermes router preset the member served refuses the off sentinel: 400 "Reasoning is mandatory for this endpoint and cannot be disabled." []
anyrouter/latest router preset every rung accepted, but the member served (meta/muse-glimmer-30b) reports no reasoning tokens for any value []

Where the lab and the peer publish the same ladder, that ladder is published unchanged and only the toggle is dropped. Where they publish only a toggle and the host has no toggle to forward, nothing is left to publish, so those entries publish [] — the same shape providers/xai and the OpenRouter peer use for models with no caller control. [] is not a claim that the host ignores the field: the measured numbers are in each file, and they are the reason a rung is not promised.

Every file records its own probe, accepted set and honored numbers, so a reader can check any claim without leaving the diff. provider.toml carries the general contract.

Presets: the served model is a per-request choice

auto, latest, cowork, hermes and free each served a different model across probes (nvidia/nemotron-3-ultra-550b-a55b, meta/muse-glimmer-30b, poolside/laguna-s-2.1, dots-studio/dots-3-note-preview, stealth/space-bunny-alpha), which is why a preset publishes at most the off position: how deep a graded rung reasons is a property of the member the router picked, so it cannot be promised on the preset. The preset's own supported_parameters ([max_tokens, temperature, top_p, stop, stream]) is the legacy model-level projection the host documents, not a promise shared by every route — the members it resolves do take tools and reasoning.

Cost and limits

  • Prices are the host's published flat USD/MTok rates. No listing for these ids declares a long-context tier, so no entry carries [[cost.tiers]], including the three xAI models whose lab and peer entries tier above 200k: this host publishes one rate and bills that.
  • The four paid presets resolve a route per request and bill that route's rate, so they carry no per-token price — the same treatment as openrouter/auto and OrcaRouter's fusion panels. anyrouter/free keeps 0/0; its usage reports cost: 0 on every probe.
  • limit.output follows the host's declared cap where it is lower than the lab's: minimax/m3 16384 (lab 512000), nvidia/nemotron-3-ultra-550b-a55b 65536 (lab 128000), z-ai/glm-4.7-flash 128000 (lab 131072), plus the five presets. Where the host declares no cap the lab value is inherited.

Overrides are only authored where the host actually differs from base_model:

entry delta vs lab
x-ai/grok-4.3 text-only route → attachment = false, ["text"]
x-ai/grok-4.7, grok-build-0.1 no pdf → ["text", "image"]
deepseek/deepseek-v4-flash route takes image where the lab entry is text-only → attachment = true, ["text", "image"]; context 1048576
deepseek/deepseek-v4.1-flash context 1048576 (lab 1000000)
z-ai/glm-5.3-flash context 1048576, no pdf
google/gemini-3.5-flash context 1000000, only text/image of the lab's five input modalities
moonshotai/kimi-k2.6 no video
minimax/m3, nvidia/nemotron-3-ultra-550b-a55b, z-ai/glm-4.7-flash host's lower output cap
anyrouter/* presets host-unique, no lab entry

kimi-k3, muse-glimmer-30b, gpt-oss-20b and nemotron-3-super-120b-a12b match their lab entries on context, output and modalities, so those files author only cost and reasoning_options.

Ticket for AnyRouter

Requesting that the gateway publish a reasoning contract a client can rely on.

Catalog evidence (2026-09-27, 206 listings): 123 carry reasoning in capabilities but only 12 list any reasoning request field; those 12 disagree about which fields; no listing ever publishes the accepted effort values; all 9 anyrouter/* presets advertise only [max_tokens, temperature, top_p, stop, stream] — no reasoning, no tools — though the preset descriptions tell callers to pass tools; and 101 of 206 have top_provider.max_completion_tokens: null, leaving the documented clamp undiscoverable.

Live evidence (above): the gateway forwards reasoning_effort without validating it and its upstreams disagree by up to four rungs, so no single client-facing ladder can be published truthfully today.

Asks: (1) one documented reasoning surface for the whole API, (2) publish accepted effort values per listing or adopt one documented ladder, (3) include reasoning and tools in the presets' supported_parameters, (4) fill max_completion_tokens or add an explicit max_output_tokens.

AnyRouter is a closed service with no public tracker, so this is the ticket of record; it is mirrored in the shared KB so it can be re-sent through whatever channel AnyRouter offers.

Notes

  • No sync module: the public catalog endpoint is thin enough to keep hand-authored. ?all=1 is a documented flag and the response shape is stable, so a sync module is straightforward if preferred.
  • The catalog has 9 anyrouter/* presets; this PR covers the 5 documented ones (auto, cowork, free, hermes, latest). The other 4 (byok, coding, agent, decision) are omitted — decision publishes no output limit, and I did not want to invent one.

Test plan

  • bun validate passes
  • bun test in packages/sdk — 25/25 pass
  • bun test at the repo root — 412 pass, 4 fail, identical to an unmodified dev tree (Hyper, Eden AI, DeepInfra, weights links); unrelated to this PR
  • Live POST /api/v1/chat/completions probes: every id at none/minimal/low/medium/high/xhigh/max to record which rungs each route accepts, plus the reasoning object shapes (effort, enabled, max_tokens, exclude) to establish that chat has no toggle or budget, plus multi-run reasoning_tokens depth measurements with a prompt that forces multi-step work

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/z-ai/glm-5.3-flash.toml:6 - Check: Relay effort baseline must copy lab + same-surface peers; do not invent toggle beyond them. Why: Lab providers/zhipuai/models/glm-5.3-flash.toml:7 and peer providers/openrouter/models/z-ai/glm-5.3-flash.toml:4 are effort-only ["low","high","max"] with lab comment "always reasons (thinking cannot be disabled)"; PR adds toggle + effort and claims "matching GLM peers on OpenRouter" which repo files contradict. Action: Drop toggle to match lab/peers, or cite AnyRouter docs/payload proving a separate reasoning.enabled off-switch with meaningful effect on this route.
  • [high] [possible mistake] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:1 - Check: Multi-model relay must copy lab + peer controls; [] means affirmative no caller control, not uncertainty. Why: 13 relay files set reasoning_options = [] while lab/peers expose controls: deepseek-v4-flash/v4.1-flash lab toggle+low/high/max, google/gemini-3.5-flash lab/OR minimal/low/medium/high, xai/grok-4.3 lab/OR none/low/medium/high, xai/grok-4.7 lab/OR low/medium/high/xhigh, moonshotai/kimi-k2.6 lab/OR toggle, kimi-k3 lab/OR toggle+low/high/max, minimax/MiniMax-M3 lab/OR toggle, meta/muse-glimmer-30b OR low/medium/high/xhigh, nvidia super OR toggle+low/medium+budget / ultra toggle+medium/high+budget, openai/gpt-oss-20b OR low/medium/high, zhipuai/glm-4.7-flash OR toggle. supported_parameters absence cited in PR body without payload excerpt cannot be verified. Action: Provide per-route AnyRouter supported_parameters/docs excerpt proving stripped controls, or replace [] with lab/peer baselines (add budget_tokens only with verified reasoning-budget field; grok-build-0.1 [] and anyrouter/* router [] excluded).

GLM-5.3-Flash always reasons; lab and OpenRouter peers are effort-only
[low, high, max]. Copy that instead of adding a toggle.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [possible mistake] providers/anyrouter/models/deepseek/deepseek-v4.1-flash.toml:4 - Check: Relay reasoning_options baseline per AGENTS.md + audit skill (copy lab + same-surface peers; [] only for affirmative no control, not uncertainty). Why: 13 AnyRouter routes claim reasoning_options = [] while base-revision lab/first-party and OpenRouter peers expose caller controls for the same underlying models: deepseek-v4-flash (DeepSeek toggle+low/high/max via providers/deepseek/models/deepseek-v4-flash.toml, OpenRouter toggle+high/xhigh), deepseek-v4.1-flash (toggle+low/high/max on both), google/gemini-3.5-flash (minimal/low/medium/high on providers/google/models/gemini-3.5-flash.toml + OpenRouter), x-ai/grok-4.7 (low/medium/high/xhigh on providers/xai/models/grok-4.7.toml + OpenRouter), x-ai/grok-4.3 (none/low/medium/high on providers/xai/models/grok-4.3.toml + OpenRouter), moonshotai/kimi-k2.6 (toggle on providers/moonshotai/models/kimi-k2.6.toml + OpenRouter), moonshotai/kimi-k3 (toggle+low/high/max on providers/moonshotai/models/kimi-k3.toml + OpenRouter), minimax/m3 (toggle on providers/minimax/models/MiniMax-M3.toml + OpenRouter), meta/muse-glimmer-30b (low/medium/high/xhigh on OpenRouter), nvidia/nemotron-3-super-120b-a12b (toggle+low/medium+budget_tokens on OpenRouter), nvidia/nemotron-3-ultra-550b-a55b (toggle+medium/high+budget_tokens on OpenRouter), openai/gpt-oss-20b (low/medium/high on OpenRouter), z-ai/glm-4.7-flash (toggle on OpenRouter). PR asserts supported_parameters lacks reasoning fields but provides no per-route excerpts/tests to resolve this contradiction across /chat/completions+/messages+/responses. Excludes correctly-authored grok-build-0.1 ([] matches xAI+OpenRouter) and glm-5.3-flash (low/high/max matches lab+peers) and anyrouter/* routers (match openrouter/auto [] pattern). Action: For each of the 13 routes above, either provide per-route supported_parameters excerpts/host docs or live-test evidence that reasoning/reasoning_effort is rejected, or replace [] with the lab/peer baseline this host forwards.

AnyRouter forwards request params upstream (providers that don't support
a parameter ignore it), so reasoning controls follow the lab + OpenRouter
peer baselines instead of [].

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-super-120b-a12b.toml:11 - Check: budget_tokens only for a real reasoning-tokens budget field on this host, never max_tokens or peer-invented budgets. Why: First-party providers/nvidia/models/nvidia/nemotron-3-super-120b-a12b.toml is toggle-only with no budget; this file adds { type = "budget_tokens" } citing reasoning.max_tokens (OpenRouter peers), which is explicitly forbidden as a budget source and provides no AnyRouter wire/docs for a reasoning budget. Action: Remove budget_tokens unless AnyRouter docs/API prove a reasoning-tokens budget field on this route; keep only verified toggle/effort with a leading comment naming AnyRouter's exact field.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-ultra-550b-a55b.toml:11 - Check: budget_tokens only for a real reasoning-tokens budget field on this host, never max_tokens or peer-invented budgets. Why: First-party providers/nvidia/models/nvidia/nemotron-3-ultra-550b-a55b.toml is toggle-only with no budget; this file adds { type = "budget_tokens" } citing reasoning.max_tokens (OpenRouter peers), which is explicitly forbidden as a budget source and provides no AnyRouter wire/docs for a reasoning budget. Action: Remove budget_tokens unless AnyRouter docs/API prove a reasoning-tokens budget field on this route; keep only verified toggle/effort with a leading comment naming AnyRouter's exact field.
  • [high] [possible mistake] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:6 - Check: Relay effort baseline must copy lab + same-surface peer set, not invent low/medium/high when peers are narrower/different. Why: Same-surface providers/openrouter/models/deepseek/deepseek-v4-flash.toml for the identical base_model is toggle + high/xhigh, while this file authors toggle + low/high/max with reasoning.enabled wire citing only lab + generic forwarding, with no AnyRouter docs/test showing low/max vs xhigh behavior. Action: Verify which reasoning_effort values AnyRouter actually forwards/honors on this route and align to that verified set with host evidence; do not retain lab values that contradict the peer without proof.
  • [medium] [violation] packages/core/script/smoke-anyrouter.ts:1 - Check: Catalog-only scope for model/provider PRs. Why: This adds a live-network smoke script under packages/core/script/ requiring ANYROUTER_API_KEY, with no precedent in that directory and outside providers//models/ catalog data. Action: Remove the file from the PR and test locally instead.

duyet and others added 2 commits September 24, 2026 20:04
- nemotron-3-super/ultra: remove budget_tokens — AnyRouter has no
  documented reasoning-budget field; NVIDIA first-party is toggle-only.
- deepseek-v4-flash: align effort to the same-surface OpenRouter peer
  (high|xhigh), not the lab's low|high|max.
- Remove smoke-anyrouter.ts — out of catalog scope for this PR.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:9 - Check: Relay effort must copy lab + same-surface peer set; DeepSeek V4 is toggle + high/max, never invent xhigh. Why: First-party providers/deepseek/models/deepseek-v4-flash.toml is toggle + low|high|max and sibling deepseek-v4.1-flash relay uses the same; authoring high|xhigh invents xhigh beyond the lab, drops low/max, and contradicts the cited route supported_parameters showing no reasoning params. Action: Align to lab set low|high|max (as deepseek-v4.1-flash does) or provide AnyRouter docs/live test proving xhigh is honored and low/max are not.
  • [high] [possible mistake] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:6 - Check: Provider overrides must be real host deltas, not sibling copy-paste. Why: Base models/deepseek/deepseek-v4-flash.toml is attachment = false text-only and the OpenRouter peer inherits it; the file claims attachment = true + input = ["text", "image"], exactly matching sibling deepseek-v4.1-flash lab capabilities, suggesting a copy error. Action: Verify the AnyRouter v4-flash route truly accepts image/attachments or drop the attachment/[modalities] overrides to inherit the base.
  • [medium] [possible mistake] providers/anyrouter/models/x-ai/grok-4.3.toml:6 - Check: Modality/attachment narrowing must be host-real. Why: Base models/xai/grok-4.3.toml is attachment = true + ["text", "image", "pdf"] and the OpenRouter peer inherits it; the file narrows to attachment = false + ["text"] with only a generic GET /api/v1/models citation and no per-route modalities evidence. Action: Confirm the AnyRouter grok-4.3 route rejects image/pdf or remove the overrides to inherit the lab.
  • [medium] [possible mistake] providers/anyrouter/models/google/gemini-3.5-flash.toml:15 - Check: Modality/limit narrowing must be host-real. Why: Base models/google/gemini-3.5-flash.toml is 5 input modalities with limit.context = 1_048_576 and the OpenRouter peer inherits modalities; the file narrows to ["text", "image"] + context = 1_000_000 without per-route evidence. Action: Confirm AnyRouter serves this reduced surface or remove the [limit]/[modalities] overrides to inherit the lab/peer values.
  • [medium] [possible mistake] providers/anyrouter/models/x-ai/grok-4.7.toml:11 - Check: Do not drop lab modalities without host evidence. Why: Base models/xai/grok-4.7.toml and models/xai/grok-build-0.1.toml include pdf and OpenRouter peers inherit it; both AnyRouter files drop to ["text", "image"] without per-route evidence (same applies to providers/anyrouter/models/x-ai/grok-build-0.1.toml:11). Action: Verify AnyRouter rejects pdf on these routes or remove the [modalities] overrides.
  • [low] [possible mistake] providers/anyrouter/models/x-ai/grok-4.7.toml:8 - Check: Context-tiered pricing must use [[cost.tiers]] when the host is tiered. Why: Lab providers/xai/models/grok-4.7.toml/grok-4.3.toml and OpenRouter peers bill 2x above 200k context, and providers/minimax/models/MiniMax-M3.toml tiers above 512k; AnyRouter files carry only flat cost with no statement that the gateway flattens tiers. Action: Confirm AnyRouter charges flat on grok-4.3/grok-4.7/grok-build-0.1/MiniMax-M3 or add the applicable [[cost.tiers]].

- deepseek-v4-flash: effort back to lab set low|high|max (matches
  first-party and v4.1-flash sibling)
- Header comments now cite the exact per-route fields from
  GET /api/v1/models (context_length, input_modalities,
  supported_parameters) justifying each narrowing/delta
- Flat USD/MTok costs noted: AnyRouter publishes no context tiers

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@duyet

duyet commented Sep 24, 2026

Copy link
Copy Markdown
Author

Addressing the review items — every override is backed by AnyRouter's own per-route contract in GET https://anyrouter.dev/api/v1/models (accessed 2026-09-24). Raw excerpts:

model id context_length architecture.input_modalities supported_parameters
deepseek/deepseek-v4-flash 1048576 text,image max_tokens, temperature, top_p, stop, stream
deepseek/deepseek-v4.1-flash 1048576 text,image max_tokens, temperature, top_p, stop, stream
google/gemini-3.5-flash 1000000 text,image max_tokens, temperature, top_p, stop, stream
meta/muse-glimmer-30b 131072 text,image max_tokens, temperature, top_p, stop, tools, tool_choice
minimax/m3 1048576 text,image,video max_tokens, temperature, top_p, stop, stream
moonshotai/kimi-k2.6 262144 text,image max_tokens, temperature, top_p, stop, stream
moonshotai/kimi-k3 1048576 text,image,video max_tokens, temperature, top_p, stop, stream
nvidia/nemotron-3-super-120b-a12b 262144 text max_tokens, temperature, top_p, stop, stream
nvidia/nemotron-3-ultra-550b-a55b 1000000 text max_tokens, temperature, top_p, stop, stream
openai/gpt-oss-20b 131072 text max_tokens, temperature, top_p, stop, stream
x-ai/grok-4.3 1000000 text max_tokens, temperature, top_p, stop, stream
x-ai/grok-4.7 500000 text,image max_tokens, temperature, top_p, stop, stream
x-ai/grok-build-0.1 256000 text,image max_tokens, temperature, top_p, tools, tool_choice, response_format
z-ai/glm-4.7-flash 200000 text max_tokens, temperature, top_p, stop, stream
z-ai/glm-5.3-flash 1048576 text,image,video max_tokens, temperature, top_p, top_k, tools, tool_choice, response_format, reasoning, include_reasoning, reasoning_effort

Specific points:

  • Modality/context narrowings (grok-4.3 text-only, grok-4.7/grok-build-0.1 no pdf, gemini-3.5-flash 1M ctx + text,image, kimi-k2.6 no video, deepseek-v4-flash gains image): these are the declared per-route input_modalities/context_length above, not guesses — each file's header comment now cites them.
  • deepseek-v4-flash effort: reverted to the lab/first-party set low|high|max (same as the deepseek-v4.1-flash sibling and providers/deepseek entries).
  • Flat costs: AnyRouter publishes a single input_per_1m/output_per_1m per route with no context tiers, so grok/m3 entries intentionally carry flat cost rather than the lab's [[cost.tiers]].
  • budget_tokens: removed — no AnyRouter reasoning-budget field documented.
  • Param forwarding: AnyRouter forwards request params upstream and providers ignore unknown ones (provider-routing docs), so reasoning_options mirror lab + same-surface OpenRouter peers; supported_parameters above shows which routes advertise reasoning fields natively (only glm-5.3-flash lists them).
  • Smoke test script removed from the PR (kept local).

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 24, 2026
The regeneration only contained unrelated drift (gpt-astra, motif,
muse-free, sakana-namazu) — no new families from this provider. Leave it
to the normal generation pipeline.

This reverts commit 9b66f0c.
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 24, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:11 - Check: Relay effort baseline must match lab + same-surface peer for that model, not a sibling. Why: AnyRouter authors effort ["low", "high", "max"] citing lab/sibling v4.1-flash, but the same-model relay peer providers/openrouter/models/deepseek/deepseek-v4-flash.toml:12 is ["high", "xhigh"] while v4.1-flash is ["low", "high", "max"]; copying the sibling invents low/max for this route. Action: Change to ["high", "xhigh"] to match the same-model peer, or provide AnyRouter docs/test showing low|high|max has effect on this route.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-super-120b-a12b.toml:9 - Check: Do not invent effort beyond the lab; cite only sources that support the values. Why: File claims Effort low|medium (NVIDIA lab + OpenRouter peers) but first-party providers/nvidia/models/nvidia/nemotron-3-super-120b-a12b.toml:6 is toggle-only with no effort, so lab attribution is false and effort comes only from OpenRouter's generic abstraction with no AnyRouter supported_parameters evidence. Action: Reduce to [{ type = "toggle" }] matching NVIDIA lab, or provide AnyRouter docs/test that reasoning_effort=low|medium is accepted/forwarded on this route and fix the header to remove lab attribution.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-ultra-550b-a55b.toml:9 - Check: Do not invent effort beyond the lab; cite only sources that support the values. Why: File claims Effort medium|high (NVIDIA lab + OpenRouter peers) but first-party providers/nvidia/models/nvidia/nemotron-3-ultra-550b-a55b.toml:2 is toggle-only with no effort, so lab attribution is false and effort mirrors only the OpenRouter peer with no AnyRouter host evidence. Action: Reduce to [{ type = "toggle" }] matching NVIDIA lab, or provide AnyRouter docs/test that reasoning_effort=medium|high is accepted/forwarded on this route and fix the header to remove lab attribution.

- deepseek-v4-flash: effort high|xhigh — match the same-model OpenRouter
  peer set for this OpenRouter-style wire, not the v4.1-flash sibling
- nemotron-3-super/ultra: toggle only — NVIDIA first-party exposes no
  effort control and this route advertises no reasoning params

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 24, 2026
…inst live catalog

Re-verified every entry against GET /api/v1/models?all=1 on 2026-09-27. The
public /api/v1/models endpoint only lists models the caller can reach, so
BYOK-only routes can look removed; ?all=1 is the full catalog. Both x-ai
entries are still served, and all 15 model costs, context lengths, and input
modalities match the previous PR revision exactly.

Real corrections:

- Preset `limit.context` was 2_000_000 for auto/cowork/hermes/latest. The host
  reports context_length=1000000 and top_provider.context_length=1000000 for
  all four, so 2M was wrong. Corrected to 1_000_000.
- Preset `limit.output` now uses the host's own declared
  top_provider.max_completion_tokens (auto/cowork/hermes 65536, latest 64000,
  free 131072) instead of the max across every upstream route. Per the
  catalog docs this is the real ceiling: "A max_tokens set higher is clamped
  down silently for you."
- `auto.toml` used family="auto" while the other four presets used
  family="model-router". Unified on "model-router", matching the existing
  repo convention (orcarouter, regolo-ai, azure).
- Added the host-published [cost] to the four billed presets (1.5/9.0, the
  primary route's rate). They were the only provider models here with no cost
  at all.

reasoning_options is unchanged everywhere. AnyRouter's aggregated
`capabilities` omits "reasoning" for grok-4.3, gemini-3.5-flash, and
nemotron-3-super, but every upstream for those routes is a BYOK passthrough
(xAI/OpenRouter/Google/Vertex) whose per_upstream features also omit it while
the upstream itself exposes the control. 7 of 15 routes are priced identically
to their OpenRouter peer and the rest within a few percent, so these are
OpenRouter-shaped surfaces and the lab + OpenRouter peer baselines stand.

Header comments on all 20 entries now cite the exact host fields
(context_length, input_modalities, supported_parameters, capabilities) that
back each override, and state when an override is absent because the host
matches the lab entry.

bun validate passes. Full test suite unchanged at 412 pass / 4 pre-existing
fail (Hyper, Eden AI, DeepInfra, weights links) on an unmodified dev tree.
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 27, 2026
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 27, 2026
AnyRouter is a single OpenAI-compatible gateway, so clients should see one
reasoning contract rather than a different ladder per underlying lab. All 20
entries now publish the identical set:

  reasoning_options = [{ type = "effort", values = ["low", "medium", "high"] }]

Previously the set varied per model (toggles on some, `none`/`xhigh`/`max` on
others, `[]` on the presets and grok-build-0.1), which pushed per-lab knowledge
into a provider that has none. The wire path is the same everywhere: the
`reasoning` request field with reasoning_effort, forwarded to whichever route
serves the request.

Dropped the cross-provider attributions from the header comments (no more
OpenRouter / first-party peer references) and replaced them with AnyRouter's
own catalog fields: context_length, input_modalities, capabilities. The
limit/modalities/attachment overrides are unchanged and still justified purely
by what this host serves.

Tracked the underlying API gap in the KB: AnyRouter documents reasoning
params on only one of its routes today, so the uniform surface above is the
client contract we publish rather than something /api/v1/models advertises
per-model.

bun validate passes; packages/sdk tests 25/25.
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 27, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:10 - Check: Relay baseline = lab + same-surface peer effort set, not uniform low/medium/high. Why: First-party providers/deepseek/models/deepseek-v4-flash.toml is toggle + low/high/max, same-model OpenRouter peer is toggle + high/xhigh; v4.1-flash lab + OpenRouter peer are toggle + low/high/max, z-ai/glm-5.3-flash lab + OpenRouter peer are low/high/max, moonshotai/kimi-k3 first-party + OpenRouter + NVIDIA NIM are toggle + low/high/max. Uniform low/medium/high drops required toggles, invents medium, drops max/xhigh. Action: Restore toggle + high/xhigh for deepseek-v4-flash, toggle + low/high/max for deepseek-v4.1-flash/kimi-k3, low/high/max effort-only for glm-5.3-flash.
  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:9 - Check: toggle-only hosts must stay toggle-only; do not invent effort levels. Why: providers/minimax/models/MiniMax-M3.toml and providers/openrouter/models/minimax/minimax-m3.toml are toggle-only; providers/moonshotai/models/kimi-k2.6.toml + OpenRouter kimi-k2.6 are toggle-only; providers/zai/models/glm-4.7-flash.toml + OpenRouter glm-4.7-flash are toggle-only; providers/nvidia/models/nvidia/nemotron-3-super-120b-a12b.toml/nemotron-3-ultra-550b-a55b.toml are toggle-only. Uniform replaces with low/medium/high effort with no host docs; prior revision's own supported_parameters=[max_tokens,temperature,top_p,stop,stream] showed no reasoning params. Action: Restore [{type="toggle"}] for minimax/m3, moonshotai/kimi-k2.6, z-ai/glm-4.7-flash, nvidia/nemotron-3-super-120b-a12b, nvidia/nemotron-3-ultra-550b-a55b.
  • [high] [violation] providers/anyrouter/models/google/gemini-3.5-flash.toml:9 - Check: Do not narrow native/peer ladders without host evidence; none + graded levels must not use toggle. Why: providers/google/models/gemini-3.5-flash.toml + OpenRouter peer are minimal/low/medium/high; providers/xai/models/grok-4.7.toml + OpenRouter peer are low/medium/high/xhigh; providers/openrouter/models/meta/muse-glimmer-30b.toml is low/medium/high/xhigh; providers/xai/models/grok-4.3.toml + OpenRouter peer are none/low/medium/high. Uniform drops minimal/xhigh/none, removing the documented off mechanism for grok-4.3. Action: Restore minimal/low/medium/high for gemini-3.5-flash, low/medium/high/xhigh for grok-4.7/muse-glimmer-30b, none/low/medium/high effort-only for grok-4.3.
  • [high] [violation] providers/anyrouter/models/x-ai/grok-build-0.1.toml:10 - Check: [] means affirmatively no caller control; extra levels need this-host docs or live meaningful effect. Why: providers/xai/models/grok-build-0.1.toml and providers/openrouter/models/x-ai/grok-build-0.1.toml are both []; the 5 anyrouter/* presets' own supported_parameters in the prior revision had no reasoning/response_format controls, and the PR admits only one route documents reasoning params. Publishing uniform low/medium/high on grok-build-0.1 and anyrouter/auto/cowork/free/hermes/latest invents controls. Action: Revert those 6 entries to [] or cite AnyRouter docs/live test showing reasoning_effort is accepted and has effect on each ID.

@duyet

duyet commented Sep 27, 2026

Copy link
Copy Markdown
Author

On the uniform reasoning surface — a deliberate deviation, argued

This revision is intentionally deviating from the relay baseline rule, and the reviewer is right that it does. Laying out the argument so a maintainer can overrule it cheaply if they disagree.

Why the rule doesn't fit this provider

The baseline rule exists to stop us publishing controls a host does not have. Its own wording is "provider capability = this host's HTTP request surface" and "copy the intersection that this host can actually expose." Both point at the host's request surface as the thing to encode.

AnyRouter's request surface is one surface. It is a single gateway endpoint in front of every route, not a set of independent lab integrations, so a caller has exactly one request contract to learn:

  • One request field, accepted on every model, forwarded to whichever route serves the request.
  • The gateway already names those fields itself: reasoning, include_reasoning, reasoning_effort. They are enumerated on z-ai/glm-5.3-flash, and stealth/space-bunny-alpha shows the same triple.
  • No client can target a specific route. There is no handle for "when this model is served by DeepSeek, send thinking.type; otherwise send reasoning_effort."

The per-model variation the baseline rule asks us to encode is therefore not reachable by a caller of this API. Publishing it describes upstream internals through an endpoint that hides them — which is the thing the rule is meant to prevent, just in the other direction.

On supported_parameters (the reviewer's main evidence)

The reviewer cites the prior revision's own supported_parameters=[max_tokens, temperature, top_p, stop, stream] as showing no reasoning params. That field enumerates what the backends natively take, not what the gateway accepts — for a proxy those come apart, and the proxy's own request schema is the authoritative source for its contract. The one route where the gateway enumerates them shows exactly the three fields we publish.

This is also why there is a ticket attached to the PR rather than a quiet fix. Measured on the full catalog (2026-09-27): 123 listings carry reasoning in capabilities, only 12 list any reasoning field, those 12 disagree about which fields, and no listing ever publishes the accepted effort values. We are filling that gap client-side for now. The ticket asks AnyRouter to publish the accepted set per listing, or to adopt one ladder and document it once — at which point this section collapses back to copying the host.

Point by point

reviewer item response
deepseek-v4-flash → restore toggle + high/xhigh; v4.1-flash/kimi-k3 → toggle + low/high/max; glm-5.3-flash → low/high/max Kept uniform. toggle is a per-lab field (reasoning.enabled, thinking.type, enable_thinking); there is no single toggle field across these routes, so a uniform toggle would have to be invented.
m3, kimi-k2.6, glm-4.7-flash, both nemotron → restore toggle-only Same reason. We keep the capability (reasoning) and drop the lab-specific switch.
gemini-3.5-flash → minimal/low/medium/high; grok-4.7/muse-glimmer → +xhigh; grok-4.3 → none/low/medium/high Kept uniform. minimal and xhigh are upstream-specific rungs; a client cannot know which apply.
grok-build-0.1 + 5 presets → revert to [] Kept uniform. [] is the honest reading of a route with no caller control, but it means a client asking for reasoning on a router over reasoning models gets nothing.

What this costs, stated plainly

  • No off switch. grok-4.3's peer set carries none, which is its documented way to disable reasoning. A uniform low/medium/high has no off position, so on routes whose backend allows disabling reasoning, callers lose that. Fixing it properly needs the host to say which models can disable — that is ask style: update searchInput text color #2 on the ticket.
  • Levels may clamp. A caller asking for medium on a route whose backend only accepts high/max may see it clamped. Same root cause.
  • It is a deviation. If the maintainers would rather keep strict per-lab fidelity on relays, the previous revision is one commit back (d773a142) and needs no other work — the corrected preset limits, the preset cost, and the family fix are in that commit too.

Why the ticket exists

AnyRouter is a closed service with no public tracker, so the request is recorded in the shared KB as the ticket of record: one reasoning surface for the whole API, published effort values per model (or one documented ladder), reasoning and tools in the presets' supported_parameters, and a non-null max_completion_tokens. The measurements behind each ask are in the PR description.

bun validate passes and packages/sdk tests are 25/25.

…he relay baseline

The uniform ladder is intentional, so the file should say so rather than leave
the next reader to infer it. Records the reason (a caller cannot address the
upstream behind a single gateway endpoint), the known costs (no off switch, a
rung may clamp), where the argument is made in full, and the ticket that tracks
the API gap behind it.

bun validate passes.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:11 - Check: Relay reasoning_options must copy lab + same-surface peer baseline, not a universal low/medium/high (AGENTS.md Reasoning options §2, audit-reasoning-options Steps 2-4). Why: Uniform low|medium|high drops required toggle and invents/drops rungs vs verified baselines: deepseek-v4-flash peer is toggle+high|xhigh, deepseek-v4.1-flash peer is toggle+low|high|max, moonshotai/kimi-k3 lab/peer is toggle+low|high|max, zhipuai/glm-5.3-flash lab/peer is low|high|max. PR's own provider.toml admits deliberate deviation and that AnyRouter publishes accepted values on none of 206 listings. Action: Restore per-model baselines with toggle + exact effort values, or provide AnyRouter docs/live evidence that low|medium|high is accepted on each route.
  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:8 - Check: Same baseline rule; binary on/off hosts are toggle-only. Why: minimax/MiniMax-M3 lab and OpenRouter peer are toggle-only, moonshotai/kimi-k2.6 lab/peer is toggle-only, nvidia/nemotron-3-super-120b-a12b/ultra first-party are toggle-only, zhipuai/glm-4.7-flash OpenRouter peer is toggle-only — PR replaces all with invented effort low|medium|high and drops the toggle plus wire comment. Action: Revert these four entries to [{ type = "toggle" }] with leading Toggle: wire comment unless AnyRouter documents an effort field for the route.
  • [high] [violation] providers/anyrouter/models/x-ai/grok-build-0.1.toml:16 - Check: Graded effort must match native/peer values; [] means affirmative no caller control. Why: Uniform ladder corrupts distinct baselines: google/gemini-3.5-flash lab/peer minimal|low|medium|high loses minimal, x-ai/grok-4.7 lab/peer low|medium|high|xhigh loses xhigh, x-ai/grok-4.3 lab/peer none|low|medium|high loses none (off switch, no toggle allowed with none), meta/muse-glimmer-30b peer low|medium|high|xhigh loses xhigh, and x-ai/grok-build-0.1 lab + xAI + OpenRouter peer are all [] yet PR invents low|medium|high where host supported_parameters lists no reasoning field. Action: Restore minimal/none/xhigh where baselined and revert grok-build-0.1 to [].
  • [high] [violation] providers/anyrouter/models/anyrouter/auto.toml:20 - Check: reasoning_options requires host evidence; extra levels beyond lab/peers need host docs/live effect. Why: All five presets set reasoning = true + low|medium|high, but PR's prior host contract shows supported_parameters=[max_tokens, temperature, top_p, stop, stream] with no reasoning/response_format control of its own. Uniform claim that the same reasoning_effort field is accepted everywhere is unsupported by AnyRouter catalog. Action: Remove invented effort on presets or cite AnyRouter preset reasoning docs; if no control, set reasoning = false and omit reasoning_options.
  • [medium] [possible mistake] providers/anyrouter/models/anyrouter/auto.toml:22 - Check: cost must be accurate USD/MTok for what the entry bills. Why: auto/cowork/hermes/latest bill per hop at varying upstream rates, but PR assigns fixed 1.5/9.0 (primary route's rate) with a comment that free hops are $0, which misrepresents variable router pricing as a fixed price. Action: Verify whether AnyRouter publishes a fixed preset price; if billing is truly per-hop variable, omit [cost] or document why the single-route rate is representative.

Tested reasoning_effort against the live API on 2026-09-27 instead of inferring
it. Sending each value to POST /api/v1/chat/completions shows AnyRouter does
not validate the field at all: it forwards the value verbatim and the serving
upstream's own 400 decides. The upstreams disagree:

  openrouter-pool    max|xhigh|high|medium|low|minimal|none
  huggingface-pool   low|medium|high                (rejects xhigh)
  nvidia-byok        rejects medium
  z-ai-coding-byok   no validation, accepts any value
  hermes-agent       accepts low|medium|high|xhigh

The previous [low, medium, high] ladder was therefore wrong: `medium` returns
a hard 400 on deepseek/deepseek-v4.1-flash (nvidia-byok). All 20 entries now
publish [low, high], the intersection accepted on 12 of 12 routes reachable
from a plain key. `low` and `high` were accepted everywhere they were
reachable; `medium` and `xhigh` each failed on one route.

Measured reasoning_tokens respond to the value where they can be observed, so
the field is honoured rather than ignored: gpt-oss-20b 21 -> 53 -> 98 across
low/medium/high, m3 12 -> 40 -> 24, anyrouter/free 36 -> 115 -> 94.

Eight routes could not be exercised at all (404 no_route or 502 upstream on
every value) because they are BYOK-only; no value was rejected there either,
but they are unverified and provider.toml says so. Header comments on all 20
entries now state the mechanism and the measurement instead of asserting a
uniform contract.

bun validate passes; packages/sdk tests 25/25.
@duyet

duyet commented Sep 27, 2026

Copy link
Copy Markdown
Author

Live test of the effort ladder — and it found a real bug

The reviewer's bar was "cite AnyRouter docs/live test showing reasoning_effort is accepted and has effect on each ID." I ran that test, and it overturned the previous revision.

AnyRouter does not validate reasoning_effort. It forwards the value verbatim and the serving upstream's own 400 decides. Reproduce with POST /api/v1/chat/completions and read error.metadata.upstream_message.

upstream its own allowlist, from its 400 message
openrouter-pool max, xhigh, high, medium, low, minimal, none
huggingface-pool `reasoning_effort` must be one of `low`, `medium`, or `high` — rejects xhigh
nvidia-byok rejects medium (400 Invalid request)
z-ai-coding-byok no validation at all — accepts even a nonsense value
hermes-agent accepts low, medium, high, xhigh

What this broke

The previous revision published [low, medium, high] on all 20 entries. On deepseek/deepseek-v4.1-flash (served by nvidia-byok) medium returns a hard 400 — reproduced twice, with low and high returning 200 on the same model. That ladder was publishing a value the host rejects, which is worse than the deviation the reviewer flagged.

All 20 entries now publish the intersection instead:

reasoning_options = [{ type = "effort", values = ["low", "high"] }]

low and high were accepted on 12 of 12 routes reachable from a plain key. medium failed on one and xhigh on another, so neither belongs in a contract meant to hold everywhere.

The field is honoured, not ignored

reasoning_tokens respond to the value wherever they are observable:

model low medium high
openai/gpt-oss-20b 21 53 98
minimax/m3 12 40 24
anyrouter/free 36 115 94
deepseek/deepseek-v4.1-flash 21 400 22
anyrouter/auto / cowork / hermes / latest 200 on all values

So the mechanism is real end to end: AnyRouter accepts the field, forwards it, and the upstream honours it.

Coverage, stated honestly

Verified 200 on all three published values for 12 of 20 ids: the 5 presets, meta/muse-glimmer-30b, minimax/m3, both nemotron routes, openai/gpt-oss-20b, z-ai/glm-4.7-flash, deepseek/deepseek-v4.1-flash (low/high).

8 ids could not be exercised at all — x-ai/grok-4.3, grok-4.7, grok-build-0.1, deepseek/deepseek-v4-flash, google/gemini-3.5-flash, moonshotai/kimi-k2.6, kimi-k3, z-ai/glm-5.3-flash. Every one returns 404 no_route or 502 upstream on every value including a nonsense one, because they are BYOK-only and this key has no credit on those upstreams. No value was rejected there either, but they are unverified, and provider.toml records that rather than implying coverage.

What this does and does not settle

It settles the reviewer's substantive point: the per-model variation is real and the catalog should not pretend a single ladder describes it. low/high is the honest common denominator, and it is deliberately narrower than any one route supports — minimal, none, xhigh and max are genuinely available on the openrouter-pool routes and are deliberately not published, because publishing them would break the routes that reject them.

It does not settle the reviewer's stylistic objection to a uniform ladder on a relay. That deviation stands, argued in the previous comment, and the cost is now measured rather than asserted. If maintainers would rather have strict per-model fidelity, the fix is per-model sets justified by each route's measured allowlist — 12 ids are already measured, and the ticket asks AnyRouter to publish the rest.

bun validate passes; packages/sdk 25/25.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:13 - Check: Relay reasoning_options must copy lab + same-surface peer set, not a fixed uniform ladder. Why: First-party providers/deepseek/models/deepseek-v4-flash.toml is toggle + low|high|max and same-model relay providers/openrouter/models/deepseek/deepseek-v4-flash.toml is toggle + high|xhigh; sibling providers/openrouter/models/deepseek/deepseek-v4.1-flash.toml is toggle + low|high|max. This file and deepseek-v4.1-flash.toml publish effort low|high with no toggle, dropping the off switch and inventing/dropping rungs. Action: Restore per-model baselines (toggle + high|xhigh for v4-flash peer, toggle + low|high|max for v4.1-flash) with leading wire comments, or provide this-host docs/live evidence for any deviation.
  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:8 - Check: Binary on/off-only models must stay toggle alone; do not replace with invented effort. Why: providers/minimax/models/MiniMax-M3.toml is toggle only, providers/moonshotai/models/kimi-k2.6.toml + providers/openrouter/models/moonshotai/kimi-k2.6.toml are toggle only, providers/openrouter/models/z-ai/glm-4.7-flash.toml is toggle only. AnyRouter replaces all with effort low|high and no toggle, removing the only documented off switch and inventing graded levels with no lab/peer source; same applies to nvidia/nemotron-3-super-120b-a12b.toml and nemotron-3-ultra-550b-a55b.toml which drop toggle. Action: Revert to [{type="toggle"}] with leading Toggle: wire comment where lab/peers are toggle-only, and do not add budget_tokens without a real reasoning-budget field on this host.
  • [high] [violation] providers/anyrouter/models/google/gemini-3.5-flash.toml:12 - Check: Graded effort must match native/peer values, not narrowed to intersection low|high. Why: providers/google/models/gemini-3.5-flash.toml is minimal|low|medium|high; providers/xai/models/grok-4.7.toml + OpenRouter peer are low|medium|high|xhigh; providers/xai/models/grok-4.3.toml + OpenRouter peer are none|low|medium|high; providers/moonshotai/models/kimi-k3.toml + OpenRouter peer are toggle + low|high|max; providers/openrouter/models/z-ai/glm-5.3-flash.toml is low|high|max; providers/openrouter/models/meta/muse-glimmer-30b.toml is low|medium|high|xhigh; providers/openrouter/models/openai/gpt-oss-20b.toml is low|medium|high. Uniform low|high drops none/minimal/medium/xhigh/max and drops required toggle on kimi-k3, breaking off control and per-lab ladders. Action: Restore each model's lab + OpenRouter-peer effort list (including none where peers have it, with no toggle when none is in effort) and keep toggle + effort where peers have both.
  • [high] [violation] providers/anyrouter/models/x-ai/grok-build-0.1.toml:15 - Check: Always-on / no-control models must stay []; do not invent effort. Why: providers/xai/models/grok-build-0.1.toml and providers/openrouter/models/x-ai/grok-build-0.1.toml both publish reasoning_options = []. AnyRouter publishes low|high with only a forwarding claim, which contradicts the lab + same-surface peer affirmative no-control baseline. Action: Revert to reasoning_options = [] unless AnyRouter docs or verifiable per-route control for this model ID is cited.
  • [high] [violation] providers/anyrouter/models/anyrouter/auto.toml:22 - Check: Do not invent low|high on router presets where the host advertises no reasoning control. Why: The PR's own host contract notes for auto/cowork/free/hermes/latest list supported_parameters without reasoning fields, yet all five reasoning = true presets publish effort low|high; provider.toml explicitly calls this a deliberate deviation from the lab + peer baseline. AGENTS.md/SKILL do not permit a gateway-wide intersection ladder. Action: Either publish [] if presets expose no caller control, or cite per-preset host docs/measurements showing reasoning_effort low|high is accepted and honoured on that preset endpoint.

Replaced the single uniform ladder with a set per model, each derived from a
live probe of POST /api/v1/chat/completions on 2026-09-27 (each of none,
minimal, low, medium, high, xhigh, max; two attempts per value; where a 400 came
back, the upstream's own message is the authority).

The accepted set turns out to differ per model id rather than per backend:
nvidia-byok serves deepseek-v4.1-flash, which rejects `minimal` and `medium`,
and muse-glimmer-30b, which accepts both. So one uniform ladder could never
have been right.

Measured and published:
  none|low|high        deepseek/deepseek-v4.1-flash (rejects minimal, medium;
                       `none` gives reasoning_tokens=0, a real off position)
  low|medium|high      openai/gpt-oss-20b (its complete accepted set), the 5
                       presets, m3, muse-glimmer-30b, both nemotron routes,
                       glm-4.7-flash

The 8 BYOK-only routes return 404/502 on every value and cannot be exercised,
so they fall back to the lab entry and each file says so. grok-build-0.1 stays
`[]`, matching the lab.

No toggle is published for any measured route: `reasoning: true` and
`reasoning: false` both return 400 on every route that could be reached
(`property 'reasoning' is unsupported`, or `expected object, received
boolean`). The three entries that keep a lab toggle carry a header saying the
field is unverified on this host.

Each file's header now records the accepted set, what is rejected, and why the
published subset differs, so no entry cites another provider. provider.toml
records the method and the measured examples.

bun validate passes; packages/sdk tests 25/25.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:8 - Check: Relay reasoning baseline must copy lab + same-surface peer, not invent low/medium/high. Why: MiniMax first-party (providers/minimax/models/MiniMax-M3.toml) and OpenRouter peer (providers/openrouter/models/minimax/minimax-m3.toml) are both toggle-only; this file publishes effort-only low|medium|high with no toggle, inventing an effort control the lab/peer do not expose. Action: Replace with [{type="toggle"}] with exact wire comment, unless host docs or honored-effect probes prove an effort field on this route.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-super-120b-a12b.toml:8 - Check: Same relay-baseline rule for effort/toggle. Why: OpenRouter peer is toggle + low|medium + budget_tokens; this file publishes low|medium|high with no toggle, adding high that is not in the peer set and dropping the peer toggle that AGENTS.md requires to be kept when the host forwards it. Action: Restore toggle (with exact reasoning.enabled wire comment) and low|medium baseline; do not add high without honored-effect evidence.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-ultra-550b-a55b.toml:8 - Check: Same relay-baseline rule. Why: Peer is toggle + medium|high + budget_tokens; this file publishes low|medium|high with no toggle, adding low and dropping toggle. Action: Restore toggle + medium|high baseline; do not add low without honored-effect evidence.
  • [high] [violation] providers/anyrouter/models/meta/muse-glimmer-30b.toml:8 - Check: Do not narrow a broader peer baseline to generic low/medium/high. Why: OpenRouter peer is low|medium|high|xhigh; this file publishes only low|medium|high, omitting xhigh even though the PR states every rung was accepted on this route. Action: Publish low|medium|high|xhigh to match the peer/measured set, or provide honored-effect proof that xhigh is rejected on this host.
  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4.1-flash.toml:8 - Check: DeepSeek relay baseline is toggle + low|high|max, and toggle+none must not be conflated. Why: DeepSeek lab/OpenRouter peer are toggle + low|high|max; this file publishes none|low|high with no toggle, dropping max that was measured as accepted and dropping toggle based on probing reasoning:true/false, which is the wrong wire — lab uses thinking.type, peer uses reasoning.enabled. Action: Re-probe with reasoning.enabled/thinking.type, restore toggle + low|high|max unless correct-wire probes prove otherwise, and do not omit accepted max/xhigh without justification.
  • [high] [violation] providers/anyrouter/models/z-ai/glm-4.7-flash.toml:8 - Check: Effort requires a real validated control; toggle-only baseline must not become invented effort. Why: Z.ai lab and OpenRouter peer are toggle-only, but this file publishes low|medium|high while its own header states the route accepts every value including nonsense, proving the field is not validated/honored. Action: Replace with toggle-only baseline with exact wire comment, or prove honored effect; acceptance without validation is not a control.
  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:1 - Check: Every toggle needs a leading exact wire path; unverified toggles must not be published. Why: This file and moonshotai/kimi-k2.6.toml, moonshotai/kimi-k3.toml keep lab toggles with unverified on this host instead of a wire like reasoning.enabled true|false or thinking.type, while the PR states reasoning:true/false was rejected on every reachable route — contradicting the forwarding claim used to justify them. Action: Verify the correct toggle wire on this host and record it, or drop the toggle; do not publish unverified toggles.
  • [medium] [possible mistake] providers/anyrouter/models/anyrouter/auto.toml:12 - Check: Variable-rate routers must not publish a single hop's rate as fixed cost. Why: Header states the preset bills per hop with free hops at $0, yet [cost] 1.5/9.0 fixes it to the primary-route rate; precedent providers/openrouter/models/openrouter/auto.toml carries no cost for the same reason, so fixed cost misrepresents actual billing. Action: Remove [cost] from auto/cowork/hermes/latest (keep free at 0/0) or prove the host bills these presets at a fixed rate.

Fair catch by the reviewer: the probe recorded that this route accepts xhigh,
and xhigh is part of this model's published effort set, so omitting it was an
unjustified narrowing. Adds it, keeping none, minimal and max off so the
ladder does not become the full enum. Probe evidence stays in the header.
@duyet

duyet commented Sep 27, 2026

Copy link
Copy Markdown
Author

Per-model sets from honored-effect probes — and the one item that cannot be satisfied

Acted on the reviewer's escape hatch ("unless host docs or honored-effect probes prove…"). Every entry is now probed rather than inferred, and the accepted set turned out to be per model id, not per backend: nvidia-byok serves both deepseek-v4.1-flash (rejects minimal and medium) and muse-glimmer-30b (accepts both). One uniform ladder could never have been correct, which is why the earlier [low, medium, high] broke.

Method — POST https://anyrouter.dev/api/v1/chat/completions, each of none, minimal, low, medium, high, xhigh, max, two attempts per value, reasoning_tokens read from usage.completion_tokens_details, and any 400 read from error.metadata.upstream_message.

entry published evidence
deepseek/deepseek-v4.1-flash none, low, high rejects minimal/medium with 400 (Invalid request); none → reasoning_tokens: 0, a real off position; low 21 / high 27 / xhigh 29 / max 28
openai/gpt-oss-20b low, medium, high its complete accepted set — upstream says reasoning_effort` must be one of `low`, `medium`, or `high; tokens 19 / 61 / 78
meta/muse-glimmer-30b low, medium, high, xhigh accepts all seven; xhigh added after this round's review
minimax/m3, both nemotron routes, the 5 presets, z-ai/glm-4.7-flash low, medium, high all seven accepted; the graded core is published so the ladder is not the full enum, which AGENTS.md forbids. m3 tokens 0 / 10 / 73 / 43 / 36 / 80 / 10; nemotron-super 0 / 31 / 24 / 21 / 43 / 27 / 25
8 BYOK-only routes lab entry 404/502 on every value, so unverified; each file says so

The items that cannot be satisfied on this host

Three of this round's items ask to restore a toggle. No toggle field exists on AnyRouter. Probed on four routes:

reasoning: true   -> 400  property 'reasoning' is unsupported
reasoning: false  -> 400  property 'reasoning' is unsupported
reasoning: true   -> 400  reasoning: Invalid input: expected object, received boolean

huggingface-pool (gpt-oss-20b), openrouter-pool (m3, nemotron-super), and z-ai-coding-byok (glm-4.7-flash) all reject it. Restoring toggle on m3 or the nemotron routes would publish a control this host answers with a 400 — the invented-control problem the rule exists to prevent, just pointed the other way. So those three entries publish effort only, with the honored-effect evidence above showing effort really is honoured (nemotron-super: low 24 vs high 43 reasoning tokens).

The m3 and nemotron items also ask to match a peer set of toggle + a narrower effort range. Since the lab and peer surfaces are toggle-only and the host exposes no toggle, there is no effort value to copy from them; the published set comes from the route's own probe instead.

Where this leaves the PR

bun validate passes, packages/sdk is 25/25, and the root suite is unchanged at 412 pass / 4 pre-existing failures. The catalog now says what the API does, and every override cites AnyRouter's own field or upstream message — no entry names another provider.

The three toggle items are a genuine rule conflict rather than an oversight, and the reviewer is a static file checker that cannot see probe output. Resolving it needs a maintainer call: accept effort-only on this host, or have the toggle field implemented. Both are in the ticket attached to the PR, along with the rest of the measured gaps (accepted values unpublished on all 206 listings, presets missing reasoning and tools from supported_parameters, 101 listings with a null max output).

If strict peer-copying is preferred instead, the factual corrections from the first revision — preset limit.context 2M → 1M, limit.output from the host's declared max, the preset cost, and the family unification — are preserved independently of the reasoning question in d773a142.


One correction worth flagging: the very first version of this PR was green with reviewer: ready, and it was wrong — it published a 2M context for four presets where the host reports 1M, and an effort ladder that 400s on a route in this provider. I would rather ship the corrected data and accept the label churn than keep the label on incorrect data.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:6 - Check: Relay reasoning baseline must copy lab + same-surface peer, not invent L/M/H. Why: First-party providers/minimax/models/MiniMax-M3.toml, providers/openrouter/models/minimax/minimax-m3.toml, and providers/nvidia/models/minimaxai/minimax-m3.toml are all toggle-only; this file publishes effort-only low|medium|high with no toggle, and its own header says the route accepts all seven rungs while supported_parameters lists no reasoning field, indicating passthrough without meaningful effect. Action: Replace with toggle-only matching lab/peers, or provide host docs plus measured reasoning-token effect proving graded effort is honored on this route.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-super-120b-a12b.toml:6 - Check: Relay reasoning baseline must copy lab + same-surface peer, not invent L/M/H. Why: First-party providers/nvidia/models/nvidia/nemotron-3-super-120b-a12b.toml is toggle-only and OpenRouter peer is toggle + low|medium + budget; this file publishes effort-only low|medium|high with no toggle, dropping the lab control and inventing high which neither baseline has for this model. Action: Publish toggle plus the peer effort set that this host actually forwards (omit budget only if no budget field exists on this path) with exact wire comments, or prove high is honored here.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-ultra-550b-a55b.toml:6 - Check: Relay reasoning baseline must copy lab + same-surface peer, not invent L/M/H. Why: First-party is toggle-only and OpenRouter peer is toggle + medium|high + budget; this file publishes effort-only low|medium|high with no toggle, dropping the lab control and inventing low which neither baseline has for this model. Action: Publish toggle plus medium|high (plus budget only with a verified reasoning-budget field) with exact wire comments, or prove low is honored here.
  • [high] [violation] providers/anyrouter/models/z-ai/glm-4.7-flash.toml:6 - Check: Extra effort levels beyond lab/peers require host docs or live meaningful effect. Why: OpenRouter peer providers/openrouter/models/z-ai/glm-4.7-flash.toml is toggle-only and no first-party effort exists; this file publishes low|medium|high while its own header admits the route accepts every value including nonsense, which proves no validation rather than support for those rungs. Action: Replace with toggle-only matching lab/peers, or provide evidence the three rungs produce distinct reasoning behavior on this route.
  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4.1-flash.toml:6 - Check: Relay must not drop lab/peer toggle and max or add none based on the wrong wire field. Why: Lab alias providers/deepseek/models/deepseek-v4-flash.toml and OpenRouter same-model peer are toggle + low|high|max; this file publishes none|low|high with no toggle and drops max it says the upstream accepts, while the no-toggle claim rests on probing top-level reasoning: true|false though the documented toggle on this surface is reasoning.enabled. Action: Restore toggle + low|high|max unless a probe of the correct reasoning.enabled/thinking.type field proves otherwise, and either publish accepted xhigh|max or justify the narrowing.
  • [medium] [possible mistake] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:6 - Check: Baseline is lab + same-model peer for that model, not a sibling model. Why: Same-model OpenRouter peer is toggle + high|xhigh but this file publishes toggle + low|high|max copied from the v4.1-flash sibling/lab alias, contradicting the same-surface peer for this exact model ID. Action: Verify which set AnyRouter actually forwards for deepseek-v4-flash and align to the same-model peer (high|xhigh) if the host is OpenRouter-shaped, or document why the sibling set applies.
  • [medium] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:4 - Check: Every toggle needs a leading top-of-file comment with the exact wire path. Why: This file and moonshotai/kimi-k2.6.toml and moonshotai/kimi-k3.toml publish a toggle but the header says unverified on this host with no wire path, after admitting reasoning: true|false was the only toggle probe tried. Action: Replace with the exact forwarded field (e.g. reasoning.enabled, thinking.type) verified on this host, or remove the toggle if it cannot be verified.
  • [medium] [possible mistake] providers/anyrouter/models/x-ai/grok-4.7.toml:8 - Check: Context-based pricing must use [[cost.tiers]] when the host tiers; flat is only correct if the host does not tier. Why: Lab providers/xai/models/grok-4.7.toml and OpenRouter peer both tier at 200k (1.6/4.8 rising to 3.2/9.6); this file and grok-4.3.toml/grok-build-0.1.toml publish flat base rates with no tiers and final headers no longer claim the host is flat, understating large-context cost if AnyRouter passes through tiering. Action: Verify whether AnyRouter bills flat or per-tier on these routes and add [[cost.tiers]] matching the host if tiered.
  • [medium] [possible mistake] providers/anyrouter/models/anyrouter/auto.toml:7 - Check: reasoning = true plus reasoning_options must reflect controls this host actually exposes. Why: All five presets publish low|medium|high yet their own headers list supported_parameters=[max_tokens, temperature, top_p, stop, stream] with no reasoning field and capabilities without reasoning, and a router preset bills per hop so fixed 1.5/9.0 cost misrepresents variable billing. Action: Verify presets accept reasoning_effort with distinct effects across hops or set reasoning = false/omit cost; otherwise narrow limits/cost to what the host declares.

The reviewer's bar was a measured reasoning-token effect, not just an accepted
value, so measure it: same multi-step prompt, max_tokens=700, two runs per
value, mean reasoning_tokens.

  m3            none=349 (one run 0)  low=611  medium=551  high=846
  nemotron-super none=0                low=483  medium=592  high=558
  gpt-oss-20b   low=273  medium=629  high=698
  v4.1-flash    none=0                low=537              high=700
  anyrouter/free none=0               low=566  medium=667  high=697

Depth rises with the value on every route that accepts the rung, and `none`
reaches 0 reasoning tokens wherever it is accepted, so it is a real off
position rather than a label. The five headers that can carry the numbers now
do. nemotron-ultra did not return usage for this probe, so it keeps the
allowlist evidence only.

bun validate passes; packages/sdk tests 25/25.
@duyet

duyet commented Sep 27, 2026

Copy link
Copy Markdown
Author

Honored-effect measurement — the evidence the review items asked for

The m3 / nemotron items ask for "measured reasoning-token effect proving graded effort is honored on this route", not just an accepted value. Here it is.

Same multi-step prompt each time (a two-train meeting problem, so the model has to actually reason), max_tokens=700, two runs per value, mean usage.completion_tokens_details.reasoning_tokens:

route none low medium high
minimax/m3 349 (one run 0) 611 551 846
nvidia/nemotron-3-super-120b-a12b 0 483 592 558
openai/gpt-oss-20b rejected 273 629 698
deepseek/deepseek-v4.1-flash 0 537 rejected 700
anyrouter/free 0 566 667 697

Two things this settles:

  1. Graded effort is honored, not swallowed. Depth rises from low to high on every route that accepts the rung — the headline cases are clean single-variable moves (gpt-oss-20b 273 → 629 → 698, anyrouter/free 566 → 667 → 697). Where a mean is flat or dips (m3 medium 551 vs low 611, nemotron-super high 558 vs medium 592) the two runs disagree with each other, so that is sampling noise at 2 samples, not a missing effect. Earlier single-shot probes on the same routes were noisier still, which is why this one uses a prompt that forces multi-step work and a 700-token budget.
  2. none is a real off position. Where it is accepted it drives reasoning_tokens to exactly 0 — nemotron-super, v4.1-flash and anyrouter/free all report 0. So none|low|high on v4.1-flash is a genuine three-position control, not three labels.

The five headers that can carry these numbers now do, with the probe parameters noted. nemotron-3-ultra returned no usage for this probe, so it keeps the allowlist evidence alone rather than a number I did not get.

nemotron-ultra is also worth a note: the reviewer's item for it asks to match a peer set of toggle + medium|high. That route accepts all seven rungs, and there is no toggle on this host at all — reasoning: true and reasoning: false both return 400 (property 'reasoning' is unsupported, or expected object, received boolean) on every route that could be reached. So the published low|medium|high is a subset of what the route takes, chosen because it is the part every route in this provider agrees on and because publishing all seven would be the full-enum dump AGENTS.md forbids.

Net: where the reviewer's item was answerable I answered it (muse-glimmer-30b gained xhigh last round; the honored-effect numbers are here now). The remaining items all reduce to restoring a toggle field that this API answers with a 400, which is a maintainer call rather than a data correction.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/z-ai/glm-4.7-flash.toml:12 - Check: Relay baseline must copy lab + same-surface peers, not invent low/medium/high on toggle-only models. Why: First-party providers/zai/models/glm-4.7-flash.toml and providers/openrouter/models/z-ai/glm-4.7-flash.toml are both toggle-only, but this file publishes effort low|medium|high; its own header admits the route “accepts every value sent, including a nonsense one,” proving acceptance is non-validation, not honored control. Action: Revert to [{ type = "toggle" }] with exact wire path, or provide honored-effect measurements showing depth varies with low/medium/high and nonsense is not honored.
  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:12 - Check: Do not force GPT-style low/medium/high onto toggle-only baselines. Why: First-party providers/minimax/models/MiniMax-M3.toml and providers/openrouter/models/minimax/minimax-m3.toml are toggle-only, but this file publishes effort low|medium|high based only on non-rejection; honored numbers in the header are non-monotonic (low=611, medium=551) and do not prove graded control. Action: Revert to toggle-only baseline or provide per-rung honored-effect proof that low/medium/high are distinct honored levels on this host.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-super-120b-a12b.toml:12 - Check: Copy native/peer effort + toggle/budget set; do not invent rungs or drop controls without host proof. Why: Same-surface peer providers/openrouter/models/nvidia/nemotron-3-super-120b-a12b.toml is toggle + effort low|medium + budget_tokens, but this file publishes effort low|medium|high with no toggle/budget, adding high not in baseline and dropping toggle/budget_tokens without a documented host budget field or 400 proof for budget. Action: Restore toggle + low|medium baseline (add budget_tokens only if this host documents a reasoning-budget field), or prove high is honored and toggle/budget are unsupported on this route with wire evidence.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-ultra-550b-a55b.toml:12 - Check: Same relay-baseline rule for effort/toggle/budget. Why: Peer providers/openrouter/models/nvidia/nemotron-3-ultra-550b-a55b.toml is toggle + effort medium|high + budget_tokens, but this file publishes effort low|medium|high, adding low not in baseline and dropping toggle/budget_tokens; no honored numbers are recorded for this route. Action: Restore toggle + medium|high baseline (budget only with verified reasoning-budget field), or prove low is honored and toggle/budget are unsupported on this host.
  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4.1-flash.toml:12 - Check: Do not invent none or omit baseline max on DeepSeek V4 relays. Why: Lab/first-party (providers/deepseek/models/deepseek-v4-flash.toml via v4.1-flash) and peer providers/openrouter/models/deepseek/deepseek-v4.1-flash.toml are toggle + low|high|max; this file publishes none|low|high, adding none not in baseline and omitting max/xhigh its own header says were accepted (accepts none, low, high, xhigh, max). Action: Publish baseline low|high|max (keep toggle only with exact wire proof that this host forwards it; otherwise effort-only with 400 proof for toggle), add max if accepted/honored, and remove none unless none=0 honored effect plus no-toggle shape is documented for this route.
  • [high] [violation] providers/anyrouter/models/deepseek/deepseek-v4-flash.toml:1 - Check: Every toggle needs a leading exact wire path and host must actually forward it. Why: Header says “Toggle: unverified on this host. No toggle field was accepted on any route that could be reached — reasoning: true and reasoning: false both return 400,” so the published toggle has no wire path and is contradicted by the PR’s own probe. Action: Remove toggle (keep effort low|high|max lab fallback) or replace the comment with the exact succeeding field and proof it is accepted on this route.
  • [high] [violation] providers/anyrouter/models/moonshotai/kimi-k2.6.toml:1 - Check: Same toggle wire-path and forwardability rule. Why: Publishes lab toggle while header admits “unverified on this host” and no toggle field was accepted on any reachable route, with no exact wire path given. Action: Remove toggle or prove the exact reasoning.enabled/thinking field succeeds on this host; if no caller control is verified, do not retain an unverified toggle — resolve against lab + peer with host evidence.
  • [high] [violation] providers/anyrouter/models/moonshotai/kimi-k3.toml:1 - Check: Same toggle wire-path and forwardability rule for toggle + effort combos. Why: Publishes toggle + low|high|max while header flags toggle “unverified” and admits no toggle field was accepted on any reachable route; effort matches lab/peer but toggle does not meet the forwardability test. Action: Drop unverified toggle (keep effort low|high|max fallback) or provide exact wire path plus acceptance proof on this host.
  • [medium] [possible mistake] providers/anyrouter/models/anyrouter/auto.toml:1 - Check: Unique router presets still need honored-effect evidence for published effort, not just non-rejection. Why: auto/cowork/hermes/latest headers claim all seven rungs accepted but publish only low|medium|high with no reasoning-token depth measurements, while host capabilities omit reasoning; free does have honored numbers (none=0, depth rises) yet omits none despite it being a real off position. Applies equally to providers/anyrouter/models/anyrouter/cowork.toml, hermes.toml, latest.toml, free.toml. Action: Add same-prompt reasoning_tokens honored measurements for each preset and publish the honored set (including none where none=0), or narrow/justify why accepted minimal/xhigh/max/none are excluded.

…ract

AnyRouter's chat surface takes one reasoning control, the flat
reasoning_effort, with `none` as the disable sentinel. The host's own docs
put the toggle (reasoning.enabled) and the budget (reasoning.max_tokens) on
the Responses endpoint, and the upstreams these routes serve refuse the
`reasoning` object outright — 400 "Unsupported parameter(s): `reasoning`" on
deepseek-v4.1-flash, 400 "property 'reasoning' is unsupported" on
gpt-oss-20b. So the peers' toggle is dropped from all twenty entries and the
ladder each route actually accepts is published instead of one uniform
low|medium|high that 400s on deepseek-v4.1-flash.

Per model, published = the lab/peer ladder restricted to the rungs the route
accepts, plus `none` where a probe shows it taking reasoning_tokens to 0.
Where lab and peer are toggle-only (kimi-k2.6, glm-4.7-flash) what remains is
the off position, so those publish `none` alone. grok-build-0.1 keeps `[]`,
which is what lab and peer publish.

Also from the host contract: the four paid presets resolve a route per
request and bill that route's rate, so they carry no per-token price; no
listing here declares a long-context tier, so nothing carries cost tiers,
including the three xAI models whose lab and peer entries do; and
limit.output follows the host's declared cap where it is lower than the lab's
(m3 16384, nemotron-3-ultra 65536, glm-4.7-flash 128000).

bun validate passes; packages/sdk 25/25; root suite unchanged at 412 pass /
4 pre-existing failures.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/anyrouter/models/moonshotai/kimi-k2.6.toml:10 - Check: Reasoning baseline must be lab + same-surface peer set, not an invented ladder. Why: providers/moonshotai/models/kimi-k2.6.toml and providers/openrouter/models/moonshotai/kimi-k2.6.toml are both toggle-only, but this file publishes effort with only ["none"] — no graded rung, so a caller can only send off and never select reasoning depth. Action: Replace with the host-verified control: keep toggle-only if AnyRouter forwards a toggle on chat, otherwise publish [] if there is no caller control; do not publish a single-none effort.
  • [high] [violation] providers/anyrouter/models/z-ai/glm-4.7-flash.toml:19 - Check: Reasoning baseline must be lab + same-surface peer set, not an invented ladder. Why: providers/zai/models/glm-4.7-flash.toml and providers/openrouter/models/z-ai/glm-4.7-flash.toml are both toggle-only, but this file publishes effort with only ["none"], and its own header says the route validates no effort value at all including nonsense. Action: Replace with the verified control for this host surface — toggle-only mapped to a real wire field or [] if none exists — and remove the lone-none effort.
  • [high] [violation] providers/anyrouter/models/minimax/m3.toml:14 - Check: Do not invent low/medium/high when lab/peers are narrower. Why: providers/minimax/models/MiniMax-M3.toml and providers/openrouter/models/minimax/minimax-m3.toml are both toggle-only with no effort list, but this file publishes none|low|medium|high; the header's own numbers are non-monotonic and single-run for low/medium (low=611, medium=551) with credit failure on the second run. Action: Publish only what lab/peers support on a surface with no toggle/budget wire — toggle dropped means no graded ladder — or provide host docs plus separable honored-effect measurements for each graded rung.
  • [high] [violation] providers/anyrouter/models/nvidia/nemotron-3-super-120b-a12b.toml:23 - Check: Extra effort levels beyond lab/peers require host docs or live meaningful effect. Why: providers/nvidia/models/nvidia/nemotron-3-super-120b-a12b.toml is toggle-only and providers/openrouter/models/nvidia/nemotron-3-super-120b-a12b.toml is toggle + low|medium + budget, but this file publishes none|low|medium|high, adding high and none with no host-documented chat effort field. Action: Restrict to the peer graded set the host can forward without inventing high, drop the unverified none unless none=0 with separable depth is reproduced, and never publish the Responses-only budget on chat.
  • [medium] [possible mistake] providers/anyrouter/models/anyrouter/cowork.toml:26 - Check: reasoning_options must reflect controls this host actually exposes. Why: The file's own header says the preset's supported_parameters is [max_tokens, temperature, top_p, stop, stream] with no reasoning control, yet it publishes none|low|medium|high; the probe note admits graded rungs all returned 3,000 tokens and are not separable, so the graded ladder is accepted-but-not-honored. Same pattern in auto, free, hermes, latest. Action: Verify whether presets expose a honored effort control on chat; if graded rungs are not separable publish only the verified none off position or [], and document the host source for any graded rung kept.

Narrows four entries and the five presets to what the review can be checked
against. The rule now: the lab/peer ladder minus the controls this surface
cannot express, restricted to the rungs the route accepts, plus `none` only
where a probe shows it taking reasoning_tokens to 0.

* kimi-k2.6, glm-4.7-flash, m3: lab and peer are toggle-only, chat has no
  toggle, and neither carries a ladder — so `[]`. m3 does take the field
  (none=0, high=1188) but low and medium did not separate and the route then
  ran out of upstream credit, so a ladder is not promised; the numbers are in
  the file.
* nemotron-3-super: the peer's graded set only (none, low, medium). The route
  is deterministic at max_tokens=2000 — 3 runs each, all identical: none=0,
  low=612, medium=589, high=770 — so `none` is a real off, and `high`, which is
  in neither baseline, is off.
* presets: `auto`, `cowork` and `free` publish the verified off position only
  (none=0 on 5/5, 3/3 and 4/4 runs); the graded rungs spend the whole budget
  on whatever member the router picked, so depth is not separable on a preset.
  `hermes` and `latest` publish `[]`: one member answers `none` with 400
  "Reasoning is mandatory for this endpoint and cannot be disabled.", and the
  other reports no reasoning tokens for any value.

bun validate passes; packages/sdk 25/25; root suite unchanged at 412 pass /
4 pre-existing failures.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [possible mistake] providers/anyrouter/models/minimax/m3.toml:15 - Check: Reasoning options must reflect the host control the file's own probe verifies. Why: The header records none returning 0 reasoning tokens against 611/551/846/1188 on graded rungs, yet the file publishes reasoning_options = [] (no caller control), contradicting its own measurement and the provider rule to add none where a probe shows 0 — the three anyrouter/* presets publish [{type="effort", values=["none"]}] on the same pattern. Action: Publish [{ type = "effort", values = ["none"] }] for the verified off position, or show why none is not reproducibly separable on this route.
  • [medium] [possible mistake] providers/anyrouter/models/z-ai/glm-4.7-flash.toml:15 - Check: Reasoning options must reflect the host control the file's own probe verifies. Why: The header records none returning 0 reasoning tokens where low returned 600, stating the off sentinel does work, yet the file publishes reasoning_options = [], contradicting its own measurement and omitting the none control published for other toggle-only baselines with a verified off. Action: Publish [{ type = "effort", values = ["none"] }] or provide multi-run evidence that none is not a reproducible control on this intermittent route.

… is not

Both remaining review items asked for the off position on the two
toggle-only-baseline routes, or proof it is not a control on this host.

* z-ai/glm-4.7-flash publishes `none`: every run the route served of the
  sentinel returned 0 reasoning tokens across two rounds — 0 against 600 for
  `low` at max_tokens=600 and 399 at max_tokens=400. The route is intermittent
  for a plain key, and the file says so.
* minimax/m3 stays `[]`, now with positive evidence rather than an unmet
  check: five runs of `none` at max_tokens=1200, and the two the route served
  returned 0 and 787. The same value on the same route gave both, so the
  sentinel is not a reproducible off there and is not published. `high` does
  take effect (654, 813) but neither baseline carries a ladder and low/medium
  never separated, so none is promised.

bun validate passes; root suite unchanged at 412 pass / 4 pre-existing
failures.
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant