Skip to content

feat(sync): replace Friendli generator with SyncProvider - #4477

Closed
Lee-Si-Yoon wants to merge 17 commits into
anomalyco:devfrom
Lee-Si-Yoon:feat/friendli-sync-v2
Closed

Lee-Si-Yoon wants to merge 17 commits into
anomalyco:devfrom
Lee-Si-Yoon:feat/friendli-sync-v2

Conversation

@Lee-Si-Yoon

@Lee-Si-Yoon Lee-Si-Yoon commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Replaces packages/core/script/generate-friendli.ts with a SyncProvider at packages/core/src/sync/providers/friendli.ts, wired into bun models:sync friendli.

Source: https://api.friendli.ai/serverless/v1/models.

Behavior

  • Resolves matching base_model metadata case-insensitively and generates override-only provider TOMLs. Friendli's HF-style IDs map to catalog labs where needed (for example zai-org → zhipuai and deepseek-ai → deepseek).
  • Skip-create for unresolvable third-party IDs (deepinfra-style gate): a brand-new remote model with no models/<lab>/ metadata to factor onto is skipped with a notice — a multi-lab relay's hourly sync never authors third-party full-inline definitions. Already-tracked files keep updating.
  • Retains dropped catalog entries (deleteMissing: false) and marks live entries deprecated from deprecation_date; clears a retained deprecated status if that model later returns active, while preserving other authored statuses such as beta.
  • Uses Friendli's real toggle, effort, and budget_tokens controls. reasoning_budget is verified as independently enforced on GLM-5.3, gemma-4-31B-it, DeepSeek-V3.2, and MiniMax-M2.5: small budgets truncate reasoning_content while completion continues. The catalog emits unbounded { type = "budget_tokens" }, not min/max bounds, because Friendli accepts values above its advertised maximum and those bounds are not safe published constraints.
  • Every authored reasoning control type gets a leading wire-path comment (Toggle / Effort / Budget, each with its docs URL), both for sync-created files and on the existing files' headers.
  • Preserves live-verified DeepSeek-V3.2 reasoning_content despite the catalog's stale interleaved: false, via a durable model-ID override.
  • Overrides limit.context from Friendli's live serving catalog. For a factored model, inherits the lab's limit.output; only a full-inline model falls back to Friendli's max_completion_tokens.
  • Pricing converts Friendli per-token USD rates to USD per million tokens, and only for unit_type = "TOKEN": any other unit (e.g. SECOND) authors no cost section rather than inventing a per-million token price.
  • Writes partial modality overrides only for modality sides Friendli supplies, so an omitted side deep-merges from the lab instead of being replaced with []. GLM-5.3-Flash's explicit text/image/video input override is retained because Friendli's catalog omits PDF and a live PDF data-URL request returns HTTP 422.

Lab metadata

  • models/lgai-exaone/k-exaone-2.0-750b-a37b.toml: 262,144 context from the official Hugging Face config; 32,768 output cited from the K-EXAONE 2.0 Technical Report evaluation setup.
  • models/lgai-exaone/k-exaone-236b-a23b.toml: 262,144 context from the official Hugging Face config, LG sources only. LG publishes no separate generation ceiling, so limit.output is deliberately omitted from the lab entry; Friendli's own serving completion cap for this route lives on the provider file instead (output = 262_144, host-cited).

Verification

  • bun models:sync friendli — idempotent (0 created, 0 updated on re-run)
  • bun validate — exit 0
  • bun run test — 253 pass / 4 fail (same known baseline failures, confirmed via git stash on this branch; no regression)
  • bunx tsc --noEmit — no errors from sync/providers/friendli.ts; an existing unrelated sync/index.ts SyncedMetadata.reasoning_options error predates this PR

Supersedes #3741.

@github-actions

Copy link
Copy Markdown
Contributor

Checking package scripts, runner header behavior, and peers for EXAONE/DeepSeek factoring.

Replace packages/core/script/generate-friendli.ts with a SyncProvider module
(packages/core/src/sync/providers/friendli.ts) wired into bun models:sync friendli.
The Friendli catalog is now synced by the same runner as every other direct provider.

Source: https://api.friendli.ai/serverless/v1/models (no auth for models listing).

Behavior:
- base_model inheritance: when the API base_model resolves to an existing
  models/<lab>/<file>.toml, emit override-only provider TOML via factorBaseModel.
- Full inline when lab metadata is missing or base_model is self-referential:
  deepseek-ai/DeepSeek-V3.2, LGAI-EXAONE/K-EXAONE-2.0-750B-A37B,
  LGAI-EXAONE/K-EXAONE-236B-A23B (deprecation_date 2026-08-20; still listed).
- id = inference model id (HF-style layout, e.g. providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml).
- reasoning_options: toggle and effort passed through from the API. budget_tokens
  is a real reasoning-budget control on Friendli (min=-1 unlimited, max=max_completion_tokens)
  but is stripped from authored TOMLs pending an AGENTS.md rule update — see PR body.
- Pricing: per-token USD to per-million-token USD (x1e6).
- deleteMissing: false — retain local files for models Friendli rotates out of its catalog.
- Case-insensitive base_model resolution (API lowercases ids that exist on disk mixed-case).

New lab metadata: models/lgai-exaone/k-exaone-2.0-750b-a37b.toml,
models/lgai-exaone/k-exaone-236b-a23b.toml (LG AI Research is the lab; full lab
files authored rather than stubs).

Models synced from live /v1/models (7): GLM-5.2, gemma-4-31B-it, DeepSeek-V3.2,
K-EXAONE-236B-A23B, GLM-5.1, MiniMax-M2.5, K-EXAONE-2.0-750B-A37B.
@Lee-Si-Yoon
Lee-Si-Yoon force-pushed the feat/friendli-sync-v2 branch from 58856ae to d7bb0a2 Compare August 22, 2026 03:13
@Lee-Si-Yoon

Lee-Si-Yoon commented Aug 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Rebased feat/friendli-sync-v2 onto current upstream/dev and force-pushed — no conflicts, this PR only touches friendli.ts, providers/friendli/**, and the two new EXAONE lab files, none of which changed on dev since.

bun validate passes. bun run test has 4 failures, but they're pre-existing on upstream/dev (confirmed against a clean worktree), not caused by this PR.

mergeable: true now. Only the automated review check is still pending.

@Lee-Si-Yoon

Lee-Si-Yoon commented Aug 22, 2026 •

Copy link
Copy Markdown
Contributor Author

budget_tokens stays out of this PR. The translator already drops it before writing any TOML, so there's nothing here that needs an AGENTS.md change — updated the PR description to reflect that instead of leaving it as an open ask. toggle and effort are unaffected.

Ready once the review check clears.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml:1 - Check: Non-lab hosts must use base_model when lab metadata exists; provider files stay override-only. Why: Friendli did not create DeepSeek V3.2, and models/deepseek/deepseek-v3.2.toml already exists (HF weights match this id). The PR keeps a full inline definition and drifts lab facts (release_date/last_updated to API created, weaker description). Action: Emit base_model = "deepseek/deepseek-v3.2" and keep only Friendli deltas (cost, reasoning_options, interleaved, real limit overrides). Fix resolveBaseModelID / aliases so self-referential deepseek-ai/DeepSeek-V3.2 factors to that lab id on every sync.
  • [high] [violation] providers/friendli/models/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B.toml:1 - Check: When lab metadata is added for a third-party host, the provider entry must base_model it (override-only). Why: The PR adds complete models/lgai-exaone/k-exaone-*.toml files, then still authors full inline Friendli copies that restate identical name/description/capabilities/modalities. That breaks the lab-vs-provider split and will fight the next sync. Action: Point both EXAONE provider TOMLs at lgai-exaone/k-exaone-2.0-750b-a37b and lgai-exaone/k-exaone-236b-a23b; leave only host fields/deltas. Ensure the resolver maps LGAI-EXAONE/... onto those lab files after they are written.
  • [high] [violation] packages/core/src/sync/providers/friendli.ts:888 - Check: Every toggle needs a leading top-of-file wire-path comment; sync must return header from translateModel. Why: newFileHeader is not part of SyncProvider (runner only uses translated.header, as in llmgateway). New/updated toggle models ship without the required comment (LGAI-EXAONE/*, zai-org/GLM-5.1, zai-org/GLM-5.2, and any future creates). Action: Return header from translateModel when options include toggle (wire path chat_template_kwargs.enable_thinking), drop dead newFileHeader, and re-sync so toggle TOMLs get the leading comment.
  • [medium] [violation] providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml:1 - Check: Models past provider deprecation must not appear as live catalog entries. Why: PR body states deprecation_date 2026-08-20; review date is 2026-08-22. isDeprecated skips updates but deleteMissing: false retains the file with no status = "deprecated", so the model stays active in the catalog. Action: When deprecation_date has passed, author/retain with status = "deprecated" (or stop adding it as a current model). Do not skip-and-retain as an unmarked live entry.
  • [medium] [possible mistake] providers/friendli/models/zai-org/GLM-5.2.toml:7 - Check: Relay reasoning_options must match this host’s real controls; do not invent effort levels. Why: The PR adds toggle + effort high/max and drops the prior hand-authored note that Friendli docs only document enable_thinking. Base providers/friendli/provider.toml still says no effort field is documented; first-party Zhipu GLM-5.2 is effort-only (no toggle). Action: Verify on Friendli’s API/docs the exact wire fields for GLM-5.2. Keep only controls that work; add a leading comment for toggle (and effort if real); update provider.toml if effort is actually supported.

…els, wire toggle header

- resolveLabModelSync now falls back to the model's own id when the API
  omits base_model (DeepSeek-V3.2, K-EXAONE-2.0-750B-A37B), factoring
  them against existing models/ metadata instead of shipping duplicate
  inline definitions.
- translateModel/buildFriendliModel mirror the deepinfra/pioneer pattern:
  a brand-new deprecated model is skipped, but a model we already track
  keeps its file and gets status = "deprecated" instead of silently
  falling out of sync with no lifecycle marker.
- K-EXAONE-236B-A23B (no longer served by the API) gets status =
  "deprecated" in place, matching the same field position used by
  providers/deepinfra and providers/greenpt.
- Removed the dead newFileHeader field (never read by the sync runner)
  and wired a toggleHeader() function into translateModel's header
  return, matching the llmgateway.ts pattern, so toggle-bearing files
  get their wire-path comment from the sync itself.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml:1 - Check: Non-lab hosts must use base_model with override-only fields when lab metadata exists. Why: This PR adds complete lab metadata at models/lgai-exaone/k-exaone-236b-a23b.toml, but the Friendli entry stays a full inline third-party definition (name, description, capabilities, modalities, limits, etc.). That duplicates lab facts and breaks the override-only rule; deprecation/deleteMissing: false does not exempt it. Action: Rewrite as base_model = "lgai-exaone/k-exaone-236b-a23b" plus only provider deltas (cost, reasoning_options, interleaved, status = "deprecated", real limit overrides).
  • [high] [violation] providers/friendli/models/zai-org/GLM-5.1.toml:1 - Check: Every toggle reasoning control requires a leading top-of-file wire-path comment. Why: Final GLM-5.1, GLM-5.2, and K-EXAONE-236B-A23B all author type = "toggle" without a leading # Toggle: … header. DeepSeek/gemma/EXAONE-2.0 got headers; these three did not. Sync sameModel compares model data only, so a later header: return will not rewrite unchanged bodies. Action: Add the Friendli wire comment (# Toggle: chat_template_kwargs.enable_thinking = true | false + docs URL) to each toggle-bearing file missing it, matching the other Friendli entries.
  • [medium] [possible mistake] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml:1 - Check: Preserve host-accurate interleaved reasoning when the API surface still exposes it. Why: The prior Friendli DeepSeek-V3.2 entry had [interleaved] field = "reasoning_content"; the regenerated file drops it entirely while other Friendli reasoners in this PR keep interleaved. Peers (DeepInfra/Chutes) also expose reasoning_content for V3.2. Action: Verify Friendli still returns reasoning_content for this model; if yes, restore [interleaved] (and ensure the sync maps API/existing interleaved so it is not silently dropped on the next run).
  • [low] [possible mistake] providers/friendli/models/zai-org/GLM-5.2.toml:7 - Check: Reasoning effort must match this host’s documented controls, not only an uncited API enum. Why: provider.toml states Friendli documents only enable_thinking (no effort field), and the previous GLM-5.2 header said the same, yet the file still authors effort = ["high", "max"] with no remaining caveat. Action: Confirm Friendli actually accepts meaningful high/max effort for this model; if not, drop effort to toggle-only and align provider.toml/comments with the verified surface.

…ore toggle headers, preserve interleaved

- K-EXAONE-236B-A23B (no longer served, status = "deprecated") now
  factors against models/lgai-exaone/k-exaone-236b-a23b.toml instead of
  staying a full inline duplicate — matches the base_model override-only
  rule that already applied to the other EXAONE entry.
- GLM-5.1 and GLM-5.2 provider files get the leading toggle wire-path
  comment that sync only writes on rewrite; sameModel treated the
  unchanged body as no-op so the header never landed. Added directly to
  match every other toggle-bearing Friendli file.
- GLM-5.2 also gets an explicit comment scoping the confirmed
  reasoning_effort=high|max control to this model only (cited to
  Friendli's own GLM-5.2 reasoning doc), since GLM-5.1 documents toggle
  only and the family-wide claim was the actual violation risk.
- translateInterleaved now falls back to the existing authored value
  when the API reports interleaved: false, instead of dropping it. The
  API is inconsistent here (DeepSeek-V3.2 reports false while peers
  DeepInfra/Chutes confirm reasoning_content is real); trusting an
  existing authored value avoids silently regressing hosts the API
  under-reports on future syncs.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml:1 - Check: Provider-authored interleaved must stay accurate for this host’s reasoning side channel. Why: The prior Friendli DeepSeek-V3.2 entry had [interleaved] field = "reasoning_content" (same as DeepInfra/other peers). Patch 2 dropped it while factoring onto base_model, and patch 3 only preserves interleaved when the existing file still has it—so a re-sync cannot restore what was already wiped. Catalog consumers lose the reasoning side-channel signal for this model. Action: Restore [interleaved] field = "reasoning_content" on this TOML (and keep the translateInterleaved fallback so future syncs do not strip it again).
  • [low] [possible mistake] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml:4 - Check: After base_model, only real host deltas should remain. Why: description = "DeepSeek chat model for instruction following, coding, and analysis" overrides the lab’s more accurate hybrid-reasoning description on models/deepseek/deepseek-v3.2.toml with a generic string. Action: Drop the provider description override unless Friendli’s API text is intentionally host-specific; prefer inheriting the lab description.
  • [low] [possible mistake] providers/friendli/provider.toml:2 - Check: Provider docs comments should match authored reasoning controls. Why: This PR authors GLM-5.2 with effort high|max (and cites Friendli’s control-reasoning doc), but provider.toml still says no effort field is documented—leaving a contradictory source of truth for the next sync/author. Action: Update the provider.toml reasoning comment to match the documented toggle (and model-specific effort where confirmed).

Called /v1/chat/completions against deepseek-ai/DeepSeek-V3.2 with
enable_thinking=true — the response includes both reasoning and
reasoning_content, confirming interleaved reasoning_content is real for
this model despite /v1/models reporting interleaved: false. Replaced
the "other hosts confirm it" comment with the actual verification.
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml:1 - Check: Provider interleaved must reflect this host’s real reasoning side channel. Why: Base already had [interleaved] field = "reasoning_content". The sync drops it (API reports interleaved: false), and the final factored file never restores it. Patch 4 only preserves an existing authored value when the API says false, but the file already lost that field, so re-sync cannot recover it. The same commit message states a live /v1/chat/completions call with enable_thinking=true returned reasoning_content. Catalog consumers will treat Friendli DeepSeek-V3.2 as non-interleaved despite a verified side channel. Action: Restore [interleaved] field = "reasoning_content" on this TOML, and make the translator seed/preserve that channel when verified (do not rely solely on a value the prior sync already wiped).
  • [medium] [possible mistake] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml (description override) - Check: After base_model, keep only real deltas vs lab metadata. Why: The file keeps description = "DeepSeek chat model for instruction following, coding, and analysis" while models/deepseek/deepseek-v3.2.toml describes the hybrid-reasoning V3.2 weights. That string matches the non-thinking deepseek-chat product copy, so this override looks stale/wrong relative to base_model = "deepseek/deepseek-v3.2". Action: Drop the description override so the lab text inherits, or replace it with a Friendli-specific description that actually describes V3.2.

@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

@rekram1-node Final follow-up pushed. Verified the DeepSeek-V3.2 interleaved question directly instead of inferring it: called /v1/chat/completions with enable_thinking=true and the response includes reasoning_content, so the metadata endpoint's interleaved: false is stale/wrong and keeping the field is correct.

Also fixed from the last review round: K-EXAONE-236B-A23B now factors against its lab metadata (base_model) instead of staying a full inline duplicate, and GLM-5.1/GLM-5.2 have their toggle wire-path headers back (sync's sameModel skipped rewriting unchanged bodies, so they never landed).

validate and review are both green. Remaining test failures are pre-existing on dev, unrelated to this PR. Ready for a final look.

- Resolve providers/index.ts conflict: keep friendli alongside the
  incoming github-copilot entry (alphabetical, both providers list).
- translateReasoningOptions no longer drops budget_tokens: the live
  GET /serverless/v1/models response and
  https://friendli.ai/docs/openapi/model-apis/chat-completions#body-reasoning-budget-one-of-0
  confirm reasoning_budget is a real, documented request field
  (min = -1 unlimited, max = max_completion_tokens), contradicting the
  prior comment claiming it "is not a real reasoning budget control on
  this host". Normalizes min to -1 to match the hand-authored GLM-5.x
  convention regardless of what the API reports.
- Re-ran bun models:sync friendli: GLM-5.1/5.2/5.3/5.3-Flash and the
  new DeepSeek/EXAONE/MiniMax entries all pick up budget_tokens.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 642cc5fcc: merged upstream/dev and corrected the budget_tokens translation. Live GET /serverless/v1/models and the chat-completions OpenAPI docs confirm reasoning_budget is a real request field on Friendli (min=-1 unlimited, max=max_completion_tokens), so the translator no longer strips it — GLM-5.1/5.2/5.3/5.3-Flash all now sync with a budget_tokens option. bun validate exit 0, bun run test 253 pass / 4 fail (no new failures vs the 248/5 baseline).

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml:1 - Check: Provider-specific interleaved must reflect this host’s real response shape when reasoning = true. Why: Base had [interleaved] field = "reasoning_content"; patch 1 dropped it (API interleaved: false), and later patches never restored it even after a live /v1/chat/completions check confirmed both reasoning and reasoning_content. The final factored file still omits interleaved, so consumers lose a verified side channel. translateInterleaved only preserves an existing authored value when the API reports false, so a re-sync from a file that already lost the field (or a clean create) permanently under-reports the capability. Action: Restore [interleaved] field = "reasoning_content" on DeepSeek-V3.2.toml, and fix the translator so a false/stale catalog flag cannot clear verified Friendli reasoning_content (e.g. seed/default for this model or for Friendli reasoners that actually emit the field).
  • [medium] [possible mistake] providers/friendli/models/google/gemma-4-31B-it.toml:4 - Check: Provider modality/attachment overrides must match what this host actually accepts, not only lab defaults. Why: Lab google/gemma-4-31b-it is attachment = true with input = ["text", "image"]. The new sync omits modality overrides when the API leaves input_modalities unset, so Friendli inherits vision. The deleted generator treated Friendli as text-only; if serverless gemma is still text-only, the catalog will advertise image input the host rejects. Action: Confirm Friendli’s gemma route accepts image input. If not, author attachment = false and text-only [modalities] (and teach the sync to emit that when the API omits modalities for text-only hosts).
  • [low] [violation] providers/friendli/provider.toml:2 - Check: Provider docs comments must match authored reasoning controls. Why: The leading comment still says Friendli documents only enable_thinking and no effort field, but this PR authors GLM-5.2 with effort high/max and cites Friendli’s own control-reasoning doc. That leaves the provider header contradictory for the next sync/reviewer. Action: Update the provider.toml reasoning comment to note model-specific reasoning_effort (at least GLM-5.2) alongside the toggle wire path.

… fix provider.toml

- translateInterleaved now consults a durable
  VERIFIED_INTERLEAVED_OVERRIDES map keyed by model id instead of relying
  solely on the fragile "existing on-disk value" fallback. The prior
  design silently forgot a live-verified capability if the authored file
  ever lost the field for any reason (which is exactly what happened to
  DeepSeek-V3.2 in this PR's own resync). Re-verified live: a real
  /chat/completions request with enable_thinking=true still returns both
  reasoning and reasoning_content for deepseek-ai/DeepSeek-V3.2 despite
  /v1/models reporting interleaved: false.
- google/gemma-4-31B-it: verified image input is real (not just catalog
  metadata) with a live request -- base64 image POST consumed 280 prompt
  tokens vs 22 for an equivalent text-only request, proving the image was
  actually processed. No modalities override needed (file correctly
  inherits the lab's text+image modalities); added a citing comment.
- provider.toml: removed the stale "no effort or budget field is
  documented" claim (same fix already applied on the friendli-glm-5.3
  branch), replaced with a comment citing reasoning_effort and
  reasoning_budget as real OpenAPI-documented request fields.

bun models:sync friendli remains idempotent after these changes.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 294acdbc0 fixing all 3 newly-flagged issues:

  1. [high] DeepSeek-V3.2 interleaved regression — real bug, not a false positive. The prior translateInterleaved only preserved a verified interleaved: reasoning_content via the on-disk existing file's value, which is fragile: if the authored file ever loses the field (which happened during this PR's own resync history), the live verification silently disappears with no trace. Fixed by adding a durable VERIFIED_INTERLEAVED_OVERRIDES map keyed by model id, consulted before any existing-file fallback. Re-verified live: POST /chat/completions with enable_thinking=true against deepseek-ai/DeepSeek-V3.2 still returns both reasoning and reasoning_content, despite /v1/models reporting interleaved: false.

  2. [medium] gemma-4-31B-it modality verification — tested with a real request instead of leaving it as a guess. GET /serverless/v1/models reports input_modalities: [text, image]; a live image POST /chat/completions (base64 PNG data URL) returned HTTP 200 consuming 280 prompt tokens vs 22 for the equivalent text-only request — the image was genuinely processed, not silently dropped or rejected. Confirmed the current no-override behavior (inheriting the lab's text+image modalities) is correct; added a citing comment so this doesn't need re-deriving.

  3. [low] provider.toml stale comment — brought in line with the same fix already applied on friendli-glm-5.3 (PR feat(friendli): add GLM-5.3 provider model #5776): reasoning_effort and reasoning_budget are documented in Friendli's chat-completions OpenAPI reference, the old "no effort or budget field is documented" claim was simply wrong.

bun models:sync friendli still idempotent after these changes. bun validate exit 0.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] packages/core/src/sync/providers/friendli.ts (translateReasoningOptions) - Check: Host reasoning_options must reflect real Friendli request controls; do not invent or drop real budget controls, and never derive budget bounds from max_completion_tokens / limit.output. Why: The translator still does if (option.type === "budget_tokens") continue, so authored TOMLs omit budget even though this PR’s provider.toml and body document reasoning_budget as a real OpenAPI control and claim GLM-5.x files carry budget_tokens. That understates the host surface and leaves the sync/docs/catalog out of sync. Action: When the API advertises budget, emit budget_tokens on those models. Keep verified bounds only (e.g. documented min = -1); do not set max from max_completion_tokens. Re-sync so GLM (and any other budget-capable) files match, and align the PR body with the actual authored options.
  • [medium] [possible mistake] providers/friendli/models/zai-org/GLM-5.2.toml / providers/friendli/models/google/gemma-4-31B-it.toml / providers/friendli/models/LGAI-EXAONE/*.toml ([limit].output) - Check: Provider limit.output must be the host completion cap, not a restated context window. Why: Sync sets output = model.max_completion_tokens, producing output equal to context (1_048_576, 262_144) far above lab defaults (GLM-5.2 131_072, gemma/EXAONE 32_768) and tighter peer caps. If Friendli’s field is effectively “context-sized max tokens,” the catalog overstates completion capacity. Action: Verify against Friendli docs or a live completion whether these values are real output caps; if not, keep lab/peer completion limits (or omit the override) instead of copying context-length max_completion_tokens.

Friendli's max_completion_tokens equals context_length for several
models it serves (GLM-5.2/5.3/5.3-Flash, gemma-4-31B-it, both EXAONE
entries, DeepSeek-V3.2), and the sync was blindly asserting that value
as limit.output on every generated file -- including ones with a
base_model to inherit from. That silently overwrote lab-verified,
genuinely tighter completion caps: DeepSeek-V3.2's lab file documents
output=64_000 out of a 128_000 context, not "same as context".

Per user direction: when a base_model exists, defer to its own
limit.output via factorBaseModel's normal inheritance instead of forcing
Friendli's completion-length report onto it. Only fall back to
Friendli's own max_completion_tokens for full-inline entries that have
no base_model to inherit a real completion-cap policy from.

context override is unaffected -- Friendli's live serving context is a
real, host-specific fact worth overriding regardless of base_model.

Verified resolved catalog after the fix:
  deepseek-ai/DeepSeek-V3.2   -> output=64_000  (was 163_840)
  zai-org/GLM-5.1/5.2/5.3/5.3-Flash -> output=131_072 (was 202_752/1_048_576)
  google/gemma-4-31B-it       -> output=32_768  (was 262_144)
  LGAI-EXAONE/K-EXAONE-2.0-750B-A37B -> output=32_768 (was 262_144)

bun models:sync friendli remains idempotent; bun run test unchanged at
253 pass / 4 fail (same pre-existing failures as prior rounds).
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 0cc8f2caf addressing the [medium] finding on limit.output.

You were right to flag it: max_completion_tokens coincides with context_length on several Friendli-served models (GLM-5.2/5.3/5.3-Flash, gemma-4-31B-it, both EXAONE entries, DeepSeek-V3.2), and the sync was blindly asserting that as limit.output on every file — including ones with a base_model to inherit from. That silently clobbered lab-verified, genuinely tighter completion caps: DeepSeek-V3.2's lab file documents output = 64_000 out of a 128_000 context, not "same as context".

Fix: when a base_model exists, limit.output is left undefined so factorBaseModel's normal inheritance takes over and pulls the lab's real completion cap. Only full-inline entries (no lab to inherit from) still fall back to Friendli's own max_completion_tokens. context overrides are untouched — Friendli's live serving context is a real, host-specific fact worth keeping regardless of base_model.

Verified resolved catalog after the fix (bun validate output):

  • deepseek-ai/DeepSeek-V3.2 → output=64_000 (was 163_840)
  • zai-org/GLM-5.1/5.2/5.3/5.3-Flash → output=131_072 (was 202_752/1_048_576)
  • google/gemma-4-31B-it → output=32_768 (was 262_144)
  • LGAI-EXAONE/K-EXAONE-2.0-750B-A37B → output=32_768 (was 262_144)

bun models:sync friendli still idempotent, bun validate exit 0, bun run test unchanged at 253 pass / 4 fail (same pre-existing failures).

On the [high] finding from the previous round (bare budget_tokens strip in translateReasoningOptions): checked the exact pushed commit (294acdbc0) directly and that code path doesn't exist there — it pushes {type: "budget_tokens", min: -1, max: option.max}, not a bare continue. That pattern only ever existed in the very first commit on this branch (d7bb0a284), long since fixed. Happy to dig further if the next review re-flags it with a specific line reference.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] packages/core/src/sync/providers/friendli.ts / providers/friendli/models/** - Check: budget_tokens bounds must be real reasoning-budget limits, not derived from completion/context caps (AGENTS.md Reasoning options; audit skill: never derive min/max from limit.output or context). Why: The PR authors budget_tokens with max taken from Friendli max_completion_tokens (hunks show max = 262_144 next to gemma/EXAONE, and the same pattern on GLM-5.x). That is the completion-length field the PR itself later stopped treating as a trustworthy limit.output, so cataloging it as a reasoning budget max invents bounds. Action: Keep budget_tokens only if reasoning_budget is a real host control; drop max (and any min) unless independently verified as a reasoning-budget range. Do not set max from max_completion_tokens, context, or lab limit.output. Update the translator and regenerate the affected TOMLs the same way.
  • [high] [violation] packages/core/src/sync/providers/friendli.ts (translateReasoningOptions) - Check: Sync code and authored TOMLs must agree; a later sync must be idempotent with the committed catalog. Why: Patch 1 still strips every budget_tokens option (if (option.type === "budget_tokens") continue), and no later commit removes that filter, while the committed Friendli TOMLs and PR body claim pass-through of budget_tokens. The next bun models:sync friendli would rewrite those files and drop the budgets (or keep strip forever if TOMLs were hand-edited). Action: Align translator and files: either pass budget_tokens through (without bogus max-from-completion) and regenerate, or strip it consistently and remove the authored [[reasoning_options]] type = "budget_tokens" blocks.
  • [medium] [violation] providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml - Check: After base_model, do not keep a host limit.output that only restates Friendli’s completion=context report when the lab has a tighter real cap. Why: Patch 6 correctly stopped forcing max_completion_tokens onto factored models and removed the inflated output overrides elsewhere, but this retained/deprecated file still has output = 262_144 while models/lgai-exaone/k-exaone-236b-a23b.toml documents output = 32_768. Because the model is absent from the live catalog (deleteMissing: false), sync will not rewrite it. Action: Drop the provider limit.output override (or set a verified Friendli-specific completion cap) so the resolved model inherits the lab 32_768.

…236B

K-EXAONE-236B-A23B.toml is retained (deleteMissing: false) after
Friendli dropped it from the live catalog, so bun models:sync friendli
never touches it and the file kept a stale limit.output = 262_144 from
before the base_model-inheritance fix landed on the other 8 files.
Removed the override so the resolved model correctly inherits the lab
file's output = 32_768, matching every other factored Friendli entry.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed `53ee2cc17` fixing the `[medium]` and responding to both `[high]` findings with evidence.

`[medium]` K-EXAONE-236B-A23B stale `limit.output` — fixed. This model was retained after Friendli dropped it from the live catalog (`deleteMissing: false`), so `bun models:sync friendli` never touches it and it kept the pre-fix `output = 262_144` override from before the base_model-inheritance change landed on the other 8 files. Removed the override; resolved catalog now correctly shows `output = 32_768` (inherited from `models/lgai-exaone/k-exaone-236b-a23b.toml`), matching every other factored entry.

`[high]` #1 "budget_tokens max invents a reasoning-budget bound from completion capacity" — re-tested this claim directly instead of assuming either side is right. Ran live `reasoning_budget=30` (generous `max_tokens=3000`) requests against three different models with different reasoning architectures:

  • `google/gemma-4-31B-it`: `reasoning_content` cut to 106 chars, `completion_tokens=972`
  • `deepseek-ai/DeepSeek-V3.2`: `reasoning_content` cut to 106 chars, `completion_tokens=597`
  • `zai-org/GLM-5.3` (prior round): `reasoning_content` cut to 130 chars, `completion_tokens=1553`

All three show the same pattern: a small `reasoning_budget` truncates only the reasoning phase while the completion phase continues normally to a full answer. That's consistent, cross-model evidence that `reasoning_budget` is a real, independently-enforced control — not a value we invented from `max_completion_tokens`. The `max` we author is Friendli's own reported `reasoning_options.budget_tokens.max` from `GET /serverless/v1/models`, read verbatim; it coincidentally equals `max_completion_tokens` for these particular models on Friendli's backend, but we aren't deriving one from the other.

`[high]` #2 "translator still does a bare `budget_tokens` continue/strip" — checked the exact pushed commit (`0cc8f2caf`, the SHA this review ran against) directly: `translateReasoningOptions` at that commit pushes `{type: "budget_tokens", min: -1, max: option.max}`, not a bare `continue`. Re-ran `bun models:sync friendli` fresh in a clean worktree from that exact commit — 0 created/updated, fully idempotent, `budget_tokens` still present in every regenerated file. I can't find the bare-strip code path this finding describes anywhere in the current branch; it only ever existed in the very first commit (`d7bb0a284`) and was removed in `642cc5fcc`. If there's a specific line/file the next review can point to, happy to look again.

`bun validate` exit 0, `bun models:sync friendli` idempotent, `bun run test` unchanged at 253 pass / 4 fail.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B.toml / packages/core/src/sync/providers/friendli.ts - Check: budget_tokens bounds must be real reasoning-budget limits, never derived from limit.output / context / completion caps (AGENTS.md → Reasoning options; audit skill Step 4). Why: Authored Friendli entries set budget_tokens.max from live max_completion_tokens (e.g. max = 262_144 on EXAONE/gemma, matching context-sized completion caps the PR itself treats as unreliable for limit.output). That publishes a bogus reasoning budget equal to the full completion window, not a verified thinking-token range. Action: Stop deriving max (and any min/max) from max_completion_tokens/context. Either author budget_tokens without fabricated bounds, or only with bounds proven as Friendli’s reasoning-budget range; re-sync so every GLM/EXAONE/gemma file matches.
  • [medium] [possible mistake] providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml - Check: After base_model, keep only provider-specific fields or real deltas — do not restated weaker generic metadata. Why: The file keeps description = "DeepSeek chat model for instruction following, coding, and analysis" while models/deepseek/deepseek-v3.2.toml has the fuller hybrid-reasoning lab description. That shadows lab metadata with a generic API/old-generator blurb rather than a Friendli-specific fact. Action: Drop the description override (and stop forcing API description through factorBaseModel when it only degrades the lab text), unless Friendli truly serves a different model identity.
  • [medium] [possible mistake] packages/core/src/sync/providers/friendli.ts - Check: Sync translator and authored TOMLs must agree on reasoning_options (pass-through vs strip). Why: Patch 1 still strips budget_tokens in translateReasoningOptions, while the PR body and later TOMLs claim pass-through with authored budget_tokens. If the strip remains, the next bun models:sync friendli will delete those options and break the claimed idempotence; if pass-through was added off-diff, the strip comments/path are stale. Action: Make the translator match the intended final policy in one place, re-run sync, and ensure generated files stay stable.

…iptions

reasoning_budget is a real independently-enforced Friendli control
(verified live on GLM-5.3, gemma-4-31B-it, and DeepSeek-V3.2), but the
catalog's reported min/max are not safe published constraints: GLM-5.3
accepted reasoning_budget=1_048_577 despite reporting max=1_048_576.
Emit budget_tokens as an unbounded capability rather than publishing a
range that the API does not enforce.

For entries with a base_model, stop overriding the canonical lab
description with generic API text. Factored provider entries now inherit
the richer lab description; full-inline entries keep the API description.

Re-synced 8 active Friendli models. budget_tokens remains present on all
budget-capable active models without fabricated min/max bounds, and the
sync is idempotent.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 383903b7f resolving all three findings from the latest pass:

  1. Unverified budget_tokens bounds — kept the verified capability but removed every emitted min/max. We confirmed reasoning_budget is real and independently enforced across GLM-5.3, gemma-4-31B-it, and DeepSeek-V3.2, but also confirmed Friendli accepts GLM-5.3 reasoning_budget=1_048_577 despite reporting max=1_048_576. The field is therefore a real control, while its catalog range is not a safe published constraint. translateReasoningOptions now emits { type = "budget_tokens" } only, and re-sync regenerated all 8 active models that expose it. No fabricated range remains.

  2. Generic description override — dropped description: model.description for factored entries. They now inherit the lab's canonical description; API description remains only for full-inline entries. DeepSeek-V3.2 now resolves to its lab's Hybrid-reasoning DeepSeek model with thinking and non-thinking modes, sparse attention, and tool-use description rather than the generic old-generator text.

  3. Sync consistency — translator now has one clear policy (keep budget_tokens as an unbounded capability), all active generated TOMLs match it, and bun models:sync friendli is idempotent after re-sync (0 created/updated on second run).

Validation: bun validate exit 0; bun run test unchanged at 253 pass / 4 known pre-existing failures.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [violation] packages/core/src/sync/providers/friendli.ts - Check: Deprecation lifecycle must clear status = "deprecated" when the live catalog no longer marks the model deprecated (same pattern as deepinfra). Why: status is set as deprecated ? "deprecated" : existing?.status, so a retained Friendli file (e.g. K-EXAONE-236B-A23B with deleteMissing: false) that later reappears without deprecation_date stays permanently deprecated. The comment claims to mirror DeepInfra, but DeepInfra explicitly clears a prior deprecated status when the model is active again. Action: When isDeprecated(model) is false, clear an existing deprecated status (keep non-deprecated hand-authored values like beta); e.g. deprecated ? "deprecated" : existing?.status === "deprecated" ? undefined : existing?.status.
  • [low] [possible mistake] providers/friendli/models/MiniMaxAI/MiniMax-M2.5.toml - Check: Relay reasoning_options must match this host’s real controls; lab/peer baseline for MiniMax-M2.5 is always-on with no caller control ([]). Why: This file moves from reasoning_options = [] to budget_tokens only. First-party providers/minimax/models/MiniMax-M2.5.toml and established relays keep []. The PR’s live budget checks cite GLM-5.3, gemma-4-31B-it, and DeepSeek-V3.2, not MiniMax-M2.5, so a catalog-wide budget_tokens pass-through may be over-claiming control on this model. Action: Confirm Friendli actually accepts/enforces reasoning_budget for MiniMaxAI/MiniMax-M2.5; if not, keep reasoning_options = [] (and teach the sync to drop advertised budget when it is not a real MiniMax control).

- Retained deleteMissing:false entries no longer remain permanently
  deprecated if they reappear active in Friendli's catalog. Clear only a
  prior deprecated status; preserve intentional lifecycle values such as
  beta.
- Verified MiniMaxAI/MiniMax-M2.5 reasoning_budget with a live request:
  reasoning_budget=30 truncates reasoning_content at 135 chars while the
  completion continues for 682 tokens. MiniMax therefore genuinely has
  the independently-enforced budget control, and keeps the unbounded
  budget_tokens policy used for other verified Friendli models.

bun models:sync friendli idempotent; bun validate exits 0; tests unchanged
at 253 pass / 4 known failures.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed feac5bfb2 resolving both latest findings:

  1. Retained deprecated lifecycle — fixed the real lifecycle bug. deleteMissing:false is intentionally used because Friendli rotates entries in and out of the live catalog. A retained file that later returns active must not stay permanently status = "deprecated". The translator now sets deprecated when the live model's date has passed, clears only a prior deprecated status when the live model is active again, and preserves other hand-authored lifecycle states such as beta.

  2. MiniMax-M2.5 budget control — verified with a direct live request rather than assuming catalog metadata applies. POST /chat/completions to MiniMaxAI/MiniMax-M2.5 with reasoning_budget=30, max_tokens=3000 returned HTTP 200; reasoning_content truncated at 135 characters mid-sentence while completion_tokens=682 continued to a full answer. This is the same independent reasoning-cap behavior verified earlier on GLM-5.3, gemma, and DeepSeek. MiniMax therefore correctly retains the unbounded { type = "budget_tokens" } capability.

bun models:sync friendli remains idempotent; bun validate exit 0; bun run test unchanged at 253 pass / 4 known failures.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [possible mistake] models/lgai-exaone/k-exaone-2.0-750b-a37b.toml:75 - Check: New lab metadata must carry complete, evidence-backed limits (providers inherit them). Why: Both EXAONE lab files set limit.output = 32_768 with no source, while Friendli’s own max_completion_tokens for these models equals context (262_144). After this PR stops overriding provider limit.output, the catalog resolves to the uncited lab cap. Action: Cite a first-party/HF/docs source for EXAONE completion limits (both k-exaone-2.0-750b-a37b and k-exaone-236b-a23b), or set lab limit.output to the verified value and keep any host-specific delta only on the Friendli files.
  • [medium] [violation] packages/core/src/sync/providers/friendli.ts:1034 - Check: Provider modality overrides must not wipe inherited lab modalities. Why: Factored entries build modalities as { input: apiInput ?? [], output: apiOutput ?? [] } whenever either API field is present. If Friendli sends only one side, the other becomes [] and replaces the lab array on merge (e.g. dropping text output or vision inputs). Action: Override only defined sides (omit the missing key so deep-merge inherits), or require both sides before writing a modalities override.
  • [medium] [possible mistake] providers/friendli/models/zai-org/GLM-5.3-Flash.toml:1920 - Check: Host modality overrides must be real deltas, not incomplete API copies. Why: The file overrides lab zhipuai/glm-5.3-flash input (text/image/video/pdf) to ["text", "image", "video"], dropping pdf, with no verification note (unlike the gemma image comment). Action: Confirm Friendli rejects PDF for this model and keep the override only if true; otherwise remove the modalities override so pdf inherits, and cite the check in a leading comment.
  • [low] [possible mistake] .pr-review/pull-request.json - Check: PR body should match final authored data for reviewable evidence. Why: The summary still describes full-inline DeepSeek/EXAONE and budget_tokens max from max_completion_tokens, but later commits factor those models and emit unbounded budget_tokens. Action: Update the PR body to the final behavior (base_model usage, unbounded budget, deferred limit.output, live verification notes per model).

- Factored modality overrides now include only API-defined sides, so an
  omitted side deep-merges from lab metadata instead of being overwritten
  with an empty array.
- Cite the official K-EXAONE 2.0 report's 32,768 generation setup for
  750B; use Friendli's documented 262,144 fallback for 236B because LG's
  public config supplies context but no lower generation ceiling.
- Document GLM-5.3-Flash's real host modality delta: live catalog omits
  PDF and a minimal PDF data URL is rejected by chat completions (422).
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed f3d1e8403 and follow-up 5ec577437 addressing the latest reviewer findings:

  • EXAONE limits: cited the official 750B HF config plus K-EXAONE 2.0 Technical Report (32,768 generated tokens in 256K context). The official 236B config states 262,144 context but no lower generation ceiling; per the agreed fallback rule it uses Friendli's live 262,144 max_completion_tokens, documented in a leading comment instead of retaining the uncited 32,768 value.
  • Partial modalities: factored entries now build an override from only sides the API actually supplies. Omitted input/output sides remain omitted and deep-merge from lab metadata; they cannot become []. Full-inline entries default any API-omitted side to text instead of an empty modality array.
  • GLM-5.3-Flash PDF: Friendli's live catalog lists text/image/video only. A minimal PDF supplied via a data URL to chat completions was rejected with HTTP 422 Failed to load media; the explicit text/image/video host override is retained with that leading evidence comment.
  • PR body: rewritten to describe the final factored model handling, unbounded verified budget capability, output inheritance/fallback policy, and live verifications.

Fresh bun models:sync friendli: 0 created/updated; bun validate: exit 0; tests: 253 pass / 4 known existing failures.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] models/lgai-exaone/k-exaone-236b-a23b.toml:19 - Check: Lab metadata must be provider-agnostic; limit.output is a lab fact, not a host serving detail. Why: The file sets output = 262_144 and cites Friendli’s live max_completion_tokens as the source. That contaminates shared models/ with a host-specific completion length (the same field this PR already treats as unreliable on factored Friendli entries because it often equals context_length). Any other provider that base_models this lab file will inherit Friendli’s cap. Action: Restore a lab-grounded limit.output (or omit a fabricated ceiling until LG documents one). If Friendli’s served completion length differs, put that only on providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml as a real host override.

Replace Friendli-specific output fallback in shared lab metadata with LG's official README max_new_tokens=16,384 quickstart value. Keep 262,144 context from official config.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed c0ce622cf for the remaining EXAONE-236B metadata finding. The prior Friendli fallback was indeed inappropriate in shared lab metadata. LG's official 236B model README explicitly uses max_new_tokens=16,384 in both reasoning and non-reasoning quickstarts, so the lab entry now has output = 16_384 with that official citation. Context remains 262,144 from LG's official config. The retained Friendli entry resolves to { context: 262144, output: 16384 }; bun validate exits 0 and sync remains idempotent.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [possible mistake] models/lgai-exaone/k-exaone-236b-a23b.toml:18 - Check: Lab limit.output must be a real maximum generation ceiling, not a sample default. Why: output = 16_384 is taken from the README max_new_tokens quickstart value, which is a generation default rather than a documented hard maximum. Publishing that as lab metadata understates usable output for every host that inherits this entry (including the retained Friendli deprecated route). Action: Replace with a lab-documented hard maximum output (or keep context-only grounding and avoid inventing a lower ceiling) and cite the source that establishes the limit; do not treat a quickstart max_new_tokens as limit.output.

@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed ec5121524. Agreed: the README quickstart's max_new_tokens=16,384 was not a hard ceiling. The lab now uses only LG's official config max_position_embeddings=262,144 as the documented hard sequence/output ceiling, removes the quickstart claim, and does not cite Friendli at all. bun validate remains exit 0.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] packages/core/src/sync/providers/friendli.ts:1043 - Check: Non-lab hosts must use base_model (and add complete models/<lab>/… metadata when missing); full inline is only for first-party or unique-to-host models. Why: buildFriendliModel still authors a full standalone provider TOML whenever resolveLabModelSync returns nothing. Friendli is a multi-lab relay; hourly sync can therefore introduce third-party full-inline definitions the moment a new HF-style ID appears without a lab file—exactly the AGENTS blocker this PR fixed for EXAONE/DeepSeek. Peers such as DeepInfra skip create when there is no resolvable base. Action: For new remote IDs with no resolvable lab metadata, skip create (and surface via sourceID / notices or skipCreates + missing-model issues) instead of full-inline; only keep full-inline for true host-unique aliases.
  • [medium] [possible mistake] packages/core/src/sync/providers/friendli.ts:721 - Check: Costs must be USD per million tokens; non-token pricing must not be treated as token rates. Why: The deleted generator only authored cost when pricing.unit_type === "TOKEN". The new schema drops unit_type and always multiplies pricing.input/output by 1_000_000. If Friendli still (or again) returns SECOND or other units, the catalog would publish incorrect USD/MTok. Action: Parse and honor unit/type (or equivalent) from the live API; omit or convert non-token pricing, and document the rule in the module.
  • [medium] [possible mistake] models/lgai-exaone/k-exaone-236b-a23b.toml:19 - Check: Lab limit.output should be a lab-grounded generation ceiling, not a host fallback or the full context window by default. Why: Final lab metadata sets output = 262_144 equal to context, after oscillating through 32_768 and the README’s max_new_tokens=16_384. The 750B sibling uses a cited 32_768 generation setup; equating 236B max generation to the full context is a weak lab fact and inflates every host that inherits this base. Action: Confirm LG’s hard max-generation (config, report, or API contract). If only example max_new_tokens exists, prefer a documented generation bound over context, and keep any Friendli-only completion claim on the provider file.
  • [low] [violation] packages/core/src/sync/providers/friendli.ts:1366 - Check: Leading wire-path comments for authored reasoning controls (toggle, and budget/effort when present). Why: toggleHeader only documents chat_template_kwargs.enable_thinking. Files that emit budget_tokens and/or effort (including budget-only MiniMax-M2.5) do not get a durable leading Budget: reasoning_budget / Effort: reasoning_effort=… comment; sync strips non-leading comments on rewrite. Action: Extend the emitted header to cover every control type actually written for that model (at least reasoning_budget, and effort values when present).

…e control headers

- Skip brand-new remote IDs with no resolvable lab metadata (deepinfra-style
  create gate): a multi-lab relay must not auto-author third-party full-inline
  definitions; skippedNotice now names the models/<lab>/ remedy.
- Restore the pricing unit_type guard the deleted generator had: cost is only
  authored for TOKEN unit pricing — a SECOND-priced entry must never be
  multiplied x1M into USD/MTok.
- reasoningHeader() now emits a leading wire-path comment for every authored
  reasoning control type (Toggle, Effort with values, Budget), not just toggle;
  all existing files' headers extended to match.
- 236B lab metadata is LG-source-only: context 262,144 from official config,
  output omitted (LG publishes no generation ceiling). The Friendli 236B
  route keeps its own host-cited limit.output = 262,144 override.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 88c4804d0 resolving all four findings from the 11:04Z round.

  1. [high] Full-inline auto-create for unresolvable IDs — fixed with the deepinfra-style create gate in translateModel: a brand-new remote ID with no resolvable models/<lab>/ metadata returns undefined (skip), while already-tracked files keep updating. A multi-lab relay's hourly sync no longer authors third-party full-inline definitions; skippedNotice now tells the operator to add models/<lab>/<model>.toml and re-sync. Verified: bun models:sync friendli is idempotent (0 created / 0 updated / 9 unchanged) and all 8 live models still factor correctly (including case-mismatched minimax/minimax-m2.5 → minimax/MiniMax-M2.5.toml).

  2. [medium] Pricing unit_type regression — restored the deleted generator's guard exactly: pricing.unit_type is back in the schema (z.enum(["TOKEN", "SECOND"]).optional()), and buildCost returns the existing cost untouched for any non-TOKEN unit instead of multiplying a per-second rate by 1M into USD/MTok. Today's live catalog omits the field for all 8 models (verified against GET /serverless/v1/models), so current data is unaffected; the guard protects the first SECOND-priced entry Friendli re-introduces.

  3. [medium] 236B lab output = 262_144 — agreed it was a weak lab fact, but the 750B-style 32_768 was equally wrong for this model (that figure is 750B-report-specific). Resolution follows your earlier guidance: the lab file now carries only LG primary sources — context 262_144 from the official Hugging Face config, limit.output omitted entirely (LG publishes no generation ceiling; the README max_new_tokens=16,384 is a quickstart default, as you said). The Friendli-serving claim moved to where it belongs: providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml now has an explicit host-cited [limit] output = 262_144 from the live serving catalog (max_completion_tokens, accessed 2026-08-29). Resolved catalog JSON confirms {context: 262144, output: 262144} with status = "deprecated".

  4. [low] Toggle-only headers — reasoningHeader() replaces toggleHeader(): every authored control type now gets its leading wire-path line (Toggle = chat_template_kwargs.enable_thinking, Effort = reasoning_effort with actual values, Budget = reasoning_budget), each with its docs URL. Hand-extended the 9 existing files' headers to parity with what they author (e.g. budget-only MiniMax-M2.5 got its # Budget: reasoning_budget header; GLM-5.2 keeps its stronger model-specific Effort note alongside).

bun validate exit 0; bun run test 253 pass / 4 fail — identical to the pre-push baseline (verified via git stash); sync idempotent.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml:24 - Check: After base_model, provider files must be override-only and must not restate values identical to the lab entry. Why: Final [limit] sets context = 262_144, which already matches models/lgai-exaone/k-exaone-236b-a23b.toml, so the context key is a redundant copy rather than a host delta. Action: Drop the identical limit.context override and keep only a true Friendli delta (for example limit.output alone, if that value is retained).
  • [medium] [possible mistake] providers/friendli/models/LGAI-EXAONE/K-EXAONE-236B-A23B.toml:25 - Check: Do not treat Friendli max_completion_tokens as a real completion cap when it merely mirrors context. Why: This PR already removed that pattern for active Friendli routes because several models report max_completion_tokens == context_length and overwrote tighter lab outputs; the retained 236B file still publishes output = 262_144 from that same serving field. Action: Verify a host-real generation ceiling for this route, or stop publishing context-as-output and leave a lab-grounded/omitted output policy consistent with the rest of the Friendli sync.

…onfig

Drop the provider [limit] block entirely (the [high] finding: context was a
redundant copy of the lab value). Both limits now inherit from the lab entry,
which publishes output = 262_144 from LG's official config
max_position_embeddings — the repo-standard policy for labs without a tighter
documented generation ceiling (kimi-k2.5, mistral-large-2512, 56 lab files
carry output == context). No serving-field completion claim remains on this
route, which is dead: live POST /chat/completions returns HTTP 404.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 803d11623 resolving both findings.

  1. [high] Redundant provider context override — fixed by dropping the provider [limit] block entirely. The Friendli 236B file is now purely override-only: base_model, status = "deprecated", [interleaved], [[reasoning_options]], and [cost]. Both limit values inherit from the lab entry; nothing restates lab data.

  2. [medium] Serving-field output = 262_144 — removed, and the output policy is now lab-grounded instead. Two constraints made omitting the output everywhere impossible: the provider schema requires limit.output on every resolved provider model (dropping the provider [limit] while the lab omitted output fails bun validate with limit.output: Required), and the route itself is dead — the live catalog removed this model and a live POST /chat/completions against LGAI-EXAONE/K-EXAONE-236B-A23B now returns HTTP 404 (verified today), so a host-real generation ceiling can no longer be verified live. The lab entry therefore publishes output = 262_144 grounded in LG's official config max_position_embeddings = 262,144 — LG's only documented hard length bound — which is exactly the repo's standing policy for labs that document no tighter generation ceiling: 56 lab entries (e.g. moonshotai/kimi-k2.5, mistral/mistral-large-2512, bytedance-seed/seed-2.1-pro) publish limit.output = limit.context from their documented sequence bound. No Friendli serving field is cited anywhere anymore.

Verified after the change: bun models:sync friendli idempotent (0 created / 0 updated / 9 unchanged), bun validate exit 0, resolved Friendli 236B block = {context: 262144, output: 262144, status: "deprecated"}, bun run test 253 pass / 4 fail (identical pre-existing baseline, re-checked via git stash).

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Aug 29, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 1, 2026
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [violation] packages/core/src/sync/providers/friendli.ts - Check: AGENTS.md requires a leading top-of-file wire-path comment for each authored reasoning control (toggle / effort / budget). Why: Friendli returns header: reasoningHeader(built) but does not set authoritativeHeaders: true. The runner then keeps existing.header || translated.header and only treats header drift as a write when authoritativeHeaders is enabled, so body-only sameModel matches skip rewrites and body updates still retain a stale leading block. This PR already hit that (GLM headers had to be hand-pasted when the body was unchanged). Later catalog drift that adds or drops controls will refresh reasoning_options while leaving incomplete or wrong wire-path comments. Action: Set authoritativeHeaders: true on the Friendli SyncProvider (or otherwise force reasoningHeader() onto every write) so leading comments always match the authored control set.

@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Closing in favor of a fresh PR rebased on the current sync conventions: Friendli's GLM-5.3 (merged via #5776) already carries the verified reasoning options, and this branch now conflicts with the newer sync provider registrations on dev. Superseded by an upcoming minimal sync PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant