feat(sync): replace Friendli generator with SyncProvider - #4477
Lee-Si-Yoon wants to merge 17 commits into
Conversation
|
Checking package scripts, runner header behavior, and peers for EXAONE/DeepSeek factoring. |
Replace packages/core/script/generate-friendli.ts with a SyncProvider module (packages/core/src/sync/providers/friendli.ts) wired into bun models:sync friendli. The Friendli catalog is now synced by the same runner as every other direct provider. Source: https://api.friendli.ai/serverless/v1/models (no auth for models listing). Behavior: - base_model inheritance: when the API base_model resolves to an existing models/<lab>/<file>.toml, emit override-only provider TOML via factorBaseModel. - Full inline when lab metadata is missing or base_model is self-referential: deepseek-ai/DeepSeek-V3.2, LGAI-EXAONE/K-EXAONE-2.0-750B-A37B, LGAI-EXAONE/K-EXAONE-236B-A23B (deprecation_date 2026-08-20; still listed). - id = inference model id (HF-style layout, e.g. providers/friendli/models/deepseek-ai/DeepSeek-V3.2.toml). - reasoning_options: toggle and effort passed through from the API. budget_tokens is a real reasoning-budget control on Friendli (min=-1 unlimited, max=max_completion_tokens) but is stripped from authored TOMLs pending an AGENTS.md rule update — see PR body. - Pricing: per-token USD to per-million-token USD (x1e6). - deleteMissing: false — retain local files for models Friendli rotates out of its catalog. - Case-insensitive base_model resolution (API lowercases ids that exist on disk mixed-case). New lab metadata: models/lgai-exaone/k-exaone-2.0-750b-a37b.toml, models/lgai-exaone/k-exaone-236b-a23b.toml (LG AI Research is the lab; full lab files authored rather than stubs). Models synced from live /v1/models (7): GLM-5.2, gemma-4-31B-it, DeepSeek-V3.2, K-EXAONE-236B-A23B, GLM-5.1, MiniMax-M2.5, K-EXAONE-2.0-750B-A37B.
58856ae to
d7bb0a2
Compare
|
Rebased
|
|
Ready once the |
Action items
|
…els, wire toggle header - resolveLabModelSync now falls back to the model's own id when the API omits base_model (DeepSeek-V3.2, K-EXAONE-2.0-750B-A37B), factoring them against existing models/ metadata instead of shipping duplicate inline definitions. - translateModel/buildFriendliModel mirror the deepinfra/pioneer pattern: a brand-new deprecated model is skipped, but a model we already track keeps its file and gets status = "deprecated" instead of silently falling out of sync with no lifecycle marker. - K-EXAONE-236B-A23B (no longer served by the API) gets status = "deprecated" in place, matching the same field position used by providers/deepinfra and providers/greenpt. - Removed the dead newFileHeader field (never read by the sync runner) and wired a toggleHeader() function into translateModel's header return, matching the llmgateway.ts pattern, so toggle-bearing files get their wire-path comment from the sync itself.
Action items
|
…ore toggle headers, preserve interleaved - K-EXAONE-236B-A23B (no longer served, status = "deprecated") now factors against models/lgai-exaone/k-exaone-236b-a23b.toml instead of staying a full inline duplicate — matches the base_model override-only rule that already applied to the other EXAONE entry. - GLM-5.1 and GLM-5.2 provider files get the leading toggle wire-path comment that sync only writes on rewrite; sameModel treated the unchanged body as no-op so the header never landed. Added directly to match every other toggle-bearing Friendli file. - GLM-5.2 also gets an explicit comment scoping the confirmed reasoning_effort=high|max control to this model only (cited to Friendli's own GLM-5.2 reasoning doc), since GLM-5.1 documents toggle only and the family-wide claim was the actual violation risk. - translateInterleaved now falls back to the existing authored value when the API reports interleaved: false, instead of dropping it. The API is inconsistent here (DeepSeek-V3.2 reports false while peers DeepInfra/Chutes confirm reasoning_content is real); trusting an existing authored value avoids silently regressing hosts the API under-reports on future syncs.
Action items
|
Called /v1/chat/completions against deepseek-ai/DeepSeek-V3.2 with enable_thinking=true — the response includes both reasoning and reasoning_content, confirming interleaved reasoning_content is real for this model despite /v1/models reporting interleaved: false. Replaced the "other hosts confirm it" comment with the actual verification.
Action items
|
|
@rekram1-node Final follow-up pushed. Verified the DeepSeek-V3.2 Also fixed from the last review round: K-EXAONE-236B-A23B now factors against its lab metadata (
|
- Resolve providers/index.ts conflict: keep friendli alongside the incoming github-copilot entry (alphabetical, both providers list). - translateReasoningOptions no longer drops budget_tokens: the live GET /serverless/v1/models response and https://friendli.ai/docs/openapi/model-apis/chat-completions#body-reasoning-budget-one-of-0 confirm reasoning_budget is a real, documented request field (min = -1 unlimited, max = max_completion_tokens), contradicting the prior comment claiming it "is not a real reasoning budget control on this host". Normalizes min to -1 to match the hand-authored GLM-5.x convention regardless of what the API reports. - Re-ran bun models:sync friendli: GLM-5.1/5.2/5.3/5.3-Flash and the new DeepSeek/EXAONE/MiniMax entries all pick up budget_tokens.
|
Pushed |
Action items
|
… fix provider.toml - translateInterleaved now consults a durable VERIFIED_INTERLEAVED_OVERRIDES map keyed by model id instead of relying solely on the fragile "existing on-disk value" fallback. The prior design silently forgot a live-verified capability if the authored file ever lost the field for any reason (which is exactly what happened to DeepSeek-V3.2 in this PR's own resync). Re-verified live: a real /chat/completions request with enable_thinking=true still returns both reasoning and reasoning_content for deepseek-ai/DeepSeek-V3.2 despite /v1/models reporting interleaved: false. - google/gemma-4-31B-it: verified image input is real (not just catalog metadata) with a live request -- base64 image POST consumed 280 prompt tokens vs 22 for an equivalent text-only request, proving the image was actually processed. No modalities override needed (file correctly inherits the lab's text+image modalities); added a citing comment. - provider.toml: removed the stale "no effort or budget field is documented" claim (same fix already applied on the friendli-glm-5.3 branch), replaced with a comment citing reasoning_effort and reasoning_budget as real OpenAPI-documented request fields. bun models:sync friendli remains idempotent after these changes.
|
Pushed
|
Action items
|
Friendli's max_completion_tokens equals context_length for several models it serves (GLM-5.2/5.3/5.3-Flash, gemma-4-31B-it, both EXAONE entries, DeepSeek-V3.2), and the sync was blindly asserting that value as limit.output on every generated file -- including ones with a base_model to inherit from. That silently overwrote lab-verified, genuinely tighter completion caps: DeepSeek-V3.2's lab file documents output=64_000 out of a 128_000 context, not "same as context". Per user direction: when a base_model exists, defer to its own limit.output via factorBaseModel's normal inheritance instead of forcing Friendli's completion-length report onto it. Only fall back to Friendli's own max_completion_tokens for full-inline entries that have no base_model to inherit a real completion-cap policy from. context override is unaffected -- Friendli's live serving context is a real, host-specific fact worth overriding regardless of base_model. Verified resolved catalog after the fix: deepseek-ai/DeepSeek-V3.2 -> output=64_000 (was 163_840) zai-org/GLM-5.1/5.2/5.3/5.3-Flash -> output=131_072 (was 202_752/1_048_576) google/gemma-4-31B-it -> output=32_768 (was 262_144) LGAI-EXAONE/K-EXAONE-2.0-750B-A37B -> output=32_768 (was 262_144) bun models:sync friendli remains idempotent; bun run test unchanged at 253 pass / 4 fail (same pre-existing failures as prior rounds).
|
Pushed You were right to flag it: Fix: when a Verified resolved catalog after the fix (
On the |
Action items
|
…236B K-EXAONE-236B-A23B.toml is retained (deleteMissing: false) after Friendli dropped it from the live catalog, so bun models:sync friendli never touches it and the file kept a stale limit.output = 262_144 from before the base_model-inheritance fix landed on the other 8 files. Removed the override so the resolved model correctly inherits the lab file's output = 32_768, matching every other factored Friendli entry.
|
Pushed `53ee2cc17` fixing the `[medium]` and responding to both `[high]` findings with evidence. `[medium]` K-EXAONE-236B-A23B stale `limit.output` — fixed. This model was retained after Friendli dropped it from the live catalog (`deleteMissing: false`), so `bun models:sync friendli` never touches it and it kept the pre-fix `output = 262_144` override from before the base_model-inheritance change landed on the other 8 files. Removed the override; resolved catalog now correctly shows `output = 32_768` (inherited from `models/lgai-exaone/k-exaone-236b-a23b.toml`), matching every other factored entry. `[high]` #1 "budget_tokens max invents a reasoning-budget bound from completion capacity" — re-tested this claim directly instead of assuming either side is right. Ran live `reasoning_budget=30` (generous `max_tokens=3000`) requests against three different models with different reasoning architectures:
All three show the same pattern: a small `reasoning_budget` truncates only the reasoning phase while the completion phase continues normally to a full answer. That's consistent, cross-model evidence that `reasoning_budget` is a real, independently-enforced control — not a value we invented from `max_completion_tokens`. The `max` we author is Friendli's own reported `reasoning_options.budget_tokens.max` from `GET /serverless/v1/models`, read verbatim; it coincidentally equals `max_completion_tokens` for these particular models on Friendli's backend, but we aren't deriving one from the other. `[high]` #2 "translator still does a bare `budget_tokens` continue/strip" — checked the exact pushed commit (`0cc8f2caf`, the SHA this review ran against) directly: `translateReasoningOptions` at that commit pushes `{type: "budget_tokens", min: -1, max: option.max}`, not a bare `continue`. Re-ran `bun models:sync friendli` fresh in a clean worktree from that exact commit — 0 created/updated, fully idempotent, `budget_tokens` still present in every regenerated file. I can't find the bare-strip code path this finding describes anywhere in the current branch; it only ever existed in the very first commit (`d7bb0a284`) and was removed in `642cc5fcc`. If there's a specific line/file the next review can point to, happy to look again. `bun validate` exit 0, `bun models:sync friendli` idempotent, `bun run test` unchanged at 253 pass / 4 fail. |
Action items
|
…iptions reasoning_budget is a real independently-enforced Friendli control (verified live on GLM-5.3, gemma-4-31B-it, and DeepSeek-V3.2), but the catalog's reported min/max are not safe published constraints: GLM-5.3 accepted reasoning_budget=1_048_577 despite reporting max=1_048_576. Emit budget_tokens as an unbounded capability rather than publishing a range that the API does not enforce. For entries with a base_model, stop overriding the canonical lab description with generic API text. Factored provider entries now inherit the richer lab description; full-inline entries keep the API description. Re-synced 8 active Friendli models. budget_tokens remains present on all budget-capable active models without fabricated min/max bounds, and the sync is idempotent.
|
Pushed
Validation: |
Action items
|
- Retained deleteMissing:false entries no longer remain permanently deprecated if they reappear active in Friendli's catalog. Clear only a prior deprecated status; preserve intentional lifecycle values such as beta. - Verified MiniMaxAI/MiniMax-M2.5 reasoning_budget with a live request: reasoning_budget=30 truncates reasoning_content at 135 chars while the completion continues for 682 tokens. MiniMax therefore genuinely has the independently-enforced budget control, and keeps the unbounded budget_tokens policy used for other verified Friendli models. bun models:sync friendli idempotent; bun validate exits 0; tests unchanged at 253 pass / 4 known failures.
|
Pushed
|
Action items
|
- Factored modality overrides now include only API-defined sides, so an omitted side deep-merges from lab metadata instead of being overwritten with an empty array. - Cite the official K-EXAONE 2.0 report's 32,768 generation setup for 750B; use Friendli's documented 262,144 fallback for 236B because LG's public config supplies context but no lower generation ceiling. - Document GLM-5.3-Flash's real host modality delta: live catalog omits PDF and a minimal PDF data URL is rejected by chat completions (422).
|
Pushed
Fresh |
Action items
|
Replace Friendli-specific output fallback in shared lab metadata with LG's official README max_new_tokens=16,384 quickstart value. Keep 262,144 context from official config.
|
Pushed |
Action items
|
|
Pushed |
Action items
|
…e control headers - Skip brand-new remote IDs with no resolvable lab metadata (deepinfra-style create gate): a multi-lab relay must not auto-author third-party full-inline definitions; skippedNotice now names the models/<lab>/ remedy. - Restore the pricing unit_type guard the deleted generator had: cost is only authored for TOKEN unit pricing — a SECOND-priced entry must never be multiplied x1M into USD/MTok. - reasoningHeader() now emits a leading wire-path comment for every authored reasoning control type (Toggle, Effort with values, Budget), not just toggle; all existing files' headers extended to match. - 236B lab metadata is LG-source-only: context 262,144 from official config, output omitted (LG publishes no generation ceiling). The Friendli 236B route keeps its own host-cited limit.output = 262,144 override.
|
Pushed
|
Action items
|
…onfig Drop the provider [limit] block entirely (the [high] finding: context was a redundant copy of the lab value). Both limits now inherit from the lab entry, which publishes output = 262_144 from LG's official config max_position_embeddings — the repo-standard policy for labs without a tighter documented generation ceiling (kimi-k2.5, mistral-large-2512, 56 lab files carry output == context). No serving-field completion claim remains on this route, which is dead: live POST /chat/completions returns HTTP 404.
|
Pushed
Verified after the change: |
|
No actionable findings. |
Action items
|
|
Closing in favor of a fresh PR rebased on the current sync conventions: Friendli's GLM-5.3 (merged via #5776) already carries the verified reasoning options, and this branch now conflicts with the newer sync provider registrations on dev. Superseded by an upcoming minimal sync PR. |
Summary
Replaces
packages/core/script/generate-friendli.tswith aSyncProvideratpackages/core/src/sync/providers/friendli.ts, wired intobun models:sync friendli.Source:
https://api.friendli.ai/serverless/v1/models.Behavior
base_modelmetadata case-insensitively and generates override-only provider TOMLs. Friendli's HF-style IDs map to catalog labs where needed (for examplezai-org→zhipuaianddeepseek-ai→deepseek).models/<lab>/metadata to factor onto is skipped with a notice — a multi-lab relay's hourly sync never authors third-party full-inline definitions. Already-tracked files keep updating.deleteMissing: false) and marks live entries deprecated fromdeprecation_date; clears a retaineddeprecatedstatus if that model later returns active, while preserving other authored statuses such asbeta.toggle,effort, andbudget_tokenscontrols.reasoning_budgetis verified as independently enforced on GLM-5.3, gemma-4-31B-it, DeepSeek-V3.2, and MiniMax-M2.5: small budgets truncatereasoning_contentwhile completion continues. The catalog emits unbounded{ type = "budget_tokens" }, notmin/maxbounds, because Friendli accepts values above its advertised maximum and those bounds are not safe published constraints.DeepSeek-V3.2reasoning_contentdespite the catalog's staleinterleaved: false, via a durable model-ID override.limit.contextfrom Friendli's live serving catalog. For a factored model, inherits the lab'slimit.output; only a full-inline model falls back to Friendli'smax_completion_tokens.unit_type = "TOKEN": any other unit (e.g.SECOND) authors no cost section rather than inventing a per-million token price.[]. GLM-5.3-Flash's explicit text/image/video input override is retained because Friendli's catalog omits PDF and a live PDF data-URL request returns HTTP 422.Lab metadata
models/lgai-exaone/k-exaone-2.0-750b-a37b.toml: 262,144 context from the official Hugging Face config; 32,768 output cited from the K-EXAONE 2.0 Technical Report evaluation setup.models/lgai-exaone/k-exaone-236b-a23b.toml: 262,144 context from the official Hugging Face config, LG sources only. LG publishes no separate generation ceiling, solimit.outputis deliberately omitted from the lab entry; Friendli's own serving completion cap for this route lives on the provider file instead (output = 262_144, host-cited).Verification
bun models:sync friendli— idempotent (0 created, 0 updatedon re-run)bun validate— exit 0bun run test— 253 pass / 4 fail (same known baseline failures, confirmed viagit stashon this branch; no regression)bunx tsc --noEmit— no errors fromsync/providers/friendli.ts; an existing unrelatedsync/index.tsSyncedMetadata.reasoning_optionserror predates this PRSupersedes #3741.