Repository navigation
feat(friendli): add GLM-5.3 provider model - #5776
Conversation
Friendli serverless now serves zai-org/GLM-5.3, base_model'd onto the existing zhipuai/glm-5.3 lab entry. Effort-only reasoning options per Friendli's documented reasoning_effort=low|high|max, limit/modalities inherited from base.
## Action items
- **[high] [violation]** `providers/friendli/models/zai-org/GLM-5.3.toml:27` - **Check:** Relay `reasoning_options` must match this host’s real request surface (lab/peers only where Friendli actually exposes the same controls); do not invent effort. **Why:** Friendli is classified as a multi-model relay. `providers/friendli/provider.toml` documents only `chat_template_kwargs.enable_thinking` and states no effort field is documented. This file’s own comment (lines 22–24) and the PR body admit the same, yet `reasoning_options` still authors `effort` `low|high|max` and omits the documented toggle. That misrepresents caller controls (clients will send unsupported `reasoning_effort` / miss `enable_thinking`). Lab/peers (`providers/zai`, `providers/zhipuai`) use `low|high|max` on Z.AI’s API with always-on thinking; that does not authorize undocumented effort on Friendli. Same-host peers that only have the thinking switch correctly use `toggle` (e.g. DeepSeek-V3.2, gemma-4-31B-it). **Action:** Align options with Friendli’s wire path for `zai-org/GLM-5.3`: if only `chat_template_kwargs.enable_thinking` works, use `[{ type = "toggle" }]` with a leading top-of-file wire comment; if thinking is always on and there is no caller control, use `[]`; only keep `effort` `low|high|max` if Friendli docs or a verified request show `reasoning_effort` (or equivalent) is accepted and effective—cite that in the PR body and a leading comment.
- **[low] [violation]** `providers/friendli/models/zai-org/GLM-5.3.toml:21` - **Check:** After `base_model`, provider files must be override-only (no identical restated fields). **Why:** `name = "GLM-5.3"` matches `models/zhipuai/glm-5.3.toml` exactly, so it is a redundant override. Putting `name` first also pushes the source/rationale comments below the first key. **Action:** Remove the identical `name` so the file stays override-only and source comments sit in a leading block above the first key. |
- Removed duplicate name field (matched models/zhipuai/glm-5.3.toml exactly; base_model already inherits it). - provider.toml: removed the stale "no effort or budget field is documented" claim. Friendli's chat-completions endpoint documents both reasoning_effort (enum) and reasoning_budget (integer) as real request fields — https://friendli.ai/docs/openapi/model-apis/chat-completions#body-reasoning-effort-one-of-0 and #body-reasoning-budget-one-of-0. - Added budget_tokens (min=-1, max=1_048_576) to GLM-5.3's reasoning_options, matching live GET /serverless/v1/models (max_completion_tokens=1_048_576). - Added [limit] override (context=1_048_576, output=1_048_576): live context_length/max_completion_tokens exceed the inherited lab file's rounded 1_000_000 figure.
|
Pushed
|
Action items
|
…M-5.3 Verified via two live POST /chat/completions requests against zai-org/GLM-5.3: reasoning_budget=30 and =50 (with generous max_tokens= 3000/4000) both cut reasoning_content mid-sentence at the requested cap, while completion_tokens continued to 1553/177 respectively -- proof the reasoning phase is capped independently of completion/context capacity. Also clarified that min=-1/max=1_048_576 are Friendli's own reported values from GET /serverless/v1/models, not values models.dev derived from max_completion_tokens itself.
|
Pushed Ran two live requests against
Both requests confirm reasoning and completion are capped independently — On the "never derived from max_tokens/limit.output/context" point: the authored
|
|
No actionable findings. |
Summary
Friendli serverless now serves
zai-org/GLM-5.3(per https://api.friendli.ai/serverless/v1/models). Adds the provider model, usingbase_modelto inherit everything from the existingmodels/zhipuai/glm-5.3.tomllab entry (name, description, limits, modalities, dates, capability flags).What's overridden (provider-only)
reasoning_options:effortonly, valueslow | high | max— Friendli only documentschat_template_kwargs.enable_thinkingas a toggle, not a separate control, so per this repo's convention only the effort levels are exposed (matches siblingGLM-5.2entry style).interleaved:reasoning_content(same field GLM-5.1/5.2 use on this host).cost: input $1.40 / output $4.40 / cache_read $0.26 per MTok, from the API's serverless pricing.limit(context 1,000,000 / output 131,072) andmodalitiesare inherited from the base model, not restated.Checklist
bun validatepasses