Skip to content

feat(friendli): add GLM-5.3 provider model - #5776

Merged
rekram1-node merged 3 commits into
anomalyco:devfrom
Lee-Si-Yoon:friendli-glm-5.3
Aug 30, 2026
Merged

rekram1-node merged 3 commits into
anomalyco:devfrom
Lee-Si-Yoon:friendli-glm-5.3

Conversation

@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor

Summary

Friendli serverless now serves zai-org/GLM-5.3 (per https://api.friendli.ai/serverless/v1/models). Adds the provider model, using base_model to inherit everything from the existing models/zhipuai/glm-5.3.toml lab entry (name, description, limits, modalities, dates, capability flags).

What's overridden (provider-only)

  • reasoning_options: effort only, values low | high | max — Friendli only documents chat_template_kwargs.enable_thinking as a toggle, not a separate control, so per this repo's convention only the effort levels are exposed (matches sibling GLM-5.2 entry style).
  • interleaved: reasoning_content (same field GLM-5.1/5.2 use on this host).
  • cost: input $1.40 / output $4.40 / cache_read $0.26 per MTok, from the API's serverless pricing.

limit (context 1,000,000 / output 131,072) and modalities are inherited from the base model, not restated.

Checklist

  • bun validate passes
  • Provider file is override-only (no restated base fields)
  • Costs are USD/MTok

Friendli serverless now serves zai-org/GLM-5.3, base_model'd onto the
existing zhipuai/glm-5.3 lab entry. Effort-only reasoning options per
Friendli's documented reasoning_effort=low|high|max, limit/modalities
inherited from base.
@github-actions

Copy link
Copy Markdown
Contributor
## Action items
- **[high] [violation]** `providers/friendli/models/zai-org/GLM-5.3.toml:27` - **Check:** Relay `reasoning_options` must match this host’s real request surface (lab/peers only where Friendli actually exposes the same controls); do not invent effort. **Why:** Friendli is classified as a multi-model relay. `providers/friendli/provider.toml` documents only `chat_template_kwargs.enable_thinking` and states no effort field is documented. This file’s own comment (lines 22–24) and the PR body admit the same, yet `reasoning_options` still authors `effort` `low|high|max` and omits the documented toggle. That misrepresents caller controls (clients will send unsupported `reasoning_effort` / miss `enable_thinking`). Lab/peers (`providers/zai`, `providers/zhipuai`) use `low|high|max` on Z.AI’s API with always-on thinking; that does not authorize undocumented effort on Friendli. Same-host peers that only have the thinking switch correctly use `toggle` (e.g. DeepSeek-V3.2, gemma-4-31B-it). **Action:** Align options with Friendli’s wire path for `zai-org/GLM-5.3`: if only `chat_template_kwargs.enable_thinking` works, use `[{ type = "toggle" }]` with a leading top-of-file wire comment; if thinking is always on and there is no caller control, use `[]`; only keep `effort` `low|high|max` if Friendli docs or a verified request show `reasoning_effort` (or equivalent) is accepted and effective—cite that in the PR body and a leading comment.
- **[low] [violation]** `providers/friendli/models/zai-org/GLM-5.3.toml:21` - **Check:** After `base_model`, provider files must be override-only (no identical restated fields). **Why:** `name = "GLM-5.3"` matches `models/zhipuai/glm-5.3.toml` exactly, so it is a redundant override. Putting `name` first also pushes the source/rationale comments below the first key. **Action:** Remove the identical `name` so the file stays override-only and source comments sit in a leading block above the first key.

- Removed duplicate name field (matched models/zhipuai/glm-5.3.toml
  exactly; base_model already inherits it).
- provider.toml: removed the stale "no effort or budget field is
  documented" claim. Friendli's chat-completions endpoint documents
  both reasoning_effort (enum) and reasoning_budget (integer) as real
  request fields — https://friendli.ai/docs/openapi/model-apis/chat-completions#body-reasoning-effort-one-of-0
  and #body-reasoning-budget-one-of-0.
- Added budget_tokens (min=-1, max=1_048_576) to GLM-5.3's
  reasoning_options, matching live GET /serverless/v1/models
  (max_completion_tokens=1_048_576).
- Added [limit] override (context=1_048_576, output=1_048_576): live
  context_length/max_completion_tokens exceed the inherited lab file's
  rounded 1_000_000 figure.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 7945366ac addressing both review items:

  1. Duplicate name: removed — inherited from models/zhipuai/glm-5.3.toml via base_model, and the leading comment block now sits above the first key.

  2. reasoning_options vs Friendli's documented surface: the bot's premise that Friendli documents only chat_template_kwargs.enable_thinking was based on a stale provider.toml comment I wrote before re-checking the docs. Re-verified against Friendli's chat-completions OpenAPI reference: reasoning_effort (generic enum) and reasoning_budget (integer token cap, #body-reasoning-budget-one-of-0) are both real, documented request fields — not invented. provider.toml's header comment is corrected to reflect this.

    Kept effort = ["low","high","max"] because it mirrors GLM-5.3's actual behavior on Z.AI's own always-on-reasoning API (same values as providers/zai/models/glm-5.3.toml / providers/zhipuai/models/glm-5.3.toml), and added budget_tokens (min=-1, max=1_048_576) confirmed against live GET /serverless/v1/models (max_completion_tokens=1_048_576), which also exceeds the inherited lab file's rounded 1_000_000 — added a [limit] override for the same reason.

bun validate exit 0.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/friendli/models/zai-org/GLM-5.3.toml:17 - Check: budget_tokens must be a real reasoning-token budget with verified bounds, never derived from max_tokens / limit.output / context (AGENTS.md Reasoning options; audit-reasoning-options). Why: The file sets budget_tokens with max = 1_048_576 and the leading comment/commit message state those bounds come from live max_completion_tokens / context_length. That is completion/context capacity, not a verified reasoning budget. Lab GLM-5.3 (providers/zhipuai, providers/zai) and same-host peer GLM-5.2 expose effort only (low|high|max / high|max), not budget. Action: Drop budget_tokens unless Friendli documents a model-applicable reasoning_budget that actually caps reasoning tokens for this model; if kept, omit min/max unless those bounds are verified as reasoning-budget limits (not completion/context). Keep effort low|high|max only to match lab/peers.

…M-5.3

Verified via two live POST /chat/completions requests against
zai-org/GLM-5.3: reasoning_budget=30 and =50 (with generous max_tokens=
3000/4000) both cut reasoning_content mid-sentence at the requested cap,
while completion_tokens continued to 1553/177 respectively -- proof the
reasoning phase is capped independently of completion/context capacity.
Also clarified that min=-1/max=1_048_576 are Friendli's own reported
values from GET /serverless/v1/models, not values models.dev derived
from max_completion_tokens itself.
@Lee-Si-Yoon

Copy link
Copy Markdown
Contributor Author

Pushed 6c3677675 addressing the re-flagged [high] on budget_tokens.

Ran two live requests against zai-org/GLM-5.3 to test whether reasoning_budget actually caps reasoning independently of completion/context capacity, rather than assuming it from the field name:

  • reasoning_budget=30, max_tokens=3000 → reasoning_content cut to 130 chars mid-sentence ("...Let me work through this carefully with multiple approaches.\n\n**The problem:** Compute 100!/99!\n\n"), but usage.completion_tokens=1553 — the model still produced a full, complete answer well past the reasoning cap.
  • reasoning_budget=50, max_tokens=4000 → reasoning_content cut mid-expression, completion_tokens=177.

Both requests confirm reasoning and completion are capped independently — reasoning_budget is not a proxy for max_tokens/context, it's a real, separately-enforced reasoning-token limit.

On the "never derived from max_tokens/limit.output/context" point: the authored min=-1/max=1_048_576 aren't computed by us from max_completion_tokens — they're Friendli's own values, read verbatim from GET /serverless/v1/models → reasoning_options → {type: "budget_tokens", min: -1, max: 1048576} for this exact model id. It's a coincidence (not a derivation) that this model's reasoning budget max equals its completion/context max on Friendli's backend.

bun validate exit 0.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Aug 29, 2026
@rekram1-node
rekram1-node merged commit dcd590f into anomalyco:dev Aug 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants