Skip to content

Update MLX support for mlx-swift-lm 3.32.3 - #274

Merged
mattt merged 4 commits into
mainfrom
mattt/mlx-swift-lm-3.32
Oct 2, 2026
Merged

mattt merged 4 commits into
mainfrom
mattt/mlx-swift-lm-3.32

Conversation

@mattt

@mattt mattt commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

mlx-swift-lm 3.32.3 came out on September 30, and from: "3.31.4" resolves to it. It makes source-breaking changes, so the MLX trait doesn't compile against it: newCache(parameters:) now throws, prepare(_:cache:windowSize:) is replaced by prepare(_:cache:state:prefill:), and Generation has a new rejectedToolCall case.

This PR updates MLXLanguageModel for those changes and requires mlx-swift-lm 3.32.x. mlx-swift-lm's minor and patch versions follow mlx-swift's, so the requirement now uses .upToNextMinor to keep the next MLX update from breaking builds the same way. That version uses mlx-swift 0.32.3, which vendors MLX 0.32.2 and includes the fix for the fused attention causal mask (ml-explore/mlx#3271).

mlx-swift 0.32.3 declares swift-tools-version 6.3, so SwiftPM can't resolve it with earlier toolchains, even when the MLX trait is off. This PR raises AnyLanguageModel's requirement to Swift 6.3 (Xcode 26.4 or later), and updates the CI matrix and README to match.

A rejected tool call (for example, one that is malformed or incomplete) now throws RejectedToolCallError. Before, mlx-swift-lm dropped that output or passed it through as text. A call to a tool that the session doesn't declare still gets a "Tool not found" result, as with the other providers.

Models that place tokens with M-RoPE (Qwen3-VL, Qwen3.5, GLM-OCR) now throw ContinuationStateError when they continue a reused cache without the model state that goes with it. The session cache doesn't keep that state, so for those models generation falls back to an empty cache.

The structured generation path now synchronizes the current GPU stream instead of a new stream, which had no pending work.

mlx-swift-lm 3.32 makes newCache(parameters:) throw, replaces prepare(_:cache:windowSize:) with prepare(_:cache:state:prefill:), and adds Generation.rejectedToolCall. AnyLanguageModel's MLX trait no longer compiled against it, and from: "3.31.4" resolves to 3.32.3. Require 3.32.3, which also brings in MLX 0.32.2 with the fused attention causal-mask fix (ml-explore/mlx#3271).

A rejected tool call now throws RejectedToolCallError. Synchronize the current GPU stream instead of a new, empty one.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Warm-cache paths do not preserve the model state required by mlx-swift-lm 3.32.3.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Updates MLX integration for mlx-swift-lm 3.32.3 compatibility.

Changes:

  • Adapts cache and prefill APIs.
  • Handles rejected tool calls.
  • Synchronizes the active GPU stream.
File Description
Package.swift Raises the mlx-swift-lm minimum version.
Sources/​AnyLanguageModel/​Models/​MLXLanguageModel.swift Updates MLX generation, caching, tool handling, and synchronization.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +860 to +862
) throws -> (
cache: [MLXLMCommon.KVCache], input: MLXLMCommon.LMInput, fullTokens: [Int32], cachedTokenCount: Int
) {
mlx-swift 0.32.3 declares swift-tools-version 6.3 and mlx-swift-lm 3.32.3 declares 6.2, so the package no longer resolves with earlier toolchains. Raise the tools version, test the Swift 6.3 floor with Xcode 26.4.1, and drop the Swift 6.1 and 6.2 CI jobs. Remove the README note about building with Xcode 16; #15 was fixed by #59.
@mattt
mattt requested a balanced review from Copilot October 1, 2026 21:44

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Warm-cache reuse drops model state required by state-producing MLX models, causing subsequent generation to fail.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Prewarm discards model state required with cached KV

Sources/​AnyLanguageModel/​Models/​MLXLanguageModel.swift:1471

This prewarm call discards the LMOutput.State returned by prepare and then stores the warmed KV cache alone. State-producing models in mlx-swift-lm 3.32.3 require that paired state when the cache is reused, so a prewarmed Qwen3-VL or GLM-OCR session can fail with ContinuationStateError.missingState. Store the returned state with newCache, or do not retain this prewarm cache when state cannot be preserved.

Low severity Incorrect upstream link for fused-attention causal-mask fix

Package.swift:44

The PR description attributes a fused-attention causal-mask fix to ml-explore/mlx#3271, but that linked pull request is titled “Nax Refactor” and does not describe such a fix. Please correct or remove the link so the dependency-upgrade rationale points to the intended upstream change.

In mlx-swift-lm 3.32.3, models that place tokens with M-RoPE (Qwen3-VL, Qwen3.5, GLM-OCR) throw ContinuationStateError when they continue a warm cache without the LMOutput.State that goes with it. The session cache keeps only the KV cache, so drop it and start from an empty cache when that happens.
@mattt
mattt requested a balanced review from Copilot October 1, 2026 21:54

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Generation omits declared tool schemas, preventing undeclared calls from producing the intended rejection error.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Pass declared tool schemas to generate to reject undeclared tools

Sources/​AnyLanguageModel/​Models/​MLXLanguageModel.swift:934

Pass the declared tool schemas to generate. In mlx-swift-lm 3.32.3, the default tools: nil accepts any parsed function name, so an undeclared tool is emitted as .toolCall rather than .rejectedToolCall; this path then returns a synthetic “Tool not found” result instead of throwing the promised RejectedToolCallError. Use an empty array when the session declares no tools so authorization still rejects tool-shaped output, and apply it to both the initial and retry calls.

mlx-swift-lm's minor and patch versions follow mlx-swift's, and 3.32 made source-breaking changes in a minor release. Require .upToNextMinor so a later MLX update doesn't break builds again.
@mattt
mattt merged commit 934d352 into main Oct 2, 2026
7 checks passed
@mattt
mattt deleted the mattt/mlx-swift-lm-3.32 branch October 2, 2026 09:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants