You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
mlx-swift-lm 3.32.3 came out on September 30, and from: "3.31.4" resolves to it. It makes source-breaking changes, so the MLX trait doesn't compile against it: newCache(parameters:) now throws, prepare(_:cache:windowSize:) is replaced by prepare(_:cache:state:prefill:), and Generation has a new rejectedToolCall case.
This PR updates MLXLanguageModel for those changes and requires mlx-swift-lm 3.32.x. mlx-swift-lm's minor and patch versions follow mlx-swift's, so the requirement now uses .upToNextMinor to keep the next MLX update from breaking builds the same way. That version uses mlx-swift 0.32.3, which vendors MLX 0.32.2 and includes the fix for the fused attention causal mask (ml-explore/mlx#3271).
mlx-swift 0.32.3 declares swift-tools-version 6.3, so SwiftPM can't resolve it with earlier toolchains, even when the MLX trait is off. This PR raises AnyLanguageModel's requirement to Swift 6.3 (Xcode 26.4 or later), and updates the CI matrix and README to match.
A rejected tool call (for example, one that is malformed or incomplete) now throws RejectedToolCallError. Before, mlx-swift-lm dropped that output or passed it through as text. A call to a tool that the session doesn't declare still gets a "Tool not found" result, as with the other providers.
Models that place tokens with M-RoPE (Qwen3-VL, Qwen3.5, GLM-OCR) now throw ContinuationStateError when they continue a reused cache without the model state that goes with it. The session cache doesn't keep that state, so for those models generation falls back to an empty cache.
The structured generation path now synchronizes the current GPU stream instead of a new stream, which had no pending work.
mlx-swift-lm 3.32 makes newCache(parameters:) throw, replaces prepare(_:cache:windowSize:) with prepare(_:cache:state:prefill:), and adds Generation.rejectedToolCall. AnyLanguageModel's MLX trait no longer compiled against it, and from: "3.31.4" resolves to 3.32.3. Require 3.32.3, which also brings in MLX 0.32.2 with the fused attention causal-mask fix (ml-explore/mlx#3271).
A rejected tool call now throws RejectedToolCallError. Synchronize the current GPU stream instead of a new, empty one.
mlx-swift 0.32.3 declares swift-tools-version 6.3 and mlx-swift-lm 3.32.3 declares 6.2, so the package no longer resolves with earlier toolchains. Raise the tools version, test the Swift 6.3 floor with Xcode 26.4.1, and drop the Swift 6.1 and 6.2 CI jobs. Remove the README note about building with Xcode 16; #15 was fixed by #59.
This prewarm call discards the LMOutput.State returned by prepare and then stores the warmed KV cache alone. State-producing models in mlx-swift-lm 3.32.3 require that paired state when the cache is reused, so a prewarmed Qwen3-VL or GLM-OCR session can fail with ContinuationStateError.missingState. Store the returned state with newCache, or do not retain this prewarm cache when state cannot be preserved.
Incorrect upstream link for fused-attention causal-mask fix
Package.swift:44
The PR description attributes a fused-attention causal-mask fix to ml-explore/mlx#3271, but that linked pull request is titled “Nax Refactor” and does not describe such a fix. Please correct or remove the link so the dependency-upgrade rationale points to the intended upstream change.
In mlx-swift-lm 3.32.3, models that place tokens with M-RoPE (Qwen3-VL, Qwen3.5, GLM-OCR) throw ContinuationStateError when they continue a warm cache without the LMOutput.State that goes with it. The session cache keeps only the KV cache, so drop it and start from an empty cache when that happens.
Pass the declared tool schemas to generate. In mlx-swift-lm 3.32.3, the default tools: nil accepts any parsed function name, so an undeclared tool is emitted as .toolCall rather than .rejectedToolCall; this path then returns a synthetic “Tool not found” result instead of throwing the promised RejectedToolCallError. Use an empty array when the session declares no tools so authorization still rejects tool-shaped output, and apply it to both the initial and retry calls.
mlx-swift-lm's minor and patch versions follow mlx-swift's, and 3.32 made source-breaking changes in a minor release. Require .upToNextMinor so a later MLX update doesn't break builds again.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
mlx-swift-lm 3.32.3 came out on September 30, and
from: "3.31.4"resolves to it. It makes source-breaking changes, so theMLXtrait doesn't compile against it:newCache(parameters:)now throws,prepare(_:cache:windowSize:)is replaced byprepare(_:cache:state:prefill:), andGenerationhas a newrejectedToolCallcase.This PR updates
MLXLanguageModelfor those changes and requires mlx-swift-lm 3.32.x. mlx-swift-lm's minor and patch versions follow mlx-swift's, so the requirement now uses.upToNextMinorto keep the next MLX update from breaking builds the same way. That version uses mlx-swift 0.32.3, which vendors MLX 0.32.2 and includes the fix for the fused attention causal mask (ml-explore/mlx#3271).mlx-swift 0.32.3 declares
swift-tools-version6.3, so SwiftPM can't resolve it with earlier toolchains, even when theMLXtrait is off. This PR raises AnyLanguageModel's requirement to Swift 6.3 (Xcode 26.4 or later), and updates the CI matrix and README to match.A rejected tool call (for example, one that is malformed or incomplete) now throws
RejectedToolCallError. Before, mlx-swift-lm dropped that output or passed it through as text. A call to a tool that the session doesn't declare still gets a "Tool not found" result, as with the other providers.Models that place tokens with M-RoPE (Qwen3-VL, Qwen3.5, GLM-OCR) now throw
ContinuationStateErrorwhen they continue a reused cache without the model state that goes with it. The session cache doesn't keep that state, so for those models generation falls back to an empty cache.The structured generation path now synchronizes the current GPU stream instead of a new stream, which had no pending work.