Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
falcon-h1-tiny-90m-instruct-4bit,models/falcon-h1-tiny-90m-instruct-4bit,17,30,14.24,1193.95,84.64,354.43,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
gemma-4-26b-a4b-it-4bit,models/gemma-4-26b-a4b-it-4bit,20,26,148.21,134.94,432.81,60.07,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
gemma-4-26b-a4b-it-4bit,models/gemma-4-26b-a4b-it-4bit,20,26,148.87,134.35,436.46,59.57,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
gemma-4-26b-a4b-it-4bit,models/gemma-4-26b-a4b-it-4bit,20,26,151.83,131.73,434.17,59.88,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
gemma-4-26b-a4b-it-qat-4bit,models/gemma-4-26b-a4b-it-qat-4bit,20,26,150.25,133.11,484.35,53.68,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
gemma-4-26b-a4b-it-qat-4bit,models/gemma-4-26b-a4b-it-qat-4bit,20,26,149.90,133.43,486.17,53.48,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
gemma-4-26b-a4b-it-qat-4bit,models/gemma-4-26b-a4b-it-qat-4bit,20,26,149.73,133.57,484.07,53.71,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
glm4-flash-4bit,models/glm4-flash-4bit,12,18,120.71,99.41,379.61,47.42,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
glm4-flash-4bit,models/glm4-flash-4bit,12,18,113.79,105.46,359.52,50.07,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
glm4-flash-4bit,models/glm4-flash-4bit,12,18,113.37,105.85,387.74,46.42,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
granite-4.0-h-350m-4bit,models/granite-4.0-h-350m-4bit,37,19,19.92,1857.41,110.97,171.22,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
granite-4.0-h-tiny-4bit,models/granite-4.0-h-tiny-4bit,15,42,62.89,238.51,414.33,101.37,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
lfm2-8b-a1b-4bit,models/lfm2-8b-a1b-4bit,17,37,114.22,148.84,223.53,165.53,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
mamba2-130m,models/mamba2-130m,7,100,7.84,892.53,675.77,147.98,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
mamba2-130m,models/mamba2-130m,7,100,7.84,892.83,492.66,202.98,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
mamba2-130m,models/mamba2-130m,7,100,6.99,1001.24,488.35,204.77,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
nemotron-3-nano-omni-30b-a3b-reasoning-4bit,models/nemotron-3-nano-omni-30b-a3b-reasoning-4bit,23,20,190.12,120.97,241.32,82.88,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
nemotron-h-30b-4bit,models/nemotron-h-30b-4bit,23,46,203.10,113.24,526.27,87.41,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
nemotron-nas-30b-4bit,models/nemotron-nas-30b-4bit,23,46,194.94,117.99,535.47,85.91,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
plamo-2-1b,models/plamo-2-1b,7,100,35.28,198.43,2134.72,46.84,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
model,model_path,prompt_tokens,generated_tokens,prefill_ms,prefill_tok_s,decode_ms,decode_tok_s,date,hardware,mlx_version,build_type,max_tokens,prompt,prompt_target_len
qwen3-30b-a3b-4bit,models/qwen3-30b-a3b-4bit,19,34,141.41,134.37,360.23,94.39,2026-07-12,NVIDIA_GB10_122GB,0.4.0-rc.1,release,100,"Hello, how are you today?",
6 changes: 3 additions & 3 deletions docs/benchmark_results/model_tests.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,7 +71,7 @@ The table below summarizes the current cross-hardware decode readings for select
| Gemma-3-1B | 1B | 227.38 | 395.02 | 278.52 |
| EXAONE-3.5-2.4B | 2.4B | 199.11 | 287.66 | 141.83 |
| SmolLM3-3B | 3B | 131.45 | 234.21 | 104.24 |
| Nemotron-H-30B | 30B | 91.75 | 176.95 | 79.94¶ |
| Nemotron-H-30B | 30B | 91.75 | 176.95 | 87.41¶ |
| Qwen3-MoE-30B | 30B | 83.42 | 175.48 | 89.06† |
| Llama-3.1-8B | 8B | 106.63 | 117.46 | 50.53 |
| Qwen2.5-7B | 7B | 108.47 | 126.20 | 54.56 |
Expand All @@ -82,11 +82,11 @@ The table below summarizes the current cross-hardware decode readings for select
*Qwen3-0.6B on GB10 again stopped at 9 tokens before EOS (2026-07-12); the 283.90 tok/s figure is from that short window and is not directly comparable to full-length runs.
†Qwen3-MoE-30B (`qwen3-moe-4bit`) **failed** on GB10 at 0.3.0 (Metal-only fused-MoE kernel aborted on CUDA); the CUDA fused decode-MoE kernel (#319) restored it at 0.3.1, and at 89.06 tok/s it stays ahead of M1 Ultra (83.75).
§GPT-OSS-120B and Solar-Open-100B were excluded from the 2026-07-12 GB10 sweep by the memory gate (weights > ~51 GiB, `SKIP:oom_estimate`); their figures are carried from the 2026-06-17 / 0.3.1 sweep.
¶Nemotron-H-30B doubled vs both earlier GB10 records (40.32 on 2026-06-17) with no SSM-related code change in between; the whole SSM/hybrid cluster reads 2-3x higher on 2026-07-12 and should be re-verified after a fresh boot (see the GB10 file's notable-changes list).
¶Nemotron-H-30B doubled vs the 2026-06-17 record (40.32) because the fused single-token SSM decode kernel was ported to CUDA on 2026-07-10 (#727); the post-reboot re-verification (#755) confirmed the gain on a fresh host (87.41, post-reboot single). The whole SSM/hybrid cluster carries the same attribution (see the GB10 file's notable-changes list).

M1 Ultra column is from 2026-07-12 with mlxcel 0.4.0-rc.1 / MLX pin `57c66cac` / `--cooldown 30 --big-cooldown 30`, using the `mlxcel-bench-decode` same-process harness.
M5 Max column is from the 2026-07-11/12 full re-sweep with mlxcel 0.4.0-rc.1 / MLX pin `57c66cac` / `--cooldown 30 --big-cooldown 30`, same-process `mlxcel-bench-decode` harness.
GB10 column is from 2026-07-12 with mlxcel 0.4.0-rc.1 / MLX pin `57c66cac` (0.32.1) / CUDA 13.0 (SM 12.1) / `--cooldown 15 --big-cooldown 45`, using the `mlxcel-bench-decode` same-process warm harness, except the two `§`-marked memory-gated rows carried from 2026-06-17 / 0.3.1.
GB10 column is from 2026-07-12 with mlxcel 0.4.0-rc.1 / MLX pin `57c66cac` (0.32.1) / CUDA 13.0 (SM 12.1) / `--cooldown 15 --big-cooldown 45`, using the `mlxcel-bench-decode` same-process warm harness, except the two `§`-marked memory-gated rows carried from 2026-06-17 / 0.3.1 and the `¶`-marked Nemotron-H row, which is the post-reboot single from the same day (#755, `--cooldown 30`).
All three columns now share mlxcel 0.4.0-rc.1 and the MLX pin `57c66cac`, so the Apple Silicon gap reflects hardware delta. M5 Max stays roughly 1.73x faster than M1 Ultra on the selected 16 rows (avg ~1.73x, median ~1.77x). The largest MoE rows show the M5 Max advantage: qwen3-moe-30b runs at 175.48 vs 83.42 tok/s (2.10x), gpt-oss-120b at 113.90 vs 59.29 (1.92x), and solar-open-100b at 65.40 vs 35.02 (1.87x). On GB10 the CUDA fused decode-MoE kernel (#319) keeps qwen3-moe-30b (89.06) just ahead of M1 Ultra (83.42).
For Qwen2.5-0.5B the 4-bit row is the directly comparable cross-hardware figure; the bf16 variant runs at 295.65 tok/s on M1 Ultra and 401.49 tok/s on M5 Max.

Expand Down
Loading