Leaderboard - TokenMark Hardware Rankings
See how devices rank for local LLM inference. Compare Apple M4, RTX 4090, Snapdragon X Elite, AMD Strix, and more on the TokenMark leaderboard.
Snapshot 2026-09-28 · 207 configs · 94 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
AMD Strix Halo (Ryzen AI Max+ 395) · 145 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3 0.6B | Q8_0 | llama.cpp/vulkan | 266 tok/s (prefill 13112) | 512 | Structured table | source |
| Qwen3 0.6B | Q8_0 | llama.cpp/rocm | 208.7 tok/s (prefill 4666.1) | 512 | Structured table | source |
| LFM2.5 8B-A1B | Q4_K_M | llama.cpp/vulkan | 176.5 tok/s (prefill 3398.4) | 512 | Structured table | source |
| Qwen3-30B-A3B-Instruct-2507 | IQ4_XS | llama.cpp/vulkan | 103.2 tok/s (prefill 1438.1) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | Q4_K_S | llama.cpp/vulkan | 98 tok/s (prefill 1406.5) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | UD-Q4_K_XL | llama.cpp/vulkan | 97.1 tok/s (prefill 1400) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | IQ4_XS | llama.cpp/vulkan | 90.4 tok/s (prefill 1372.3) | 512 | Structured table | source |
| Qwen3 30B-A3B NEO-MAX | IQ4_XS | llama.cpp/vulkan | 87.4 tok/s (prefill 1396.1) | 512 | Structured table | source |
| Qwen3.6 35B-A3B | Q4_0 | llama.cpp/vulkan | 81.3 tok/s (prefill 1243.5) | 512 | Structured table | source |
| Qwen3-30B-A3B | UD-Q4_K_XL | llama.cpp/rocm | 79.1 tok/s (prefill 687.9) | Structured table | source | |
| Nemotron Cascade 2 30B-A3B | IQ4_XS | llama.cpp/vulkan | 79 tok/s (prefill 1325.3) | 512 | Structured table | source |
| Gemma-4-26B-A4B | UD-Q4_K_XL | llama.cpp | 78 tok/s | Owner-submitted | source | |
| Ornith 1.5 35B-A3B | FP4 | 76.9 tok/s | Owner-submitted | source | ||
| Nemotron 3 Nano 30B-A3B | IQ4_XS | llama.cpp/vulkan | 76 tok/s (prefill 1312.5) | 512 | Structured table | source |
| Qwen3.5 35B-A3B | IQ4_XS | llama.cpp/vulkan | 75.2 tok/s (prefill 1170.3) | 512 | Structured table | source |
AMD Gorgon Halo (Ryzen AI Max+ PRO 495) · 0 single-stream configs
No single-stream configs recorded yet.
NVIDIA DGX Spark (GB10) · 39 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-heretic | NVFP4 | vllm | 96 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth) | NVFP4 | vllm | 92.6 tok/s | Structured table | source | |
| GPT-OSS-20B | Q4_K_XL | llama.cpp | 90.5 tok/s | Structured table | source | |
| Ornith-1.0-35B-AEON | NVFP4 | vllm | 84.8 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B | NVFP4 | vllm | 77.4 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B | Q4_K_XL | llama.cpp | 62.7 tok/s | Structured table | source | |
| Ornith-1.0-35B | Q8_0 | llama.cpp | 56.1 tok/s | Structured table | source | |
| DeepSeek-V4-Flash-0731-JA-REAP-K216 | EXL3-3BPW | 55 tok/s | 256000 | Extracted, source-linked | source | |
| Qwen3-Coder-Next-80B-A3B | Q4_K_XL | llama.cpp | 50.1 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B | llama.cpp | 48.2 tok/s | Structured table | source | ||
| Nemotron-3-Nano-30B-A3B | llama.cpp | 47.9 tok/s | Structured table | source | ||
| Kimi-Linear-48B-A3B | Q8_0 | llama.cpp | 47.2 tok/s | Structured table | source | |
| Qwen3.8-Flash-Next | NVFP4 | vllm (speculative decode, mtp-2) | 44.2 tok/s | 262144 | Extracted, source-linked | source |
| Nemotron-3-Nano-Omni-30B-A3B | NVFP4 | vllm | 41.7 tok/s | Structured table | source | |
| Nemotron-3-Nano-30B-A3B | FP8 | vllm | 40.7 tok/s | Structured table | source |
Apple M-series Max (MacBook Pro / Mac Studio Max) · 2 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3.8-Flash-Next-REAP-288 | MLX-4BIT | mlx/pmlx | 37 tok/s | Extracted, source-linked | source | |
| Qwen3.8-Flash-Next-REAP-288 | MLX-4BIT | mlx | 28 tok/s | Extracted, source-linked | source |
Apple Mac Studio (M-series Ultra) · 3 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| GLM-5.3-Flash | MLX-4BIT | mlx | 34.2 tok/s | Extracted, source-linked | source | |
| DeepSeek-V4-Flash-0731-MXFP4-MLX-Abliterated | MXFP4 | mlx | 29.6 tok/s (prefill 454) | 1048576 | Extracted, source-linked | source |
| GLM-5.3-Flash Abliterated | MLX-4BIT | mlx (speculative decode, mtp) | 24 tok/s (prefill 365) | 16384 | Extracted, source-linked | source |
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.