TokenMark: Local AI Config Tracker for Strix Halo & DGX Spark
Measured tok/s configs for local LLMs on AMD Strix Halo, NVIDIA DGX Spark and Apple Mac, source-linked and updated daily. One searchable leaderboard.
Snapshot 2026-10-04 · 254 configs · 116 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
AMD Strix Halo (Ryzen AI Max+ 395) · 159 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3 0.6B | Q8_0 | llama.cpp/vulkan | 266 tok/s (prefill 13112) | 512 | Structured table | source |
| Qwen3 0.6B | Q8_0 | llama.cpp/rocm | 208.7 tok/s (prefill 4666.1) | 512 | Structured table | source |
| LFM2.5 8B-A1B | Q4_K_M | llama.cpp/vulkan | 176.5 tok/s (prefill 3398.4) | 512 | Structured table | source |
| Ornith-1.5-35B-A3B | FP4 | (speculative decode, mtp) | 105.6 tok/s | Owner-submitted | source | |
| Qwen3-30B-A3B-Instruct-2507 | IQ4_XS | llama.cpp/vulkan | 103.2 tok/s (prefill 1438.1) | 512 | Structured table | source |
AMD Gorgon Halo (Ryzen AI Max+ PRO 495) · 0 single-stream configs
No single-stream configs recorded yet.
NVIDIA DGX Spark (GB10) · 46 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3 1.7B | Q4_K_M | llama.cpp | 146.1 tok/s | 2048 | Extracted, source-linked | source |
| Qwen3.6-35B-heretic | NVFP4 | vllm | 96 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth) | NVFP4 | vllm | 92.6 tok/s | Structured table | source | |
| GPT-OSS-20B | Q4_K_XL | llama.cpp | 90.5 tok/s | Structured table | source | |
| Ministral 3B | Q4_K_M | llama.cpp | 86.6 tok/s | 2048 | Extracted, source-linked | source |
Apple M-series Max (MacBook Pro / Mac Studio Max) · 23 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| llama-3.2-1b | MLX-4BIT | mlx | 527.7 tok/s (prefill 17477) | 512 | Extracted, source-linked | source |
| llama-3.2-1b | Q4_K_M | llama.cpp | 439.6 tok/s (prefill 16968) | 512 | Extracted, source-linked | source |
| qwen3-4b | MLX-4BIT | mlx | 186 tok/s (prefill 5506) | 512 | Extracted, source-linked | source |
| qwen3-4b | Q4_K_M | llama.cpp | 159.7 tok/s (prefill 4973) | 512 | Extracted, source-linked | source |
| qwen3.6-35b-a3b | MLX-4BIT | mlx | 151.1 tok/s (prefill 3333) | 512 | Extracted, source-linked | source |
Apple Mac Studio (M-series Ultra) · 7 single-stream configs
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3.6 35B-A3B | MLX-4BIT | mlx | 109.1 tok/s (prefill 6012) | 32000 | Extracted, source-linked | source |
| Qwen3.8-Flash-Next | MLX-4BIT | mlx (speculative decode, mtp) | 76.4 tok/s | Extracted, source-linked | source | |
| Qwen3 14B | MLX-4BIT | mlx | 55.4 tok/s (prefill 2278) | 32000 | Extracted, source-linked | source |
| Qwen3.8 27B | MLX-4BIT | mlx | 42.6 tok/s (prefill 1526) | 32000 | Extracted, source-linked | source |
| GLM-5.3-Flash | MLX-4BIT | mlx | 34.2 tok/s | Extracted, source-linked | source |
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.