Leaderboard - TokenMark Hardware Rankings

See how devices rank for local LLM inference. Compare Apple M4, RTX 4090, Snapdragon X Elite, AMD Strix, and more on the TokenMark leaderboard.

Snapshot 2026-09-28 · 207 configs · 94 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.

AMD Strix Halo (Ryzen AI Max+ 395) · 145 single-stream configs

ModelQuantBackendDecodeContextTrustSource
Qwen3 0.6BQ8_0llama.cpp/vulkan266 tok/s (prefill 13112)512Structured tablesource
Qwen3 0.6BQ8_0llama.cpp/rocm208.7 tok/s (prefill 4666.1)512Structured tablesource
LFM2.5 8B-A1BQ4_K_Mllama.cpp/vulkan176.5 tok/s (prefill 3398.4)512Structured tablesource
Qwen3-30B-A3B-Instruct-2507IQ4_XSllama.cpp/vulkan103.2 tok/s (prefill 1438.1)512Structured tablesource
Qwen3-Coder 30B-A3BQ4_K_Sllama.cpp/vulkan98 tok/s (prefill 1406.5)512Structured tablesource
Qwen3-Coder 30B-A3BUD-Q4_K_XLllama.cpp/vulkan97.1 tok/s (prefill 1400)512Structured tablesource
Qwen3-Coder 30B-A3BIQ4_XSllama.cpp/vulkan90.4 tok/s (prefill 1372.3)512Structured tablesource
Qwen3 30B-A3B NEO-MAXIQ4_XSllama.cpp/vulkan87.4 tok/s (prefill 1396.1)512Structured tablesource
Qwen3.6 35B-A3BQ4_0llama.cpp/vulkan81.3 tok/s (prefill 1243.5)512Structured tablesource
Qwen3-30B-A3BUD-Q4_K_XLllama.cpp/rocm79.1 tok/s (prefill 687.9)Structured tablesource
Nemotron Cascade 2 30B-A3BIQ4_XSllama.cpp/vulkan79 tok/s (prefill 1325.3)512Structured tablesource
Gemma-4-26B-A4BUD-Q4_K_XLllama.cpp78 tok/sOwner-submittedsource
Ornith 1.5 35B-A3BFP476.9 tok/sOwner-submittedsource
Nemotron 3 Nano 30B-A3BIQ4_XSllama.cpp/vulkan76 tok/s (prefill 1312.5)512Structured tablesource
Qwen3.5 35B-A3BIQ4_XSllama.cpp/vulkan75.2 tok/s (prefill 1170.3)512Structured tablesource

AMD Gorgon Halo (Ryzen AI Max+ PRO 495) · 0 single-stream configs

No single-stream configs recorded yet.

NVIDIA DGX Spark (GB10) · 39 single-stream configs

ModelQuantBackendDecodeContextTrustSource
Qwen3.6-35B-hereticNVFP4vllm96 tok/sStructured tablesource
Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth)NVFP4vllm92.6 tok/sStructured tablesource
GPT-OSS-20BQ4_K_XLllama.cpp90.5 tok/sStructured tablesource
Ornith-1.0-35B-AEONNVFP4vllm84.8 tok/sStructured tablesource
Qwen3.6-35B-A3BNVFP4vllm77.4 tok/sStructured tablesource
Qwen3.6-35B-A3BQ4_K_XLllama.cpp62.7 tok/sStructured tablesource
Ornith-1.0-35BQ8_0llama.cpp56.1 tok/sStructured tablesource
DeepSeek-V4-Flash-0731-JA-REAP-K216EXL3-3BPW55 tok/s256000Extracted, source-linkedsource
Qwen3-Coder-Next-80B-A3BQ4_K_XLllama.cpp50.1 tok/sStructured tablesource
Qwen3.6-35B-A3Bllama.cpp48.2 tok/sStructured tablesource
Nemotron-3-Nano-30B-A3Bllama.cpp47.9 tok/sStructured tablesource
Kimi-Linear-48B-A3BQ8_0llama.cpp47.2 tok/sStructured tablesource
Qwen3.8-Flash-NextNVFP4vllm (speculative decode, mtp-2)44.2 tok/s262144Extracted, source-linkedsource
Nemotron-3-Nano-Omni-30B-A3BNVFP4vllm41.7 tok/sStructured tablesource
Nemotron-3-Nano-30B-A3BFP8vllm40.7 tok/sStructured tablesource

Apple M-series Max (MacBook Pro / Mac Studio Max) · 2 single-stream configs

ModelQuantBackendDecodeContextTrustSource
Qwen3.8-Flash-Next-REAP-288MLX-4BITmlx/pmlx37 tok/sExtracted, source-linkedsource
Qwen3.8-Flash-Next-REAP-288MLX-4BITmlx28 tok/sExtracted, source-linkedsource

Apple Mac Studio (M-series Ultra) · 3 single-stream configs

ModelQuantBackendDecodeContextTrustSource
GLM-5.3-FlashMLX-4BITmlx34.2 tok/sExtracted, source-linkedsource
DeepSeek-V4-Flash-0731-MXFP4-MLX-AbliteratedMXFP4mlx29.6 tok/s (prefill 454)1048576Extracted, source-linkedsource
GLM-5.3-Flash AbliteratedMLX-4BITmlx (speculative decode, mtp)24 tok/s (prefill 365)16384Extracted, source-linkedsource

More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.