TokenMark: Local AI Config Tracker for Strix Halo & DGX Spark

Measured tok/s configs for local LLMs on AMD Strix Halo, NVIDIA DGX Spark and Apple Mac, source-linked and updated daily. One searchable leaderboard.

Snapshot 2026-10-04 · 254 configs · 116 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.

AMD Strix Halo (Ryzen AI Max+ 395) · 159 single-stream configs

ModelQuantBackendDecodeContextTrustSource
Qwen3 0.6BQ8_0llama.cpp/vulkan266 tok/s (prefill 13112)512Structured tablesource
Qwen3 0.6BQ8_0llama.cpp/rocm208.7 tok/s (prefill 4666.1)512Structured tablesource
LFM2.5 8B-A1BQ4_K_Mllama.cpp/vulkan176.5 tok/s (prefill 3398.4)512Structured tablesource
Ornith-1.5-35B-A3BFP4(speculative decode, mtp)105.6 tok/sOwner-submittedsource
Qwen3-30B-A3B-Instruct-2507IQ4_XSllama.cpp/vulkan103.2 tok/s (prefill 1438.1)512Structured tablesource

AMD Gorgon Halo (Ryzen AI Max+ PRO 495) · 0 single-stream configs

No single-stream configs recorded yet.

NVIDIA DGX Spark (GB10) · 46 single-stream configs

ModelQuantBackendDecodeContextTrustSource
Qwen3 1.7BQ4_K_Mllama.cpp146.1 tok/s2048Extracted, source-linkedsource
Qwen3.6-35B-hereticNVFP4vllm96 tok/sStructured tablesource
Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth)NVFP4vllm92.6 tok/sStructured tablesource
GPT-OSS-20BQ4_K_XLllama.cpp90.5 tok/sStructured tablesource
Ministral 3BQ4_K_Mllama.cpp86.6 tok/s2048Extracted, source-linkedsource

Apple M-series Max (MacBook Pro / Mac Studio Max) · 23 single-stream configs

ModelQuantBackendDecodeContextTrustSource
llama-3.2-1bMLX-4BITmlx527.7 tok/s (prefill 17477)512Extracted, source-linkedsource
llama-3.2-1bQ4_K_Mllama.cpp439.6 tok/s (prefill 16968)512Extracted, source-linkedsource
qwen3-4bMLX-4BITmlx186 tok/s (prefill 5506)512Extracted, source-linkedsource
qwen3-4bQ4_K_Mllama.cpp159.7 tok/s (prefill 4973)512Extracted, source-linkedsource
qwen3.6-35b-a3bMLX-4BITmlx151.1 tok/s (prefill 3333)512Extracted, source-linkedsource

Apple Mac Studio (M-series Ultra) · 7 single-stream configs

ModelQuantBackendDecodeContextTrustSource
Qwen3.6 35B-A3BMLX-4BITmlx109.1 tok/s (prefill 6012)32000Extracted, source-linkedsource
Qwen3.8-Flash-NextMLX-4BITmlx (speculative decode, mtp)76.4 tok/sExtracted, source-linkedsource
Qwen3 14BMLX-4BITmlx55.4 tok/s (prefill 2278)32000Extracted, source-linkedsource
Qwen3.8 27BMLX-4BITmlx42.6 tok/s (prefill 1526)32000Extracted, source-linkedsource
GLM-5.3-FlashMLX-4BITmlx34.2 tok/sExtracted, source-linkedsource

More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.