# TokenMark, every single-stream config Snapshot: 2026-09-10T06:44:15.420Z. Generated from the tracker snapshot; each line links to where it was measured. Overview and how to cite: https://tokenmark.app/llms.txt ## AMD Strix Halo (Ryzen AI Max+ 395) (/hardware/strix-halo) - Qwen3 0.6B · Q8_0 · llama.cpp/vulkan · 266 tok/s (prefill 13112) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3 0.6B · Q8_0 · llama.cpp/rocm · 208.7 tok/s (prefill 4666.1) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - LFM2.5 8B-A1B · Q4_K_M · llama.cpp/vulkan · 176.5 tok/s (prefill 3398.4) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3-30B-A3B-Instruct-2507 · IQ4_XS · llama.cpp/vulkan · 103.2 tok/s (prefill 1438.1) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3-Coder 30B-A3B · Q4_K_S · llama.cpp/vulkan · 98 tok/s (prefill 1406.5) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3-Coder 30B-A3B · UD-Q4_K_XL · llama.cpp/vulkan · 97.1 tok/s (prefill 1400) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3-Coder 30B-A3B · IQ4_XS · llama.cpp/vulkan · 90.4 tok/s (prefill 1372.3) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3 30B-A3B NEO-MAX · IQ4_XS · llama.cpp/vulkan · 87.4 tok/s (prefill 1396.1) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6 35B-A3B · Q4_0 · llama.cpp/vulkan · 81.3 tok/s (prefill 1243.5) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3-30B-A3B · UD-Q4_K_XL · llama.cpp/rocm · 79.1 tok/s (prefill 687.9) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/Qwen3-30B-A3B-UD-Q4_K_XL/results.jsonl - Nemotron Cascade 2 30B-A3B · IQ4_XS · llama.cpp/vulkan · 79 tok/s (prefill 1325.3) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Nemotron 3 Nano 30B-A3B · IQ4_XS · llama.cpp/vulkan · 76 tok/s (prefill 1312.5) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.5 35B-A3B · IQ4_XS · llama.cpp/vulkan · 75.2 tok/s (prefill 1170.3) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Gemma 4 26B-A4B IT QAT · UD-Q4_K_XL · llama.cpp/vulkan · 74.8 tok/s (prefill 1432) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/vulkan · 74.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3-Coder 30B-A3B · UD-Q4_K_XL · llama.cpp/rocm · 73.7 tok/s (prefill 1285.3) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/vulkan · 72.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/rocm · 68.3 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3.6-35B-A3B-MTP · UD-Q4_K_XL · vulkan · 66 tok/s · ctx 262144 · trust: Verified (ran it) · source: https://github.com/sypherin/strix-halo-setup/blob/HEAD/README.md - Qwen AgentWorld 35B-A3B · IQ4_XS · llama.cpp/vulkan · 65.7 tok/s (prefill 1182.8) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.5 35B-A3B · Q4_K_M · llama.cpp/vulkan · 64.9 tok/s (prefill 1080) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/rocm · 64.5 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3.6 35B-A3B · Q4_K_M · llama.cpp/vulkan · 63.8 tok/s (prefill 1064) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/vulkan-radv · 62.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/vulkan-radv-performance · 62.2 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3-Coder-Next 80B-A3B · IQ4_XS · llama.cpp/vulkan · 61.9 tok/s (prefill 739) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6 35B-A3B · Q4_K_M · ollama/vulkan · 60.6 tok/s (prefill 793.1) · ctx 2048 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/vulkan · 58.7 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3-Next 80B-A3B · UD-Q4_K_XL · llama.cpp/vulkan · 54.9 tok/s (prefill 657) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.5 35B-A3B · Q4_K_M · llama.cpp/rocm · 54.7 tok/s (prefill 1047) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Gemma 4 26B-A4B IT · Q4_K_M · llama.cpp/vulkan · 54.2 tok/s (prefill 1323.4) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/rocm-7.14 · 53.5 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - llama-2-7b · Q4_0 · llama.cpp/rocm · 53 tok/s (prefill 1327) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/llama-2-7b.Q4_0/results.jsonl - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/rocm-7.14-pr26592 · 53 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6 35B-A3B · Q4_K_M · llama.cpp/rocm · 52.7 tok/s (prefill 1186.2) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/rocm-7.2.4 · 51.7 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/ryzen-ai-halo-results.json - llama-2-7b · Q4_K_M · llama.cpp/rocm · 50.4 tok/s (prefill 1083) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/llama-2-7b.Q4_K_M/results.jsonl - Qwen3-Next 80B-A3B · UD-Q4_K_XL · llama.cpp/rocm · 49.6 tok/s (prefill 800.4) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/vulkan-radv · 49.2 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/vulkan-radv-performance · 48.9 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-35B-A3B · UD-Q4_K_XL · llama.cpp/rocm · 48.7 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Gemma 4 26B-A4B · Q4_K_M · llama.cpp/vulkan · 48.5 tok/s (prefill 1142) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/rocm-7.14 · 48.1 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/rocm-7.14-pr26592 · 47.9 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - gpt-oss-20b-F16 · unquantised/unknown · llama.cpp/rocm · 47.8 tok/s (prefill 1213.3) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/gpt-oss-20b-F16/results.jsonl - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/rocm-7.2.4 · 46.5 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/ryzen-ai-halo-results.json - shisa-v2-llama3.1-8b.i1 · Q4_K_M · llama.cpp/rocm · 43.5 tok/s (prefill 878.2) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/shisa-v2-llama3.1-8b.i1-Q4_K_M/results.jsonl - gemma-4-26B-A4B-it · UD-Q8_K_XL · llama.cpp/rocm-7.2.4 · 41.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/ryzen-ai-halo-results.json - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/vulkan · 41 tok/s · trust: Verified (ran it) · source: https://github.com/sypherin/strix-halo-setup/blob/HEAD/README.md - Qwen3.6-35B-A3B · UD-Q8_K_XL · llama.cpp/rocm · 39 tok/s · trust: Verified (ran it) · source: https://github.com/sypherin/strix-halo-setup/blob/HEAD/README.md - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/vulkan · 35.7 tok/s (prefill 226.8) · ctx 524288 · trust: Extracted, source-linked · source: https://github.com/pepuscz/strix-halo-deepseek-v4-flash/blob/HEAD/README.md - Qwen3.8-Flash-Next · IQ3_XXS · llama.cpp · 35.7 tok/s · ctx 156000 · trust: Extracted, source-linked · source: https://huggingface.co/EasiiX/Qwen3.8-Flash-Next-MTP-Strix-Halo-GGUF/blob/main/README.md - DeepSeek-V4-Flash-0731 · 2.58bpw-mix · ? · 34.2 tok/s · trust: Extracted, source-linked · source: https://huggingface.co/otheru/DeepSeek-V4-Flash-Strix-Halo-GGUF/blob/main/README.md - gpt-oss-120b-F16 · unquantised/unknown · llama.cpp/rocm · 33.9 tok/s (prefill 534.1) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/gpt-oss-120b-F16/results.jsonl - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/rocm · 29.6 tok/s (prefill 142.8) · ctx 131072 · trust: Extracted, source-linked · source: https://github.com/pepuscz/strix-halo-deepseek-v4-flash/blob/HEAD/README.md - Gemma 4 12B IT QAT · UD-Q4_K_XL · llama.cpp/vulkan · 29.3 tok/s (prefill 816.3) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.8-Flash-Next · IQ4_XS · llama.cpp/vulkan · 27.2 tok/s (prefill 394.7) · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.8-27B · UD-Q4_K_XL · llama.cpp/vulkan · 24.6 tok/s · ctx 131072 · trust: Verified (ran it) · source: https://github.com/sypherin/strix-halo-setup/blob/HEAD/README.md - Qwen3.8-Flash-Next · Q4_K_XL · llama.cpp/vulkan · 24 tok/s · trust: Extracted, source-linked · source: https://github.com/routhjim/fn-expert-swap/blob/HEAD/README.md - Qwen3.8-Flash-Next · IQ3_XXS · llama.cpp · 23.5 tok/s · ctx 156000 · trust: Extracted, source-linked · source: https://huggingface.co/EasiiX/Qwen3.8-Flash-Next-MTP-Strix-Halo-GGUF/blob/main/README.md - GLM-4.5-Air · UD-Q4_K_XL · llama.cpp/vulkan · 23.4 tok/s (prefill 179.9) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/GLM-4.5-Air-UD-Q4_K_XL/results.jsonl - Qwen3.8-27B · UD-Q5_K_XL · llama.cpp · 23 tok/s · trust: Extracted, source-linked · source: https://github.com/PieBru/Qwen-3.8-27B_Strix-Halo_gfx1151/blob/HEAD/README.md - dots.llm1.inst · UD-Q4_K_XL · llama.cpp/rocm · 22.7 tok/s (prefill 182) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/dots.llm1.inst-UD-Q4_K_XL/results.jsonl - Qwen3.5-122B-A10B · UD-Q4_K_XL · llama.cpp/rocm-7.2.4 · 21.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/ryzen-ai-halo-results.json - Qwen3.8 27B · Q4_K_M · ollama/vulkan · 20.4 tok/s (prefill 292.5) · ctx 4096 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.8 27B · Q4_K_M · ollama · 20.4 tok/s (prefill 292.5) · trust: Extracted, source-linked · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/README.md - Llama-4-Scout-17B-16E-Instruct · UD-Q4_K_XL · llama.cpp/rocm · 20.2 tok/s (prefill 306.4) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/Llama-4-Scout-17B-16E-Instruct-UD-Q4_K_XL/results.jsonl - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/vulkan-radv-performance · 19.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Nemotron 3 Super 120B-A12B · IQ4_XS · llama.cpp/vulkan · 18.9 tok/s (prefill 297.1) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Hunyuan-A13B-Instruct · UD-Q6_K_XL · llama.cpp/rocm · 18.4 tok/s (prefill 297.5) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/Hunyuan-A13B-Instruct-UD-Q6_K_XL/results.jsonl - Llama 4 Scout 109B · Q4_K_M · llama.cpp/vulkan · 18.3 tok/s (prefill 331) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.8-27B · UD-Q6_K_XL · llama.cpp · 17 tok/s (prefill 346) · ctx 131072 · trust: Extracted, source-linked · source: https://github.com/PieBru/Qwen-3.8-27B_Strix-Halo_gfx1151/blob/HEAD/README.md - DeepSeek-V4-Flash-0731 · UD-IQ2_XXS · llama.cpp/rocm-7.14 · 16.2 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - DeepSeek-V4-Flash-0731 · UD-IQ2_XXS · llama.cpp/rocm-7.2.4 · 16.1 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/rocm-7.14 · 16 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3-235B-A22B-Instruct-2507 · UD-Q3_K_XL · llama.cpp/vulkan · 15.9 tok/s (prefill 117.1) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/Qwen3-235B-A22B-Instruct-2507-UD-Q3_K_XL/results.jsonl - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/rocm-7.2.4 · 15.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - DeepSeek-V4-Flash-0731 · UD-IQ2_XXS · llama.cpp/rocm-7.14-pr26592 · 14.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Mistral-Small-3.1-24B-Instruct-2503 · UD-Q4_K_XL · llama.cpp/rocm · 14.7 tok/s (prefill 368.5) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/Mistral-Small-3.1-24B-Instruct-2503-UD-Q4_K_XL/results.jsonl - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/rocm-7.14-pr26592 · 14.5 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/rocm · 13.5 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/vulkan · 13.3 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - DeepSeek V4 Flash 284B · UD-IQ2_XXS · llama.cpp/vulkan · 13.3 tok/s (prefill 155.6) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6 27B MTP NVFP4 v3 · NVFP4 · llama.cpp/vulkan · 13.2 tok/s (prefill 374) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6 27B MTP NVFP4 v3 · NVFP4 · llama.cpp/vulkan · 13.2 tok/s (prefill 374) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - DeepSeek-V4-Flash-0731 · UD-IQ2_XXS · llama.cpp/vulkan-radv-performance · 13.1 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/rocm · 12.4 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - gemma-3-27b-it · UD-Q4_K_XL · llama.cpp/rocm · 12.1 tok/s (prefill 302.2) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/gemma-3-27b-it-UD-Q4_K_XL/results.jsonl - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/vulkan · 11.7 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Gemma 4 31B IT QAT · Q4_0 · llama.cpp/vulkan · 11.4 tok/s (prefill 308.3) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - DeepSeek-V4-Flash-0731 · UD-IQ2_XXS · llama.cpp/vulkan-radv · 9.1 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - DeepSeek-V4-Flash-0731 · UD-IQ3_XXS · llama.cpp/vulkan-radv · 9 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · Q8_0 · llama.cpp/rocm-7.2.4 · 7.8 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/ryzen-ai-halo-results.json - Qwen3.6 27B MTP · Q8_0 · llama.cpp/vulkan · 7.7 tok/s (prefill 341.9) · ctx 512 · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/rocm-7.14 · 6.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/rocm-7.14-pr26592 · 6.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/rocm-7.2.4 · 6.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/vulkan-radv · 6.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/vulkan-radv-performance · 6.6 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/docs/toolbox-performance-results.json - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/rocm · 6.5 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - Qwen3-32B · Q8_0 · llama.cpp/rocm · 6.4 tok/s (prefill 226.1) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/Qwen3-32B-Q8_0/results.jsonl - Qwen3.6-27B · UD-Q8_K_XL · llama.cpp/vulkan · 6.3 tok/s · trust: Structured table · source: https://github.com/kyuz0/amd-strix-halo-toolboxes/blob/HEAD/benchmark/results-mtp/summary.json - shisa-v2-llama3.3-70b.i1 · Q4_K_M · llama.cpp/rocm · 5.1 tok/s (prefill 94.7) · trust: Structured table · source: https://github.com/lhl/strix-halo-testing/blob/HEAD/llm-bench/shisa-v2-llama3.3-70b.i1-Q4_K_M/results.jsonl - Llama 3.1 70B · Q4_K_M · ollama/vulkan · 4.7 tok/s (prefill 79.6) · trust: Structured table · source: https://github.com/hogeheer499-commits/strix-halo-guide/blob/HEAD/data/benchmarks.csv ## NVIDIA DGX Spark (GB10) (/hardware/dgx-spark) - Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth) · NVFP4 · vllm · 92.6 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3.6-35b-nvfp4-unsloth-fast.json - Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth) · NVFP4 · vllm · 92.6 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3.6-35b-nvfp4-unsloth-fast.json - GPT-OSS-20B · Q4_K_XL · llama.cpp · 90.5 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/gpt-oss-20b-llamacpp.json - Qwen3.6-35B-A3B · Q4_K_XL · llama.cpp · 62.7 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3.6-35b-q4-llamacpp.json - Ornith-1.0-35B · Q8_0 · llama.cpp · 56.1 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/ornith-35b-q8-llamacpp.json - DeepSeek-V4-Flash-0731-JA-REAP-K216 · EXL3-3BPW · ? · 55 tok/s · ctx 256000 · trust: Extracted, source-linked · source: https://huggingface.co/Laplace1313/DeepSeek-V4-Flash-0731-JA-REAP-K216-EXL3-3bpw-DGX-Spark/blob/main/README.md - Qwen3-Coder-Next-80B-A3B · Q4_K_XL · llama.cpp · 50.1 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3-coder-next-q4-llamacpp.json - Qwen3.6-35B-A3B · unquantised/unknown · llama.cpp · 48.2 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3.6-35b-q8-llamacpp.json - Nemotron-3-Nano-30B-A3B · unquantised/unknown · llama.cpp · 47.9 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/nemotron3-nano-30b-q8-llamacpp.json - Kimi-Linear-48B-A3B · Q8_0 · llama.cpp · 47.2 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/kimi-linear-48b-q8-llamacpp.json - Qwen3.8-Flash-Next · NVFP4 · vllm · 44.2 tok/s · ctx 262144 · trust: Extracted, source-linked · source: https://huggingface.co/YSLAB-ai/Qwen3.8-Flash-Next-NVFP4-BF16PLE-DGX-Spark/blob/main/README.md - Nemotron-3-Nano-30B-A3B · FP8 · vllm · 40.7 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/nemotron3-nano-30b-fp8.json - Gemma-4-12B-IT · Q4_K_M · llama.cpp · 28.4 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/gemma-4-12b-it-llamacpp.json - Qwen3.8-Flash-Next · NVFP4 · vllm · 27.6 tok/s · trust: Extracted, source-linked · source: https://huggingface.co/sayyidfareed/Qwen3.8-Flash-Next-DGX-Spark-1M-Recipe/blob/main/README.md - Qwen3.8-Flash-Next · NVFP4 · vllm · 27.3 tok/s · ctx 262144 · trust: Extracted, source-linked · source: https://huggingface.co/YSLAB-ai/Qwen3.8-Flash-Next-NVFP4-BF16PLE-DGX-Spark/blob/main/README.md - Qwen-AgentWorld-35B-A3B · BF16 · vllm · 26.9 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen-agentworld-35b-bf16.json - Qwen3.8-Flash-Next · NVFP4 · vllm · 26.7 tok/s · ctx 1000000 · trust: Extracted, source-linked · source: https://huggingface.co/sayyidfareed/Qwen3.8-Flash-Next-DGX-Spark-1M-Recipe/blob/main/README.md - SuperQwen3.8-27b-abliterated · NVFP4 · vllm · 25.8 tok/s · ctx 262043 · trust: Extracted, source-linked · source: https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark/blob/main/README.md - Qwen3.6-27B-NVFP4 (Unsloth) · NVFP4 · vllm · 23.3 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3.6-27b-nvfp4-unsloth.json - Qwen3.6-27B-NVFP4 (Unsloth) · NVFP4 · vllm · 23.3 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwen3.6-27b-nvfp4-unsloth.json - Gemma-4-26B-A4B-IT · BF16 · vllm · 22.3 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/gemma-4-26b-a4b-bf16.json - Step-3.7-Flash · IQ4_XS · llama.cpp · 19.9 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/step-3.7-flash-llamacpp.json - Nex-N2-Pro-397B-A17B · unquantised/unknown · llama.cpp · 18.9 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/nex-n2-pro-iq1m-llamacpp.json - Qwopus3.6-27B-Coder · Q8_0 · llama.cpp · 7.6 tok/s · trust: Structured table · source: https://github.com/jvr0x/dgx-spark-bench/blob/HEAD/results/qwopus3.6-27b-coder-q8-llamacpp.json ## Apple M-series Max (MacBook Pro / Mac Studio Max) (/hardware/mac-max) - Qwen3.8-Flash-Next-REAP-288 · MLX-4BIT · mlx/pmlx · 37 tok/s · trust: Extracted, source-linked · source: https://huggingface.co/sh0wie/Qwen3.8-Flash-Next-REAP-288-MLX-4bit/blob/main/README.md - Qwen3.8-Flash-Next-REAP-288 · MLX-4BIT · mlx · 28 tok/s · trust: Extracted, source-linked · source: https://huggingface.co/sh0wie/Qwen3.8-Flash-Next-REAP-288-MLX-4bit/blob/main/README.md ## Apple Mac Studio (M-series Ultra) (/hardware/mac-ultra) - GLM-5.3-Flash · MLX-4BIT · mlx · 34.2 tok/s · trust: Extracted, source-linked · source: https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4/blob/HEAD/README.md - DeepSeek-V4-Flash-0731-MXFP4-MLX-Abliterated · MXFP4 · mlx · 29.6 tok/s (prefill 454) · ctx 1048576 · trust: Extracted, source-linked · source: https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-DeepSeekV4F-0731-MXFP4-Abliterated-Dual-ANE-CPU/blob/HEAD/README.md - GLM-5.3-Flash Abliterated · MLX-4BIT · mlx · 24 tok/s (prefill 365) · ctx 16384 · trust: Extracted, source-linked · source: https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4/blob/HEAD/README.md