Contributors › hogeheer's local-LLM benchmark configs
hogeheer's local-LLM benchmark configs
Every tracked benchmark config measured by hogeheer: model, quant, backend, tokens per second, source-linked and trust-tiered.
Snapshot 2026-09-10 · 148 configs · 69 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
37 configs · 25 models · 323 stars · 266 tok/s best single-stream · hogeheer on github.
Configs measured by hogeheer
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3 0.6B | Q8_0 | llama.cpp/vulkan | 266 tok/s (prefill 13112) | 512 | Structured table | source |
| Qwen3 0.6B | Q8_0 | llama.cpp/rocm | 208.7 tok/s (prefill 4666.1) | 512 | Structured table | source |
| LFM2.5 8B-A1B | Q4_K_M | llama.cpp/vulkan | 176.5 tok/s (prefill 3398.4) | 512 | Structured table | source |
| Qwen3-30B-A3B-Instruct-2507 | IQ4_XS | llama.cpp/vulkan | 103.2 tok/s (prefill 1438.1) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | Q4_K_S | llama.cpp/vulkan | 98 tok/s (prefill 1406.5) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | UD-Q4_K_XL | llama.cpp/vulkan | 97.1 tok/s (prefill 1400) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | IQ4_XS | llama.cpp/vulkan | 90.4 tok/s (prefill 1372.3) | 512 | Structured table | source |
| Qwen3 30B-A3B NEO-MAX | IQ4_XS | llama.cpp/vulkan | 87.4 tok/s (prefill 1396.1) | 512 | Structured table | source |
| Qwen3.6 35B-A3B | Q4_0 | llama.cpp/vulkan | 81.3 tok/s (prefill 1243.5) | 512 | Structured table | source |
| Nemotron Cascade 2 30B-A3B | IQ4_XS | llama.cpp/vulkan | 79 tok/s (prefill 1325.3) | 512 | Structured table | source |
| Nemotron 3 Nano 30B-A3B | IQ4_XS | llama.cpp/vulkan | 76 tok/s (prefill 1312.5) | 512 | Structured table | source |
| Qwen3.5 35B-A3B | IQ4_XS | llama.cpp/vulkan | 75.2 tok/s (prefill 1170.3) | 512 | Structured table | source |
| Gemma 4 26B-A4B IT QAT | UD-Q4_K_XL | llama.cpp/vulkan | 74.8 tok/s (prefill 1432) | 512 | Structured table | source |
| Qwen3-Coder 30B-A3B | UD-Q4_K_XL | llama.cpp/rocm | 73.7 tok/s (prefill 1285.3) | 512 | Structured table | source |
| Qwen AgentWorld 35B-A3B | IQ4_XS | llama.cpp/vulkan | 65.7 tok/s (prefill 1182.8) | 512 | Structured table | source |
| Qwen3.5 35B-A3B | Q4_K_M | llama.cpp/vulkan | 64.9 tok/s (prefill 1080) | 512 | Structured table | source |
| Qwen3.6 35B-A3B | Q4_K_M | llama.cpp/vulkan | 63.8 tok/s (prefill 1064) | 512 | Structured table | source |
| Qwen3-Coder-Next 80B-A3B | IQ4_XS | llama.cpp/vulkan | 61.9 tok/s (prefill 739) | 512 | Structured table | source |
| Qwen3.6 35B-A3B | Q4_K_M | ollama/vulkan | 60.6 tok/s (prefill 793.1) | 2048 | Structured table | source |
| Qwen3-Next 80B-A3B | UD-Q4_K_XL | llama.cpp/vulkan | 54.9 tok/s (prefill 657) | 512 | Structured table | source |
| Qwen3.5 35B-A3B | Q4_K_M | llama.cpp/rocm | 54.7 tok/s (prefill 1047) | 512 | Structured table | source |
| Gemma 4 26B-A4B IT | Q4_K_M | llama.cpp/vulkan | 54.2 tok/s (prefill 1323.4) | 512 | Structured table | source |
| Qwen3.6 35B-A3B | Q4_K_M | llama.cpp/rocm | 52.7 tok/s (prefill 1186.2) | 512 | Structured table | source |
| Qwen3-Next 80B-A3B | UD-Q4_K_XL | llama.cpp/rocm | 49.6 tok/s (prefill 800.4) | 512 | Structured table | source |
| Gemma 4 26B-A4B | Q4_K_M | llama.cpp/vulkan | 48.5 tok/s (prefill 1142) | 512 | Structured table | source |
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.