Contributors › lhl's local-LLM benchmark configs
lhl's local-LLM benchmark configs
Every tracked benchmark config measured by lhl: model, quant, backend, tokens per second, source-linked and trust-tiered.
Snapshot 2026-09-10 · 148 configs · 69 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
15 configs · 14 models · 252 stars · 79.1 tok/s best single-stream · lhl on github.
Configs measured by lhl
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3-30B-A3B | UD-Q4_K_XL | llama.cpp/rocm | 79.1 tok/s (prefill 687.9) | Structured table | source | |
| llama-2-7b | Q4_0 | llama.cpp/rocm | 53 tok/s (prefill 1327) | Structured table | source | |
| llama-2-7b | Q4_K_M | llama.cpp/rocm | 50.4 tok/s (prefill 1083) | Structured table | source | |
| gpt-oss-20b-F16 | llama.cpp/rocm | 47.8 tok/s (prefill 1213.3) | Structured table | source | ||
| shisa-v2-llama3.1-8b.i1 | Q4_K_M | llama.cpp/rocm | 43.5 tok/s (prefill 878.2) | Structured table | source | |
| gpt-oss-120b-F16 | llama.cpp/rocm | 33.9 tok/s (prefill 534.1) | Structured table | source | ||
| GLM-4.5-Air | UD-Q4_K_XL | llama.cpp/vulkan | 23.4 tok/s (prefill 179.9) | Structured table | source | |
| dots.llm1.inst | UD-Q4_K_XL | llama.cpp/rocm | 22.7 tok/s (prefill 182) | Structured table | source | |
| Llama-4-Scout-17B-16E-Instruct | UD-Q4_K_XL | llama.cpp/rocm | 20.2 tok/s (prefill 306.4) | Structured table | source | |
| Hunyuan-A13B-Instruct | UD-Q6_K_XL | llama.cpp/rocm | 18.4 tok/s (prefill 297.5) | Structured table | source | |
| Qwen3-235B-A22B-Instruct-2507 | UD-Q3_K_XL | llama.cpp/vulkan | 15.9 tok/s (prefill 117.1) | Structured table | source | |
| Mistral-Small-3.1-24B-Instruct-2503 | UD-Q4_K_XL | llama.cpp/rocm | 14.7 tok/s (prefill 368.5) | Structured table | source | |
| gemma-3-27b-it | UD-Q4_K_XL | llama.cpp/rocm | 12.1 tok/s (prefill 302.2) | Structured table | source | |
| Qwen3-32B | Q8_0 | llama.cpp/rocm | 6.4 tok/s (prefill 226.1) | Structured table | source | |
| shisa-v2-llama3.3-70b.i1 | Q4_K_M | llama.cpp/rocm | 5.1 tok/s (prefill 94.7) | Structured table | source |
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.