Contributors › lhl's local-LLM benchmark configs

lhl's local-LLM benchmark configs

Every tracked benchmark config measured by lhl: model, quant, backend, tokens per second, source-linked and trust-tiered.

Snapshot 2026-09-10 · 148 configs · 69 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.

15 configs · 14 models · 252 stars · 79.1 tok/s best single-stream · lhl on github.

Configs measured by lhl

ModelQuantBackendDecodeContextTrustSource
Qwen3-30B-A3BUD-Q4_K_XLllama.cpp/rocm79.1 tok/s (prefill 687.9)Structured tablesource
llama-2-7bQ4_0llama.cpp/rocm53 tok/s (prefill 1327)Structured tablesource
llama-2-7bQ4_K_Mllama.cpp/rocm50.4 tok/s (prefill 1083)Structured tablesource
gpt-oss-20b-F16llama.cpp/rocm47.8 tok/s (prefill 1213.3)Structured tablesource
shisa-v2-llama3.1-8b.i1Q4_K_Mllama.cpp/rocm43.5 tok/s (prefill 878.2)Structured tablesource
gpt-oss-120b-F16llama.cpp/rocm33.9 tok/s (prefill 534.1)Structured tablesource
GLM-4.5-AirUD-Q4_K_XLllama.cpp/vulkan23.4 tok/s (prefill 179.9)Structured tablesource
dots.llm1.instUD-Q4_K_XLllama.cpp/rocm22.7 tok/s (prefill 182)Structured tablesource
Llama-4-Scout-17B-16E-InstructUD-Q4_K_XLllama.cpp/rocm20.2 tok/s (prefill 306.4)Structured tablesource
Hunyuan-A13B-InstructUD-Q6_K_XLllama.cpp/rocm18.4 tok/s (prefill 297.5)Structured tablesource
Qwen3-235B-A22B-Instruct-2507UD-Q3_K_XLllama.cpp/vulkan15.9 tok/s (prefill 117.1)Structured tablesource
Mistral-Small-3.1-24B-Instruct-2503UD-Q4_K_XLllama.cpp/rocm14.7 tok/s (prefill 368.5)Structured tablesource
gemma-3-27b-itUD-Q4_K_XLllama.cpp/rocm12.1 tok/s (prefill 302.2)Structured tablesource
Qwen3-32BQ8_0llama.cpp/rocm6.4 tok/s (prefill 226.1)Structured tablesource
shisa-v2-llama3.3-70b.i1Q4_K_Mllama.cpp/rocm5.1 tok/s (prefill 94.7)Structured tablesource

More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.