Contributors › julianmb's local-LLM benchmark configs
julianmb's local-LLM benchmark configs
Every tracked benchmark config measured by julianmb: model, quant, backend, tokens per second, source-linked and trust-tiered.
Snapshot 2026-09-16 · 158 configs · 75 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
10 configs · 4 models · 211 stars · 76.9 tok/s best single-stream · julianmb on github.
Configs measured by julianmb
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Ornith 1.5 35B-A3B | FP4 | 76.9 tok/s | Owner-submitted | source | ||
| Ornith 1.5 35B-A3B | Q4_K_M | 71.6 tok/s | Owner-submitted | source | ||
| Qwen 3.8 27B | FP4 | 33.8 tok/s | Owner-submitted | source | ||
| DeepSeek V4 Flash 284B | IQ2_XXS | 32 tok/s | Owner-submitted | source | ||
| Qwen 3.8 27B | FP4 | 26.9 tok/s | 32768 | Owner-submitted | source | |
| Qwen3.8-27B-DFlash2 | Q4_K_M | vulkan | 21.2 tok/s (prefill 211) | 32768 | Owner-submitted | source |
| Qwen 3.8 27B | FP8 | llama.cpp | 19 tok/s | Owner-submitted | source | |
| Qwen 3.8 27B | Q4_K_M | 12.4 tok/s | Owner-submitted | source | ||
| Qwen 3.8 27B | Q4_K_M | llama.cpp | 12.3 tok/s | 32768 | Owner-submitted | source |
| Qwen 3.8 27B | FP16 | llama.cpp | 5 tok/s | Owner-submitted | source |
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.