Contributors › jvr0x's local-LLM benchmark configs
jvr0x's local-LLM benchmark configs
Every tracked benchmark config measured by jvr0x: model, quant, backend, tokens per second, source-linked and trust-tiered.
Snapshot 2026-09-10 · 148 configs · 69 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
20 configs · 16 models · 4 stars · 92.6 tok/s best single-stream · jvr0x on github.
Configs measured by jvr0x
| Model | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth) | NVFP4 | vllm | 92.6 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B-NVFP4-Fast (Unsloth) | NVFP4 | vllm | 92.6 tok/s | Structured table | source | |
| GPT-OSS-20B | Q4_K_XL | llama.cpp | 90.5 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B | Q4_K_XL | llama.cpp | 62.7 tok/s | Structured table | source | |
| Ornith-1.0-35B | Q8_0 | llama.cpp | 56.1 tok/s | Structured table | source | |
| Qwen3-Coder-Next-80B-A3B | Q4_K_XL | llama.cpp | 50.1 tok/s | Structured table | source | |
| Qwen3.6-35B-A3B | llama.cpp | 48.2 tok/s | Structured table | source | ||
| Nemotron-3-Nano-30B-A3B | llama.cpp | 47.9 tok/s | Structured table | source | ||
| Kimi-Linear-48B-A3B | Q8_0 | llama.cpp | 47.2 tok/s | Structured table | source | |
| Nemotron-3-Nano-30B-A3B | FP8 | vllm | 40.7 tok/s | Structured table | source | |
| Gemma-4-12B-IT | Q4_K_M | llama.cpp | 28.4 tok/s | Structured table | source | |
| Qwen-AgentWorld-35B-A3B | BF16 | vllm | 26.9 tok/s | Structured table | source | |
| Qwen3.6-27B-NVFP4 (Unsloth) | NVFP4 | vllm | 23.3 tok/s | Structured table | source | |
| Qwen3.6-27B-NVFP4 (Unsloth) | NVFP4 | vllm | 23.3 tok/s | Structured table | source | |
| Gemma-4-26B-A4B-IT | BF16 | vllm | 22.3 tok/s | Structured table | source | |
| Step-3.7-Flash | IQ4_XS | llama.cpp | 19.9 tok/s | Structured table | source | |
| Nex-N2-Pro-397B-A17B | llama.cpp | 18.9 tok/s | Structured table | source | ||
| Qwopus3.6-27B-Coder | Q8_0 | llama.cpp | 7.6 tok/s | Structured table | source |
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.