Contributors › Metrale's local-LLM benchmark configs
Metrale's local-LLM benchmark configs
Every tracked benchmark config measured by Metrale: model, quant, backend, tokens per second, source-linked and trust-tiered.
Snapshot 2026-10-06 · 258 configs · 118 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
2 configs · 2 models · 2 stars · 85 tok/s best single-stream · Metrale on github.
Configs measured by Metrale
| Model | Hardware | Quant | Backend | Decode | Context | Trust | Source |
|---|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | DGX Spark | FP8 | metrale (speculative decode, mtp-1) | 85 tok/s | 2048 | Extracted, source-linked | source |
| Qwen3.8-27B | DGX Spark | NVFP4 | metrale (speculative decode, mtp) | 25.1 tok/s | 2048 | Extracted, source-linked | source |
More: all configs · hardware · news · contributors · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.