Community Intel: daily local-AI model and hardware releases, explained
What changed for single-box local AI, rebuilt every morning from the model and hardware releases the tracker harvested plus pinned reports. Same feed as the home page column.
Snapshot 2026-09-28 · 207 configs · 94 models. Every number is generated from the tracker snapshot and linked to where it was measured; nothing here is typed by hand.
News that matters to a single-box owner
- Faster prompt lookup drafting in llama.cpp (2026-09-26): llama.cpp is the primary inference engine for local LLMs on these devices, and faster drafting directly improves generation speed.
Notable releases and pinned reports, explained
- ji-farthing/Qwen3.8-Flash-Next-Uncensored-ik-llama-GGUF (2026-09-28): This is a fictional, non-existent model, so it is completely irrelevant to single-box local inference.
- ClearFracture/dgx-spark-ubuntu-setup (2026-09-28): This is a setup script, not a model, so it doesn't add new inference capabilities but merely automates the vLLM environment on your 128GB DGX Spark.
- gufo-org/gufo (2026-09-28): Gufo enables Strix Halo users to run Qwen 27B at 70 tok/s, proving high-performance local inference on 128GB hardware.
- mlx-community/YOLO26s-OptiQ-6bit (2026-09-28): YOLO26s-OptiQ-6bit is a tiny object detection model that easily fits in 128GB, offering fast local inference without needing massive VRAM.
- ggml-org/llama.cpp (2026-09-28): This release adds native DGX Spark support, enabling optimized local inference for enthusiasts running models on their 128GB NVIDIA box.
- Edge0: a 35B MoE model on-device in ~3 GB RAM by streaming experts from storage (2026-09-10): Open-source (Apache-2.0) streaming-MoE runtime: a 35B 4-bit model (256 experts, ~23 GB on disk) runs at ~2.9 GB peak RAM and 15-18 tok/s, and an 8B at ~1 GB and 24-25 tok/s, by keeping only the active experts in memory (SSD offload + a prerouter that predicts routing a step ahead + Recover-LoRA). Ships on macOS/Apple Silicon today; the '35B on an iPhone' headline is still a demo, with iOS and CUDA on the roadmap.
- Google AI Edge Gallery: run Gemma-class models fully on-device on Android and iOS (2026-09-10): Google's open app to try on-device GenAI locally on Android and iOS with no internet after the model download: chat, image understanding and audio transcription on mobile NPUs via LiteRT, running Gemma-class models on the phone itself.
- ExecuTorch: PyTorch's on-device runtime for LLMs on phones, wearables and microcontrollers (2026-09-10): Meta/PyTorch's edge inference runtime runs LLMs and vision models on Android, iOS and embedded hardware, benchmarked head-to-head against llama.cpp, ONNX Runtime, LiteRT and CoreML. The PyTorch-native path from a trained model to a phone.
- MLC LLM: compile and run LLMs on mobile CPUs and NPUs (Snapdragon Hexagon) (2026-09-10): A TVM-based compiler and deployment engine that targets mobile CPUs and NPUs directly: on a Snapdragon 8 Elite (Galaxy S25 Ultra) MLCChat runs Qwen3 1.7B at ~40 tok/s and Phi-4-mini at ~22 tok/s, roughly 2-3x faster than CPU-only on the same phone.
- LiteRT-LM: Google's on-device LLM runtime that powers Gemini Nano (2026-09-10): Google's on-device LLM stack that superseded the now maintenance-only MediaPipe LLM Inference API; it powers Gemini Nano in Chrome and on Pixel, and is reached on Android through the ML Kit GenAI APIs.
- Intelligence per Watt (Stanford): local models answer 88.7% of 1M real queries, IPW up 5.3x since 2023 (2026-09-06): Stanford measured task accuracy per watt across 20+ local models and 8 accelerators on 1 million real queries: the best local model of 20B active parameters or fewer answers 88.7%, intelligence per watt rose 5.3x from 2023 to 2025, and a local-first router that is right 80% of the time cuts energy 64% and cost 59% against cloud only. The only local box they profiled is an Apple M4 Max; the Apache-2.0 harness reads AMD power, so Strix Halo numbers are ours to add.
- NVIDIA IFA 2026: llama.cpp up to 1.9x, vLLM 1.4x on two DGX Sparks, one-click local setup (2026-09-03): llama.cpp hits up to 1.9x on an RTX 5090 (kernel opts, better speculative decoding, faster prefill) and vLLM 1.4x on two DGX Sparks; one-click Windows setup ships via Ollama and LM Studio, and NVIDIA PAIR spreads inference across your local-network PCs. Same box, bigger and faster local models, so a lot of current bench numbers are now stale.
Latest releases tracked
- ji-farthing/Qwen3.8-Flash-Next-Uncensored-ik-llama-GGUF (2026-09-28)
- ClearFracture/dgx-spark-ubuntu-setup (2026-09-28): Canonical Ubuntu 24.04 setup for NVIDIA DGX Spark with Docker, vLLM, and a redacted STIG assessment
- gufo-org/gufo (2026-09-28): Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_
- mlx-community/YOLO26s-OptiQ-6bit (2026-09-28)
- ggml-org/llama.cpp (2026-09-28): LLM inference in C/C++
- Osmantic/ODS (2026-09-28): ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into
- nero-/deepseek-v41-flash-gb10-ring (2026-09-28): DeepSeek-V4.1-Flash TP4 on a switchless ring of 2x DGX Spark + 2x ASUS GX10 (Mia's SGLang line + sparkring transport + L
- rlindsey2/sunkcost (2026-09-28): How long until local AI pays for itself? A break-even calculator for running open models on a Mac, DGX Spark or Strix Ha
- jschmied/qwen38-flash-next-gb10 (2026-09-28): Getting Qwen's Qwen4-architecture preview (Qwen3.8-Flash-Next) to run on a single DGX Spark GB10 — 125.9 GiB of weights
- nv-drollins/nous-research-spark (2026-09-28): One-shot installer: Hermes Agent + local vLLM (Qwen3.6-35B-A3B-NVFP4) on NVIDIA DGX Spark
- pom11/hscc (2026-09-28): Turn a DGX Spark GPU cluster into a self-running team of specialized AI agents — cluster control, role-specialized worke
- styles01/sparkrun-recipes (2026-09-28): Custom inference recipes for NVIDIA DGX Spark — Qwen 122B DFlash hybrid
- KitsuneOff/Qwen3-14B-GGUF (2026-09-28)
- nightmedia/Qwen3.5-9B-Seven-q8-hi-mlx (2026-09-28)
- 1bit-MONSTER/engine (2026-09-28): 1bit engine: local LLM inference server for AMD Ryzen AI (Strix Halo). XDNA 2 NPU, Radeon GPU via Vulkan/HRX/ROCm, CPU r
- shcherbakov22/yet-another-halo-engine (2026-09-28): Modern custom inference engine for Strix Halo targeting hybrid NPU+GPU inference
- 1bit-MONSTER/Qwen3-Coder-30B-A3B-Q4_0-H32-GGUF (2026-09-28)
- DuoNeural/Qwen2.5-7B-Instruct-GTAP-Q4_K_M-GGUF (2026-09-28)
- championswimmer/kev-local (2026-09-28): Run jaredpalmer/kev (small decision-model LLM) locally on AMD Strix Halo via ROCm
- adrienbrault/qwen3.8-27b-rtx5090 (2026-09-28): Qwen3.8-27B on RTX 5090s — 262K ctx, 1.4M-token KV pool, ~220 t/s code decode. NVFP4 + vLLM + sm120 patches, reproducibl
- headpiece747/ninfer-5090-windows (2026-09-28): Native Windows port of NInfer engine for RTX 5090. Features Qwen3.8-27B with QUASAR and NInfer models, MTP/DFlash2 with
- kyuz0/amd-strix-halo-toolboxes (2026-09-28)
- BenGamliel/NInferEZ-Engine (2026-09-28): Unofficial Windows distribution of iamwavecut/ninfer-all with target-specific builds and an integration contract for NIn
- Luce-Org/lucebox (2026-09-28): LLM speculative inference server for heterogeneous hardware & consumer GPUs
- logic65/Whittle-Qwen-3.8-35B-A3B-GGUF (2026-09-28)
- dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx (2026-09-28)
- mradermacher/Qwen-Writer-9B-GGUF (2026-09-28)
- mradermacher/Whittle-Qwen-3.8-35B-A3B-GGUF (2026-09-28)
- pwl1987-dev/unified-ai-platform (2026-09-28): Qwen3.8-27B on 8×RTX 4090: 4-replica 256K inference + OpenResty dynamic LB, QLoRA post-training, LoRA→GGUF patches, A/B
- OliviaRossi/Qwen3.5-9B-C3SM-SDM-Agentic-Coder-Abliterated-GGUF (2026-09-28)
More: all configs · hardware · methodology · llms.txt · llms-full.txt · API (OpenAPI) · JSON snapshot. Built by Altronis, private on-prem AI, Singapore.