Strata Runs 125B Qwen3.8-Flash-Next on a 12GB Gaming GPU at Up to 94 Tokens/s
TL;DR
Strata is an open-source (MIT) inference engine with a one-click installer that runs Qwen3.8-Flash-Next, a 125B-parameter mixture-of-experts model, on an ordinary gaming PC with a 12 GB or larger NVIDIA or AMD card and 32-64 GB of RAM. The author's own benchmark on an RTX 5070 (12 GB) with 64 GB of DDR5 shows 53 to 94 output tokens per second depending on the quantization, and 1,600 to 2,650 tokens per second reading a 32K prompt. The repo was created on September 24, has nearly 11,000 stars and 962 forks, shipped 30 tagged releases between September 27 and October 4, and hit 570 points on Hacker News on Sunday. It exposes OpenAI, Anthropic and Responses-style APIs on localhost, so Claude Code, Codex CLI and friends can point at it directly.
What it is
Strata is one opinionated thing done well: get one specific model running fast on hardware you already own. Download the zip, double-click START-HERE.bat on Windows or run ./setup.sh on Linux, answer three questions (model size, context length, whether you want image input), and wait for a ~70 GB download. A browser app opens at 127.0.0.1:8080 with chat, a live monitor of GPU, CPU and RAM, and the settings.
The model is Qwen's own. According to the Qwen3.8-Flash-Next model card, it has 125B parameters with 6B activated per token, plus a 51B n-gram embedding and a 4B multi-token prediction (MTP) layer, spread across 48 layers of 512 experts each, with 10 routed experts and one shared expert firing per token. Native context is 262,144 tokens.
Strata does not invent the compression. It ships quantized builds from ISTA-DASLab, a fine-tune called Swift 1.5 from UkisAI, and an experimental 4-bit build from Unsloth, and it borrows parts of llama.cpp / ggml. The engine and the scheduling around those weights are the new part.
The numbers, measured by the author
The README publishes a measured table for two machines. The NVIDIA box is an RTX 5070 (12 GB), a six-core Ryzen 5 7600 and 64 GB of DDR5-5200. Output speed was measured on short chats, prompt speed on a 32K-token prompt:
- Prompt reading at 32K: 2,650 tokens/s (Q2_0), 2,090 (IQ2_XS), 1,750 (IQ3_XXS), 1,620 (IQ3_S), 2,180 (Coder).
- AMD: on an RX 9070 XT (16 GB) with a Ryzen 9 3900X and 47 GB of RAM, Q2_0 writes 60 tokens/s, IQ2_XS 52 and the Coder 44.
- Long context: the details page lists Q2_0 at 73.7 tokens/s output with a 128K prompt and 60.3 at the full 262K window on the same 12 GB card.
The Hacker News thread added anecdotes that are worth exactly what anecdotes are worth: the submitter reported 124 tokens/s on an RTX 4090 with 128 GB of DDR5, one commenter reported about 30 tokens/s from the Coder build on a 10 GB RTX 3080 with 48 GB of RAM, and another reported 33 tokens/s decode on an 8 GB RTX 2060 at 128K context. Two contributors have filed formal reports in the repo's community benchmark folder, including a pair of 2018-era AMD Instinct MI50 cards on a 32 GB Xeon.
How a 125B model fits in 12 GB
It doesn't, and that is the trick. A mixture-of-experts model only touches a small slice of its weights per token: here 10 of 512 experts per layer, across 48 layers, or 480 of 24,576 expert blocks. Strata splits the machine into tiers based on how often each expert gets called:
- GPU: attention and DeltaNet layers, routers, shared experts, the output head, the MTP draft layer, the hot part of the KV cache, and an expert cache that fills the rest of VRAM. The cache adapts to the conversation as you chat. Per the how-it-works doc, each extra GB of VRAM holds about 700 more experts, so more VRAM beats a faster GPU.
- RAM: every expert, pinned. When a token routes to an expert the card doesn't hold, the CPU computes it in place with AVX-512 or AVX2 kernels, at the same time as the GPU handles the cached ones.
- SSD: the 28.8 GB n-gram embedding table, read a few rows per token through the OS file cache.
Think of it as a restaurant line. The ten dishes everyone orders are prepped at the pass (VRAM), the full menu is in the walk-in (RAM) with a second cook (the CPU) plating the odd orders, and the wine cellar (SSD) only gets a trip when someone asks for something specific. Nobody waits for a delivery truck per plate.
Two more tricks carry the speed. The model's built-in MTP layer drafts up to three tokens, and one pass over all 48 layers verifies them, averaging 2.4-3.2 tokens per pass for a 1.6-1.8x speedup with identical output. And from 64K context up, Strata keeps most of the KV cache in RAM and only the part attention reads on the card, which freed room for 3,872 cached experts instead of 1,589 at 262K and lifted Q2_0 output there from 50.9 to 62.6 tokens/s.
The Coder build: half the experts, most of the skill
If you have 32 GB of RAM, the installer points you at the Coder build. ISTA-DASLab kept 256 of each layer's 512 experts, chosen with RCO (paper) by minimizing KL divergence against the full model on code, agentic and image calibration data. Its own card reports 91.3% of the full model's SWE-bench Verified score and 98.7% of its LiveCodeBench v6 score. The first shard is 29.6 GB.
The cost is everything that isn't code. Strata's model guide says plainly that the Coder is weaker outside code and in non-English text, citing an issue where Chinese answers came out wrong or looping.
Built for agents, installed by agents
The integration surface is the part builders should study:
- OpenAI-compatible at
http://127.0.0.1:8080/v1, any key, any model name. - Anthropic Messages at
/v1/messages, so Claude Code works withANTHROPIC_BASE_URL=http://127.0.0.1:8080. - Responses API at
/v1/responsesfor Codex CLI. - An MCP server so an assistant can install, start and stop Strata itself.
The recommended install path is to paste one sentence into your coding agent telling it to follow AI_SETUP.md, which checks your GPU, RAM and disk and picks a size. Yes, the installer for a local model is a prompt for a cloud model. Hacker News noticed: the top security subthread compared it to curl | bash, and the submitter said they read setup.py and setup.sh before trusting it. You should too.
The maker story
The project is essentially one person. The GitHub account Niko1221 has a single public repo, and its owner accounts for 520 of the commits; the next-largest of the 30 contributors has 21. The release cadence is the tell of a maker working in public with a hungry user base: 30 tagged releases in eight days, each landing fixes and features users reported (a mapped low-RAM mode for 32 GB machines, AMD HIP support, a CJK-aware draft vocabulary, multi-GPU). The only monetization is a Buy Me a Coffee button, which, against 962 forks, is a very generous exchange rate.
The lesson for anyone building on open models: the frontier was never just the weights. Qwen released a model that "usually needs a server"; a single developer built the scheduler, packaging and installer that turned it into something you double-click. Distribution and UX around an open model are a product, and this one found its audience in about a week. The flip side is visible too: the LocalLLaMA subreddit has already produced a "Strata is good now please stop" thread complaining about the volume of praise posts.
Caveats
- 2-bit is 2-bit. The fastest builds are 2-3 bit quantizations. One HN commenter pointed to research finding 2-bit quantization often causes broad degradation, and another reported severe quality loss from their own pruned 2-bit experiment. The README recommends IQ3_S when RAM allows.
- Self-reported speed. The headline tables come from the author's machines, and the community reports are a handful of users. Your RAM speed, PCIe link and CPU matter a lot in a design where the CPU computes the cache misses.
- Startup hurts. The README warns your PC may be slow or unresponsive for 1-3 minutes while it loads 35-55 GB into RAM and locks part of it for the GPU.
- One request at a time by default; parallel slots exist but slow each answer on a 12 GB card.
- Not bit-for-bit deterministic at temperature 0 by default, because CPU and GPU kernels round differently; there is an opt-in flag set for reproducible output.
- Full quality is slow. The near-lossless Unsloth UD-Q4_K_XL build streams from SSD and manages 7-8.5 tokens/s on a 64 GB PC.
Key Takeaways
- Strata runs the 125B-parameter (6B active) Qwen3.8-Flash-Next on 12 GB+ NVIDIA or AMD cards with 32-64 GB of RAM, under an MIT license.
- The author measured 53-94 output tokens/s and 1,600-2,650 prompt tokens/s at 32K on an RTX 5070, depending on the quant.
- The method is tiering: hot experts cached in VRAM, all experts in RAM with the CPU computing misses in parallel, the n-gram table on SSD, plus MTP speculative decoding.
- It speaks OpenAI, Anthropic and Responses APIs locally, so existing coding agents can use it without code changes.
- A largely one-person repo reached nearly 11,000 stars and 30 releases in its first ten days, which says a lot about demand for packaging, not just weights.
- The fast builds are aggressively quantized or pruned; test quality on your own tasks before you cancel anything.
Sources: Strata on GitHub, Strata technical details, How Strata works, Qwen3.8-Flash-Next model card, ISTA-DASLab Coder GGUF model card, ISTA-DASLab GSQ-RCO GGUF, Hacker News discussion, Strata community MI50 benchmark