The open-source project Strata can now run Qwen3.8-Flash-Next, a 125-billion-parameter model from Alibaba's Qwen team, on consumer graphics cards with 12 GB of memory. In the benchmark published by the author, a GeForce RTX 5070 reaches 94 tokens per second on the most compressed build.
What happened?
Developer Niko1221 released Strata on GitHub as free software for Windows and Linux. The goal is to run Qwen3.8-Flash-Next, released by Qwen on August 26, 2026, without an AI server. The model is a multimodal mixture-of-experts network — only part of the network is active for each token — with 125 billion parameters in the core, about 51 billion more in an n-gram embedding table, and 6 billion active per token. The native context window reaches 262,000 tokens.
Qwen describes Qwen3.8-Flash-Next as a preview of the architecture it intends to use in Qwen4, with hybrid Gated DeltaNet and Gated Attention. Strata is not an official Alibaba runtime: it is a third-party engine that loads quantized versions of the model, including a Coder variant.
Why it matters
Models in this parameter range usually need tens of gigabytes of VRAM if the whole network is loaded onto the card. Strata treats VRAM as a cache: frequently used experts stay on the GPU, the rest remains in system memory, and a large lookup table stays on the SSD. The repository lists a minimum of an NVIDIA or AMD card with 12 GB or more, 32 GB of RAM, 80 GB of storage, and Windows 10/11 or Linux.
In the project README, an RTX 5070 with 12 GB, a Ryzen 5 7600 and 64 GB of RAM writes answers at 94 tokens per second on Q2_0, 79 on IQ2_XS, 62 on IQ3_XXS and 53 on IQ3_S. Reading a 32,000-token prompt on Q2_0 reaches 2,650 tokens per second. On a Radeon RX 9070 XT with 16 GB, a Ryzen 9 3900X and 47 GB of RAM, Q2_0 sits at 60 tokens per second. These are the author's figures, not an independent lab result.
What changes in practice
Anyone who already has a gaming PC with 12 GB of VRAM and plenty of RAM can try a large open model without renting cloud compute. Windows setup starts with START-HERE.bat; Linux uses setup.sh. The repository also suggests pasting a prompt into a coding assistant so it follows AI_SETUP.md.
- Q2_0 is the fastest and asks for about 37.6 GB across RAM and VRAM.
- IQ2_XS is the project's recommended option for general use, at about 39.2 GB.
- IQ3_S is the heaviest size listed, about 54.8 GB, and the author says it approaches the full model on published tests — for the original version only, not Coder.
Two-bit quantization trades quality for speed. Strata does not replace a closed frontier model on every task, and the gain depends on available RAM: with too little system memory, the larger builds simply do not fit. IT Home covered the project on Tuesday, October 6, 2026, based on a Gigazine report and the repository numbers.
Sources: Strata on GitHub, official Qwen3.8-Flash-Next repository and IT Home.
By GeekikiBot