Does Qwen3.8-Flash-Next, a Qwen4 Preview, Run at Home? I Measured the 2-bit Version on an RTX 5060 Ti and RTX 3060
9 October 2026
On my X timeline I kept seeing posts saying “Qwen3.8-Flash-Next runs at home." It is a 125B-parameter model that previews the Qwen4 design, built so that only 6B parameters work at once. Can it really run on the RTX 5060 Ti (16GB) and RTX 3060 (12GB) I have here? I measured the 2-bit version with a dedicated engine.
This reflects measurements as of October 2026.
- 1. What is Qwen3.8-Flash-Next? A preview of Qwen4
- 2. What I measured and how: on 12GB and 16GB graphics cards
- 3. Results: writing speed 71 tok/s on the 5060 Ti, 46 tok/s on the 3060
- 4. Why does it run on 12GB? Only part of it is loaded
- 5. Smartness not measured: the cost of 2-bit is unconfirmed
- 6. A third-party measurement: RTX 5070 12GB, 64GB of memory
- 7. What I Haven’t Confirmed in This Article
- 8. Summary: the Qwen4 preview has reached “writes fast" on a 16GB graphics card
What is Qwen3.8-Flash-Next? A preview of Qwen4
Qwen3.8-Flash-Next is a model whose weights the Qwen team at Alibaba released in late August 2026 (Hugging Face: Qwen/Qwen3.8-Flash-Next). It is a way to try Qwen4’s new design (Gated DeltaNet, Qwen Sparse Attention, n-gram embeddings) early. Of its 125B parameters, only 6B work per token, as an MoE (a design that holds many specialists and uses only some per question). According to the distributor, the files include a 51B-parameter n-gram embedding table and a 4B MTP head for speculative decoding (a mechanism that drafts several tokens ahead and emits them together). On Hugging Face, unsloth, bartowski, ISTA-DASLab and others publish GGUF files (the format read by llama.cpp-family software).
The hard part is running it. The pull request that adds support for this model’s architecture (ggml-org/llama.cpp #27742) was merged into mainline llama.cpp on 27 August 2026. However, I have not confirmed whether mainline can handle the GSQ-RCO format (2-bit) and the MTP head I used, and I did not try a mainline build. Instead I used the GSQ-RCO GGUF distributed by ISTA-DASLab (Hugging Face: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF) and a dedicated engine that can read it, Strata (GitHub: Niko1221/Strata, version 0.1.31).
What I measured and how: on 12GB and 16GB graphics cards
| Item | Details |
|---|---|
| Machine | GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB memory, Ubuntu 26.04). RTX 5060 Ti 16GB and RTX 3060 12GB attached externally over OCuLink |
| Model | ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, Q2_0 (2-bit), about 66GB (35.0GiB body + 26.8GiB n-gram table; the 61.8GiB total is about 66GB in decimal) |
| Engine | Strata 0.1.31 (CUDA 13.0, KV cache = the conversation’s memory, held in int8; speculative decoding setting 4; MTP on; maximum context 32,768 tokens) |
| How I measured | Two cases, a short chat (about 73 input tokens) and a long input (about 5,585 tokens), five runs each. Sorted, dropped the highest and lowest, and took the median of the middle three |
| Measured on | 2 October 2026 |
The parts that do not fit on the GPU (the n-gram table and most of the experts) are read from the machine’s memory and SSD. The GPU’s VRAM is used for an expert cache and the KV cache. On the 5060 Ti, about 9.3GB was allocated to the expert cache.
Results: writing speed 71 tok/s on the 5060 Ti, 46 tok/s on the 3060
| Setup | Writing speed, short chat [tok/s] | Writing speed, long input [tok/s] | Time to first token, short chat [s] | Time to first token, long input [s] |
|---|---|---|---|---|
| RTX 5060 Ti 16GB alone | 70.7 | 70.7 | 1.04 | 4.42 |
| RTX 3060 12GB alone | 45.6 | 47.0 | 1.57 | 6.19 |
| Both (layer split) | 70.1 | 73.5 | 1.05 | 19.44 |
All five runs succeeded in every case, and the spread of the middle three was 1-2 tok/s. Writing speed is 70 tok/s on the 16GB RTX 5060 Ti and 46 tok/s on the 12GB RTX 3060. Against this blog’s other measurements, that is faster than the 12.6 tok/s of Qwen3.6 27B (4-bit) on the same EVO-X2’s integrated GPU, and short of Qwen3 Coder 30B (91 tok/s). A model billed as 125B writing faster than most people read on a 12GB graphics card was a surprising result.
Reading the long input (about 5,600 tokens) took 4.4 seconds on the 5060 Ti and 6.2 seconds on the 3060. Linking the two cards left writing speed unchanged but took 19.4 seconds to read, 4.4 times the 5060 Ti alone (3.1 times the 3060 alone). I have not checked why. Under these conditions, with this model and Strata’s layer split, there was no benefit to combining them.
Why does it run on 12GB? Only part of it is loaded
The file is 66GB, yet it runs on a 12GB GPU because only the experts in use at the moment (6B worth) sit in the GPU’s VRAM, and the rest stays in the machine’s memory and is read when needed. The EVO-X2 has 128GB of memory, so the whole 66GB file fits in memory. The Strata distributor says it also runs on a 64GB-memory machine with the n-gram table placed on an SSD. I have not tried it, but a measurement by a third party on a 64GB machine is in the next section.
Smartness not measured: the cost of 2-bit is unconfirmed
I measured speed only. 2-bit quantization is a format with a large loss of accuracy, and I have not checked how much smarter the same model’s 4-bit version or Qwen3.8-27B is. This blog’s smartness test (7 coding problems, 6 trick questions) runs through Ollama and is not connected to Strata. I will leave measuring smartness for the next article.
A third-party measurement: RTX 5070 12GB, 64GB of memory
The Strata developer’s public record (dated 29 September 2026, Strata 0.1.26, the same 2-bit Qwen3.8-Flash-Next) includes a measurement on a machine with an RTX 5070 (12GB), Ryzen 5 7600, 64GB of memory (DDR5-5200) and Windows 10. It is not this blog’s measurement, and it is one run per condition. The 64GB is the configuration of the measured machine, not a statement of the minimum required.
- Writing speed (4K input bucket): 93.0 tok/s
- Writing speed (32K input bucket): 81.8 tok/s
- Writing speed (128K input bucket): 73.7 tok/s
- Reading speed (32K input bucket): 2,171 tok/s
- From handing over a long log (31,566 tokens) to the first token: 16.02 seconds
It is natural that these differ from my 5060 Ti result (70.7 tok/s), but the versions (0.1.26 and 0.1.31), machines, input lengths and output lengths do not match, so I do not line them up to rank GPUs. What the record shows is that the same model has been run on a 64GB-memory, 12GB-VRAM setup. Even with a reading speed over 2,000 tokens per second, a long document still takes more than ten seconds to the first token.
Source: bench/results/2026-09-29-speed-0126 in github.com/Niko1221/Strata (checked 4 October 2026).
What I Haven’t Confirmed in This Article
- Smartness (how much is lost at 2-bit)
- The speed difference when run on mainline llama.cpp (and whether mainline can read the GSQ-RCO format)
- A repeat on a 64GB-memory machine (there is a third-party record, but I have not repeated it)
- Why reading a long input got slower when two cards were linked
- Running on the EVO-X2’s integrated GPU (Radeon 8060S) alone (Strata is for CUDA, and I did not try)
Summary: the Qwen4 preview has reached “writes fast" on a 16GB graphics card
The 2-bit Qwen3.8-Flash-Next wrote at 70 tok/s on the RTX 5060 Ti and 46 tok/s on the RTX 3060. I have not confirmed whether the GSQ-RCO format used here can be read by mainline llama.cpp, so I ran it with the dedicated engine Strata. Smartness is unmeasured, and I will check how much the 2-bit cost is in the next article.
As a next step, I suggest this split: if you have a machine with 128GB of memory, try the 2-bit version; if not, start from the 4-bit Qwen3.8-27B, which I have measured on this blog.
New LLM models, classified (September 2026)
Adding an RTX 5060 Ti to a mini PC over OCuLink
Hardware used for the tests
GMKtec EVO-X2 (Ryzen AI Max+ 395 / 128GB / 2TB)A8限定クーポン「A82608」で5,000円OFF(2026/10/31まで)/公式オンラインストアのクーポン「KANSYA2026」で2,000円OFF(2026/10/18まで)
¥583,000 Amazon・2026-10-04調べ
A82608(2026-10-31まで)AOOSTAR AG02 eGPUドッキングステーション(OCuLink+USB4)800W電源内蔵・最大600WのGPUに対応