How Many GB Does DeepSeek V4.1-Flash Need to Run? — I Checked Every Quantization Being Distributed
DeepSeek V4.1-Flash is getting a lot of attention. If you’re trying to run it on your own PC, the thing you probably want to know first is: how many GB of memory does it take to run? I went back through every quantization (a format that shrinks the numbers inside the model) being distributed on Hugging Face and recounted all of them, as of September 28, 2026. Is there one that actually fits on a PC you’d have at home?
This is what I found as of September 28, 2026.
I tried loading it onto this blog’s GMKtec EVO-X2 (128GB of memory) and forcing it to run even on the assumption that it would spill over onto the SSD, and it didn’t work — I wrote up that attempt here.
- 1. What Is DeepSeek V4.1-Flash?
- 2. What I Checked, and How Far
- 3. How Many GB Are the Quantizations Being Distributed?
- 4. Which Machines Can Fit It?
- 5. Before Capacity — Is There Software to Run It?
- 6. Does the Previous V4-Flash Run?
- 7. If It Doesn’t Run on Your Own PC, Where Can You Use It?
- 8. What This Article Hasn’t Confirmed
- 9. Summary — Can You Run V4.1-Flash on Your Own PC?
- 10. Sources
What Is DeepSeek V4.1-Flash?
DeepSeek V4.1-Flash is a generative AI model that China’s DeepSeek published on Hugging Face on September 10, 2026. It’s released under the MIT license, and anyone can download the weights (the model’s core data).
According to the official documentation (the model card and technical report), there are two main goals. One is to get better at “agent" work, where the AI uses tools to carry out multiple steps on its own. The other is to shrink the area that remembers long conversations (the KV cache). Measured per token, this area is described as about a quarter the size of the previous V4-Flash, and about 1/437th the size of the original DeepSeek V1. Handling long contexts of up to 1 million tokens, and being able to read images, are also new.
“Flash" is the name DeepSeek gives to the models in its lineup that prioritize speed and lightness. In the V4 generation, two versions came out: the lighter “V4-Flash" and the larger “V4-Pro." In the V4.1 generation, as of September 28, only Flash has been released; V4.1-Pro hasn’t come out.
| Model | Released | Total size | Amount used per calculation (active parameters) | Context length | Images |
|---|---|---|---|---|---|
| DeepSeek V4-Flash | April 22, 2026 (updated version on July 31) | 284B | 13B | — | — |
| DeepSeek V4-Pro | April 22, 2026 (updated version on August 13) | 1.6T | 49B | — | — |
| DeepSeek V4.1-Flash | September 10, 2026 | 552B (plus a 196B Engram module for memory) | 8B / 16B | 1 million tokens | Yes |
B stands for 1 billion. The “Total size" and “Amount used per calculation" columns in the table are the values from the official model card. Despite the name “Flash," V4.1-Flash’s total size ended up about twice that of V4-Flash. On the other hand, the amount it uses per calculation is 8B (when reading input) and 16B (when generating text), which is about the same as V4-Flash’s 13B. Even though only part of it is used in each calculation, though, running it on your own PC requires putting the whole thing into memory. It’s the total size that determines how much memory you need. The reason the file is so large comes down to that total size, plus the fact that it carries an Engram memory module (196B).
On the basic benchmarks the developer publishes, V4.1-Flash scores 74.1 on MMLU-Pro (a broad knowledge test) — about the same as V4-Pro (73.5), which is roughly three times its total size. That said, this is the developer’s own evaluation, and this blog hasn’t verified it.
What I Checked, and How Far
I gathered every Hugging Face repository with “DeepSeek-V4.1" in its name, plus every repository registered as “based on DeepSeek-V4.1-Flash," and totaled up the GGUF files (the format used by tools like llama.cpp) by quantization tier. I only counted cases where all of the split files were fully uploaded. I excluded repositories that held only the image-reading component, or that turned out to be empty.
The official DeepSeek release doesn’t include a GGUF. Everything in the table below is a volunteer’s own conversion, distributed independently.
How Many GB Are the Quantizations Being Distributed?
| Quantization | Size [GiB] | Distributor | llama.cpp patch files |
|---|---|---|---|
| DSV41-mixedq2 (custom mix) | 157.3 | apetersson | — |
| Q2_K | 246.3 | AMAImedia / vcruz305 | Included |
| EngramQ5-Q2_K (custom) | 312.3 | smalinin | — |
| Q3_K_M | 323.4 | Solstice-AI / vcruz305 | Included |
| Q2 (custom, single file) | 340.6 | antirez (plus two more with safety guardrails removed) | — |
| EngramQ8-IQ2_XXS family (custom) | 345.9 | smalinin | — |
| ROCMFP2S / MIX10 / ROCMFP23 (custom) | 356.4-377.5 | Lucebox | — |
| MXFP4 | 375.8-473.1 | mxxm-t / kernelpool / JigSawPT / smalinin | — |
| EngramQ8-Q2_K_Protected (custom) | 385.0 | smalinin | — |
| GSQ-RCO 3.0bit (custom) | 387.2 | pfeifferj / taurusduan | Included |
| EngramQ8-IQ3_XS (custom) | 404.2 | smalinin | — |
| Q4_K_M | 414.2 | AMAImedia / Solstice-AI / vcruz305 | Included |
| Q8_0 | 473.1 | Solstice-AI / vcruz305 | Included |
| Q4 (custom) | 483.0 | antirez | — |
The lightest is MixedQ2 at 157.3GiB, and even Q2_K, the lightest among the standard tiers, comes to 246.3GiB. When I checked on September 13, there was also a Q1_0 tier at 98.6GiB, but as of September 28 I couldn’t find it. It appears to have been taken down.
“Included" in the right-hand column means files for patching llama.cpp (patches) are bundled alongside the model. Read straightforwardly, that means these aren’t meant to load in the official llama.cpp as-is — a modified llama.cpp is assumed.
Which Machines Can Fit It?
GiB and GB are slightly different units. 1GiB is about 1.07GB, so a machine with “192GB of memory" works out to about 178.8GiB. The comparisons below use that conversion.
Even on a machine where the CPU and GPU share memory (unified memory), you can’t hand all of it over to the GPU. You need to leave a few GiB aside for the OS.
| Machine (memory) | Rough amount assignable to GPU | Quantization that fits | Assessment |
|---|---|---|---|
| GMKtec EVO-X2 (128GB) | About 112-120GiB (this blog’s own measurement, Linux) | None | Doesn’t work |
| GMKtec EVO-X5 Pro (192GB), Windows | Up to 160GB (official) = about 149GiB | None (even the lightest, MixedQ2, is 157.3GiB, which exceeds about 149GiB) | Tough |
| GMKtec EVO-X5 Pro (192GB), Linux | About 170GiB (estimate, assuming the same settings as EVO-X2) | MixedQ2 (157.3GiB) | Looks like it fits, capacity-wise |
| Mac Studio (256GB) | About 230GiB (estimate, with the ceiling raised via settings) | No GGUF. A volunteer-pruned MLX version exists (198.3GiB) | Possible with a volunteer-pruned version |
| Two NVIDIA DGX Sparks connected together (128GB x 2 = 256GB) | About 230GiB combined (estimate, assuming about 115GiB per unit) | MixedQ2 (157.3GiB). Assumes software that can split the model across two units | Looks like it fits, capacity-wise |
| 512GB class | About 470GiB | Q2_K through Q4_K_M | Plenty of headroom, capacity-wise |
DGX Spark is NVIDIA’s small AI-focused computer. Each unit has 128GB of memory, and two can be connected directly via the ConnectX-7 port on the back (up to 200Gbps). NVIDIA states that “connecting two units lets you handle models with up to 405 billion parameters at 4-bit (FP4)." V4.1-Flash has 552 billion parameters in total (plus 196 billion for memory), so it exceeds that official guideline. However, at MixedQ2, which works out to roughly 2.45 bits, the combined memory of two units comes out to enough capacity. Here too, this assumes software capable of splitting a single model across two machines.
The MLX version in the Mac Studio row isn’t a GGUF — it’s a format for Mac. rapid-mlx made it by trimming out some of the model’s internal experts (a method called REAP) and reducing it to 2 bits, so it may be less capable than the original model.
Before Capacity — Is There Software to Run It?
Even with enough capacity, it won’t run if the loading software doesn’t support it. V4.1-Flash uses a new architecture called Causal Encoder-Decoder. On this blog’s llama.cpp (build b10605), it stopped at the loading stage, before capacity even came into play.
The distributor of the lightest one, MixedQ2, writes on its description page that “this is just the weights placed here — there’s no software yet that’s been shown to run it all the way through." Most distributors of Q2_K and Q4_K_M bundle patch files for llama.cpp. As of September 28, I couldn’t find a form of it that runs on the official llama.cpp alone.
This blog has a policy of not running programs found from outside sources. I haven’t tried the modified llama.cpp.
Does the Previous V4-Flash Run?
For the previous generation, DeepSeek V4-Flash (284B), many GGUFs that fit on a 128GB machine have been released.
| Distributor | Quantization | Size [GiB] |
|---|---|---|
| takanori-ishikawa (a pruned 155B version, experts removed) | UD-IQ3_XXS | 55.0 |
| unsloth | UD-IQ1_S | 76.9 |
| unsloth | UD-IQ2_XXS | 84.6 |
| unsloth (July 31 updated version) | UD-IQ3_S | 108.1 |
| ggml-org | Q2_K | 109.3 |
ggml-org is the team behind llama.cpp. The fact that they’ve released a GGUF suggests V4-Flash is built to load in the official llama.cpp. This blog hasn’t run it yet.
Given that the GGUF and the software support for it came together first for V4-Flash, there’s a chance V4.1-Flash could become runnable on a home machine too, once official llama.cpp support arrives, or once a more aggressively pruned version shows up.
If It Doesn’t Run on Your Own PC, Where Can You Use It?
DeepSeek V4.1-Flash can be used right away through internet-based services. Here are the main ways to use it that I could confirm as of September 28, 2026.
| How to use it | Pricing model | Rough price (per 1 million tokens) |
|---|---|---|
| DeepSeek’s official site/app (chat) | Free | — |
| DeepSeek’s official API (model name deepseek-flash) | Pay-as-you-go | Input $0.15 / output $0.60 (off-peak hours). Double during peak hours |
| Ollama’s cloud (deepseek-v4.1-flash:cloud) | Monthly plan (Pro / Max / Team), or pay-as-you-go on a free account | Ollama states its pay-as-you-go pricing matches DeepSeek’s official API |
| OpenRouter (a gateway aggregating multiple providers) | Pay-as-you-go | Varies by provider (DeepSeek’s own listing matches its official $0.15 / $0.60) |
| Fireworks | Pay-as-you-go | Input $0.22 / output $0.66 (as listed on OpenRouter) |
“Peak hours" for DeepSeek’s official API are 10:00-13:00 and 15:00-19:00 Japan time on weekdays (excluding Chinese public holidays). When part of the input matches previous input (when the cache kicks in), the input price drops to as low as $0.003. Pricing changes often, so check each service’s pricing page before you use it.
If your reason for wanting to try it on your own PC is “I don’t want to send information outside," V4.1-Flash isn’t a fit yet. You’d need to use one of the previous V4-Flash’s lighter GGUFs, or wait for a smaller version of V4.1 to appear.
What This Article Hasn’t Confirmed
- I haven’t measured startup, speed, or capability on real hardware for any quantization (this blog doesn’t have a machine it fits on)
- The “amount assignable to the GPU" for EVO-X5 Pro (Linux) and Mac Studio is an estimate based on EVO-X2’s actual measurements and typical settings
- Whether it actually runs on a modified llama.cpp or on SGLang hasn’t been tested
- The reason the Q1_0 (98.6GiB) that was visible on September 13 was removed is unknown
- The official performance figures are the developer’s own evaluation, and this blog hasn’t verified them
Summary — Can You Run V4.1-Flash on Your Own PC?
DeepSeek V4.1-Flash is a large model DeepSeek released on September 10, 2026, aimed at agent work and long contexts. Despite the name Flash, its total size is about twice that of V4-Flash, and the lightest GGUF being distributed for it comes to 157.3GiB.
To sum up: it doesn’t work on an EVO-X2 (128GB); on an EVO-X5 Pro (192GB) it’s tough on Windows but looks like it fits, capacity-wise, on Linux; at the 256GB class, it’s possible with a volunteer-pruned version; and at the 512GB class, there’s plenty of headroom, capacity-wise. On every machine, this assumes software support for running it.
In conclusion, V4.1-Flash is a model you can’t run on your own PC without expensive hardware. In the V4 generation, volunteer-pruned versions and a GGUF readable by the official llama.cpp eventually came together. A smaller-sized version of V4.1 might show up too. Until then, waiting is the only option.
If you want to try it now, the easy route is the official chat or API. If you insist on your own PC, the previous V4-Flash’s lighter GGUFs (77-109GiB) are a realistic option on a 128GB machine. If you’re looking for models that run at a given amount of memory, this article may help too.
Sources
- DeepSeek V4.1-Flash model card (official)
- apetersson: MixedQ2 distribution page
- vcruz305: Q2_K through Q8_0 distribution page
- smalinin: Engram-family / MXFP4 distribution page
- rapid-mlx: expert-pruned MLX version
- ggml-org: DeepSeek V4-Flash GGUF
- unsloth: DeepSeek V4-Flash (July 31 version) GGUF
- DeepSeek API: models and pricing (official)
- Ollama: deepseek-v4.1-flash
- OpenRouter: DeepSeek V4.1 Flash
- Leadtek: NVIDIA DGX Spark Founders Edition (up to 405B when connecting two units)











Discussion
New Comments
No comments yet. Be the first one!