DeepSeek V4.1-Flash Won’t Run — I Tried to Get It Working Locally
DeepSeek V4.1-Flash has been getting a lot of attention. Its total parameter count is a huge 552B, but the design is efficient enough that only 8B (while reading) or 16B (while generating) actually run per token, and its context window goes up to 1 million tokens. It was also posting strong numbers on coding benchmarks.
Looking into it, it seemed like it could be run through Ollama, so I tried it on the machine I have on hand, a mini PC with 128GB of unified memory, the EVO-X2. As things stand, I haven’t managed to get it running on the EVO-X2. That said, this is the result for one distributed file that I actually tried. As I’ll get into below, there are several other distributions out there, and I haven’t been able to try all of them.
This is research as of September 2026. I tried loading it on my own machine and confirmed the failure myself. I haven’t measured the quality of the text it generates.
- 1. What I Looked Into, and How Far
- 2. What Was Generating the Buzz
- 3. Five Months After V4 — What Changed?
- 4. How Large Are the Distributed Files?
- 5. Capacity Wasn’t the Problem — Loading Itself Stopped
- 6. The Model Listed on Ollama Is the “Cloud" Version
- 7. Is There Any Way to Run It Right Now?
- 8. What I Haven’t Confirmed in This Article
- 9. Summary — It Stopped Short of the Capacity Question
What I Looked Into, and How Far
I looked into four things: what’s generating the buzz, when it was released, how large the distributed files are, and whether my own machine can load it.
What Was Generating the Buzz
I checked the model card on the distribution site (Hugging Face) along with the technical report published alongside it. Every number in this section is self-reported by the developer, DeepSeek.
An Efficient Design
- Out of 552B total parameters, only 8B (while reading) or 16B (while generating) actually run per token (a design called Causal Encoder-Decoder. The previous generation, DeepSeek-V4-Pro, ran 49B out of 1.6T parameters every time, so the amount of active compute has been cut sharply)
- An MoE (Mixture of Experts) design with 384 “experts" on hand, where only 1 shared expert plus 6 selected ones run for each token
- It also carries an extra mechanism for holding onto conversation memory over long stretches (Engram, equivalent to 196B) and a mechanism that speeds up generation by looking ahead (DSpark)
It Can Hold a Long Conversation in a Light Memory Footprint
The technical report’s title is “KV Cache Compression" (compressing the memory area used to hold a conversation), and that appears to be this release’s main highlight. The context window tops out at 1 million tokens. The memory footprint per token of conversation is given as 890 bytes, which the report states is about a quarter of the previous version’s (DeepSeek V4-Flash) and about 1/437th compared with an earlier generation (DeepSeek-V1).
It Goes Toe-to-Toe With Well-Known Models on Benchmarks
On several benchmarks, numbers were listed side by side with other well-known models (all figures below are the developer’s own published values).
| Benchmark | DeepSeek V4.1-Flash | Comparison |
|---|---|---|
| Terminal-Bench 2.1 (agentic terminal-operation tasks) | 90.6% | Opus-5.0: 89.1% / GPT-5.6 Sol: 88.8% |
| DeepSWE v1.1 (code-fix resolution rate) | 74.2% | Opus-5.0: 74.0% / GPT-5.6 Sol: 73.0% / Previous version (V4-Flash): 54.4% |
| CyberGym (security-related tasks) | 88.1% | GPT-5.6 Sol: 84.5% |
| GPQA Diamond (graduate-level expert knowledge) | 90.9% | Opus-5.0: 93.4% / GPT-5.6 Sol: 94.1% |
On several coding and terminal-operation benchmarks, it matches or beats well-known models like Opus-5.0 and GPT-5.6 Sol, while on GPQA Diamond, which tests specialist knowledge, it comes in a bit lower. It doesn’t win across the board.
It Handles Both Text and Images
This is a model that handles images as well as text, posting 95.6% on a benchmark that tests understanding of images (DocVQA) and 86.0% on one that tests locating objects (RefCOCO). A feature that lets you continuously adjust “how deeply it thinks" on a scale from 1 to 100 was also listed as a new addition aimed at agentic use (using tools to carry out a task).
The model card also stated that, at the time of release, 59 different quantized versions were already prepared for llama.cpp, Ollama, LM Studio, and Jan. Going by that number alone, support looks solid, but where I actually got stuck was one step further than that: whether the tools themselves could read the new structure.
I also checked where it was actually being run. What stood out most were large-scale, data-center-grade GPU setups (reports of it running on 8 RTX PRO 6000 cards, for example), but among those I also found a record of someone trying to run it on a consumer machine with a large amount of unified memory (a 192GB Mac Studio). My EVO-X2 (128GB unified memory) sits closer to that second category. That’s what motivated me to try it myself.
Five Months After V4 — What Changed?
| Model | Released |
|---|---|
| DeepSeek V4-Flash | April 22, 2026 |
| DeepSeek V4.1-Flash | September 10, 2026 |
That’s about 5 months after the previous version. Checking the model architecture listed on the distributor’s Hugging Face page, V4.1-Flash is registered under a new type, DeepseekV41ForCausalLM (deepseek41).
What Changed
I compared the previous version (DeepSeek V4-Flash) against the distributor’s config files and model card. Everything below is published on Hugging Face for both versions.
| Item | V4-Flash (April 2026) | V4.1-Flash (September 2026) |
|---|---|---|
| Total parameters | 284B | 552B |
| Active per token | 13B | 8B (reading) / 16B (generating) |
| What it handles | Text only | Text and images |
| Context window | 1 million tokens | 1 million tokens (unchanged) |
| Number of layers | 43 layers | 40 layers (2 stages of 20 + 20) |
| Number of experts | 256 | 384 (still 6 active per token) |
| Training data used | Over 32 trillion tokens | 45 trillion tokens |
What caught my eye is that the total parameter count roughly doubled, while the amount active per token went down. At the reading stage specifically, it dropped from 13B to 8B. The documentation explains this as a design aimed at handling long inputs.
The other change is that it can now handle images. V4-Flash was a text-only model. V4.1-Flash adds an image-handling component, and the documentation states that text and images were trained together from the start.
The context window stays at 1 million tokens, unchanged. The previous version already supported 1 million tokens, so this isn’t something that grew with the new release. What changed instead is the amount of memory needed to hold onto that full 1 million tokens.
There’s also a new setting. A mechanism was added that lets you specify how much it thinks before answering as an integer from 1 to 100. It puts control over how much thinking (and cost) goes into an answer in the user’s hands.
The Scores Didn’t All Go Up
The model card includes a benchmark table measured under the same conditions as the previous version. I pulled out the comparison between the two base models. All figures are the developer’s own published values.
| Benchmark | V4-Flash-Base | V4.1-Flash-Base | Difference |
|---|---|---|---|
| HumanEval (writing code) | 69.5 | 79.4 | +9.9 |
| SimpleQA-Verified (knowing facts) | 30.1 | 42.3 | +12.2 |
| MMLU-Pro (broad knowledge) | 68.3 | 74.1 | +5.8 |
| BigCodeBench (coding tasks) | 56.8 | 60.6 | +3.8 |
| GSM8K (word-problem math) | 90.8 | 93.0 | +2.2 |
| MGSM (multilingual math) | 85.7 | 80.2 | −5.5 |
| DROP (reading comprehension) | 88.6 | 87.9 | −0.7 |
| BBH (a mix of reasoning tasks) | 86.9 | 86.1 | −0.8 |
| AGIEval (exam-style questions) | 83.9 | 83.4 | −0.5 |
Coding and factual knowledge improved. On the other hand, multilingual math dropped by 5.5 points. Reading comprehension and reasoning benchmarks also slipped slightly. Switching to the new version doesn’t come out as an improvement across the board.
There’s no equivalent number for V4-Flash on the image-handling benchmarks. It was a text-only model, so there’s nothing to compare against.
How Large Are the Distributed Files?
Multiple distributors have published GGUF versions (the format used to run a model locally) at different quantization levels. Some of these were only partially uploaded, with the full set of split files missing, so I counted how many split files were actually present for each and confirmed the smallest complete one.
| Quantization | Size | Status |
|---|---|---|
| Q1_0 | 98.6GiB | The smallest one with a complete set of split files |
| MixedQ2 | 157.3GiB | Complete |
| Q2_K | 246.3GiB | Complete |
My EVO-X2 has 128GB of unified memory. The actual upper limit assigned to the GPU side comes to about 112GiB (I confirmed the boot-time settings amdgpu.gttsize and ttm.pages_limit directly from /proc/cmdline on the machine. Both are set to the equivalent of 112GiB).
The smallest, Q1_0 (98.6GiB), fits within that 112GiB ceiling. Going by size alone, it should load.
I Worked Out Which Quantization Levels the EVO-X2 Could Plausibly Run
I re-sorted the three distributed quantization levels by how many bits per parameter each one compresses down to (working backward from the 552B total parameter count).
| Quantization | Size | Per parameter | Fits in 112GiB? |
|---|---|---|---|
| Q1_0 | 98.6GiB | About 1.53 bits | Fits (about 13GiB of margin) |
| MixedQ2 | 157.3GiB | About 2.45 bits | Doesn’t fit |
| Q2_K | 246.3GiB | About 3.83 bits | Doesn’t fit |
Converting the 112GiB ceiling into bits per parameter comes out to about 1.74 bits. The lightest quantization currently available, Q1_0 (1.53 bits), already clears that bar. Put another way, from a pure capacity standpoint, there’s no need to wait for an even lighter quantization to show up.
That said, there’s a large gap between Q1_0 and MixedQ2, spanning 1.53 to 2.45 bits, and nothing has been distributed yet to fill that gap. Memory is also needed separately for the KV cache used to hold a long-running conversation, so that 13GiB of margin can’t really be called generous.
Capacity Wasn’t the Problem — Loading Itself Stopped
Capacity wasn’t the issue, but llama.cpp has no entry for this model. Just to check, I fed the GGUF file alone into the llama.cpp build I use for testing on this machine (build b10605), and it failed right at the loading stage. As expected, since llama.cpp doesn’t support it, it doesn’t run.
For reference, a different new model (Spark-X2.5-4B, a 4B-class model at 2.42GiB) failed in the exact same way. That one had plenty of capacity to spare, but llama.cpp doesn’t support it either, so it didn’t run.
The Model Listed on Ollama Is the “Cloud" Version
DeepSeek V4.1-Flash showed up in Ollama’s model listing, so I tried ollama run deepseek-v4.1-flash. The result: a complete failure. It didn’t run.
Digging further, it turns out this isn’t a model that runs locally through Ollama — it’s an Ollama Cloud model. It had racked up 12,200 pulls (a download-like metric) within 3 days of release, so I’d assumed it was something that ran on my own machine. Just to be sure, I checked the full list of tags too.
There was only one tag available: deepseek-v4.1-flash:cloud. That cloud tag is a feature where Ollama forwards the request to a model running in the cloud — the weights never actually land on your own computer. This is one example of how “usable through Ollama" and “runs on your own GPU" turned out to be two different things.
I also went back and checked what Ollama Cloud actually is on its official page. It requires an account, and beyond the free tier there are $20/month and $100/month plans. The actual computation runs on Ollama’s own servers (mainly in the US, and in some cases Europe or Singapore), meaning whatever you type gets sent to Ollama. The official page states that “input and output are not logged or used for training," but that’s their own operating policy, and it isn’t something I can verify myself.
What this blog calls a “local LLM" is a way of using a model entirely inside your own computer, without sending any data outside. Ollama Cloud uses the same Ollama tool, but the model actually runs on an outside server, which is a different question from the one I set out to check: whether it runs on my own EVO-X2.
Is There Any Way to Run It Right Now?
Once I understood why it wouldn’t load, I looked into whether there was any way to get something running right away.
- The pull request for support in upstream llama.cpp (
convert : add DeepSeek V4.1) was submitted on September 10, 2026, and as of September 14 it still hasn’t been merged (checked its status on GitHub) - I found a community fork of llama.cpp built specifically for the Ryzen AI Max family (the same chip family as the EVO-X2), but a request there asking for DeepSeek V4.1 support was also still open and unresolved, so it isn’t usable right now either
- The previous version (DeepSeek V4-Flash) is registered under a different architecture type (
DeepseekV4ForCausalLM), and upstream llama.cpp already has code that handles it. I haven’t tried it myself, but until V4.1-Flash becomes loadable, the previous version looks likely to load in the meantime - While searching, I also found that community members have already distributed a lightweight quantization roughly equivalent to Q1. I couldn’t judge how trustworthy that distributor is, though, so I decided not to try it this time. That means I don’t know whether these versions actually load
Even when there’s enough capacity, if the tool you’re running it with (llama.cpp) doesn’t support that model’s architecture, it stops right at the loading stage. Newer models sometimes adopt newer architectures, and there’s a lag before tool support catches up. That said, this applies to the specific distributed file I tried. Other community-distributed files may be a different story.
What I Haven’t Confirmed in This Article
- It may well run through a different tool, such as Ollama. I only checked with llama.cpp (b10605)
- I’m guessing a newer llama.cpp build won’t change this, going by the status of the pending pull request (not yet merged), but I haven’t actually tried one
- I haven’t tried whether the previous version (DeepSeek V4-Flash) actually loads on my own machine
- I haven’t tried whether the community-distributed Q1-equivalent quantization actually loads. That’s because I couldn’t verify how trustworthy the distributor is
- I haven’t measured answer quality or speed. I only checked whether it loads
Summary — It Stopped Short of the Capacity Question
DeepSeek V4.1-Flash comes to at least 98.6GiB even at its smallest quantization, but that’s a size that fits, capacity-wise, on my 128GB machine (whose actual GPU allocation ceiling is about 112GiB). What the distributed file I tried got stuck on wasn’t capacity — it was the loading stage before capacity even came into play. That said, this isn’t a final verdict of “it doesn’t run" — it’s “it didn’t run within what I tried." A separate, roughly Q1-equivalent community distribution already exists, and I haven’t been able to try that one.
Setting that aside, it’s worth writing up the capacity question on its own. Given that the smallest quantization currently distributed is 98.6GiB, running this locally requires at least about 99GB of memory. Even on the EVO-X2, with its 128GB of unified memory, the gap to the actual usable ceiling (about 112GiB) is only about 13GiB — not much room to spare. For an ordinary computer (typically in the 16-64GB range), this rules it out before capacity even becomes a question.
Judging only by “does it have enough capacity" would have missed a loading failure like this one. At the same time, this is also a model with a genuinely high capacity bar to clear. When trying a new model, it seems worth checking both the capacity math and whether the tool you’re using supports that model’s architecture. Until support lands, trying the previous version remains an option.
I’ll keep an eye out for whether a lighter quantization, or a smaller version that runs on less memory, shows up, and wait it out for now.