Too Many New AI Models, Can’t Tell Which Is Which? — I Sorted the Latest 17 by Size, Design, and Use
New generative AI models are showing up almost every day. Many of the names are ones I’ve never heard of, and it gets hard to tell which ones actually matter to me. What exactly are these recently released models?
This is what I found as of September 28, 2026.
- 1. The Generative AI Models I Looked At
- 2. The 17 Models, Ranked by Popularity
- 3. Which Sizes Actually Fit on Your Own PC?
- 4. What Is MoE, and Which Models Use It?
- 5. What Is Each Model Good At?
- 6. Who Released Each Model?
- 7. Where Is the Popularity Concentrated?
- 8. How Many Days After Release Did the GGUF Appear?
- 9. How Fast Are They on the Same PC?
- 10. What This Article Hasn’t Confirmed
- 11. Summary — If It’s Your Own PC, Which One Should You Try First?
The Generative AI Models I Looked At
I gathered the models from Hugging Face, the site where AI models get distributed. Almost every model used for local LLMs is published there.
The scope is 17 text models released between August 5 and September 21, 2026. I picked them from the top 60 on Hugging Face’s trending list and the top 30 as of September 13, keeping only the original releases from each developer. Versions remade by other people — shrunk versions, further fine-tuned versions — aren’t counted as separate models; their download counts are folded into the original’s total instead.
For each one, I checked the following:
- Popularity: the download count Hugging Face shows for the last 30 days — the total of the original file and the GGUF (explained below)
- Size: parameter count (the number of internal values in the model; 1B = 1 billion of them) and the file size once shrunk to 4-bit (Q4_K_M)
- Design: whether it’s “dense," using everything every time, or “MoE," using only part of it
- Stated strengths: what the developer writes on the model’s description page (the model card)
- GGUF release lag: how many days passed between the original release and the first GGUF appearing
“Stated strengths" summarizes what the developer says — this blog hasn’t verified any of it.
What Is GGUF?
GGUF (pronounced “gee-guff") is a file format for running generative AI models on your own PC.
Most of the original files developers publish come in a format called safetensors. Bigger models split into more files — among the 17 here, that ranged from 1 file up to 644. The internal values are also kept at full, unshrunk precision, so the files are large and mainly assume they’ll be run on dedicated server software.
GGUF is what you get when you convert that original file so llama.cpp — the leading free software for running generative AI on your own PC — can load it. It bundles the model’s values and settings into a single file, and most are shrunk (quantized) down to 4-bit or so before being distributed. Ollama and LM Studio, apps for running generative AI on your own PC, can also load GGUF.
Labels in filenames like “Q4_K_M" or “Q8_0" indicate the kind of shrinking used. The smaller the number, the smaller the file, but answer quality drops a bit each step. This article uses Q4_K_M as the baseline, since it balances size and quality well and is the most commonly used.
GGUF files are often made and published by volunteer converters rather than the developers themselves. Of the 16 models here (excluding Ternary-Bonsai-2-27B, which was distributed as GGUF from the start), only 3 — Xing4.0-29B-A4B, MiniCPM5-2B, and ZDTaichu5.0-9B — had a GGUF released by the developer itself. That’s why the gap between the original release and the GGUF appearing varies by model.
The 17 Models, Ranked by Popularity
I sorted them by download count and listed what kind of model each one is.
| No. | Model | Publisher | Released | Downloads (30-day) | Parameters | Design | 4-bit size | Stated strengths |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8-27B | Qwen (Alibaba) | Aug 5 | 38.81 million | 27.8B | Dense | 16.5GB | Good at coding and multi-step agent tasks |
| 2 | Qwen3.8-Flash-Next | Qwen (Alibaba) | Aug 24 | 5.18 million | 180.0B | MoE (10 of 512 experts) | 111.3GB (UD-Q4_K_XL) | An experimental new design built to speed up long-context agent processing |
| 3 | Ternary-Bonsai-2-27B | PrismML (startup) | Sep 16 | 3.56 million | — | — | 5.9GB (1.75-bit) | Keeps 27B-class reasoning ability well despite being small |
| 4 | DeepSeek-V4.1-Flash | DeepSeek | Sep 10 | 1.83 million | 763.2B | MoE (6 of 384 experts) | 444.7GB | Good at running efficiently and saving memory even with long context |
| 5 | MiniCPM5-2B | OpenBMB (Tsinghua-affiliated) | Sep 6 | 880,000 | 2.5B | Dense | 1.6GB | A lightweight 2B model for resource-constrained devices, with top-tier performance for its class |
| 6 | Nex-N2.5-mini | Nex AGI (startup) | Sep 8 | 170,000 | 35.1B | MoE (8 of 256 experts) | 22.3GB | Good at long-running real-world agent work |
| 7 | MiMo-V2.6-Distill-Qwen-9B | Xiaomi | Sep 21 | 130,000 | 9.4B | Dense | 5.8GB | Good at coding and tool-using agent work |
| 8 | Edge0-35B-A3B-preview | Edge0 (startup) | Sep 8 | 80,000 | 34.7B | MoE (8 of 256 experts) | — | A 35B-class model that runs on smartphone-level memory |
| 9 | MiMo-V2.6-Pro-RL | Xiaomi | Sep 21 | 77,000 | 1,024.2B | MoE (8 of 384 experts) | 528.9GB (MXFP4) | Good at large-scale reinforcement learning that keeps improving itself |
| 10 | Hemmingway-1 | Altworld (startup) | Sep 20 | 67,000 | 26.9B | Dense | 17.4GB | Good at writing everyday emails and messages in natural, human-sounding prose |
| 11 | Xing4.0-29B-A4B | XingChen-AGI (startup) | Sep 16 | 62,000 | 31.2B | MoE (4 of 64 experts) | 19.0GB | Good at agent work that plans ahead and uses tools |
| 12 | MiMo-V2.6-Flash-RL | Xiaomi | Sep 21 | 51,000 | 310.8B | MoE (8 of 256 experts) | 167.4GB (MXFP4) | A reinforcement-learning model that balances lightness and performance |
| 13 | Nex-N2.5-Pro | Nex AGI (startup) | Sep 8 | 41,000 | 396.8B | MoE (10 of 512 experts) | 250.8GB | Good at long-running real-world agent work |
| 14 | ZDTaichu5.0-9B | TaichuAI (China-research-institute-affiliated) | Sep 4 | 22,000 | 9.8B | Dense | 5.6GB | Good at spatial understanding and tool use with images and video |
| 15 | Nex-N2.5-Max | Nex AGI (startup) | Sep 7 | 13,000 | 1,600.8B | MoE (6 of 384 experts) | 897.8GB (mix of 4-bit and 8-bit) | Good at long-running real-world agent work |
| 16 | AliceAI-Foundation-80B-A3B-Base | Yandex (major Russian search company) | Sep 12 | 7,743 | 81.3B | MoE (10 of 512 experts) | 48.3GB | A base model especially strong on Russian-language factual knowledge questions |
| 17 | LensVLM-9B | Apple | Sep 21 | 6,207 | 9.4B | Dense | 5.8GB | Can unpack just the needed part of a compressed document image to read it |
The download counts are for the last 30 days, not a cumulative total since release — newer models naturally show lower numbers. Sizes are based on Q4_K_M; where Q4_K_M isn’t distributed, I noted the closest roughly-4-bit format in parentheses. UD-Q4_K_XL and MXFP4 are different roughly-4-bit shrinking methods from Q4_K_M. Ternary-Bonsai-2-27B was distributed from the start as a GGUF in its own format, shrunk to 1.75-bit — less than half of 4-bit. “—" means no GGUF exists.
Which Sizes Actually Fit on Your Own PC?
Whether a model runs on your own PC comes down to file size first — it won’t run unless the file fits in your graphics card’s memory (VRAM) or your PC’s main memory. In practice you need a few extra GB beyond the file size, for the space that holds the conversation so far (the KV cache). Macs and mini PCs built with shared memory can use their main memory for graphics too, so they can handle bigger models than a typical PC with the same amount of memory.
| Size range | Use | Model | 4-bit size |
|---|---|---|---|
| 8GB or less | Small, agent | MiniCPM5-2B | 1.6GB |
| 8GB or less | Reads images too, agent | ZDTaichu5.0-9B | 5.6GB |
| 8GB or less | Agent, programming | MiMo-V2.6-Distill-Qwen-9B | 5.8GB |
| 8GB or less | Reads images too | LensVLM-9B | 5.8GB |
| 8GB or less | Reasoning, small | Ternary-Bonsai-2-27B | 5.9GB (1.75-bit) |
| 8-24GB | Programming, agent | Qwen3.8-27B | 16.5GB |
| 8-24GB | Writing | Hemmingway-1 | 17.4GB |
| 8-24GB | Agent, programming | Xing4.0-29B-A4B | 19.0GB |
| 8-24GB | Agent, reads images too | Nex-N2.5-mini | 22.3GB |
| 24-100GB | Base model | AliceAI-Foundation-80B-A3B-Base | 48.3GB |
| Over 100GB | Agent, reads images too | Qwen3.8-Flash-Next | 111.3GB (UD-Q4_K_XL) |
| Over 100GB | Agent, reads images too | MiMo-V2.6-Flash-RL | 167.4GB (MXFP4) |
| Over 100GB | Agent, reads images too | Nex-N2.5-Pro | 250.8GB |
| Over 100GB | Agent, reads images too | DeepSeek-V4.1-Flash | 444.7GB |
| Over 100GB | Agent, reads images too | MiMo-V2.6-Pro-RL | 528.9GB (MXFP4) |
| Over 100GB | Agent | Nex-N2.5-Max | 897.8GB (mix of 4-bit and 8-bit) |
| No GGUF | Small | Edge0-35B-A3B-preview | — |
Of the 17 models, 9 fit within 24GB of VRAM, and 6 exceed 100GB. Even shrunk to 4-bit, DeepSeek-V4.1-Flash comes to about 445GB, and Nex-N2.5-Max to about 898GB.
What Is MoE, and Which Models Use It?
MoE (Mixture of Experts) splits the inside of a model into many “experts" and uses only some of them each time it produces a word. Xing4.0-29B-A4B, for example, has 64 experts and uses only 4 of them at a time. Even though the whole model is large, running only a small slice of it each time means it can write answers faster than a dense model of the same overall size. That said, the experts you’re not using still have to sit in memory — the memory you need is determined by the model’s total size, not the active slice.
10 of the 17 models here use MoE. The “A3B" or “A4B" in a model’s name means roughly 3 billion or 4 billion parameters are active at a time.
What Is Each Model Good At?
| Use | VRAM needed | Model |
|---|---|---|
| Agent | 8GB or less | MiniCPM5-2B, ZDTaichu5.0-9B, MiMo-V2.6-Distill-Qwen-9B |
| Agent | 8-24GB | Qwen3.8-27B, Xing4.0-29B-A4B, Nex-N2.5-mini |
| Agent | Over 100GB | Qwen3.8-Flash-Next, MiMo-V2.6-Flash-RL, Nex-N2.5-Pro, DeepSeek-V4.1-Flash, MiMo-V2.6-Pro-RL, Nex-N2.5-Max |
| Programming | 8GB or less | MiMo-V2.6-Distill-Qwen-9B |
| Programming | 8-24GB | Qwen3.8-27B, Xing4.0-29B-A4B |
| Reads images too | 8GB or less | ZDTaichu5.0-9B, LensVLM-9B |
| Reads images too | 8-24GB | Nex-N2.5-mini |
| Reads images too | Over 100GB | Qwen3.8-Flash-Next, MiMo-V2.6-Flash-RL, Nex-N2.5-Pro, DeepSeek-V4.1-Flash, MiMo-V2.6-Pro-RL |
| Writing | 8-24GB | Hemmingway-1 |
| Reasoning | 8GB or less | Ternary-Bonsai-2-27B |
| Small | 8GB or less | MiniCPM5-2B, Ternary-Bonsai-2-27B |
| Small | No GGUF | Edge0-35B-A3B-preview |
| Base model | 24-100GB | AliceAI-Foundation-80B-A3B-Base |
Models that fit more than one use appear in more than one row. These uses are broken out from the full model-card description, so they include uses that didn’t fit into the one-line summary in the earlier table. “Agent" means the AI judges for itself and carries out multi-step tasks using tools such as web search or running code. “Base model" means a model that isn’t practical to use as a chat partner as-is — it’s published for other developers to further train on top of.
Who Released Each Model?
I sorted the 17 models by who released them and what they were built on.
| Publisher | Category | Model | Based on |
|---|---|---|---|
| Apple | Major company / research institute | LensVLM-9B | Qwen3.5-9B |
| DeepSeek | Major company / research institute | DeepSeek-V4.1-Flash | — |
| OpenBMB (Tsinghua-affiliated) | Major company / research institute | MiniCPM5-2B | — |
| Qwen (Alibaba) | Major company / research institute | Qwen3.8-27B | — |
| Qwen (Alibaba) | Major company / research institute | Qwen3.8-Flash-Next | — |
| Xiaomi | Major company / research institute | MiMo-V2.6-Distill-Qwen-9B | Qwen3.5-9B |
| Xiaomi | Major company / research institute | MiMo-V2.6-Flash-RL | — |
| Xiaomi | Major company / research institute | MiMo-V2.6-Pro-RL | — |
| Yandex (major Russian search company) | Major company / research institute | AliceAI-Foundation-80B-A3B-Base | — |
| Altworld | Startup | Hemmingway-1 | Qwen3.8-27B |
| Edge0 | Startup | Edge0-35B-A3B-preview | Qwen3.6-35B-A3B |
| Nex AGI | Startup | Nex-N2.5-Max | — |
| Nex AGI | Startup | Nex-N2.5-Pro | Nex-N2 (an earlier version from the same developer) |
| Nex AGI | Startup | Nex-N2.5-mini | Nex-N2 (an earlier version from the same developer) |
| PrismML | Startup | Ternary-Bonsai-2-27B | Qwen3.8-27B |
| TaichuAI (China-research-institute-affiliated) | Startup | ZDTaichu5.0-9B | Qwen3.5-9B |
| XingChen-AGI | Startup | Xing4.0-29B-A4B | — |
Even though the names look new, 6 of the 17 models were built on top of a Qwen model. Models released by major companies, like Apple’s LensVLM-9B and Xiaomi’s MiMo-V2.6-Distill-Qwen-9B, are also built on Qwen3.5-9B underneath. What a model is based on is written on its model card, so for an unfamiliar name, checking what it’s built on first is the quickest way to get a feel for it.
Where Is the Popularity Concentrated?
Adding up the 30-day download counts for all 17 models comes to about 50.99 million. About 76% of that is concentrated in a single model, Qwen3.8-27B.
There’s also a clear split in whether the original file or the GGUF gets downloaded more. I sorted by how many times higher the GGUF downloads are than the original’s (excluding Ternary-Bonsai-2-27B, since it was distributed as GGUF from the start).
| Model | Original downloads | GGUF downloads | GGUF ÷ original | Category |
|---|---|---|---|---|
| Nex-N2.5-mini | 9,507 | 160,000 | 17.1x | GGUF-heavy (5x or more) |
| MiMo-V2.6-Distill-Qwen-9B | 8,839 | 120,000 | 14.1x | GGUF-heavy (5x or more) |
| Hemmingway-1 | 5,904 | 61,000 | 10.3x | GGUF-heavy (5x or more) |
| Qwen3.8-27B | 6.73 million | 32.08 million | 4.8x | Middle |
| Qwen3.8-Flash-Next | 1.23 million | 3.95 million | 3.2x | Middle |
| LensVLM-9B | 1,740 | 4,467 | 2.6x | Middle |
| DeepSeek-V4.1-Flash | 650,000 | 1.18 million | 1.8x | Middle |
| AliceAI-Foundation-80B-A3B-Base | 3,456 | 4,287 | 1.2x | Middle |
| MiMo-V2.6-Flash-RL | 26,000 | 25,000 | 1.0x | Middle |
| ZDTaichu5.0-9B | 12,000 | 10,000 | 0.9x | Middle |
| Xing4.0-29B-A4B | 45,000 | 17,000 | 0.4x | Middle |
| Nex-N2.5-Pro | 33,000 | 8,387 | 0.3x | Middle |
| MiniCPM5-2B | 770,000 | 110,000 | 0.1x | Original-heavy (GGUF under 1/5) |
| MiMo-V2.6-Pro-RL | 75,000 | 1,726 | 0.0x | Original-heavy (GGUF under 1/5) |
| Nex-N2.5-Max | 13,000 | 246 | 0.0x | Original-heavy (GGUF under 1/5) |
| Edge0-35B-A3B-preview | 80,000 | 0 | 0.0x | Original-heavy (GGUF under 1/5) |
Since GGUF is the format for running a model on your own PC, a model with more GGUF downloads suggests more people are using it that way. Conversely, a model with mostly original-file downloads is either aimed at people running it on a server, or simply doesn’t have a GGUF yet.
That said, this is only an inference from download counts. If most of the GGUF downloads come from automated processes — mirroring or test fetches — rather than people, this reading doesn’t hold. Hugging Face doesn’t publish a breakdown of its download counts, so there’s no way to confirm this either way.
How Many Days After Release Did the GGUF Appear?
| Model | First GGUF | GGUF count (as of Sep 28) |
|---|---|---|
| DeepSeek-V4.1-Flash | Same day | 16 |
| Nex-N2.5-mini | Same day | 15 |
| MiMo-V2.6-Distill-Qwen-9B | Same day | 25 |
| Hemmingway-1 | Same day | 19 |
| MiniCPM5-2B | Same day (published by the developer) | 34 |
| Xing4.0-29B-A4B | Same day (published by the developer) | 7 |
| Qwen3.8-27B | 1 day later | 141 |
| Qwen3.8-Flash-Next | 1 day later | 111 |
| MiMo-V2.6-Flash-RL | 1 day later | 13 |
| MiMo-V2.6-Pro-RL | 2 days later | 3 |
| Nex-N2.5-Max | 2 days later | 1 |
| LensVLM-9B | 2 days later | 4 |
| Nex-N2.5-Pro | 4 days later | 4 |
| AliceAI-Foundation-80B-A3B-Base | 10 days later | 2 |
| ZDTaichu5.0-9B | 12 days later | 3 |
| Edge0-35B-A3B-preview | Still none (20 days elapsed) | 0 |
Of the 16 models excluding Ternary-Bonsai-2-27B (distributed as GGUF from the start), 9 had a GGUF out by the day after release. Edge0-35B-A3B-preview, on the other hand, still has none 20 days after release, even with about 80,000 downloads of the original file. According to its model card, Edge0 is distributed in its own format built to run on smartphone-level low memory, and no GGUF is provided. The lack of a GGUF looks like it comes down to how it’s distributed, not low interest.
Even once a GGUF exists, it doesn’t necessarily run right away. DeepSeek-V4.1-Flash had a GGUF out on release day, but it failed to load on llama.cpp (build b10605, as of September 2026) on this blog’s measurement PC (DeepSeek V4.1-Flash Won’t Run — I Tried to Get It Working Locally). Models with a new architecture can be unreadable until llama.cpp adds support for them.
How Fast Are They on the Same PC?
I’m in the middle of measuring how fast each runnable model is on the same PC (GMKtec EVO-X2, 128GB of memory), the same llama.cpp build, and the same format (Q4_K_M). Once the results are ready, I’ll add the read speed and write speed here (tok/s = tokens, the fragments of text a model processes, per second).
What This Article Hasn’t Confirmed
- “Stated strengths" are the developer’s own claims; this blog hasn’t tested them
- How the 17 models were selected depends on Hugging Face’s trending list, and the ranking mechanism behind that list isn’t public
- GGUF counts only include ones registered on Hugging Face against the original model; unregistered GGUFs aren’t counted
- Download counts can’t distinguish between human and automated fetches
Summary — If It’s Your Own PC, Which One Should You Try First?
Of the 17 recently released models, 9 fit comfortably within 24GB and are easy to run on your own PC, while 6 exceed 100GB and are hard to run on a single home PC. It’s easier to keep up with this space if you treat the big models everyone’s talking about and the models you can actually try on your own PC as two separate things.
If you’re trying one on your own PC, the amount of memory you have decides your candidates almost by itself.
| VRAM / memory needed | Model | Use | Note |
|---|---|---|---|
| About 8GB | MiniCPM5-2B | Small, agent | |
| About 8GB | ZDTaichu5.0-9B | Reads images too, agent | |
| About 8GB | MiMo-V2.6-Distill-Qwen-9B | Agent, programming | |
| About 8GB | LensVLM-9B | Reads images too | |
| About 8GB | Ternary-Bonsai-2-27B | Reasoning, small | Uses its own 1.75-bit format |
| 24GB | Qwen3.8-27B | Programming, agent | |
| 24GB | Hemmingway-1 | Writing | |
| 24GB | Xing4.0-29B-A4B | Agent, programming | |
| 24GB | Nex-N2.5-mini | Agent, reads images too | At 22.3GB, it likely won’t fit in 24GB |
| 64-128GB (main memory) | AliceAI-Foundation-80B-A3B-Base | Base model | A base model, not practical to use as a chat partner as-is |
If you’re not sure which to try first, Qwen3.8-27B is a safe bet — it has the most downloads and had a GGUF out the next day. If you’re short on VRAM, pick something matching your use from the “8GB or less" group in the same list above. When you spot a new model’s name, checking three things — “how many GB at Q4_K_M," “is it MoE," and “is there a GGUF" — will quickly tell you whether it’s relevant to you.










Discussion
New Comments
No comments yet. Be the first one!