Too Many New AI Models, Can’t Tell Which Is Which? — I Sorted the Latest 17 by Size, Design, and Use

This page contains advertising (affiliate links). See our Privacy Policy for details.

New generative AI models are showing up almost every day. Many of the names are ones I’ve never heard of, and it gets hard to tell which ones actually matter to me. What exactly are these recently released models?

This is what I found as of September 28, 2026.

Sponsored

The Generative AI Models I Looked At

I gathered the models from Hugging Face, the site where AI models get distributed. Almost every model used for local LLMs is published there.

The scope is 17 text models released between August 5 and September 21, 2026. I picked them from the top 60 on Hugging Face’s trending list and the top 30 as of September 13, keeping only the original releases from each developer. Versions remade by other people — shrunk versions, further fine-tuned versions — aren’t counted as separate models; their download counts are folded into the original’s total instead.

For each one, I checked the following:

  • Popularity: the download count Hugging Face shows for the last 30 days — the total of the original file and the GGUF (explained below)
  • Size: parameter count (the number of internal values in the model; 1B = 1 billion of them) and the file size once shrunk to 4-bit (Q4_K_M)
  • Design: whether it’s “dense," using everything every time, or “MoE," using only part of it
  • Stated strengths: what the developer writes on the model’s description page (the model card)
  • GGUF release lag: how many days passed between the original release and the first GGUF appearing

“Stated strengths" summarizes what the developer says — this blog hasn’t verified any of it.

What Is GGUF?

GGUF (pronounced “gee-guff") is a file format for running generative AI models on your own PC.

Most of the original files developers publish come in a format called safetensors. Bigger models split into more files — among the 17 here, that ranged from 1 file up to 644. The internal values are also kept at full, unshrunk precision, so the files are large and mainly assume they’ll be run on dedicated server software.

GGUF is what you get when you convert that original file so llama.cpp — the leading free software for running generative AI on your own PC — can load it. It bundles the model’s values and settings into a single file, and most are shrunk (quantized) down to 4-bit or so before being distributed. Ollama and LM Studio, apps for running generative AI on your own PC, can also load GGUF.

Labels in filenames like “Q4_K_M" or “Q8_0" indicate the kind of shrinking used. The smaller the number, the smaller the file, but answer quality drops a bit each step. This article uses Q4_K_M as the baseline, since it balances size and quality well and is the most commonly used.

GGUF files are often made and published by volunteer converters rather than the developers themselves. Of the 16 models here (excluding Ternary-Bonsai-2-27B, which was distributed as GGUF from the start), only 3 — Xing4.0-29B-A4B, MiniCPM5-2B, and ZDTaichu5.0-9B — had a GGUF released by the developer itself. That’s why the gap between the original release and the GGUF appearing varies by model.

Sponsored

The 17 Models, Ranked by Popularity

I sorted them by download count and listed what kind of model each one is.

No.ModelPublisherReleasedDownloads (30-day)ParametersDesign4-bit sizeStated strengths
1Qwen3.8-27BQwen (Alibaba)Aug 538.81 million27.8BDense16.5GBGood at coding and multi-step agent tasks
2Qwen3.8-Flash-NextQwen (Alibaba)Aug 245.18 million180.0BMoE (10 of 512 experts)111.3GB (UD-Q4_K_XL)An experimental new design built to speed up long-context agent processing
3Ternary-Bonsai-2-27BPrismML (startup)Sep 163.56 million——5.9GB (1.75-bit)Keeps 27B-class reasoning ability well despite being small
4DeepSeek-V4.1-FlashDeepSeekSep 101.83 million763.2BMoE (6 of 384 experts)444.7GBGood at running efficiently and saving memory even with long context
5MiniCPM5-2BOpenBMB (Tsinghua-affiliated)Sep 6880,0002.5BDense1.6GBA lightweight 2B model for resource-constrained devices, with top-tier performance for its class
6Nex-N2.5-miniNex AGI (startup)Sep 8170,00035.1BMoE (8 of 256 experts)22.3GBGood at long-running real-world agent work
7MiMo-V2.6-Distill-Qwen-9BXiaomiSep 21130,0009.4BDense5.8GBGood at coding and tool-using agent work
8Edge0-35B-A3B-previewEdge0 (startup)Sep 880,00034.7BMoE (8 of 256 experts)—A 35B-class model that runs on smartphone-level memory
9MiMo-V2.6-Pro-RLXiaomiSep 2177,0001,024.2BMoE (8 of 384 experts)528.9GB (MXFP4)Good at large-scale reinforcement learning that keeps improving itself
10Hemmingway-1Altworld (startup)Sep 2067,00026.9BDense17.4GBGood at writing everyday emails and messages in natural, human-sounding prose
11Xing4.0-29B-A4BXingChen-AGI (startup)Sep 1662,00031.2BMoE (4 of 64 experts)19.0GBGood at agent work that plans ahead and uses tools
12MiMo-V2.6-Flash-RLXiaomiSep 2151,000310.8BMoE (8 of 256 experts)167.4GB (MXFP4)A reinforcement-learning model that balances lightness and performance
13Nex-N2.5-ProNex AGI (startup)Sep 841,000396.8BMoE (10 of 512 experts)250.8GBGood at long-running real-world agent work
14ZDTaichu5.0-9BTaichuAI (China-research-institute-affiliated)Sep 422,0009.8BDense5.6GBGood at spatial understanding and tool use with images and video
15Nex-N2.5-MaxNex AGI (startup)Sep 713,0001,600.8BMoE (6 of 384 experts)897.8GB (mix of 4-bit and 8-bit)Good at long-running real-world agent work
16AliceAI-Foundation-80B-A3B-BaseYandex (major Russian search company)Sep 127,74381.3BMoE (10 of 512 experts)48.3GBA base model especially strong on Russian-language factual knowledge questions
17LensVLM-9BAppleSep 216,2079.4BDense5.8GBCan unpack just the needed part of a compressed document image to read it

The download counts are for the last 30 days, not a cumulative total since release — newer models naturally show lower numbers. Sizes are based on Q4_K_M; where Q4_K_M isn’t distributed, I noted the closest roughly-4-bit format in parentheses. UD-Q4_K_XL and MXFP4 are different roughly-4-bit shrinking methods from Q4_K_M. Ternary-Bonsai-2-27B was distributed from the start as a GGUF in its own format, shrunk to 1.75-bit — less than half of 4-bit. “—" means no GGUF exists.

Sponsored

Which Sizes Actually Fit on Your Own PC?

Whether a model runs on your own PC comes down to file size first — it won’t run unless the file fits in your graphics card’s memory (VRAM) or your PC’s main memory. In practice you need a few extra GB beyond the file size, for the space that holds the conversation so far (the KV cache). Macs and mini PCs built with shared memory can use their main memory for graphics too, so they can handle bigger models than a typical PC with the same amount of memory.

Size rangeUseModel4-bit size
8GB or lessSmall, agentMiniCPM5-2B1.6GB
8GB or lessReads images too, agentZDTaichu5.0-9B5.6GB
8GB or lessAgent, programmingMiMo-V2.6-Distill-Qwen-9B5.8GB
8GB or lessReads images tooLensVLM-9B5.8GB
8GB or lessReasoning, smallTernary-Bonsai-2-27B5.9GB (1.75-bit)
8-24GBProgramming, agentQwen3.8-27B16.5GB
8-24GBWritingHemmingway-117.4GB
8-24GBAgent, programmingXing4.0-29B-A4B19.0GB
8-24GBAgent, reads images tooNex-N2.5-mini22.3GB
24-100GBBase modelAliceAI-Foundation-80B-A3B-Base48.3GB
Over 100GBAgent, reads images tooQwen3.8-Flash-Next111.3GB (UD-Q4_K_XL)
Over 100GBAgent, reads images tooMiMo-V2.6-Flash-RL167.4GB (MXFP4)
Over 100GBAgent, reads images tooNex-N2.5-Pro250.8GB
Over 100GBAgent, reads images tooDeepSeek-V4.1-Flash444.7GB
Over 100GBAgent, reads images tooMiMo-V2.6-Pro-RL528.9GB (MXFP4)
Over 100GBAgentNex-N2.5-Max897.8GB (mix of 4-bit and 8-bit)
No GGUFSmallEdge0-35B-A3B-preview—

Of the 17 models, 9 fit within 24GB of VRAM, and 6 exceed 100GB. Even shrunk to 4-bit, DeepSeek-V4.1-Flash comes to about 445GB, and Nex-N2.5-Max to about 898GB.

Sponsored

What Is MoE, and Which Models Use It?

MoE (Mixture of Experts) splits the inside of a model into many “experts" and uses only some of them each time it produces a word. Xing4.0-29B-A4B, for example, has 64 experts and uses only 4 of them at a time. Even though the whole model is large, running only a small slice of it each time means it can write answers faster than a dense model of the same overall size. That said, the experts you’re not using still have to sit in memory — the memory you need is determined by the model’s total size, not the active slice.

10 of the 17 models here use MoE. The “A3B" or “A4B" in a model’s name means roughly 3 billion or 4 billion parameters are active at a time.

Sponsored

What Is Each Model Good At?

UseVRAM neededModel
Agent8GB or lessMiniCPM5-2B, ZDTaichu5.0-9B, MiMo-V2.6-Distill-Qwen-9B
Agent8-24GBQwen3.8-27B, Xing4.0-29B-A4B, Nex-N2.5-mini
AgentOver 100GBQwen3.8-Flash-Next, MiMo-V2.6-Flash-RL, Nex-N2.5-Pro, DeepSeek-V4.1-Flash, MiMo-V2.6-Pro-RL, Nex-N2.5-Max
Programming8GB or lessMiMo-V2.6-Distill-Qwen-9B
Programming8-24GBQwen3.8-27B, Xing4.0-29B-A4B
Reads images too8GB or lessZDTaichu5.0-9B, LensVLM-9B
Reads images too8-24GBNex-N2.5-mini
Reads images tooOver 100GBQwen3.8-Flash-Next, MiMo-V2.6-Flash-RL, Nex-N2.5-Pro, DeepSeek-V4.1-Flash, MiMo-V2.6-Pro-RL
Writing8-24GBHemmingway-1
Reasoning8GB or lessTernary-Bonsai-2-27B
Small8GB or lessMiniCPM5-2B, Ternary-Bonsai-2-27B
SmallNo GGUFEdge0-35B-A3B-preview
Base model24-100GBAliceAI-Foundation-80B-A3B-Base

Models that fit more than one use appear in more than one row. These uses are broken out from the full model-card description, so they include uses that didn’t fit into the one-line summary in the earlier table. “Agent" means the AI judges for itself and carries out multi-step tasks using tools such as web search or running code. “Base model" means a model that isn’t practical to use as a chat partner as-is — it’s published for other developers to further train on top of.

Sponsored

Who Released Each Model?

I sorted the 17 models by who released them and what they were built on.

PublisherCategoryModelBased on
AppleMajor company / research instituteLensVLM-9BQwen3.5-9B
DeepSeekMajor company / research instituteDeepSeek-V4.1-Flash—
OpenBMB (Tsinghua-affiliated)Major company / research instituteMiniCPM5-2B—
Qwen (Alibaba)Major company / research instituteQwen3.8-27B—
Qwen (Alibaba)Major company / research instituteQwen3.8-Flash-Next—
XiaomiMajor company / research instituteMiMo-V2.6-Distill-Qwen-9BQwen3.5-9B
XiaomiMajor company / research instituteMiMo-V2.6-Flash-RL—
XiaomiMajor company / research instituteMiMo-V2.6-Pro-RL—
Yandex (major Russian search company)Major company / research instituteAliceAI-Foundation-80B-A3B-Base—
AltworldStartupHemmingway-1Qwen3.8-27B
Edge0StartupEdge0-35B-A3B-previewQwen3.6-35B-A3B
Nex AGIStartupNex-N2.5-Max—
Nex AGIStartupNex-N2.5-ProNex-N2 (an earlier version from the same developer)
Nex AGIStartupNex-N2.5-miniNex-N2 (an earlier version from the same developer)
PrismMLStartupTernary-Bonsai-2-27BQwen3.8-27B
TaichuAI (China-research-institute-affiliated)StartupZDTaichu5.0-9BQwen3.5-9B
XingChen-AGIStartupXing4.0-29B-A4B—

Even though the names look new, 6 of the 17 models were built on top of a Qwen model. Models released by major companies, like Apple’s LensVLM-9B and Xiaomi’s MiMo-V2.6-Distill-Qwen-9B, are also built on Qwen3.5-9B underneath. What a model is based on is written on its model card, so for an unfamiliar name, checking what it’s built on first is the quickest way to get a feel for it.

Sponsored

Where Is the Popularity Concentrated?

Adding up the 30-day download counts for all 17 models comes to about 50.99 million. About 76% of that is concentrated in a single model, Qwen3.8-27B.

There’s also a clear split in whether the original file or the GGUF gets downloaded more. I sorted by how many times higher the GGUF downloads are than the original’s (excluding Ternary-Bonsai-2-27B, since it was distributed as GGUF from the start).

ModelOriginal downloadsGGUF downloadsGGUF ÷ originalCategory
Nex-N2.5-mini9,507160,00017.1xGGUF-heavy (5x or more)
MiMo-V2.6-Distill-Qwen-9B8,839120,00014.1xGGUF-heavy (5x or more)
Hemmingway-15,90461,00010.3xGGUF-heavy (5x or more)
Qwen3.8-27B6.73 million32.08 million4.8xMiddle
Qwen3.8-Flash-Next1.23 million3.95 million3.2xMiddle
LensVLM-9B1,7404,4672.6xMiddle
DeepSeek-V4.1-Flash650,0001.18 million1.8xMiddle
AliceAI-Foundation-80B-A3B-Base3,4564,2871.2xMiddle
MiMo-V2.6-Flash-RL26,00025,0001.0xMiddle
ZDTaichu5.0-9B12,00010,0000.9xMiddle
Xing4.0-29B-A4B45,00017,0000.4xMiddle
Nex-N2.5-Pro33,0008,3870.3xMiddle
MiniCPM5-2B770,000110,0000.1xOriginal-heavy (GGUF under 1/5)
MiMo-V2.6-Pro-RL75,0001,7260.0xOriginal-heavy (GGUF under 1/5)
Nex-N2.5-Max13,0002460.0xOriginal-heavy (GGUF under 1/5)
Edge0-35B-A3B-preview80,00000.0xOriginal-heavy (GGUF under 1/5)

Since GGUF is the format for running a model on your own PC, a model with more GGUF downloads suggests more people are using it that way. Conversely, a model with mostly original-file downloads is either aimed at people running it on a server, or simply doesn’t have a GGUF yet.

That said, this is only an inference from download counts. If most of the GGUF downloads come from automated processes — mirroring or test fetches — rather than people, this reading doesn’t hold. Hugging Face doesn’t publish a breakdown of its download counts, so there’s no way to confirm this either way.

Sponsored

How Many Days After Release Did the GGUF Appear?

ModelFirst GGUFGGUF count (as of Sep 28)
DeepSeek-V4.1-FlashSame day16
Nex-N2.5-miniSame day15
MiMo-V2.6-Distill-Qwen-9BSame day25
Hemmingway-1Same day19
MiniCPM5-2BSame day (published by the developer)34
Xing4.0-29B-A4BSame day (published by the developer)7
Qwen3.8-27B1 day later141
Qwen3.8-Flash-Next1 day later111
MiMo-V2.6-Flash-RL1 day later13
MiMo-V2.6-Pro-RL2 days later3
Nex-N2.5-Max2 days later1
LensVLM-9B2 days later4
Nex-N2.5-Pro4 days later4
AliceAI-Foundation-80B-A3B-Base10 days later2
ZDTaichu5.0-9B12 days later3
Edge0-35B-A3B-previewStill none (20 days elapsed)0

Of the 16 models excluding Ternary-Bonsai-2-27B (distributed as GGUF from the start), 9 had a GGUF out by the day after release. Edge0-35B-A3B-preview, on the other hand, still has none 20 days after release, even with about 80,000 downloads of the original file. According to its model card, Edge0 is distributed in its own format built to run on smartphone-level low memory, and no GGUF is provided. The lack of a GGUF looks like it comes down to how it’s distributed, not low interest.

Even once a GGUF exists, it doesn’t necessarily run right away. DeepSeek-V4.1-Flash had a GGUF out on release day, but it failed to load on llama.cpp (build b10605, as of September 2026) on this blog’s measurement PC (DeepSeek V4.1-Flash Won’t Run — I Tried to Get It Working Locally). Models with a new architecture can be unreadable until llama.cpp adds support for them.

Sponsored

How Fast Are They on the Same PC?

I’m in the middle of measuring how fast each runnable model is on the same PC (GMKtec EVO-X2, 128GB of memory), the same llama.cpp build, and the same format (Q4_K_M). Once the results are ready, I’ll add the read speed and write speed here (tok/s = tokens, the fragments of text a model processes, per second).

Sponsored

What This Article Hasn’t Confirmed

  • “Stated strengths" are the developer’s own claims; this blog hasn’t tested them
  • How the 17 models were selected depends on Hugging Face’s trending list, and the ranking mechanism behind that list isn’t public
  • GGUF counts only include ones registered on Hugging Face against the original model; unregistered GGUFs aren’t counted
  • Download counts can’t distinguish between human and automated fetches

Summary — If It’s Your Own PC, Which One Should You Try First?

Of the 17 recently released models, 9 fit comfortably within 24GB and are easy to run on your own PC, while 6 exceed 100GB and are hard to run on a single home PC. It’s easier to keep up with this space if you treat the big models everyone’s talking about and the models you can actually try on your own PC as two separate things.

If you’re trying one on your own PC, the amount of memory you have decides your candidates almost by itself.

VRAM / memory neededModelUseNote
About 8GBMiniCPM5-2BSmall, agent
About 8GBZDTaichu5.0-9BReads images too, agent
About 8GBMiMo-V2.6-Distill-Qwen-9BAgent, programming
About 8GBLensVLM-9BReads images too
About 8GBTernary-Bonsai-2-27BReasoning, smallUses its own 1.75-bit format
24GBQwen3.8-27BProgramming, agent
24GBHemmingway-1Writing
24GBXing4.0-29B-A4BAgent, programming
24GBNex-N2.5-miniAgent, reads images tooAt 22.3GB, it likely won’t fit in 24GB
64-128GB (main memory)AliceAI-Foundation-80B-A3B-BaseBase modelA base model, not practical to use as a chat partner as-is

If you’re not sure which to try first, Qwen3.8-27B is a safe bet — it has the most downloads and had a GGUF out the next day. If you’re short on VRAM, pick something matching your use from the “8GB or less" group in the same list above. When you spot a new model’s name, checking three things — “how many GB at Q4_K_M," “is it MoE," and “is there a GGUF" — will quickly tell you whether it’s relevant to you.

Related reading

Sponsored