What Runs on a Phone or in an App? 23 Small Generative AI Models Compared on Smartness, Speed and Size

October 11, 2026

This page contains advertising (affiliate links). See our Privacy Policy for details.

11 October 2026

On 6 September 2026, OpenBMB in China released a small generative AI model called MiniCPM5-2B. The official description says it was built to run inside devices such as smartphones and PCs. If you run a model on a device, how does it differ from other small models?

These are measurements as of 29 September 2026.

Sponsored

What kind of model is MiniCPM5-2B?

MiniCPM5-2B comes from OpenBMB, a research group linked to Tsinghua University in China. It is the second in the series after MiniCPM5-1B, released in May, with the same design scaled to 2B (2 billion parameters).

The official model card says it was built for “on-device use, local use, and resource-constrained environments". It claims top-level performance among open models of similar size and competes with 4B-class models. The license is Apache-2.0, a condition that makes it easy to embed in apps and distribute or use commercially.

However, the performance claims are the developer’s own evaluation. So I gathered models of similar size and lined them up on smartness, speed and size using this blog’s yardsticks.

Sponsored

What I compared, and how I measured

I compared 23 small models from 0.6B to 7.5B, from 2024’s Llama 3.2 up to 2026’s Qwen3.5, Gemma 4 and MiniCPM5. All are in Q4_K_M, a 4-bit format.

What I looked atHow I measured
SizeFile size in Q4_K_M (how heavy it is to put in an app or have users download)
SmartnessThis blog’s yardstick AEB L1: the number correct out of 7 programming problems and 6 trick questions. On an RTX 3090, with a setting that suppresses variation in answers (temperature 0), one run each
SpeedWriting speed (tok/s) on a GMKtec EVO-X2 in two ways: using the integrated GPU (Radeon 8060S), and running on only 4 CPU threads without the GPU. Both are the median of five runs
LicenseAs shown on the official distribution page

tok/s is the number of tokens (word fragments) written per second.

The CPU-only values are a guide for running on a machine without a GPU. A smartphone CPU differs in design and power conditions, so these are not the speed on a smartphone itself.

Sponsored

Results for the 23 models

ModelSize [GiB]Programming /7Trick questions /6Writing speed, GPU [tok/s]Writing speed, CPU [tok/s]License
Qwen3.5-4B2.547/76/665.417.7Apache-2.0
Nemotron 3 Nano 4B2.637/76/669.518.8NVIDIA custom
MiniCPM5-2B1.506/76/6123.834.9Apache-2.0
Gemma 4 E2B2.886/76/6116.134.0Apache-2.0
Gemma 4 E4B4.626/76/661.017.0Apache-2.0
Qwen3-4B2.326/75/677.521.1Apache-2.0
Gemma 3n E4B4.227/74/654.717.6Gemma custom
Llama 3.2 3B1.875/75/699.226.2Llama 3.2 custom
Gemma 3n E2B2.816/74/688.331.3Gemma custom
Gemma 3 4B2.313/76/675.720.7Gemma custom
Granite 3.3 2B1.445/74/6119.933.7Apache-2.0
LFM2-2.6B1.454/74/6122.634.1LFM custom
Granite 4.1 3B1.954/74/691.425.1Apache-2.0
SmolLM3-3B1.785/73/696.827.1Apache-2.0
LFM2-1.2B0.683/72/6257.071.1LFM custom
Qwen3-1.7B1.032/73/6167.846.0Apache-2.0
Qwen3.5-2B1.184/71/6135.837.4Apache-2.0
Phi-4-mini2.312/73/679.221.3MIT
Qwen3-0.6B0.361/72/6356.2113.3Apache-2.0
Qwen3.5-0.8B0.493/70/6245.978.0Apache-2.0
Gemma 3 1B0.741/71/6200.258.7Gemma custom
Llama 3.2 1B0.742/70/6242.264.4Llama 3.2 custom
MiniCPM5-1B0.641/70/6291.385.6Apache-2.0

The order is by the total number of correct answers on programming and trick questions, highest first.

Sponsored

How are smartness and size related?

At the top are Qwen3.5-4B and Nemotron 3 Nano 4B, both getting all 13 questions right. Both are around 2.5GiB.

MiniCPM5-2B, at 1.50GiB, got 6 programming problems and 6 trick questions right, landing in about the same place as the top of the 4B class. Its size is about 60% of Qwen3.5-4B. The official claim of “top level for 2B, competing with 4B" held up, roughly, on this blog’s yardstick too.

On the other hand, the models of 1B or less (Qwen3-0.6B, Qwen3.5-0.8B, Gemma 3 1B, Llama 3.2 1B, MiniCPM5-1B) got only 1 to 3 programming problems. The smaller you go, the more smartness drops.

Sponsored

How different is the speed?

Writing speed was set mostly by size. On the same machine under the same conditions, the less a model has to read, the faster it writes. On the GPU, MiniCPM5-2B writes at 123.8 tok/s and Qwen3.5-4B at 65.4 tok/s, a difference of about 1.9 times. On CPU alone it is 34.9 and 17.7 tok/s, about the same ratio.

The fastest is Qwen3-0.6B (356.2 tok/s on GPU, 113.3 tok/s on CPU), but its smartness is among the lowest of the 23.

Sponsored

Smartness, speed or size: which to prioritize?

PriorityCandidatesReason
SmartnessQwen3.5-4B (2.54GiB)All 13 correct. It also had the lowest rate of making things up this time. Its writing speed is about half of MiniCPM5-2B’s
BalanceMiniCPM5-2B (1.50GiB) / Gemma 4 E2B (2.88GiB)About the same smartness as the top, with writing speed about 1.8-1.9 times faster. MiniCPM5-2B is also smaller
Speed and sizeQwen3.5-2B (1.18GiB) / LFM2-1.2B (0.68GiB)130-260 tok/s on GPU. But smartness is clearly lower than the top. Suits producing short sentences of a fixed form
Ease of licenseApache-2.0 or MIT modelsIf you embed and distribute in an app, read the terms first for custom licenses (Gemma 3, Llama 3.2, LFM, NVIDIA)

Which to prioritize depends on the device’s memory and the use. As a rule of thumb, if you want short replies back quickly inside an app, go for speed and size; if you use it for summarizing text or simple programming, prioritize smartness.

I should also state when this reading would be wrong. If writing speed is set by size, models of similar size should have similar speeds. Indeed, the three at 1.44-1.50GiB (MiniCPM5-2B, LFM2-2.6B, Granite 3.3 2B) came out together at 119.9-123.8 tok/s on GPU. If a model of the same size were much faster or slower, that would mean something other than size (such as an MoE design) is deciding the speed.

Sponsored

What I Haven’t Confirmed in This Article

  • I did not measure on a real smartphone. The CPU-only values are for 4 threads of a PC CPU
  • Smartness is one run each on this blog’s yardstick (7 programming problems, 6 trick questions). I did not measure naturalness of prose or quality of Japanese
  • I did not measure other ways of shrinking than Q4_K_M (such as formats for Apple or Google devices)

Summary: for on-device use, start with MiniCPM5-2B or Qwen3.5-4B

MiniCPM5-2B, released for devices, reached about the smartness of the 4B class at a small 1.5GiB, at about twice the speed. As the official description says, it is well balanced across size, speed and smartness.

If you cannot decide on a first model: if you have memory to spare and want smartness, Qwen3.5-4B; if you also want lightness and speed, try MiniCPM5-2B. Both are Apache-2.0, a license that is easy to work with when embedding in an app. Models of 1B or less are candidates when you limit the use to short, fixed replies.

For what runs at each memory size, including larger models, see here.

What runs by memory size (VRAM)

Other articles can be found from this summary page.

Local AI guide

Sources consulted

  • OpenBMB: MiniCPM5-2B model card (official) – huggingface.co/openbmb/MiniCPM5-2B
  • Qwen: Qwen3.5-4B (official) – huggingface.co/Qwen/Qwen3.5-4B
  • Google: Gemma 4 E2B (official) – huggingface.co/google/gemma-4-E2B-it
  • NVIDIA: Nemotron 3 Nano 4B (official) – huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-GGUF
  • Liquid AI: LFM2 (official) – huggingface.co/LiquidAI/LFM2-1.2B
Sponsored