Which Local LLM to Install First in 2026: Picks by VRAM and Use, From Measured Numbers

本ページは広告(アフィリエイトプログラム)を含みます。詳しくはプライバシーポリシーをご覧ください。

8 October 2026

Open Ollama’s model list and you see hundreds of names. Which one should you install first? Over 2026 I have measured more than fifty models on the same machine. This article re-sorts those numbers by three questions — will it fit on my PC, is it fast, is it smart — and narrows down candidates by use.

This reflects measurements as of October 2026.

How to choose: three yardsticks

I think people get stuck choosing models because the yardsticks get mixed together. This article separates them into three.

YardstickWhat it looks atNumbers used in this article
Will it fit?Whether the model file fits in GPU memory (VRAM) or a mini PC’s shared memoryFile size in Q4_K_M format (4-bit quantization: a format that cuts size to about a quarter at a small cost in accuracy) [GiB]
Is it fast?How many tokens per second it can write (a token is roughly one word fragment)Writing speed tg128 [tok/s], reading speed pp512 [tok/s]
Is it smart?Whether it writes short programs correctly and avoids trick questionsAEB L1 (7 coding problems, 6 trick questions, fabrication rate on false premises)

“Smart" here is this blog’s own short test. It does not measure writing quality or taste in prose. I come back to this under “What I Haven’t Confirmed in This Article".

This blog has several related articles. “What runs on a phone or in an app?" (2 October) compares 23 models of 4B parameters or fewer on smartness, speed and size. “What runs on 8GB of VRAM?" (September) lines up 22 models by memory needed. “Gemma 4 or Qwen 3.6?" (June) compares two model families by VRAM. “Choosing a model: five axes" covers the way of thinking. This article is the entry point that links them, and it goes as far as picking a first model for each use. The table of 23 small models is in the 2 October article, so I only pull the top ones here, and for mid-size and larger I list the newer numbers measured on 3 October.

What runs on a phone or in an app? (in Japanese)

Gemma 4 vs Qwen: comparison by VRAM

Choosing a local LLM: five axes

A few short definitions. Q4_K_M is 4-bit quantization, which cuts the size to about a quarter at a small cost in accuracy. tg128 is the speed when writing 128 tokens; pp512 is the speed when reading 512 tokens. MoE (mixture of experts) is a design that holds many specialists and uses only some of them per question, so the file is large but little works at once and it runs fast. Fabrication rate is the share of answers, among those that could be scored out of the 18 false-premise questions, where the model answered with a made-up story instead of pointing out the error (answers that could not be scored are left out of the denominator).

How I measured speed and smartness

I measured speed on the integrated GPU (Radeon 8060S) of a GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB memory). I used the official llama.cpp build b11192, measured five times under the same conditions, dropped the highest and lowest, and took the median of the middle three. The 23 small models were measured on 29 September 2026, the mid-size ones (9-35B) on 3 October. The large models (100B class) include additional measurements from 3 October, but some are values from Ollama or an older build (b10605) in September, noted in the table.

For smartness, I put the same questions to each model once and scored them (RTX 3090, Ollama, temperature 0). The 23 small models are recorded as “how many of 7 problems correct", the 8 large ones as “average out of 100", and because the scoring versions differ, I do not compare small and large scores side by side.

What runs by memory size (VRAM)

Installing one first: fast and smart with 8GB of VRAM or less

On a gaming laptop or an 8GB graphics card, I suggest starting with models whose file is around 3GB. The full list of 23 small models is in the 2 October article, so here I pull only the top four. Of these, the top two got every coding problem (7 of 7) and every trick question (6 of 6) right.

Model (name in Ollama)Size [GiB]CodingTrick questionsWriting speed [tok/s]Reading speed [tok/s]
Qwen3.5-4B (hf.co/unsloth/Qwen3.5-4B-GGUF:Q4_K_M)2.547/76/665.41,987
Nemotron 3 Nano 4B (nemotron-3-nano:4b)2.637/76/669.51,940
MiniCPM5-2B (hf.co/bartowski/MiniCPM5-2B-GGUF:Q4_K_M)1.506/76/6123.83,826
Gemma 4 E2B (hf.co/unsloth/gemma-4-E2B-it-GGUF:Q4_K_M)2.886/76/6116.13,179

At 65 tok/s the text appears faster than most people read. For a first model I suggest Qwen3.5-4B. Its fabrication rate on false-premise questions (answering with a made-up story) was 0.07 (1 of the 15 scorable answers; 3 could not be scored), the lowest of the 23 models tested for smartness. It is only an 18-question test, though, so the confidence interval is wide and the ranking is not definitive. Nemotron 3 Nano 4B got the same number right but had a fabrication rate of 0.45 (5 of the 11 scorable answers; 7 could not be scored), often accepting a false premise and answering anyway. On a separate 96-problem arithmetic test, Nemotron 3 Nano 4B scored 0.479 (46 of 96) with its “thinking" feature (reasoning before answering) switched off and 0.896 (86 of 96) with thinking allowed. The median output with thinking on was 207 tokens. Arithmetic tests calculation alone and is a different test from the coding and trick-question tests above.

If you want something lighter, MiniCPM5-2B is 1.5GB and twice as fast at 124 tok/s, and it only dropped one problem.

Help with coding: the main pick at 16-24GB

Here are the candidates for writing programs, with the smartness I measured and the speed on the EVO-X2’s integrated GPU. Robustness on long instructions is not measured in this test.

Model (name in Ollama)Size [GiB]Coding averageTrick-question averageWriting speed [tok/s]Fits in
Qwen3 Coder 30B (qwen3-coder:30b)17.395.06/6 (correct count)91.124GB VRAM
Nemotron 3 Nano 30B (nemotron-3-nano:30b)22.696.76/6 (correct count)69.024GB VRAM
Ornith 1.5 35B-A3B (hf.co/deepreinforce-ai/Ornith-1.5-35B-A3B-GGUF:Q4_K_M)20.295.36/6 (correct count)75.524GB VRAM
Ornith 1.5 9B (hf.co/deepreinforce-ai/Ornith-1.5-9B-GGUF:Q4_K_M)5.487.46/6 (correct count)38.78GB VRAM
Qwen3-Coder-Next (qwen3-coder-next)48.295.296.4not measured128GB mini PC

Qwen3 Coder 30B calls itself 30B but has a design where only a small part works at once (MoE), so with a 17GB file it ran at 91.1 tok/s, faster than the 4B Qwen3.5-4B (65.4 tok/s). With a short-conversation setting it fits whole in the 24GB VRAM of an RTX 3090. The file does not fit on a 16GB graphics card. It got all 7 coding problems right, with an average of 95.0.

The three 24GB models (Qwen3 Coder 30B, Nemotron 3 Nano 30B, Ornith 1.5 35B-A3B) got every coding and trick question right, with averages in a cluster at 95-97. They reached the level where this test cannot separate them (a ceiling), so I cannot rank the three. On the 96-problem arithmetic test, Nemotron 3 Nano 30B scored 0.594 (57 of 96) with thinking off and 1.000 (96 of 96) with thinking allowed. The median output with thinking on was 312 tokens. The thinking-on result is a perfect-score ceiling, so this test cannot rank them either. If you choose by speed, Qwen3 Coder 30B at 91 tok/s is the one.

If you want to stay within 8GB, Ornith 1.5 9B: 5.4GB, 6/7 on coding, 6/6 on trick questions, average 87.4. Its 39 tok/s writing speed is about 60% of a 4B-class model.

As smart as possible: the ceiling for what fits in 24GB

These are the three models I measured for smartness that fit within 24GB of VRAM.

Model (name in Ollama)Size [GiB]Coding averageTrick-question averageFabrication rateWriting speed [tok/s]
Gemma 4 31B (hf.co/unsloth/gemma-4-31B-it-GGUF:Q4_K_M)17.195.997.20.1511.6
Qwen3.6 27B (hf.co/unsloth/Qwen3.6-27B-GGUF:Q4_K_M)15.795.194.80.1312.6
Qwen3.6 35B-A3B (hf.co/unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M)20.694.794.80.2162.7

All three average around 95, and the gaps are too small to rank by score. Choosing by speed seems best. Gemma 4 31B and Qwen3.6 27B run at about 12 tok/s on the EVO-X2’s integrated GPU, while Qwen3.6 35B-A3B runs at 63 tok/s, a fivefold difference. 35B-A3B is a MoE (many specialists, only some used per question), and only 3B worth of parameters work at once. I have not measured quality in Japanese on this blog.

With a 128GB mini PC: 100B-class models run

On a mini PC with 128GB of shared memory like the EVO-X2, models over 60GB can be loaded. Here are writing speeds on the integrated GPU. The period and software differ by row, so I added notes.

Model (name in Ollama)Size [GiB]Coding averageWriting speed [tok/s]Notes
gpt-oss 120B (hf.co/ggml-org/gpt-oss-120b-GGUF)59.085.855.2MXFP4 format. llama.cpp b11192 (37.2 tok/s on Ollama, September 2026)
Nemotron 3 Super 120B-A12B (nemotron-3-super:120b-a12b)80.087.321.2Measured 5 times on Ollama (September 2026)
GLM-4.5-Air (Q4_K_M)67.8not measured25.6llama.cpp b10605
Qwen3 235B (Q2_K)79.8not measured21.82-bit quantization. llama.cpp b10605
MiniMax M2.7 (minimax-m2.7)101.0not measured29.1llama.cpp b10605

Even 100B-class models run at 21-55 tok/s, well above reading speed. The average smartness scores of gpt-oss 120B and Nemotron 3 Super are 85-87, which on paper is lower than the three 27-31B models that fit in 24GB (around 95). But each model was measured once, and from 9B up the test is near its ceiling, so I do not draw a better-or-worse conclusion from that gap. The test is short coding and trick questions, so summarizing long documents or breadth of knowledge could look different.

Small and fast: what can a 1B-class model do?

Models of 1B or less, aimed at phones and old laptops, sat at the bottom of the 23. Qwen3-0.6B was the fastest at 356 tok/s, but got only 1 of 7 coding problems, and Llama 3.2 1B, gemma-3-1b and MiniCPM5-1B are in the same band. For the detailed table and how to read it, see the 2 October article. I also measured Qwen3.5-0.8B-Japanese-SFT, which was further trained for Japanese, but it scored 0 on coding and trick questions. It measures something different from the Japanese common-sense scores its author publishes, so this result does not show how good its Japanese is.

What I Haven’t Confirmed in This Article

  • I did not measure the quality of Japanese prose. Smartness here is the result of 7 coding problems, 6 trick questions and 18 false-premise questions
  • Smartness is one measurement per model, and the scoring version differs between small and large. From 9B up, the coding and trick questions are near a perfect score (a ceiling), so this test cannot rank the top models against each other. I do not compare small (correct out of 7) and large (average out of 100). The fabrication rate comes from 18 questions (only the scorable answers form the denominator) with a wide confidence interval, so I do not use it as grounds for a ranking
  • “Fits" is a guide for a short-conversation setting. With a long context you need extra memory for the KV cache (the conversation’s memory)
  • Speeds are for the EVO-X2’s integrated GPU. Gemma 4 31B, the two Qwen3.6 models and gpt-oss 120B were measured with Hugging Face’s official GGUF rather than the Ollama distribution (the Ollama builds are in a format llama.cpp cannot read). On NVIDIA graphics cards the numbers will differ (RTX 3090 values are in another article)
  • Some large-model rows mix in an older llama.cpp build (b10605) or Ollama measurements. I noted this in the table
  • Model versions are as of September-October 2026. Ollama tags get updated, so the same name can change in content

Summary: a first model, and a second

With 8GB of VRAM or less, Qwen3.5-4B. For coding, Qwen3 Coder 30B on 24GB, or Ornith 1.5 9B on 8GB. For as smart as possible, any of Gemma 4 31B, Qwen3.6 27B or Qwen3.6 35B-A3B. With a 128GB mini PC, gpt-oss 120B. That is how the measurements sort out.

As a next step, I suggest checking your own VRAM first and installing just one model from the tables above. How to install Ollama and have a first conversation is covered in the next article.

Getting started with Ollama

If you already know the model you want to run and want to choose hardware, start from the GPU list and the local AI guide.

GPU spec list

Local AI guide

Hardware used for the tests

I measured speed on a GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB memory) and smartness on a desktop with an RTX 3090.

GMKtec EVO-X2 (Ryzen AI Max+ 395 / 128GB / 2TB)A8限定クーポン「A82608」で5,000円OFF(2026/10/31まで)/公式オンラインストアのクーポン「KANSYA2026」で2,000円OFF(2026/10/18まで)

¥583,000 Amazon・2026-10-04調べ

公式サイトのみ ¥5,000引きクーポン配布中 クーポンコード A82608(2026-10-31まで)