Slower Despite the Name “Lightning”? I Measured the New and Old Nemotron Side by Side
10 October 2026
A new version of NVIDIA’s Nemotron series is out, and the name includes “Lightning". Is it really faster than the previous version?
These are measurements from September 2026.
How I measured the new and old Nemotron
| Model | Released | Size (Q4_K_M) | Architecture |
|---|---|---|---|
| nemotron-3-nano:30b | December 2025 | 22.60GiB | Hybrid MoE |
| Nemotron-3.5-Lightning-30B-A3B | August 2026 | 23.52GiB | Hybrid MoE (NemotronHForCausalLM) |
They are about eight months apart and differ in size by about 4%. Both carry the A3B name (30B total, 3B active) and have the same model type (architecture), NemotronHForCausalLM. The measurement conditions were as follows.
| Machine | GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB unified memory) |
| GPU | Integrated Radeon 8060S |
| Software | llama.cpp (build b11192, re-measured on 28 September 2026), Vulkan |
| Metrics | pp512 (speed reading 512 tokens) / tg128 (speed writing 128 tokens) [tok/s] |
| Runs | Five per condition. Sorted, dropped the highest and lowest, and took the median of the middle three |
How reading speed and writing speed changed
| Model | Reading speed pp512 [tok/s] | Writing speed tg128 [tok/s] |
|---|---|---|
| nemotron-3-nano:30b (old) | 1,495.4 | 66.3 |
| Nemotron-3.5-Lightning-30B-A3B (new) | 1,537.4 | 57.5 |
Reading speed was about 3% faster on the new one, essentially the same (1,537.4 / 1,495.4 = 1.03). Writing speed, on the other hand, was about 13% slower on the new one (57.5 / 66.3 = 0.87). The spread over the five runs was 65.8-66.9 tok/s for the old model and 55.4-58.4 tok/s for the new one, so the difference is larger than measurement noise.
When I measured on 13 September with an older build (b10605), the reading speed came out identical for the new and old models. I suspected a copying mistake in my records, so on 28 September I re-measured both on the same machine with the newer build and replaced the values above. The direction of the writing speed, with the new one slower, did not change on re-measurement (about 16% slower on 13 September).
Why is it slow despite the name?
Reading speed (prompt processing) is about the same and only writing speed (generation) dropped. Within what I have measured on this blog, writing speed is roughly decided by this division.
Writing speed ~ memory bandwidth / amount read per token
Since I measured on the same machine, the numerator (bandwidth) is common. The difference is in the denominator, the amount read each time a token is written. Both claim a “30B, 3B active" MoE layout, but the actual design of the active part (number of layers, how they are distributed and so on) may have changed so that more is read per token. The model card from the distributor has no detailed description of whether the design changed, so I cannot say for sure.
The name “Lightning" suggests speed, but I could not confirm what benchmark the official side based the name on. At least in this blog’s environment and with these metrics, it came out slower than the previous generation.
What I Haven’t Confirmed in This Article
- I did not measure smartness. This is a comparison of speed and size only. “Lightning" may be a name about smartness or another metric
- Both models were compared at one quantization only (Q4_K_M, UD-Q4_K_M)
- The values are for one integrated GPU. Other machines and GPUs may give different numbers
- I could not confirm official release notes or the benchmark basis for the name. Where the name comes from is a guess
Summary: the name and the contents are separate matters
Compared with its predecessor nemotron-3-nano:30b, Nemotron-3.5-Lightning-30B-A3B has about the same reading speed (the new one is about 3% faster) and about 13% slower writing speed. That is the opposite of the “fast" image the name suggests.
Even for models in the same series with similar size, you cannot know the speed until you measure it. This is the third time this blog has measured a successor in the same series (Ornith 1.0 to 1.5, Nemotron 3 to 3.5). Two came out “about unchanged" and one came out “slower", and so far there is no case where the new version was clearly faster.
Other articles I have researched can be found from this summary page.