How Much Better Is Ornith 1.5? — I Measured 1.0 and 1.5 With Three Yardsticks

This page contains advertising (affiliate links). See our Privacy Policy for details.

Ornith 1.5 was released on August 18. Considering that Ornith 1.0 had only come out in June, that is a very short gap. How different are these two models, really?

The new version is said to have improved on every item, yet the file sizes are almost the same. I ran two pairs, 9B and 35B, to see whether the difference actually shows up when measured.

Measured as of September 2026.

Sponsored

What Is Ornith?

Ornith is an open-source LLM published by a company called DeepReinforce. Version 1.0 was released on June 21, 2026, and 1.5 on August 18. It is released under the MIT license, with no regional restrictions, and commercial use is allowed.

The publisher highlights three features.

PurposeBuilt for coding “agents" — meant to get work done by operating a terminal and calling tools
BaseNot built from scratch, but further trained on top of Gemma 4 and Qwen 3.5
How it’s madeIt is trained not only on answers but on the procedure for reaching the answer itself

The third is the selling point of the series. Normally, people prepare the problems and the framework for solving them, and the model solves them. Ornith is described as having the model come up with that framework too, and improving both together. Version 1.5 reportedly goes a step further, having the model create the problems themselves.

Version 1.0 came in four sizes: 9B, 31B, 35B, and 397B. In 1.5, 31B was dropped, leaving three: 9B, 35B, and 397B. The 9B compared here is a dense model; the 35B is MoE.

The 1.5 9B also comes in a version shrunk down to 1.5GB for smartphones.

According to Hugging Face, 1.0-9B has 1.01 million downloads and 1.5-35B has 360,000 (as of September 23, 2026). About a month after 1.5 came out, the old version is still being downloaded more.

Sponsored

What I Compared — Models and Three Yardsticks

I compared two pairs: 9B and 35B.

Model Released Architecture Quantization
Ornith-1.0-9B June 21, 2026 Dense Q4_K_M
Ornith-1.5-9B August 18, 2026 Dense Q4_K_M
Ornith-1.0-35B June 2026 MoE Q4_K_M
Ornith-1.5-35B August 2026 MoE Q4_K_M

I prepared three yardsticks.

① Speed How fast it reads and how fast it writes. Measured by me
② Smarts (43 questions) Code generation, spec traps, how well it avoids saying things it can’t know, and so on
③ How well its confidence holds up (20 questions) When it says “80%," is it really right 80% of the time?

The conditions for ① were as follows.

Machine GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB unified memory)
GPU Pinned to the integrated Radeon 8060S
Runtime llama.cpp (build b10941), Vulkan
Metrics pp512 (speed of reading 512 tokens) / tg128 (speed of writing 128 tokens) [tok/s]
Runs 5 runs per condition. Sorted, the top and bottom outliers dropped, median of the middle 3
Sponsored

Speed Results (Measured)

Model Read speed pp512
[tok/s]
Write speed tg128
[tok/s]
Ornith-1.0-9B 1,078.0 38.3
Ornith-1.5-9B 997.3 37.9
Ornith-1.0-35B 1,146.5 74.0
Ornith-1.5-35B 1,122.8 73.2

tok/s means tokens per second — how fast the model reads input and writes output. The new version is slightly slower. Write speed dropped by about 1%, and read speed on the 9B dropped by 7.5%.

9B and 35B show the same trend. At the very least, you cannot say “the new version got faster."

Why did it get slower? I looked inside, and 1.5 has one extra layer.

ModelLayers [layers]Parameters [count]
Ornith-1.0-9B328.95 billion
Ornith-1.5-9B339.20 billion
Ornith-1.0-35B4034.7 billion
Ornith-1.5-35B4135.5 billion

Think of a layer as one stage in the process that handles text. Incoming text passes through stage 1, stage 2, and so on, and the answer comes out at the end — much like an assembly line in a factory. More stages means it takes longer to get through.

In other words, 1.5 has been slightly rebuilt inside. It is not just extra training; the structure itself has been changed.

I also looked into what the extra layer is. It is a layer for predicting several words ahead at once — something that is there to make the model faster. But the log from running it on my machine showed this:

model has unused tensor blk.32.nextn.eh_proj.weight -- ignoring

It means “this part isn’t used, so it’s being ignored." llama.cpp loads this layer but does not use it.

This is where the speed story connects. The parameters grew by the size of a layer that goes unused, so it runs that much slower. That seems to explain the roughly 1% gap. Run it with a different tool and this layer might kick in and make it faster instead, but I have not tried that.

Sponsored

What Do the Publisher’s Benchmark Scores Say?

The publisher’s model card includes a table putting 1.0-9B and 1.5-9B side by side. The new version is ahead on all 15 items.

Just listing the item names doesn’t tell you much, though. These are tests that are widely used to measure how smart an AI is, and each one measures something different.

TestWhat it measuresType
BrowseCompHas the model browse the web on its own to research something. Can it track down information that is hard to find?Task
Toolathlon-VerifiedAgainst real software, can it finish a long job while using tools many times?Task
NL2RepoGiven only a spec, build a whole program from scratch. Starting from an empty workspace, it decides the structure, gathers the needed parts, and takes it all the way to something that can actually be installed. 104 problemsTask
MCP-AtlasGiven 36 real servers and 220 tools, it has to choose which ones to use. 1,000 problemsTask
SWE-bench ProA harder version, aimed at large and complex programs, of the test where the model fixes bugs found on GitHubTask
GPQA Diamond198 graduate-level science questions. So hard that experts score 65%, and non-experts score 34% even with web accessKnowledge
HLE (no tools)True to its name, “Humanity’s Last Exam": 2,500 of the hardest questions from more than 100 fields. Even top AIs struggle to scoreKnowledge
SWE-bench VerifiedReads real bug reports from GitHub and writes the fix. The version whose problems and tests were checked by handTask

As the right-hand column shows, only two, GPQA Diamond and HLE, are “knowledge" tests. The other six look at whether the model can finish work using tools. The former is about knowledge and reasoning; the latter is about how well you can hand it a job. This table is in the same order as the score table below.

A note on NL2Repo. According to the paper that created this test, even the strongest AI passes less than 40% of the tests, and almost never builds a whole program correctly. The failures it lists include stopping early on its own, losing track of the overall plan, and breaking the connections between files. Against that, 1.5’s score of 32.4 is a respectable showing.

Item 1.0-9B
[score]
1.5-9B
[score]
Difference
[points]
BrowseComp 44.8 56.4 +11.6
Toolathlon-Verified 33.4 41.2 +7.8
NL2Repo 27.2 32.4 +5.2
MCP-Atlas 49.4 54.2 +4.8
SWE-bench Pro 42.9 47.5 +4.6
GPQA Diamond 82.5 86.4 +3.9
HLE (no tools) 16.8 20.2 +3.4
SWE-bench Verified 69.4 70.6 +1.2

The scores are copied as-is from the publisher. No unit is given, but it reads as “higher is better."

Looking at the table again with this grouping in mind, a clear bias appears. The big gains are BrowseComp, Toolathlon, and MCP-Atlas — all on the “finish the job with tools" side. Meanwhile SWE-bench Verified, the standard code-fixing test, barely moved at +1.2. It got “smarter," but what grew was research and planning.

However, all of these are self-reported by the publisher, and I found no report of a third party reproducing the same results. That is exactly why I wanted to check on my own machine, and it is the reason for this article.

Sponsored

Results of My Own Test Set

Next, I ran my own test set, prepared to measure smarts. It consists of 7 code-generation questions, 6 spec traps, 18 questions it cannot answer (to see whether it answers anyway), and 12 questions it can answer (to see whether it answers them properly).

Metric 1.0-9B 1.5-9B 1.0-35B 1.5-35B
Code generation [correct / questions] 7/7 7/7 7/7 7/7
Spec traps [correct / questions] 6/6 6/6 6/6 6/6
Hit rate on answerable questions [ratio, 1.0 = all correct] 1.0 1.0 1.0 1.0
Code quality [points / 100] 95.57 96.79 95.37 95.96
Trap quality [points / 100] 95.63 98.48 94.63 97.71
Wrong-answer rate [ratio, closer to 0 is better] 0.188 0.214 0.235 0.154

All four models got a perfect score on the number of correct answers. Old and new did not differ by a single question.

The only differences were in the quality scores (+1 to 3 points) and the wrong-answer rate. And the wrong-answer rate is odd: it got worse on the 9B and better on the 35B. Same new version, opposite results depending on size.

The wrong-answer rate is a ratio, so 0.188 means “got 18.8% of the judgeable answers wrong." On top of that, it comes from only 18 questions; computing the confidence interval gives 0.076–0.476 for 1.5-9B and 0.043–0.422 for 1.5-35B. Both overlap with the 1.0 intervals. The gap is not large enough to call either one better.

Sponsored

Can You Trust Its Confidence? (Measured)

This is the third yardstick. Let me start by explaining the word “confidence."

When you have an AI answer something, you can also have it output how sure it is of that answer. That degree of sureness is called confidence — a number like “I think this answer is about 80% likely to be right."

Here, there is something I need to say up front.

The confidence in this test is not what the AI said when asked “how sure are you, in percent?" When an AI writes text, it picks the next word based on probabilities — “yes" at 0.8, “no" at 0.2, and so on. Reading that probability off directly is what I mean by confidence here.

So this is the probability of choosing the next word, not the probability that the answer is correct. It is a real computed number, but keep in mind that it does not guarantee accuracy.

Usable confidence is handy. You can route answers: pass anything at 90% or above automatically, and send 50% to a person. But that only works if the number can be trusted. If answers labeled 80% are actually right only 30% of the time, it is useless as a routing threshold.

So I had the models solve 20 questions with known answers. Each shows one line of a server log and asks the model to choose one of three: “critical," “watch," or “ignore." It outputs its confidence along with the answer.

There are two things to look at. The first is simply how many it got right.

ModelCorrect [correct / 20 questions]
Ornith-1.0-9B15/20 (75%)
Ornith-1.5-9B17/20 (85%)
Ornith-1.0-35B19/20 (95%)
Ornith-1.5-35B16/20 (80%)

For this number, simply higher is better. With 20 questions, one question moves it by 5%. Comparing old and new, the 9B improved from 15→17, while the 35B dropped from 19→16. Again, opposite results depending on size.

The second is whether the confidence can be trusted. Here higher is not better. What matters is whether the stated number matches reality.

I sorted the 20 questions by confidence level and counted how often each group was actually right.

ModelConfidence bandQuestions in that band
[questions]
Confidence the AI stated
(average) [%]
Share actually correct
[%]
1.0-9B50–70%1257.1%91.7%
1.5-9B50–70%1361.0%92.3%
1.0-35BUnder 50%1844.4%94.4%
1.5-35BUnder 50%1842.6%77.8%

How to read the table: the first row means of the 12 questions where it said “I’m about 57% sure," 11 (91.7%) were actually correct. Ideally the two numbers are close — saying 57% and being right 57% of the time is what “trustworthy" looks like.

In the results, all four were right far more often than their stated numbers. The 1.0-35B in particular said “less than half sure" yet got 17 of 18 right.

Generative AI is often said to be overconfident, but these four were the opposite. They undersell themselves and are actually right a lot. If you set a rule like “discard anything under 50% confidence," you would throw away a lot of correct answers.

Sponsored

So, Did It Get Smarter?

Here are the results so far, side by side.

What was measured9B35B
The publisher’s 15 itemsBetter on every item(no old-vs-new table)
SpeedSlightly slowerSlightly slower
Correct answers, 43 questionsNo change (perfect score)No change (perfect score)
Wrong-answer rate, 43 questionsWorseBetter
Correct answers, 20 questionsBetterWorse

Some things got better and some got worse. And items that improved on the 9B got worse on the 35B, and vice versa.

This pattern is most reasonably read as falling within the margin of error. With 20 questions, one question moves the result by 5%. The 9B’s “+10%" is just two questions. The wrong-answer rate also comes from 18 questions, and the confidence intervals for old and new overlap. It is too small to call either better or worse.

The same goes for speed. The write-speed gap is about 1% — not something you would notice in use.

In short, as far as these measurements go, I saw no change in either smarts or speed.

That does not mean “1.5 is the same as 1.0," though. There may be progress in areas my tests couldn’t measure. What the publisher says improved a lot is the ability to browse the web for research and to finish long jobs using tools many times. Neither my 43 questions nor my 20 questions cover those abilities.

The publisher uses tasks on the scale of hundreds or thousands of problems. Trying to see the same gap with only a few dozen questions is asking too much.

Sponsored

What I Haven’t Confirmed in This Article

  • The publisher’s 15 items are quoted as self-reported; I have not reproduced them
  • At the scale of 43 and 20 questions, a few percent of difference can’t be told apart from chance. I did not run a significance test
  • The wrong-answer rate comes from 18 questions, and the confidence intervals for old and new overlap. It is not grounds for ranking them
  • For the new 35B, some of the 20 questions could not be judged
  • All comparisons use a single quantization level, Q4_K_M. Other levels may give different results
  • The measurements are from one integrated GPU. A different machine will give different numbers
  • The way I measured confidence is my own setup and may differ from how the publisher intends it to be used

Summary — If You’re Installing Now, 1.5 Looks Fine

  • 1.0 was already smart enough. All four models got a perfect score on the 43 smarts questions, with not a single question of difference between old and new
  • Speed was only slightly slower, with a write-speed gap of about 1%. Not something you would notice in use
  • → My tests could not bring out a difference. The questions were too easy, so it ended up comparing perfect scores with perfect scores
  • On the other hand, the publisher reports 1.5 ahead on all 15 items. The gains are in web research and tool use, which my tests did not include
  • I did confirm the insides have changed. One more layer, more parameters. It is not just extra training; the structure was changed
  • The confidence behavior differed too. On the same questions, the way probability is spread across the next word has changed

Putting it all together, if you are installing now, 1.5 looks fine.

My tests showing no difference is not a case of “1.5 is not good." It is a case of my only having questions easy enough for 1.0 to already get a perfect score. There is no evidence that it got worse, and speed is the same in practice.

On top of that, the publisher says it got smarter, and the insides really have been rebuilt. The added layer is meant to be a “make it faster" mechanism, and my tool doesn’t use it yet. As support improves, 1.5 could pull ahead.

That said, I found no reason to go out of your way to replace a setup already running 1.0. The results don’t show that swapping will reliably make things faster. If you’re installing fresh, go with 1.5; if it’s already running, leave it — that looks like the conclusion this time.

When you hear “a new version is out," it’s hard to know whether to switch. But as happened here, if your own questions are too easy, you won’t see a difference even if there is one. Once perfect scores line up, that yardstick has stopped working.

Sponsored

1.0 or 1.5 — Which Should You Install?

Based on the results so far, here is how to choose.

If you want to…PickWhy
Install it fresh1.5I found nothing that got worse, and the publisher reports improvements. Choosing the newer one seems to cost you nothing
You already run 1.0Stay on 1.0No result shows that switching makes things better. No reason to spend the effort
Have it write codeEitherBoth got perfect scores on my 43 questions, with no difference. In the publisher’s numbers too, the standard SWE-bench Verified is nearly the same at +1.2
Hand it hard code fixesLean 1.5In the publisher’s numbers, the harder SWE-bench Pro shows a gap of 42.9→47.5 (self-reported)
Have it research on the web or use toolsLean 1.5This is the area the publisher reports big gains in (self-reported; not verified by me)
Run it even a little faster1.01.5 is about 1% slower because of its extra layer. But it’s not a difference you’d notice
Save storage space1.05.24GiB vs 5.38GiB for 9B; 19.71GiB vs 20.22GiB for 35B. The gap is about 0.5GiB

Honestly, within what I measured on my own machine, I couldn’t tell 1.0 and 1.5 apart. Both got a perfect score on the 43 smarts questions, and the speed gap is about 1%. Whichever you choose, you’re unlikely to notice a difference in use.

So think of the table above as “which way to lean when you’re unsure" and nothing more. My measurements don’t show a gap big enough to recommend either one strongly.

Choosing a size is a separate question from choosing a version. The 35B came out nearly twice as fast as the 9B (74.0 vs 38.3 tok/s). That’s because the 35B is a MoE design that uses only part of itself for each token it writes. If you can set aside 20GB of storage, the 35B is the more comfortable choice.

If you’re unsure, the quickest way is probably to install the 1.5 9B first and try it on your own work. As this test showed, some things you only find out by measuring on your own use case.

Speed could be measured. The difference in smarts could not. But you only learn what you can’t measure by measuring.

Sponsored