How Much Better Is Ornith 1.5? — I Measured 1.0 and 1.5 With Three Yardsticks
Ornith 1.5 was released on August 18. Considering that Ornith 1.0 had only come out in June, that is a very short gap. How different are these two models, really?
The new version is said to have improved on every item, yet the file sizes are almost the same. I ran two pairs, 9B and 35B, to see whether the difference actually shows up when measured.
Measured as of September 2026.
- 1. What Is Ornith?
- 2. What I Compared — Models and Three Yardsticks
- 3. Speed Results (Measured)
- 4. What Do the Publisher’s Benchmark Scores Say?
- 5. Results of My Own Test Set
- 6. Can You Trust Its Confidence? (Measured)
- 7. So, Did It Get Smarter?
- 8. What I Haven’t Confirmed in This Article
- 9. Summary — If You’re Installing Now, 1.5 Looks Fine
- 10. 1.0 or 1.5 — Which Should You Install?
What Is Ornith?
Ornith is an open-source LLM published by a company called DeepReinforce. Version 1.0 was released on June 21, 2026, and 1.5 on August 18. It is released under the MIT license, with no regional restrictions, and commercial use is allowed.
The publisher highlights three features.
| Purpose | Built for coding “agents" — meant to get work done by operating a terminal and calling tools |
| Base | Not built from scratch, but further trained on top of Gemma 4 and Qwen 3.5 |
| How it’s made | It is trained not only on answers but on the procedure for reaching the answer itself |
The third is the selling point of the series. Normally, people prepare the problems and the framework for solving them, and the model solves them. Ornith is described as having the model come up with that framework too, and improving both together. Version 1.5 reportedly goes a step further, having the model create the problems themselves.
Version 1.0 came in four sizes: 9B, 31B, 35B, and 397B. In 1.5, 31B was dropped, leaving three: 9B, 35B, and 397B. The 9B compared here is a dense model; the 35B is MoE.
The 1.5 9B also comes in a version shrunk down to 1.5GB for smartphones.
According to Hugging Face, 1.0-9B has 1.01 million downloads and 1.5-35B has 360,000 (as of September 23, 2026). About a month after 1.5 came out, the old version is still being downloaded more.
What I Compared — Models and Three Yardsticks
I compared two pairs: 9B and 35B.
| Model | Released | Architecture | Quantization |
|---|---|---|---|
| Ornith-1.0-9B | June 21, 2026 | Dense | Q4_K_M |
| Ornith-1.5-9B | August 18, 2026 | Dense | Q4_K_M |
| Ornith-1.0-35B | June 2026 | MoE | Q4_K_M |
| Ornith-1.5-35B | August 2026 | MoE | Q4_K_M |
I prepared three yardsticks.
| ① Speed | How fast it reads and how fast it writes. Measured by me |
| ② Smarts (43 questions) | Code generation, spec traps, how well it avoids saying things it can’t know, and so on |
| ③ How well its confidence holds up (20 questions) | When it says “80%," is it really right 80% of the time? |
The conditions for ① were as follows.
| Machine | GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB unified memory) |
| GPU | Pinned to the integrated Radeon 8060S |
| Runtime | llama.cpp (build b10941), Vulkan |
| Metrics | pp512 (speed of reading 512 tokens) / tg128 (speed of writing 128 tokens) [tok/s] |
| Runs | 5 runs per condition. Sorted, the top and bottom outliers dropped, median of the middle 3 |
Speed Results (Measured)
| Model | Read speed pp512 [tok/s] |
Write speed tg128 [tok/s] |
|---|---|---|
| Ornith-1.0-9B | 1,078.0 | 38.3 |
| Ornith-1.5-9B | 997.3 | 37.9 |
| Ornith-1.0-35B | 1,146.5 | 74.0 |
| Ornith-1.5-35B | 1,122.8 | 73.2 |
tok/s means tokens per second — how fast the model reads input and writes output. The new version is slightly slower. Write speed dropped by about 1%, and read speed on the 9B dropped by 7.5%.
9B and 35B show the same trend. At the very least, you cannot say “the new version got faster."
Why did it get slower? I looked inside, and 1.5 has one extra layer.
| Model | Layers [layers] | Parameters [count] |
|---|---|---|
| Ornith-1.0-9B | 32 | 8.95 billion |
| Ornith-1.5-9B | 33 | 9.20 billion |
| Ornith-1.0-35B | 40 | 34.7 billion |
| Ornith-1.5-35B | 41 | 35.5 billion |
Think of a layer as one stage in the process that handles text. Incoming text passes through stage 1, stage 2, and so on, and the answer comes out at the end — much like an assembly line in a factory. More stages means it takes longer to get through.
In other words, 1.5 has been slightly rebuilt inside. It is not just extra training; the structure itself has been changed.
I also looked into what the extra layer is. It is a layer for predicting several words ahead at once — something that is there to make the model faster. But the log from running it on my machine showed this:
model has unused tensor blk.32.nextn.eh_proj.weight -- ignoring |
It means “this part isn’t used, so it’s being ignored." llama.cpp loads this layer but does not use it.
This is where the speed story connects. The parameters grew by the size of a layer that goes unused, so it runs that much slower. That seems to explain the roughly 1% gap. Run it with a different tool and this layer might kick in and make it faster instead, but I have not tried that.
What Do the Publisher’s Benchmark Scores Say?
The publisher’s model card includes a table putting 1.0-9B and 1.5-9B side by side. The new version is ahead on all 15 items.
Just listing the item names doesn’t tell you much, though. These are tests that are widely used to measure how smart an AI is, and each one measures something different.
| Test | What it measures | Type |
|---|---|---|
| BrowseComp | Has the model browse the web on its own to research something. Can it track down information that is hard to find? | Task |
| Toolathlon-Verified | Against real software, can it finish a long job while using tools many times? | Task |
| NL2Repo | Given only a spec, build a whole program from scratch. Starting from an empty workspace, it decides the structure, gathers the needed parts, and takes it all the way to something that can actually be installed. 104 problems | Task |
| MCP-Atlas | Given 36 real servers and 220 tools, it has to choose which ones to use. 1,000 problems | Task |
| SWE-bench Pro | A harder version, aimed at large and complex programs, of the test where the model fixes bugs found on GitHub | Task |
| GPQA Diamond | 198 graduate-level science questions. So hard that experts score 65%, and non-experts score 34% even with web access | Knowledge |
| HLE (no tools) | True to its name, “Humanity’s Last Exam": 2,500 of the hardest questions from more than 100 fields. Even top AIs struggle to score | Knowledge |
| SWE-bench Verified | Reads real bug reports from GitHub and writes the fix. The version whose problems and tests were checked by hand | Task |
As the right-hand column shows, only two, GPQA Diamond and HLE, are “knowledge" tests. The other six look at whether the model can finish work using tools. The former is about knowledge and reasoning; the latter is about how well you can hand it a job. This table is in the same order as the score table below.
A note on NL2Repo. According to the paper that created this test, even the strongest AI passes less than 40% of the tests, and almost never builds a whole program correctly. The failures it lists include stopping early on its own, losing track of the overall plan, and breaking the connections between files. Against that, 1.5’s score of 32.4 is a respectable showing.
| Item | 1.0-9B [score] |
1.5-9B [score] |
Difference [points] |
|---|---|---|---|
| BrowseComp | 44.8 | 56.4 | +11.6 |
| Toolathlon-Verified | 33.4 | 41.2 | +7.8 |
| NL2Repo | 27.2 | 32.4 | +5.2 |
| MCP-Atlas | 49.4 | 54.2 | +4.8 |
| SWE-bench Pro | 42.9 | 47.5 | +4.6 |
| GPQA Diamond | 82.5 | 86.4 | +3.9 |
| HLE (no tools) | 16.8 | 20.2 | +3.4 |
| SWE-bench Verified | 69.4 | 70.6 | +1.2 |
The scores are copied as-is from the publisher. No unit is given, but it reads as “higher is better."
Looking at the table again with this grouping in mind, a clear bias appears. The big gains are BrowseComp, Toolathlon, and MCP-Atlas — all on the “finish the job with tools" side. Meanwhile SWE-bench Verified, the standard code-fixing test, barely moved at +1.2. It got “smarter," but what grew was research and planning.
However, all of these are self-reported by the publisher, and I found no report of a third party reproducing the same results. That is exactly why I wanted to check on my own machine, and it is the reason for this article.
Results of My Own Test Set
Next, I ran my own test set, prepared to measure smarts. It consists of 7 code-generation questions, 6 spec traps, 18 questions it cannot answer (to see whether it answers anyway), and 12 questions it can answer (to see whether it answers them properly).
| Metric | 1.0-9B | 1.5-9B | 1.0-35B | 1.5-35B |
|---|---|---|---|---|
| Code generation [correct / questions] | 7/7 | 7/7 | 7/7 | 7/7 |
| Spec traps [correct / questions] | 6/6 | 6/6 | 6/6 | 6/6 |
| Hit rate on answerable questions [ratio, 1.0 = all correct] | 1.0 | 1.0 | 1.0 | 1.0 |
| Code quality [points / 100] | 95.57 | 96.79 | 95.37 | 95.96 |
| Trap quality [points / 100] | 95.63 | 98.48 | 94.63 | 97.71 |
| Wrong-answer rate [ratio, closer to 0 is better] | 0.188 | 0.214 | 0.235 | 0.154 |
All four models got a perfect score on the number of correct answers. Old and new did not differ by a single question.
The only differences were in the quality scores (+1 to 3 points) and the wrong-answer rate. And the wrong-answer rate is odd: it got worse on the 9B and better on the 35B. Same new version, opposite results depending on size.
The wrong-answer rate is a ratio, so 0.188 means “got 18.8% of the judgeable answers wrong." On top of that, it comes from only 18 questions; computing the confidence interval gives 0.076–0.476 for 1.5-9B and 0.043–0.422 for 1.5-35B. Both overlap with the 1.0 intervals. The gap is not large enough to call either one better.
Can You Trust Its Confidence? (Measured)
This is the third yardstick. Let me start by explaining the word “confidence."
When you have an AI answer something, you can also have it output how sure it is of that answer. That degree of sureness is called confidence — a number like “I think this answer is about 80% likely to be right."
Here, there is something I need to say up front.
The confidence in this test is not what the AI said when asked “how sure are you, in percent?" When an AI writes text, it picks the next word based on probabilities — “yes" at 0.8, “no" at 0.2, and so on. Reading that probability off directly is what I mean by confidence here.
So this is the probability of choosing the next word, not the probability that the answer is correct. It is a real computed number, but keep in mind that it does not guarantee accuracy.
Usable confidence is handy. You can route answers: pass anything at 90% or above automatically, and send 50% to a person. But that only works if the number can be trusted. If answers labeled 80% are actually right only 30% of the time, it is useless as a routing threshold.
So I had the models solve 20 questions with known answers. Each shows one line of a server log and asks the model to choose one of three: “critical," “watch," or “ignore." It outputs its confidence along with the answer.
There are two things to look at. The first is simply how many it got right.
| Model | Correct [correct / 20 questions] |
|---|---|
| Ornith-1.0-9B | 15/20 (75%) |
| Ornith-1.5-9B | 17/20 (85%) |
| Ornith-1.0-35B | 19/20 (95%) |
| Ornith-1.5-35B | 16/20 (80%) |
For this number, simply higher is better. With 20 questions, one question moves it by 5%. Comparing old and new, the 9B improved from 15→17, while the 35B dropped from 19→16. Again, opposite results depending on size.
The second is whether the confidence can be trusted. Here higher is not better. What matters is whether the stated number matches reality.
I sorted the 20 questions by confidence level and counted how often each group was actually right.
| Model | Confidence band | Questions in that band [questions] | Confidence the AI stated (average) [%] | Share actually correct [%] |
|---|---|---|---|---|
| 1.0-9B | 50–70% | 12 | 57.1% | 91.7% |
| 1.5-9B | 50–70% | 13 | 61.0% | 92.3% |
| 1.0-35B | Under 50% | 18 | 44.4% | 94.4% |
| 1.5-35B | Under 50% | 18 | 42.6% | 77.8% |
How to read the table: the first row means of the 12 questions where it said “I’m about 57% sure," 11 (91.7%) were actually correct. Ideally the two numbers are close — saying 57% and being right 57% of the time is what “trustworthy" looks like.
In the results, all four were right far more often than their stated numbers. The 1.0-35B in particular said “less than half sure" yet got 17 of 18 right.
Generative AI is often said to be overconfident, but these four were the opposite. They undersell themselves and are actually right a lot. If you set a rule like “discard anything under 50% confidence," you would throw away a lot of correct answers.
So, Did It Get Smarter?
Here are the results so far, side by side.
| What was measured | 9B | 35B |
|---|---|---|
| The publisher’s 15 items | Better on every item | (no old-vs-new table) |
| Speed | Slightly slower | Slightly slower |
| Correct answers, 43 questions | No change (perfect score) | No change (perfect score) |
| Wrong-answer rate, 43 questions | Worse | Better |
| Correct answers, 20 questions | Better | Worse |
Some things got better and some got worse. And items that improved on the 9B got worse on the 35B, and vice versa.
This pattern is most reasonably read as falling within the margin of error. With 20 questions, one question moves the result by 5%. The 9B’s “+10%" is just two questions. The wrong-answer rate also comes from 18 questions, and the confidence intervals for old and new overlap. It is too small to call either better or worse.
The same goes for speed. The write-speed gap is about 1% — not something you would notice in use.
In short, as far as these measurements go, I saw no change in either smarts or speed.
That does not mean “1.5 is the same as 1.0," though. There may be progress in areas my tests couldn’t measure. What the publisher says improved a lot is the ability to browse the web for research and to finish long jobs using tools many times. Neither my 43 questions nor my 20 questions cover those abilities.
The publisher uses tasks on the scale of hundreds or thousands of problems. Trying to see the same gap with only a few dozen questions is asking too much.
What I Haven’t Confirmed in This Article
- The publisher’s 15 items are quoted as self-reported; I have not reproduced them
- At the scale of 43 and 20 questions, a few percent of difference can’t be told apart from chance. I did not run a significance test
- The wrong-answer rate comes from 18 questions, and the confidence intervals for old and new overlap. It is not grounds for ranking them
- For the new 35B, some of the 20 questions could not be judged
- All comparisons use a single quantization level, Q4_K_M. Other levels may give different results
- The measurements are from one integrated GPU. A different machine will give different numbers
- The way I measured confidence is my own setup and may differ from how the publisher intends it to be used
Summary — If You’re Installing Now, 1.5 Looks Fine
- 1.0 was already smart enough. All four models got a perfect score on the 43 smarts questions, with not a single question of difference between old and new
- Speed was only slightly slower, with a write-speed gap of about 1%. Not something you would notice in use
- → My tests could not bring out a difference. The questions were too easy, so it ended up comparing perfect scores with perfect scores
- On the other hand, the publisher reports 1.5 ahead on all 15 items. The gains are in web research and tool use, which my tests did not include
- I did confirm the insides have changed. One more layer, more parameters. It is not just extra training; the structure was changed
- The confidence behavior differed too. On the same questions, the way probability is spread across the next word has changed
Putting it all together, if you are installing now, 1.5 looks fine.
My tests showing no difference is not a case of “1.5 is not good." It is a case of my only having questions easy enough for 1.0 to already get a perfect score. There is no evidence that it got worse, and speed is the same in practice.
On top of that, the publisher says it got smarter, and the insides really have been rebuilt. The added layer is meant to be a “make it faster" mechanism, and my tool doesn’t use it yet. As support improves, 1.5 could pull ahead.
That said, I found no reason to go out of your way to replace a setup already running 1.0. The results don’t show that swapping will reliably make things faster. If you’re installing fresh, go with 1.5; if it’s already running, leave it — that looks like the conclusion this time.
When you hear “a new version is out," it’s hard to know whether to switch. But as happened here, if your own questions are too easy, you won’t see a difference even if there is one. Once perfect scores line up, that yardstick has stopped working.
1.0 or 1.5 — Which Should You Install?
Based on the results so far, here is how to choose.
| If you want to… | Pick | Why |
|---|---|---|
| Install it fresh | 1.5 | I found nothing that got worse, and the publisher reports improvements. Choosing the newer one seems to cost you nothing |
| You already run 1.0 | Stay on 1.0 | No result shows that switching makes things better. No reason to spend the effort |
| Have it write code | Either | Both got perfect scores on my 43 questions, with no difference. In the publisher’s numbers too, the standard SWE-bench Verified is nearly the same at +1.2 |
| Hand it hard code fixes | Lean 1.5 | In the publisher’s numbers, the harder SWE-bench Pro shows a gap of 42.9→47.5 (self-reported) |
| Have it research on the web or use tools | Lean 1.5 | This is the area the publisher reports big gains in (self-reported; not verified by me) |
| Run it even a little faster | 1.0 | 1.5 is about 1% slower because of its extra layer. But it’s not a difference you’d notice |
| Save storage space | 1.0 | 5.24GiB vs 5.38GiB for 9B; 19.71GiB vs 20.22GiB for 35B. The gap is about 0.5GiB |
Honestly, within what I measured on my own machine, I couldn’t tell 1.0 and 1.5 apart. Both got a perfect score on the 43 smarts questions, and the speed gap is about 1%. Whichever you choose, you’re unlikely to notice a difference in use.
So think of the table above as “which way to lean when you’re unsure" and nothing more. My measurements don’t show a gap big enough to recommend either one strongly.
Choosing a size is a separate question from choosing a version. The 35B came out nearly twice as fast as the 9B (74.0 vs 38.3 tok/s). That’s because the 35B is a MoE design that uses only part of itself for each token it writes. If you can set aside 20GB of storage, the 35B is the more comfortable choice.
If you’re unsure, the quickest way is probably to install the 1.5 9B first and try it on your own work. As this test showed, some things you only find out by measuring on your own use case.
Speed could be measured. The difference in smarts could not. But you only learn what you can’t measure by measuring.