Ornith-1.0: self-improving open-source models for agentic coding

github.com

204 points by danboarder 15 hours ago


CharlesW - 14 hours ago

Previously: https://news.ycombinator.com/item?id=48709744

https://swelljoe.com/post/will-it-mythos/: "Poor performer here, only found the one bug that almost every model found, despite its performance on other benchmarks being excellent for its size. […] It also performs poorly in a chat without tools, exhibiting an ehthusiasm for hallucination. I’m currently working on a replication of this with full tool access, including bash/Python, which may allow this model to be competitive."

ricardobayes - 12 hours ago

This is the first Qwen fine-tune that is not immediately rejected by the local LLM community, and in some cases even being recommended. Based on my limited usage, it is good, gives creative solutions to coding problems. I don't expect 9-35B models to one-click create full apps. Most people who were complaining did so .

Narew - 3 hours ago

From what I personally tested Ornith-1.0 35B is slightly better than Qwen-3.6 35B. My tests are tasks that consist of adding/modify feature in a big C++ codebase. The part that I find interesting is that the model is way faster than Qwen3.6 35B. It seems Ornith produce a smaller chain of thought. On my test it can be 3 time faster to produce the answer.

I use it via llamacpp and codex-cli.

kennywinker - 14 hours ago

Can anyone explain what’s the story here? Is this just a re-skinned qwen? Who is deepreinforce-ai and why isn’t this model listed on their website?

How does it self-improve, does the model change on disk - or just during a single context run it gets better?

S0y - 12 hours ago

These are simply benchmaxxed versions of either Qwen or Gemma 4.

giancarlostoro - 9 hours ago

> the dense 9B fits on a single 80GB GPU

Us mere mortals cannot use this.

agenticup - 2 hours ago

can the orniths self scaffolding could learn to scaffold the rlm loop?

RandyOrion - 4 hours ago

Glad to see more open models. However, where are the 31b models?

- 13 hours ago
[deleted]
v3ss0n - 11 hours ago

Self-Improving bullshit. It is just Qwen 3.5 finetune benchmaxxed . Nothing spectacular . even fails at benchmarks. Long session tool calls sucks and hallucinate a lot with that too. Just use Qwen 3.6 and 3.5 122b.

anana_ - 12 hours ago

They keep mentioning a 31B dense model, but there are no benchmarks or weights for it anywhere?

jkwang - 15 minutes ago

[flagged]

modgate - 2 hours ago

[flagged]

fratefritto - 11 hours ago

[flagged]

- 11 hours ago
[deleted]