Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

bloomberg.com

419 points by garo-pro 17 hours ago


ricardobeat - 15 hours ago

I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart.

The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had this problem was Mimo 2.5, which is quite dated at this point. As a result of this, you cannot leave it unattended / not usable for agents.

giamma - 16 hours ago

https://unwall.app/www.bloomberg.com/news/articles/2026-08-2...

WithinReason - 16 hours ago

Mixed signals, here it's performing below even GPT-5.4 Nano:

https://livebench.ai/

while here it outperforms Fable by a significant margin:

https://oxalpha.com/

but if the latter is true, will people still say it was "distilled" from Fable?

_pdp_ - 14 hours ago

Ox Alpha has been running on auto-pilot for the past 5 days on various experiments.

Very impressive model.

Here are some examples, open-source documented and the data available in HF datasets:

https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone

https://openzot.github.io/arcade/ - https://github.com/openzot/arcade

https://openzot.github.io/machinery/ - https://github.com/openzot/machinery

esskay - 16 hours ago

I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.

harlan_pdx - 15 hours ago

Releasing weights is the right move. Keeps them competitive with DeepSeek on the open side.

freakynit - 15 hours ago

It one-shotted generation of Java bindings for this project: https://github.com/jeffhajewski/latticedb

Related PR: https://github.com/jeffhajewski/latticedb/pull/5

The session used ~100K input tokens, ~60K output tokens, and ~80K thinking tokens.

I reviewed it using gpt-sol-medium, and it seems to be satisfied with it's work.

RataNova - 13 hours ago

It writes pretty clean code and holds context alright, but it starts stumbling and losing the plot on complex bash scripts with pipelines. Waiting for the weights to drop so we can dig under the hood and see what is going on there

garo-pro - 12 hours ago

https://z.ai/blog/glm-5.3-flash

garo-pro - 17 hours ago

Unfortunately I can't find sources other than this for now but this seems to be legit.

itsryanlenk - 12 hours ago

Looking forward to seeing the stats.

I gave it an abandoned repo for an Aseprite MCP someone made and told it to iterate with a laundry list of things I wanted from it to include thousands of plugins.

Came back 20 hours later and it shit out a pretty surprising little tool, will post the public repo when I get time.

seydor - 15 hours ago

Funny how all china companies are expected to release weights by default

syntaxing - 14 hours ago

I’m more curious on the size. If it’s smaller than or equal size to GLM 5.3, this would be a crazy good model. If it’s closer to deepseek pro, it would be a good model. If it’s near Kimi K3, I think it’s competitive but nothing particularly differentiating.

j_maffe - 16 hours ago

Anyone has a link to a report of its capabilities? I can't find a reliable source.

akshay_akula - 6 hours ago

Reads like the usual pattern, somewhere between Sonnet and Opus but not something you leave unattended. Curious if the weights release changes that.

amathur2k - 13 hours ago

Which harness are u folks using, I have tried opencode and claude code. Both absolutely keep hanging due to the model running into loops and becoming unavailable. Unable to do even simple things

hypfer - 15 hours ago

> The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.

Where? And "Tonight" in which timezone?

dgellow - 16 hours ago

Do we know the size of the model?

tosh - 16 hours ago

my guess is this is a small model punching way above its weight

on toy benches it made quite a few mistakes but was able to fix all of them on its own

(meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)

yipinwong - 13 hours ago

Ox Alpha was working good for me but I do not it if a trend starts where openrouter hides where the traffic is going to.

- 16 hours ago
[deleted]
SyneRyder - 15 hours ago

Rather than a pelican, for fun I showed it a couple of screenshots from Niu Lai and asked it to create an SVG inspired by the images. I explained a little about how the movie had been made by a mother & son team, initially derided but then went on to surprise cult box office success. It came up with this:

https://x.com/syneryder/status/2091978367579156569/photo/1

Created in a single turn - but technically not a "one-shot", because I gave it a tool to convert SVG to PNG so it could visualize what it had made. I asked it to keep iterating with tools during the same turn until it was happy.

I've also been using Ox Alpha for tasks that better resemble real work, and I'm really enjoying working with it. I've downgraded my Anthropic account so I can put some budget towards Ox Alpha instead, with the rumors that this one is going to be cheap. Opus & Fable are still better at getting large tasks / features done autonomously, but Ox Alpha can work autonomously too, and it's fun. I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models. (As much as I don't want to say that, as someone with Claude /stickers on their laptop.)

fen_wick - 15 hours ago

Good to see more competition in the open weights space. The more players the better.

redox99 - 14 hours ago

Ox alpha is better at UI than GPT 5.6 Sol. Not a high bar considering Sol sucks at UI, but as someone who just has a codex sub, I've used almost 1B tokens of ox alpha these last few days to complement Sol smartness.

Inference was atrocious in terms of speed and constant timeouts. If it's served fast it will be a delight to use.

ashing - 11 hours ago

I'm going to use this model hard.

respectattentio - 15 hours ago

it's for sure better than deepseek flash 07/31

glimshe - 15 hours ago

There's a lot of brand confusion among the Chinese models right now. Kimi, Qwen, GLM, Z.ai, Ox. We might know the difference (or I should say, someone does because I'm losing track already) but these models have no chance at end user penetration and loyalty until there's a single focused survivor.

It took me a year talking about it until my wife knew that ChatGPT and Gemini are two different things.

PS: some replies, especially if you do a deep dive on comment history, clearly expose the joint effort to drum up support for Chinese models. This has been clear on HN lately as anything even slightly critical of Chinese tech gets downvoted unnaturally quickly. One can just wonder what's behind the effort...

ThouYS - 14 hours ago

Calling it now: The big deal about this model is the sheer volume they were offering through openrouter and OpenCode. How? Chinese AI accelerators / nvidia-free stack

danieltk76 - 13 hours ago

to be honest I found it underwhelming.

m00dy - 14 hours ago

Yeah, it was identified as a GLM-series model quite a while ago. You can also check out this AI model fingerprinting resource [0].

[0]: https://openrating.io/blog/current-state-of-ai-model-fingerp...

xbmcuser - 14 hours ago

will we reach the singularity once the llm can be used to program the llm?

kosolam - 15 hours ago

Only reason people are interested is it’s free at the moment. I wasn’t impressed by its performance. Once the model gets a price tag it’s usage will be negligible.

13639366668 - 14 hours ago

[flagged]

lampcord - 14 hours ago

[dead]

daveyoung - 16 hours ago

[dead]

hncsiocp9x - 16 hours ago

[dead]