GLM-5.3-Flash
z.ai717 points by Philpax 7 hours ago
717 points by Philpax 7 hours ago
If you're on opencode's go $10/mo plan and want to use GLM-5.3-flash right now on pi, you can add this to models.json until pi updates to support it:
{
"providers": {
"opencode-go": {
"models": [
{
"id": "glm-5.3-flash",
"name": "GLM-5.3 Flash",
"api": "openai-completions",
"baseUrl": "https://opencode.ai/zen/go/v1",
"reasoning": true,
"input": ["text", "image"],
"cost": {
"input": 0.15,
"output": 0.5,
"cacheRead": 0.03,
"cacheWrite": 0
},
"compat": {
"supportsStore": false,
"supportsDeveloperRole": false,
"maxTokensField": "max_tokens"
},
"contextWindow": 1000000,
"maxTokens": 131072,
"thinkingLevelMap": {
"off": null,
"minimal": null,
"low": "low",
"medium": null,
"high": "high",
"xhigh": null,
"max": "max"
}
}
]
}
}
}Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash
I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.
I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.
I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
I am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...
DS4 Flash absolutely kicks ass for reverse engineering and bug hunting. Almost no point in considering paying for a bigger model, although it's possible the stuff I've fed it (wide variety of older DOS/Windows stuff and device firmwares) might be easier targets.
I've been reverse engineering LEON3-FT SPARC v8 BE code, so I wouldn't say it's common :D. When attached to Ghidra through a MCP the things you can do with this are simply crazy.
https://github.com/bethington/ghidra-mcp . Works flawlessly.
Ha - I saw this, but took one look at the slopfest README and it sorta scared me away. Will give it a go thanks.
Curious what MCP setup you use? I'm not sure which one I have wired up, but I have to restart Ghidra every time I change file. I think it's either LaurieWired's original or a fork of it
Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it.
https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...
I gave Qwen 3.8 27B and Opus 4.8 the same task in the same codebase. They both came up with the same diff. It wasn't a particularly challenging task (removing a feature flag and updating applicable specs), but it was character for character.
Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent decisions about greyer areas of good software architecture.
Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).
I’ve been doing the same thing, giving the same tasks to Qwen 3.8 27B and Opus, and the main difference is that Qwen does not consider edge cases which Opus catches. It’s good at the happy path, but even when hinting that there are uncovered edge cases and gotchas it’s oblivious to it. So I feel like I need a bigger model to do planning/review.
As a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely useless at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant style model. Zero chance I'd "work" with it, I spent days trying to get it to do something for me that was usable that I didn't have to have reviewed and refined by a frontier level model or myself. Couldn't do it. The idea that qwen 3.8 27b is _anywhere near_ Opus 4.8 is laughable. Pure benchmaxxing.
DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.
GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.
I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side.
The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.
It's the first small local model I've felt like I can do real work with.
What harness? I've had similar results as _Implicated_ said above - it's not done well in any of the tests I've tried with it. I currently have it hung off DS4Flash as a pseudo-vision tool and subagent only because of this.
I'm using pi inside a self-made harness. I've found going super lightweight with context (AGENTS.md is maybe 20 lines) and letting the model discover what it needs to gives the best results.
Fascinating that after years of for me this for me that, literally no one writes 2-3 other words like “i do react frontend” or whatever just for us to know why the results are different
Sparks don't have enough memory bandwidth, for the same 20k you're better off buying RTX or Apple M5 Ultra machines.
I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.
> get myself four sparks at a decent price
Wow, if you don't mind me asking. How and where?
I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely.
They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
> it's the only model in the whole lineup that isn't priced insanely
$4,000 isn't priced insanely? ye gads
Compare to the cost of professional-grade tools in other trades and craft hobbies.
Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies.
And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.
That's only half the reason it's expensive.
The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.
> it would likely take years to spend $4000 (plus the real cost of electricity)
Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.
As a counterpoint, my homelab/home-LLM hardware has appreciated in value by about 60% since I bought it.
Of course, it's not real unless I sell, and the value will eventually go down, but so far I have significant paper profits.
Also, DeepSeek token prices are continuing to _increase_, not decrease.
> DeepSeek token prices are continuing to _increase_
One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...