GLM-5.3: Frontier coding with emergent cyber capabilities
z.ai897 points by pella 12 hours ago
897 points by pella 12 hours ago
I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)!
I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
At my work I have a $500 monthly AI budget. I have been using the $200 Claude subscription and most of my use is with Claude code. I think I'm going to switch to either kimi or glm and use the opencode harness. Both fable 5 and opus 5 have outright refused things like security related bug fixes and making monitoring tools. I am so happy that open models are good now
Yes, I am tired of Claude and GPTs. I am ready to diversify my $300 per month on other vendors. Will try GLM. How was your rate limits and availability experience on $80 dollar plan?
It’s comparable to Anthropic usage, to be honest. 2x GLM agents ate 18% of weekly usage on this mid-tier plan within ~8 hrs (non-stop work, a lot of tool calls, appx 4 compactions each), I think. I didn’t make a proper statistics snapshot, sorry.
You should try a better harness. Try pi, or ohmypi if you want a good OOB experience
what is a harness? The comments below are mixing IDE/ADE but other suggestions are purely terminal things and I don't get what their value is over just a terminal. Is a harness like a loop where it's just a vague thing that everyone nods about but everyone is nodding at something different?
I’m in the Claude code harness for everything boat too. What are the alternatives?
What the person above is suggesting:
(no personal opinions of either, links might be useful)
I think that OpenCode is nice, their CLI version is enjoyable and their desktop/web version is okay:
I also quite like driving OpenCode through something like Kepler / Paseo and tools like that (with those I can still use my Anthropic Condition by Claude Code being treated similarly - as something that gets tasks dispatched to it, while the GUI I see is Kepler / Paseo).
On the desktop side, ZCode was surprisingly usable for something that came out of nowhere (I wasn't aware of it at all before trying out the GLM Coding Plan): https://zcode.z.ai/en
Last time I tried some of these, none of them had the "manual mode" that CC has, where it shows you change by change as diffs and you can edit them before accepting and moving on to the next change. I like that because if it's going off pattern I can spot it early on and guide it correctly, instead of having to review the whole completed diff at the end when it's too late. I should spend the weekend checking them out again to see if they added that but I assume with everyone going full agent mode they probably didn't.
The philosophy with Pi is it is minimal (but functional) out of the box and easily extensible. I'm not familiar with that feature but I would not at all be surprised if someone already coded a Pi extension that does it.
Both Pi and OpenCode let you customize them. You tell the AI you want "something like claude code manual mode", and they'll modify your configs to do the same thing, or build an extension for you
(however, it's much faster to use Plan Mode to build a plan of what it will do, and then execute the plan in Build Mode. you can also have the AI make a script that will be executed deterministically)
T3 Code has been amazing. Completely free. Really impressed with the desktop app and the mobile app experience and the way it works seamlessly has me actually accomplishing tons of stuff while I'm out on mobile that I would otherwise have to wait to come home for. First time in a while I'm actually excited to use a desktop UI instead of the terminal. Blows away the official Claude Code mobile app. I can switch between my Claude and Codex monthly subscriptions in it as well. There's a TestFlight beta SwiftUI mobile version that's so much nicer than the one in the App Store. I'm running the nightly version of the desktop app.
And this is coming from someone that's not particularly a big fan of Theo. T3 Code should get more recognition; people aren't just aware of it yet.
If you like running everything in a VM and using a web browser as your UI, Shelley is very good: https://github.com/boldsoftware/shelley
It works nicely in the browsers on my tablet and phone, too.
On exe.dev you can ask it to customize itself, and it will automatically rebase your customizations when upgrading to a new release.
ArtificialAnalysis puts out benchmarks for harnesses now as well, and OpenCode seems to be winning it. https://artificialanalysis.ai/agents/coding-agents#coding-ag...
I only found this yesterday, and it inspired me to start testing out OpenCode.
Just a warning, this is on Opus. There's not a clear harness winner. It will change depending on the models.
Can OpenCode dispatch background subagents yet? I tried it a week ago and saw nothing. This is 99% of my workflow at this point.
Having just "asked" my opencode instance-- Yes, but it's behind an experimental flag and not necessarily feature-complete.
In the new v2 beta, yes. Major QoL upgrade, so much less sitting around waiting.
The v2 branch of OpenCode has not been touched for months, if it's beeing developed then I don't know where.
this is very easy to check. i promise i'm not tricking you and just uploaded this. https://github.com/anomalyco/opencode/tree/v2
this is the integration branch for https://opencode.ai/v2 . it has been for months. it's where the Effect-based refactor has been landing.
I'm not surprised to also see Cursor above Claude code, their harness is very good.
In what scenarios?
It indexes the code efficiently, seems to find stuff quicker, it has a very nice UI (much better than Claude Codes IMO), it has a nice sub-agent UX which I find triggers more reliably, diffs render nicely. Otherwise it just seems to work in a purely vibes sense.
That said Claude Code is perfectly fine. I just prefer the integrated experience of using Cursors since I already use VSCode, but I still mostly use Claude Code because of their Max/Fable plan.
Piggybacking on this thread to ask my question: What are alternatives that are multiplayer (team oriented) by default? For example, I want my team to see all my sessions easily, vise versa. another way of stating: all the agents are running in a container that that any member of the team can view and interact with.
Mine is a WIP for automation but has similar concepts to what you're looking for: https://github.com/rush86999/atom
https://github.com/tontinton/maki is tackling the right issues IMO. not sure how they compare with the rest
I think that writing your own harness is a rite of passage now, just like writing your own search engine or database, rolling your own crypto…
Anyways, please try mine!
thank you all! got something to tinker with this weekend
i like to challenge my assumptions and try new tools
Just as a +1 anecdote. I enjoy using pi a lot. I used to h think the harness matters a lot but with the current iteration of models I am starting to sway that while it matters it’s less and less important and that CC is bloated. I did some quick tests when I switched and a task that would take $5 in tokens would be completed in $0.50 in pi. Very anecdotal and I don’t have a test framework setup to make this very official but increasingly felt like CC was spinning its wheels on the easiest of tasks.
I’ve tried a bunch of them, and I seriously do not understand these recommendations. It was a rough road and a steep hill, but right now CC is absolutely the best harness on the market, as for me, whatever top tier model is under the hood (mostly, some of them, like DeepSeek, don’t fit CC at all).
Inversely I don’t understand the praise for CC. These days it feels like bloatware. It absolutely can get the work done but when I measure on token and time use it ends up being a multiple of pi like harnesses.
CC works but for me it felt like increasingly they have zero incentive to make it a great experience. You hear folks like Boris talk about spinning up thousands of agents over night and agents chatting back and forth in GitHub issues and while I think it’s great from figuring out what the future looks like I don’t think it represents the reality of ROI today. So the folks building the tool are so disconnected I am simply not sure it’s a great experience anymore.
So is the quantitative difference in token use the only difference or do you think there's also a different qualitat? I'm on CC only and immensely happy. Very productive both at work and privately and at work I average around $250 a month which probably means nothing but it's little compared to my salary.
Is that the main concern though, cost?
For me, at least it's that the newer Claude models seem optimised for one-shotting things, which is not what I want. As the amount of code per turn increases, I have a harder job keeping up and ensuring that it's doing what I want.
That being said, I had to nope out of a similar thing from GPT 5.6 today, so it appears to be a US frontier lab issue. Claude is particularly bad though, as it produces far too much code even when I tell it not to, unlike GPT (and Kimi) which at least listen to me a little better.
More generally, I want a usable human review experience, and Claude code doesn't deliver that for me.
Quality is hard to measure and I would not say the concern is so much cost but the intersection of cost and time. Often I am jamming on something and I like being somewhat in the loop. So maybe same level of quality, I am using Anthropic modela for both harnesses, but I get to the output quicker and at a drastically lower cost.
Funnily enough, I would say almost the opposite. CC’s feature set is basically table stakes for an agent these days (does it have ACP yet? Very close to behind table stakes if not) and it has a lot of bloat powering that.
IMO part of it is that the underlying LLMs have gotten better enough that harnesses feel better even if they haven’t changed. I have a toy harness that barely implements the features you’d expect and it works surprisingly well. Like there’s literally nothing clever, it calls tools and that’s about it, and it still mostly does the right thing.
Why would anyone ever need ACP? I'm not trying to be an asshole. I just seriously don't understand the value proposition.
Edit: lol, I don't think ACP is even actively developed anymore. It seems to have been merged into another seemingly pointless standard with an even worse name, A2A. [0]
It's buggier for me than it has ever been before. I don't think that agentic coding always leads to such a buggy mess. I just don't think that the Anthropic front-end software team is very good at agentic coding.
I'm gonna shamelessly plug my own here :) https://dirge-code.github.io/
I like it.
I think creating your own agent is the Hello World of agentic coding. Instead of Rust, I used D for mine.
what is this comment based on ? vibes?
Based on the fact that Claude Code is only optimized for Anthropic models, whereas Pi and Omp are optimized for a wide variety of models, including open weights.
they are not really optimized for 'wide variety of models' . what optimization did pi do for glm 5.3?
Vibes like your low quality comment?
What’s the counter argument? pi and ohmypi are pretty fantastic. Of course like all developer tools it depends how you do your work but I am not sure what you are trying to achieve in your comment.
how would i comeup with counter argument if i dont know what original argument is. No one is disagreeing with your subjective experience, gp comment said 'better' without qualification.
“ Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.”
Cybersecurity capability might be nerfed
> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available
I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.
This is a "they have guns so we need guns" scenario.
You can't guarantee everyone else will use a neutered model.
Isn’t this essentially what anthropic is doing, albeit in a manual fashion? They work with code owners to run mythos and find issues.
Only if you're some big corporation with deep pockets. They actually accepted me into their cyber program but Fable's still locked down.
OpenAI now makes it easy to join their verified security program. Took me 5 minutes, and I was able to get GPT to do a full end-to-end pen test
Impossible with source code, possible to bypass with app/site
Don't we already do this with services like Let's Encrypt, which is arguably more sensitive? If you had the codebase you could fake it, but it would still provide some amount of protection against abuse.
With Let's Encrypt, all the verification is done on their side with them controlling the connection between themselves and whatever they're trying to verify.
In this case, you can put whatever you want between the harness you're running (or modify the harness itself), and essentially "lie" to the model. Any verification technique would be fairly trivial to bypass, while you continue to run the harness locally.
How much usage do you get out of it per week? How many millions of tokens?
Anthropic was stingy as hell with its Fable and cybersecurity nonsense, switched to OpenAI which is much better but still not enough. I'm tempted to switch again...
They're invaluable for developers to fix their code. This is definitely an area where AI decisively beats human devs in a very valuable way. It can try so much surface area so fast.
If it won't attack my stuff, it won't help me build my stuff to be secure.
Exactly my CoT! I hope z.ai won’t change this behavior after training it on our input the same way as Anthropic did (shame on you, folks, seriously)
> after training it on our input the same way as Anthropic did (shame on you, folks, seriously)
What do you mean with this? Honest question!
Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/
Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.
I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
> ... Anthropic's Project Glasswing is supposed to find them quite a while ago?
That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
> and that Apple would not attribute to GLM
That was a wtf to me, so I checked Apple’s latest iOS release security content and GLM & z.ai is mentioned once (under WebKit), Anthropic is mentioned twice, Codex is mentioned once. Not clear if there are other instances where the model did most of the work but wasn’t credited. I didn’t bother to check other releases.
> That was my thought too. For all of Anthropic's talk about their "adversaries"
It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.
Then "security through secrecy" is really bad mantra especially in the age of AI: others will find the same zero days very soon. If they attack you, then this loses the whole plot. If they propose a fix, then your arsenal becomes smaller.
This is a coherent explanation for why federal model censorship has started with cyber capabilities. But this GLM model release is an in-your-face challenge to that policy. They now have to either set models free or impose a censorship regime that will put anyone not under it at an advantage. Or muddle along in the middle as usual.
> Or muddle along in the middle as usual.
I'm not a gambling person, but if I was this would be my bet.
Who says they missed them? Could also be sitting pretty in CIA’s long list of ready to go Vault7-like exploits.
Probably Anthropic found them too and promptly got a call from Isreal to stop looking.
> I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
You have to consider that having an LLM scan for vulnerabilities is hardly infallible. It is a search guided by heuristics and given a large enough codebase, it is unlikely to identify all vulnerabilities.
Personally, I've had Fable 5, GPT 5.6 Sol, and GLM 5.2 all looking for correctness issues in an old abandoned WIP codebase of mine and all of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.
> [A]ll of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.
This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots. Frontier models should large dominate their competitors on a capability basis, but if GLM 5.2 (now 5.3) is routinely finding bugs / vulnerabilities missed by Fable and Sol then GLM might be genuinely a frontier-grade model by itself.
> This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots.
Not necessarily. Even near the frontier, we don't really have a total ordering of capabilities, but a partial order. And even frontier models make plenty of mistakes. Combined with the randomness inherent in searching large codebases for vulnerabilities or correctness issues, it is entirely plausible that even much weaker models (and GLM-5.2 isn't even weak) can stumble upon issues that stronger models missed.
My current hypothesis – for which I have only limited evidence, unfortunately – is that it is better to have multiple reasonably powerful (but not necessarily frontier) models looking for issues than just one very powerful one. And even then you're likely to miss out on some issues.
Fable 5 is just over two months old.
For normal software it would be as you say, but LLM progress is so ridiculously fast that things go from "bleeding edge" to "eh, you'll do" in about that timeframe, and "eh, you'll do" to "why even bother with this old rubbish?" in the same again.
Or, from a different perspective, we can expect some new frontier model from Anthropic in a week or two, and from OpenAI in a month or so.
I find that LLMs also generate a lot of false positives, or extremely minor issues that don't warrant a fix (that are always overstated by the LLM as very important!). Signal to noise is still not great and requires somebody to wade through and pick out the actual good findings.
> and Anthropic's Project Glasswing is supposed to find them quite a while ago?
We cannot trust a single company to report security issues, it’s good to see competition in that domain
In a similar vein, does anyone know how to classify the kinds of problems that are being found?
Is it possible to build heavier traditional linting to catch whatever is being caught in a more deterministic way? It seems to me that would be far more efficient in the long run (even if the efficiency is only for the AI to know that aspect was already checked).
this is impressive and actually matches my expectations in terms of near term AI progress. we are going to continue to seem impressive progress in coding & related, anything where verifiability is scalable in an automated way: https://transitions.substack.com/p/a-quantum-of-ai-progress?...
> but isn't the cost for such a scan getting lower by the week
Not with Anthropic's models!
If you look at the distribution of their findings in the linked post, most of theirs are issues introduced a long time ago, almost all before 2006.
Complete speculation, but I wonder if they and Anthropic are scanning very different codebases and Anthropic's skew would be in the other direction.
Interesting... So Chinese models are not so bad?
There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.
Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.
Unfortunately, those in power, pretty much all over the world, lie/deceive themselves and believe they can.
In this case depriving US companies would be the point though, so that's not necessarily a disadvantage.
> Anyone who knows anything realises banning things is a) impossible and
Maybe "It's really hard" is more accurate? We (humanity) for most part basically agreed to ban the usage of various chemical weapons in wartime, which seems to have drastically reduced the usage of it, even though it's still used by shit actors today from time to time. But it's hard to deny that usage didn't decrease after banning it, which makes "banning" maybe not completely useless for certain things.
"Banning" things that can be easily copied over cyberweb transportation pipes feels like an fool's errand though, regardless of what it is. It's just too easy to get around, compared to actual physical items I suppose.
This is different now. US labs and companies are not releasing frontier-level models openly (specially those capable of assisting cyber intelligence work), but commercializing them instead. Thus, any ban would not be symmetrical to begin with, and that is precisely what maintains the balance.
It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't actually care about the final few %.
Regardless of whether or not adversaries are using them, the US has by far the most compute available, and we've now hit the line where major providers are no longer releasing their best models. The public gets the "current" level of intelligence, while the US government gets to control access to the actual frontier of non-public AI. From their perspective, their enemies using GLM5.3 while they have GPT6 and Mythos6 or whatever is a fine trade.
I don't support a ban at all, nor the US's behavior, I'm just pointing out some facts that change the argument.
But the real bad guys will be this final few %, which defeats the purpose. The 99% will be average user which will swing to cheapest AI or easiest to access.
I don't think this really works because the Chinese government is going to be incentivised to tip off the US companies to deny the US government those exploits. I guess maybe that's what the open source patch program here is about, making sure banning the models doesn't work because they can just report the exploits without the company running the model themselves.
Do you actually believe this?
The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so?
[1] https://www.washingtonpost.com/graphics/2020/world/national-...
You should probably realise that the world has radically changed since then. This kind of thing works when you have a significant lead in the field that makes keeping vulnerabilities open sufficiently low risk for your own side. But if your adversaries have similar capabilities, then the calculation changes.
Has anything changed? Governments are hoarding undisclosed vulnerabilities, using them as they see fit instead of fixing. Every espionage, surveillance, or war campaign (see Russia v Ukraine, US/Israel v Iran etc) is followed by a ton of burned 0-days.
>This kind of thing works when you have a significant lead in the field
No? It works even if the adversary has the same capabilities. It only stops working when everything is fixed.
I believe it is unlikely. (Not because I do not believe NSA is hoarding 0-days, but for many other reasons.)
I'm curious: to any professional vulnerability researchers reading this, what do you think?
I used to call everything a conspiracy theory, but then Glenn Greenwald published "No Place to Hide: Edward Snowden, the NSA and the Surveillance State".
Now i know that reality is worse than the worst conspiracy theorist.
I don't think reasonable people post here much anymore. It's mostly galaxy brained conspiracy theorists and ignormamuses posting political garbage. Reddit-lite on the way to full blown Reddit
Why would you even believe the opposite? US spooks have been amassing vulnerabilities and relying on them for decades, they literally pioneered it in the 90's if not earlier. Everyone does it now but the US is the biggest of them all. Surely this devalues a lot of what they did. Moreover, the way the US government handled new capabilities, and OpenAI's training policy (they are in bed with the government) just scream "we want to create weapons for cyber-offence and deny them to everyone else"
It might not be the reason, but of course it's a contributing factor.
[flagged]
Good thing I said nothing of that (especially nothing about China). Reread it again to understand you built an incredible strawman and ignored my last sentence.
Well we know that the US government is pushing to restrict access to such models while the Chinese are publishing them for free, so it's mostly a matter of motivations, not the actual facts of the matter. And the USG has a documented history of unsavory behavior (including toward its own citizenry) in that area.
So we might ask if one of the reasons the US is being the bad guy is it's usual spying antics, and we're left asking why China is being the good guy.
Intelligence agencies have been known for exploiting and planting software and hardware Buga for decades, going as far as weakening cryptographic standards or intercepting hardware in transit to implant a backdoor device.
Why do you _not_ believe it's a possibility?
They've always been good enough for double digit less money. Always. Anyone thinking "Chinese models fake models built using dirty distillation scam" don't know what they're talking about.
Distillation is just forcing the model to use an exam prep workbook for training instead of generic publicly available textbooks. The models themselves has to be smart enough for that to work. It's the exact same thing as Asian tiger mom double schoolwork strategy, to paint a picture.
To carry on this analogy - do test prep workbooks make you meaningfully more competent in general, or is it benchmaxing? (Versus studying textbooks for a similar time, of course.)
Their best models are getting more and more expensive, and still aren't SOTA.
It's almost like there's an actual cost to developing these models, and the Chinese don't have magic dirt that allows them to do it at a fraction of the cost.
Looks like they're going for good PR now, to avoid smearing by the "Western" models. Smart!
I'd love to live in a society where people and corporations do good things for PR.
> I'd love to live in a society where people and corporations do good things for PR.
Maybe so, but I'm not sure I'd like to live in China of all places. (Don't get me wrong. Lotta places I'd like to visit if I ever got the chance, and China's on that list, but to live there? I don't think so.) Maybe one of the Nordic countries?
> Anthropic's Project Glasswing is supposed to find them quite a while ago?
Someone still has to run it. The analysis and fix could be someone's machine but not committed / published.
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize.
These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin.
I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.
This is going to be catastrophic.
Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.
I take it from [1] (transcript of recent DeepSeek CEO discussion with investors) that DeepSeek would disagree on the immediate catastrophic impact to the likes of OpenAI or Anthropic. The reason is even though technology parity mostly exists, only OpenAI, Anthropic et al have the inference capacity to gain market share and generate revenue. Chinese vendors don't have the chips needed to scale up inference and gain market share, and the DeepSeek CEO doesn't think this would happen in optimistic circumstances in the next 3 years, but thinks it might be possible in 5 years.
In summary, regardless of country of origin, availability of inference capacity is the moat protecting the likes of OpenAI and Anthropic, not technology superiority.
[1] https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his
That merely pushes the valuation onto the hardware makers, not the companies that have the temporary preferential access to their hardware.
That makes them at best temporary middlemen.
It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't).
Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.
The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.
If you have exponentially increasing use of your harness, then it's true that every day you capture exponentially more data, but it's also true that every day exponentially more data will slip through the cracks of your would-be monopoly and that data arrives at your competitors via various channels (competitor harnesses, subsidized reselling, etc)
The very exponential that you are relying on to give you runaway improvement is also giving exponentially increasing data to your competitors. All else being equal your competitors stay a step behind but you never develop a monopoly either. That's the best case for Anthropic/OpenAI. In reality, training data is just one variable, exponentials don't last forever, and your competitors will get better at capturing a bigger slice of training data.
If user data would become such a key ingredient (which it might, i actually remember noam shazeer talking about the importance of user data), i think chinese labs can still get it from china, as keep in mind it ahs a billion people behind the great firewall banned from using us llms. And btw broadly for any gap like this, you really gotta consider that if its becoming a bottleneck, chinese labs will find a way to buy it from one of the labs unless theres strict regulation at the government level
But is that data good? That's the question. As in, is my usage at work:
a) indicative of problems that aren't already out there in the wild? (no) b) are the responses I'm getting so good and novel that the model can improve itself? (no)
It's the garbage in garbage out idea, just scaled up. If the model gave a bad answer, and I didn't catch it, and you now train on that I/O pair (my perhaps crappy prompt, the bad output), then you're not going to improve anything.
It seems like the user response rating mechanism might be a valuable signal
Yes, RSI seems to be the new AI industry McGuffin of 2026, just as agentic capability has become table stakes and scaremongering has become a punchline.
The Chinese models are adopting licensing quite rapidly and Xi will soon enough close them for security reasons. The most widely used model, integrated across Bytedance apps and operations, has never been open and is most closely associated with the state.
I would add that it is not just capacity, but also negotiation ability. With scale comes the ability to negotiate better prices than everyone else. Even if you can find capacity for your smallish user base, your inference cost can not match these companies unless you have a technical advantage for your inference cases. Squeezing the hardware requires request batching and caching which are far easier at scale and sustained user activity.
Thanks for sharing.
Is lack of inference chips due to the trading blocks by trump administration? What if Trump agrees to sell chips to china, would they collapse then? That's not a very strong position to be at
is that releveant if people can host their own models? that activity still undermines the valuation / diminishes the US companies 'moat' ?
People can host these models is doing a lot of lifting here, these are models that depends on 5 digits on specialized installation to run on.
IMHO, this has the impact of softening the impact of data centers sitting unused in the long term if they can still serve open weight models, even if Anthropic or OAI have to scale down their expansion rate to pay the bills.
Regardless, reality has to give at some point; these valuations don't make any sense. We've been valuing GenAI as disruptive work, when in reality they're much closer to cloud providers with a beefy, one-pony-trick R&D department.
they have the capacity via market manipulation; so you know, they only have things they've bought on the governments future debt obligations.
so, you know, they're as vulnerable as utilities at this point, if only there were people who gave a shit more about society than greed.
I have already begun winding down my spend on claude and OAI to make room for infra budget. Anecdotal, but I have no doubt a lot of others are doing the same, I very much agree the US players have major issues looming. What an exciting time to be alive!
Not exciting for anyone directly or indirectly invested in a frontier lab or its partners. And that is a lot of people, including you.
No time like the present to pull out and reduce your exposure. I brought this up in my employer's forums 4 months ago and honestly it's been clear even before then. In particular, the upcoming IPOs of both oAI and Anthropic will likely be disastrous for the public - the floor is falling from under them and I don't know if they can be scrappy and work with fewer resources - their internal culture may not support this. We all knew in our hearts they're a commodity - just see how easily you can switch between the 2 of them - and now there are 10 more options costing a fraction.
When Xi Jinping did the announcement of their open weights push, they might as well cancelled their IPOs....
Yes buuuut…. I do quite a bit of day trading (maybe closer to scalping) for the first few hours the market is open, everyday. Anecdotally: despite everyone knowing its valuation was ridiculous, I rode that SpaceX train pretty hard and made a pretty penny.
I close-out all my positions by end-of-trading everyday… so when the day came when there was a very clear and very scary indicator during early trading hours, quickly followed by SpaceX’s catastrophic fall right after opening bell, that was the end of my involvement….
And I fully expect oAI and anthro to be the same way. They’re being propped up with private loans, subsidies, and other tricky bookkeeping techniques. You would think their CEOs would pivot away from their current public personas. Ironically, they are like a poor man’s Elon Musk… and that doesn’t bode well for their companies
Yes experienced investors will profit from it and leave the general public holding the bag, that's the plan I'm afraid.
The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.
I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this.
Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volume they need to sell.
Finally, there is a data wall. Sure, they can keep scaling RL on math problems and code. But with everything else, where will the supervision come from when they need several orders of magnitude more?
I agree. And even if they were able to do it for one more round, it's not a sustainable strategy. What they (Anthropic and OpenAI) need to do is build platforms and integrate verticals.
assuming the technology of model architectures does not gain any further breakthroughs that returns us back to the gains previously seen. I'm of the opinion that we still have some discoveries on the mathematical side of the fence to go that will improve models further.
> I'm of the opinion that we still have some discoveries on the mathematical side of the fence to go that will improve models further.
That's assuming the infrastructure needed to develop models stays available financially and supply wise. A lot of the services used to train and develop models are supplied and funded by people who are looking for multiple returns of investment. If/when OpenAI and Anthropic valuations fall and they inevitably get acquired, will Meta/Alphabet/Microsoft still want to spend lots of money for unclear returns in the short-term? Nvidia and co are on a one way train service to hype town. I don't think they will be happy to get on a coach to hype town Temu version. The shareholders likely won't.
Also, the backlash against LLMs is growing rapidly. AI content, data centres, etc is quickly gaining negative connotations outside of visual and music artists circles. While existing models are going nowhere, developing more advanced models is very quickly getting unpopular. LLMs Data centres increasing people's bills, Anthropic destroying old books, chat bots giving unethical advice to vulnerable people, etc. It won't be long before LLM infrastructure becoming an electoral issue.
Will a small research oriented community be big enough justify maintaining the apparatus needed to produce infra tech at a profitable level post OpenAI?
The car industry is also a trillion $$ market in the US. I don't see why that would go any differently from the Chinese cars ban.
You wouldn't download a car, would you?
I would download it if I could, no question about it. And 64GB more RAM if I am at it.
Most of the money will come from companies/corporations who will be required to buy safe AI. The public will be just banned from buying which might make it hard (ie: site/payment blocked) but not impossible. It could be good enough for the big whales.
The US will just do what they did with Chinese EVs: ban the superior technology to protect US companies.
Why is anything going to be catastrophic? Companies can go bankrupt without catastrophes for the rest of us. Happens all the time.
I read it as catastrophic for the companies trying to IPO. It'll be great for the rest of us though.
A large portion of the economy is currently tied up in the musical chairs shell game that is AI hype. When the music stops there are going to be CEOs looking for handouts and justifying it with spooky national security buzzwords. How we respond to that will depend on whether it happens in an admin that is famously captured by the industry or not.
I seriously need to start considering the scenario in which this leads to next global financial crisis.
Just keep in mind that it can take a whole for things to play out. I’m someone who believe the US AI industry is completely unsustainable and built on sand, and will crash even if the current AI itself turns out to be very successful. But that doesn’t mean everything will burn to the ground next week. In a history book things will look very sudden but at normal speed that can easily take months to years to fully play out.
Also, take in consideration that the AI trade infected a lot of other trade in the economy, if you decide at some point to move your money to a place that is safe in case of a downturn be sure to carefully evaluate that’s actually the case
These crises are manufactured by the central banks.
Compare and contrast how the dot-com bust did _not_ lead to global financial crises. Nor did Black Monday, nor the recent string of bank failures in the US.
('Manufactured' above means that central banks are responsible. I make no judgement on intent here. Around 2008 it was incompetence by the Fed and ECB as far as I can tell. The Fed started paying interest on excess reserves and the ECB even increased rates. Twice. Amongst quite a few other missteps.)
A lot of the performance of these open source models might come from distilling the closed frontier models. If those can't raise the funds anymore to train newer and better models then the whole improvement cycle might slow down.
Another interesting potential market here will be 'LLM in a box'. All the hardware and other tooling in a prebuilt, but modular, package ready to go. Pay one up-front cost, get a system running [whatever open LLM] with a token rate of [x], optionally configured to be immediately ready for distributed usage. Basically the opposite of cloud stuff: no rent, no dependency, 100% guaranteed uptime, guaranteed security/privacy (at least subject to your own actions), and so on.
Palantir already offers a "turnkey AI datacenter", i.e. a rack with "NVIDIA Blackwell Ultra systems with eight NVIDIA Blackwell Ultra GPUs and NVIDIA Spectrum-X™ Ethernet networking for AI training and inference".
It is said that it comes with all hardware and software required to run inference or training with an open weights LLM.
The existence of this product, which competes with cloud-based offerings like those of OpenAI and Anthropic, is presumably the reason why the Palantir CEO criticized very harshly some time ago the business model of OpenAI/Anthropic.
While I doubt that the ethics of Palantir is any better than of OpenAI/Anthropic, in this particular case I have to agree with Alex Karp about "Sovereign AI", i.e. that only losers will make their business completely dependent on an external entity like OpenAI or Anthropic, who are certainly not trustworthy.
I'm not sure a data center run by ... Palantir of all organizations is what people have in mind when they worry about data sovereignty.
They are selling it, not running it.
It is just a dedicated computer system, which should be managed by its owner, like any other on-prem servers.
I doubt that it has a good price/performance ratio, but it is a solution for those who feel that they do not want to search, buy, assemble, install and configure every HW/SW component.
I take it we saw different demos.
I'm under no NDA, if you actually want to know what's up.
I'm assuming you're alluding to them selling a managed solution, alongside the unmanaged solution that the GP is referring to?
Fair but the idea of "running your LLM setup" at every "need" level and corresponding cost does make sense.
For a lot of people (and orgs I'd guess) who just go and buy ≈$20 per month plans (or more for teams), they might not even need a fraction of that cost or capability. A lot of them don't even need it for coding or graphics. Even the API access based pricing aren't great from these frontier US AI houses. The distribution of "LLM being" offered will also give rise to many open-router like offering but at the end point level - direct interfaces to the customers. Pick your vendor sort.
AI shouldn't become another "search means Google".
“100% guaranteed downtime when you least can afford it and the support tickets are your problem.”
We’ve a hybrid shop, including hosting our own ML infra, and we save a ton from cloud spend with local ML. Easily one million USD over past three years. But it’s not “free”, you are shifting a lot of labor into your plate.
And with that also gain institutional knowledge, skill up your workers and attract talent that wants to work on this stuff.
All boils down to short-term/long-term thinking.
This. People WANT to work on this stuff. And having skilled workers is a precious advantage.
Still has to break even on the balance sheet, especially at a bootstrapped startup. We actually made most of the financial windfall in translation API fees oddly enough.
For our own model training we needed to do some large scale translation tasks of a large dataset (1M or so documents, 10 or so target languages), running full-size NLLB on-prem saved us an absurd amount of money vs Google Translate API.
(For reference doing 1M target docs into a single language in Google Translate API is roughly $120k list price. You can run full size NLLB on an 48GB NVIDIA A600 and the major difference for us was speed, but for this task time to completion wasn’t an issue.)
> 100% guaranteed uptime
Disagree there but I think this is an interesting idea. We would need to find some more cost-efficient hardware to run it on than Nvidia GPUs.
It will come... all big hardware players (Intel, AMD, Broadcom) and dozens of startups (Tenstorrent, etc.) are working on it...
What makes that kinda complicated is that multi-user throughput of LLMs scale well but single-user performance often stays constant at low ends. If you could saturate e.g. 16 concurrent session-month of demand, you can just go buy 16 of 32GB GPUs and start charging monthly for inference. That could work if you had e.g. over thousand total employees with hundreds of devs eager to trying it out, but only if the company is also interested in a private inference experiment.
You're talking about multi-session vs. single-session throughput. A single user can easily leverage multiple sessions via e.g. subagent swarms, especially on a lower-end setup where any single session is going to be quite slow. Saturating utilization during off-hours is harder but potentially quite feasible by assigning lower priority, unattended tasks/inference loops.
I think at this point the question is: will the US government be willing and capable to justify the trillion dollar valuation for _one_ of the companies via regulatory capture? The US has a workforce of 170m, so 1.7 trillion would come down to 10k per person, or a discounted cashflow at 3% of 25 USD per month - not including private use, students etc.
Why would you restrict to the US workforce? ChatGPT has a billion users.
It’s a common denominator if you want to do napkin-math for a whole national economy. Regulatory capture is like a tax on those people not on the beneficiary side, so if the government were to nationalize both supply (no export license for SOTA models) and demand (no foreign or self-hosted LLMs allowed), they’d end up making everyone else pay for it in some way or the other. The governmental utility function will then include only those using the services for direct economic benefit.
They have a stupid plan to buy ten to fifty percent of all the SOTA AI companies, and giving us all a fraction of the money.
Trump keeps calling his enemies “communists”… then turns around and ‘seizes the means of production’ himself.
It is impossible to justify the absurd private valuations they have given themselves in collusion with investors.
I wish they had tried to IPO because then we’d see the judgement of the market on this. But that’s why they didn’t this year. How long can they keep up the charade that their models are uniquely valuable and on the path to AGI?
> private valuations they have given themselves in collusion with investors.
What's the collusion?
Circular investment deals and investment deals at valuations which have no possible justification.
Let me ask it differently. You state the companies and their investors are colluding. Who are they colluding against?
The public that buys the stock at ipo at this inflated valuation and unwittingly buys indexes which include it (as with spacex).
its interesting, as it's typically the banks and against the public at large because the goal is to jimmy up valuations to justify IPOs then sell on opening; just like spacex.
It's what enron was doing; it's what most of crypto's offshoots were doing.
Sure you can blame the marks of the grift and say "well the public should know they're faking all this cash flow expectation".
It seems like you're either driving the grift economy or part of the collusion.
It's similar to how a cult operates, so I'll be frank: your skepticism seems biased.
> It's what enron was doing;
Enron hid billions of dollars in debt and fake profits.
Is this what you think is happening here?
Nvidia has made a lot of very suspicious circular funding deals. I suspect we’ll find fraud when the bubble bursts yes.