China’s open-weights AI strategy is winning
werd.io1053 points by benwerd 16 hours ago
1053 points by benwerd 16 hours ago
The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.
- PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to.
- PC office productivity software destroyed expensive professional products.
- Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world.
Ignoring the huge Chinese open-weight models for a moment:
- The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious.
- There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality.
- Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s.
Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.
Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends.
Personally, I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications. I think when the economics make more sense, product designers will make things that people actually want to use that will pretty transparently handle whatever model interactions are necessary, when it makes sense. As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment, and b) be obstinate, obtuse, or argumentative, or generally just be something that I have to explain things to. I think a lot of tech folks are far more biased than they realize by the “ooh, neato” factor when imagining how nontechnical people might want to use things. And the weight of these tools just feels wrong for what a lot of people use them for: the thing that plays whatever music I feel like hearing absolutely does not need to be able to generate a volumes of fanfic about the movie that song was in. It’s abstractly impressive that something could do that, but it’s just not useful.
> I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications.
I use AI chat every day, I find it endlessly useful. It’s replaced google search.
> I use AI chat every day, I find it endlessly useful. It’s replaced google search.
Extremely subsidized agentic search is very superior to Google at the moment, and of course it is. Google is a public company. The AI summary model has to work instantly, is likely as dumb as a 8T param model, and gives you incorrect details constantly. This sucks so much for Google. If you click on "AI Mode," suddenly the facts become more accurate.
Of course, if I want a real answer I happen to go to claude.ai, set it to a the best model, wait for a minute, and use many watts of energy. Slow agentic search that takes many seconds, and is greatly subsidized, is certainly better. This should not be a surprise, should it?
I think it was on a sub like r/singularity that I saw a post along the lines of "of course most people think that 'AI' sucks, as normies are interacting with 8T param models."
tone: genuinely confused about the world, not criticizing
I've found LLMs useful for surfacing popular recommendations. I also get the overwhelming feeling that it's all very early days still when the machine mixes together whatever was crawled into the weights with a web search or two and dumps it into a markdown blurb.
I totally agree with the above that a more polished and less obvious use of LLMs integrated back into search engines may be more useful, but will definitely be more usable.
I’m sure a lot of people here do. I wouldn’t exactly call this a representative sample.
It’s a sample of the forerunners
See: product adoption cycle
I encountered multiple people in HN that had Apple Vision Pro. How many people here use Linux? Have flagship model phones? Drank soylent? Microdosed LSD?
Extrapolating based on what you see on HN doesn’t make sense.
Google search still happens, it’s just your agent doing it.
Google search is basically Gemini now. You get a Gemini summation including several links.
"You get a Gemini summation including several links."
Who does "you" refer to
Me, I don't get a Gemini summation (Tested with old version of Chrome)
As such I do not believe that "Google search is basically Gemini now"
I believe Google search is still scanning through a doclist to find which documents, if any, contain words parsed from a query. These documents are pointed to by the URLs I get in the SERPs
I do not get any Gemini summation
A typical US user using a typical browser with typical settings gets an AI summary above search results.
If this doesn’t describe you, then ymmv. Talk to your government or turn down your content filtering or reset the default settings in your browser, if you want to see what we see.
>As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment
This is EXACTLY what people like/are addicted to about chatbots.
My sister-in-law bombed an interview and asked AI about her answers to the interviewer's questions, chatgpt or whatever it was told her that her answers weren't bad, but that the interviewer could not see the gold in her responses. She said she felt much better.
I see this effect with all the non-tech people in my life
The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefits accumulated. With models right now, there is just no way to do distributed training, or really, any large scale parallel way to improve them.
So whatever the short term strategy driving publicizing the model weights (e.g. potentially, to create a price war in order to put pressure on western companies and deprive them of the money they need), we can't ignore the fact that incentives and decisions could easily change in the future, and unless there is a way to truly decentralize models improvements - the party could stop at any time.
But different huge companies have different incentives. It is very much in Nvidia’s interest to have me running a powerful open source model on a $4k machine that they sell me.
Is it? When they could be having you running an even more powerful model on a $50k machine they sell by the pallet-load to enterprise consumers? We already see RAM manufacturers abandoning the low-end market in favor of server support. It's not clear to me that Nvidia sees personal GPUs as their best long term investment compared to selling millions of server-farm class machines
You mean a $500k machine, or a $15M rack... the costs have gotten unimaginably large from the lens of just a decade ago.
Selling to individuals can be a hugely more robust predictable business, the problem with selling by the pallet load is spiky revenue that can also quickly fall off a cliff if larger customers stop buying. The other problem is sales negotiations driving down margins for bulk buyers etc. Consumer hardware is a very attractive market in lots of ways, just look at Apple.
Right. It's probably important to distinguish what the hardware manufacturers' incentives are when their supply is constrained and when it isn't. (I am not an expert on anything.) As long as supply is strongly constrained they're naturally going to sell to the highest bidders, which are the big LLM SaaS players. If and when supply is no longer constrained, though, things will look very different:
1) The LLM SaaS companies are a form of vertical disintegration for the hardware providers, a middleman covering costs and taking profits out of the money that comes from customers to the hardware providers. That changes somewhat if there are no longer good models available for local use at no cost to the hardware guys, but only somewhat
2) The LLM SaaS companies are efficient users of their hardware resources. While supply is constrained this helps to make them top bidders and so attractive customers for the hardware manufacturers. When supply is not constrained this should reverse. Which is the more attractive class of customer to a hardware maker: the company full of people with higher degrees who spend their whole working day fighting to pare back resource usage, or the guy who leaves his laptop idle about 18 hours per day on average?
It's notable that nVidia, for instance, has continued to put significant emphasis on AI compact desktops and laptops. And while no doubt that's partly in the service of better developer relations and good PR in general, it's probably also nVidia eyeing the exit, and preparing for a future transition from selling shovels to the army to selling shovels at Walmart. But of course the future isn't clear and obvious. If the hardware makers, maybe the RAM guys in particular, turn out to have underbuilt future capacity starting in the present then we could be stuck in constrained supply for quite a long time. (Futher) government action could affect things etc. etc. And if the frontier labs soon find new ways to use still larger amounts of memory, GPU capacity etc. that isn't butting up against diminishing returns then they'll likely remain kings for some time, though that does not seem probable now.
They might not release such a thing, lest it puts all of their customers out of business.
Well that's kind of the point of the article. That in order to "win", the US needs an incentive structure that encourages open models.
I'm not sure what that looks like though.
Why do people always bring up state support when it comes to China? As if the U.S. doesn't provide massive tax breaks and explicit funding to industry?
It's on every tech post about China, as if it gives them some sort of "unfair" advantage.
Because the scale of the subsidization is unimaginably different.
yep. the ppp loans alone dwarf any state subsidies any other country has ever done
> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."
There is more to society than capitalism.
> Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."
> There is more to society than capitalism.
I don't read GP like that. I read it as "we should recognize a situation of unstable incentives for an important outcome, and start thinking about other solutions."
This is precisely why we need more projects like this https://github.com/bigscience-workshop/petals
There's no fundamental reason why models couldn't be developed and trained using community efforts. It might not be as fast and efficient, but it's definitely possible.
I am also confused by this point. The American government could force OpenAI and Anthropic to open their models, but then they would instantly evaporate, right? It doesn't seem like a choice that they can make, so framing it as a "winning" strategy doesn't make any sense to me. In what world could those companies have existed and opened their models?
Why did Google give Android away for free?
But they could do that because it didn’t cost them hundreds of billions of dollars to create Android and they could retain 99% of control over end users.
They could, but they don’t have to. The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.
The one thing that is sort of ironic or bad is that between Russia and the Ukraine there’s a large number of mathematically inclined people that if it wasn’t for the Putin war, their brain power working on AI models would have probably pushed open source down the road, even faster…
> The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.
This reads just like "AGI is 2 years away", I'll go set my calendar...
But I still don't get it. Like China could be come the world's leading producer of chocolate...if they started giving away chocolate for free. Would we be having this conversation saying that Switzerland lost because they were greedy and protectionist and didn't decide to give away their chocolate for free first (I realize Switzerland probably isn't actually the world's top producer of chocolate).
Long history of open source projects already have answers on how to monetize a free and complex open source product.
- Low development cost: collaborative efforts from open source contributors, innovative model training and serving for llm (Chinese models costs a fraction to train and their local chip design and manufacturing are catching up, plus cheap electricity)
- monetizing by selling hosted services, while leaving the core product free to tinker with / self host. China’s gdp is 2/3 of the US and it’s already a huge market for AI - which OAI and A\ don’t enter.
- for (the US) market that they can’t enter, let the US cloud providers to do free marketing / advocacy for them. Gaining share of mind. It costs them nothing.
The idea that Chinese models cost less to train seems to be based on that one time DeepSeek estimated the training cost for their V3 model at GPU rental rates as $5 million, and comparing this to other companies' entire R&D budgets. Yet DeepSeek raised $7 billion of fresh money last month, enough to train more than 1000 such models. What gives?
- You need to train lots of experimental models to dial in the training process just right for the one model that actually gets released in the end. Fortunately, these can be smaller.
- However, everyone is training much bigger models now, and doing a lot of RL rollouts on top.
- You can't get the GPUs for this piecemeal at rental rates because they need to be wired together using high-bandwidth interconnects.
- Nvidia GPUs are much more expensive in China, and local alternatives are still immature and not as efficient. Some companies have gotten around this using data centers in Singapore, which should tell you that electricity prices are not the primary consideration.
- The one line item where Chinese companies can probably save quite a bit of money is salaries for rank-and-file researchers.
In any case, they need to make back that money somehow. Giving away freebies isn't going to cut it.
In Russian opposition's mostly liberal discussions their school of thought connects several things together (sorry for not going directly to Marx's "General Intellect" and "Fragment on Machines" and using AI summaries instead ) - general idea of communism in China vs. techno-libertarianism of Thiel, Musk and the likes, and the Marx's thinking like:
"Fragment on Machines":
"he explores how human knowledge and collective intellect become embedded into machines, divorcing the worker from their own creativity."
"General Intellect":
"These texts are widely discussed for his concept of the General Intellect—the idea that society's shared, collective knowledge increasingly drives production rather than raw manual labor, and that this knowledge is alienated from workers and used as an instrument of capital."
(note: my point isn't to pass any political judgement here, like what real communism in China or not real, is it good or bad, i just find it interesting that pure political discussions by people with no technical credentials bring AI as a major factor today)
> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
Right... and there are two problems with this:
1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the underlying assumptions about compute and scale needed are mostly on the mark. So you can release open-weight models and they will have great use cases and applications, but ultimately similar to how you don't use an open-source phone or a budget Android phone from Wal-Mart and you buy an iPhone instead, you will see that although they "do the same thing" one product is clearly superior and you just have to pay for it. For this to not be true...
2. then it incentivizes most (all?) companies, American, Chinese, or European to halt development of models because if you spend all the CAPEX and it can just be copied and turned open-source nobody will invest in that. Given that China is not halting development of proprietary models I believe the current strategy and the subsequent approach to release open-weight models is at best a stall tactic, and at worse a sign of desperation.
Open source and the support and development models around it have been great. But folks are a little too dogmatic about it. Open-source software isn't a moral good, and closed-source software isn't a moral wrong either.
Imagine there’s a school where all the kids there are being tutored by the best. Also imagine a bunch of neighboring schools drastically falling behind that would need insane amounts of money to keep up.
This becomes a problem because all the kids from the rich school will dominate the order schools. They’ll get even more money as time goes on from their kids paying it forward to the point where all other kids are bound to work for them.
Now let’s say one other school does have the money for best tutors, BUT they know they’ll run out pretty quickly. Instead of trying to compete in a losing game, they decide to give every school in the world access to their elite lesson plan. Now, for a time, everyone will be on close to a level playing field. If the other schools improve upon their own lesson plans and keep sharing them with others, one day the elite school will wake up to find they are no longer on top. The parents have started to move their kids to other schools because the rich school is no longer attractive at the high cost they charge students
Sure and to complete your analogy here, the rich schools realize that the curriculum they develop and put a lot of time and money into creating is just used by the cheap schools, so they stop developing it because nobody loses money for long and so neither the rich or cheap schools develop any new curriculum.
Now what?
The fundamental problem here is incentives and tactics. Either the models are actually better (which I think the iPhone to cheap Android phone really speaks to, i.e. they do the same thing but one is 50x better at 5x-10x the cost) and thus they can be gate kept and like the iPhone the vast majority of profits go to a select few with high end implementations. OR the models aren't actually that much better, companies lose a fortune and then nobody can create any better commercial models or build out scale needed for open source models because it's not profitable.
We could wind up with only open-source models or something along those lines, but if the compute and scale is needed to train the models, nobody will be able to do that profitably and so AI research is either gate kept and silo'd for something like military applications or it just doesn't really happen because there's no funding for this scale of build out.
iphone is 50x better? because you get a blue box around your text rather than green?
androids and iphones are approximately the same thing
the kinda obvious direction LLM training can go is into the direction of particle physics, and the training is set up democratically and through universities and via multi-state funding
then the resulting weights end up open, the same as the particle detection data
> androids and iphones are approximately the same thing
Yet...
Android holds 70.6% of global active devices to iOS at 28.7%, but iOS captures 64.2% of consumer app spend. [1]
> and the training is set up democratically and through universities and via multi-state fundingPossible, certainly. But this case also applies to China and its "open-weights" strategy. They won't be able to form companies either or get ahead.
[1] https://www.digitalapplied.com/blog/mobile-os-market-share-2...
Try selecting a file in your 50x iphone - triple copy of the same file or ateast double copy (assuming os will post a soft link). If a large file then you are toast.
Try mmapping > 5GB file in your 50x better iPhone.
Try running any service in the background.
The list goes on and on.
Your 50x better suddenly became 50x worse compared to a much cheaper android.
> so they stop developing it because nobody loses money for long and so neither the rich or cheap schools develop any new curriculum
Why do you assume the poor schools wouldn't be smart enough to keep it going? It's very likely the can collectively beat the rich school now that the one other rich school opened access to their materials and led the charge.
> but if the compute and scale is needed to train the models, nobody will be able to do that profitably
But they would. Efficiently hosting models will be the real business and early access to models with incremental improvements will not be the moat once thought. The reason other companies don't feel they can compete is the same reason OAI and Anthropic will lose their lead. They banked too heavily on another player NOT leading the charge on open research and poured disgusting amounts of money at closed source models.
China has proved they can take the limited resources available to them and build something better than what the US is offering consumers [1]. I'm just waiting for other countries to start pitching in.
Reminds me of the NSA and their early battles with cryptographers who believed in open research.
Software being open source has many strong positive externalities. It advances human knowledge and freedom. If you don't think that counts as a moral good then I'm baffled by what you think a moral good is.
It depends on how it is applied. You can release open source software that advances knowledge and freedom that results in economic destruction or the loss of life, for example.
Open source is in the tradition of humans sharing past knowledge, long-term we just can’t keep a secret it’s a time, honored tradition…
> Open-source software isn't a moral good
Yes it is.
Prove it
I’ll take a swing at it. Open Source is a form of sharing with the wider community. Closed Source is not sharing. Moral good is based on doing good outside of your own benefit (the opposite of selfishness.)
Ergo it’s a kind of moral good.
And I’m not even an advocate for open source.
This argument boils down to X is good, therefore more of X is good. But you can see how that breaks down with even trivial examples. Not a great argument.
The second piece of this "a moral good is based on doing good outside of your own benefit" - says who? Why? This logic is also faulty. You're also cargo-cutting self-interest in here as a moral failure when many good things depend on humans acting in their own self interest. For example I completely and selfishly installed a new tree at my house. But the community benefits from carbon capture, shade, &c.
I understand the sentiment you have here and I think for everyday use and having some guiding principles it is probably fine, but don't confuse this for a principle that is actually examined. You can find contradictions rather easily, never mind solid arguments which expose cases where what you think is true is not really true and so forth.
Thanks for the discussion.
>This argument boils down to X is good, therefore more of X is good.
No, I only argued that it was a moral good, the kind of good. I actually may disagree with others about whether you should pursue a good just because it’s good.
>says who? Why?
Good question, it’s just a common framing that I see in classical discussions. I didn’t intend for it to be exclusive, I think there’s moral good outside of that.
>don't confuse this for a principle that is actually examined
I hear you, I think this is a simplified version suitable for an online comment. In particular I’m not saying that if you do something other than a moral good then you are doing something wrong. There are many actions that are morally neutral. Also it is possible to construct artificial situations where you may violate some moral good in pursuit of another.
It's important also that open-weight isn't open source. If you can't download the training data (fully labeled), source code of the NN, and follow the README to build and train it yourself assuming oyu had the hardware then it's not open source.
tldr there's no "source" in open weight models therefore they are not open source.
Exactly. AFAIK none of the popular "open Chinese" models have published the full pre- and post-training pipeline, so the models are only partially open, if at all (plus the openly shared final weights, of course).
Another important thing that made software usage and education available for most of the world was piracy. I remember as a kid growing up in a developing country, any software (windows, office, Visual Basic, flash, dreamweaver, etc.) was less than 1$. That allowed me to try out and learn so many things on my own without paying a huge amount of money for the license. And I think this is true for most of the software developers of my generation who grew up in developing countries
"I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now."
I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac.
The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today.
I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model.
I think the question is even a bit more nuanced than that. Even if frontier models can maintain a big gap that gap has to actually _matter_. If a local model satisfies my everyday use cases adequately then I may not really care that a frontier model is 5, 10, 100x better at ultra high order reasoning tasks.
I think that reality is probably not all that far off for a huge swath of use cases.
This is exactly the mainframe vs PC dynamic.
100%, I thought about writing that out explicitly. I really feel like we are extremely close to reaching that kind of breaking point for most folks LLM use cases.
Hell, Bonsai Labs 27B parameter model can run on phones with their ternary implementation which is quite efficient. Scale that up to frontier model parameters and it's quite likely we can run them on current laptops.
Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone
But isn't there the raw intelligence of a smart model and then the practical intelligence fuelled by how many parameters it has? You probably will barely be able to fit a 70 billion parameter model on a phone in 3-4 years let alone a 2+ trillion parameter model... so it depends on what you call intelligence
> I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now
10-15 years? The current rate is closer to 10-15 months.
15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.
Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index.
Here are a number of open weights models that you can run locally compared with the frontier class models from 7 to 15 months ago: https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4...
I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops.
Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024.
https://nitter.net/i/status/2079088670804767114
So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone.
Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.
> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.
There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors
12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.
Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.
Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).
Pretending progress hasn't been mindboggling is insane.
Maybe it got a lot less and I just got used to it. True.
Still feels too much for me. Breaks my workflow for no reason. Too much overhead for me, if I can't trust the output
No. We need objectively around 192 to 512gb of very fast memory to be able to run really useful models. I don't see local hardware with these specs coming in 1 to 2 years. There are a big number of initiatives currently taking place to increase ram output. But it will take another 3 years minimum to close the current supply issues. China is fast pacing forward to have its own chip baking factories with small enough nano scales to have fast chips. Will also take a few years.
> 10-15 years? The current rate is closer to 10-15 months.
The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.
The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.5 and GPT-4. And now in just over one month, major models already included Claude Fable 5, Claude Sonnet 5, the GPT-5.6 series, Kimi K3, GLM 5.2, Qwen 3.8 Max, Grok 4.5... and the official DeepSeek V4 release is coming soon.
Iteration speed is now measured in days.
> - PCs destroyed minicomputers.
What's weird is that with "store your everything in the cloud and pay a monthly recurring subscription", we have now regressed to a 1960s/1970s timesharing revenue model for individual workstation computers.
The default new factory out of box workflow for "enrollment" in google services, iCloud or Microsoft-everything on a new ios, macos, windows or android personal computing device is clearly designed to sign people up for subscriptions.
And same general idea of "move all your servers to the cloud" recurring revenue for what is effectively the same as mainframe timesharing for key business functions, by renting VMs in GCP, Azure, AWS in perpetuity.
Yes, you can still use your desktop or laptop PC in 2026 with zero external third party subscriptions (other than maybe your residential home ISP), but how many non-tech people actually do so now?
I do agree that Chinese open-source models are going to play a bigger and bigger role in the entire ecosystem moving forward, but I don't agree with you in the sense that they are going to eventually "win."
Just because they are cheapp doesn't mean they automatically win. You've picked a lot of great examples, but there is still a little bit of cherry-picking.
One clear outlier is the iPhone, which coexists with Android globally. Even though the iPhone is the leader in the US, and globally Android has the majority of the smartphone market share, they still cater to different price points and different ecosystems, and generally the iPhone has better margins.
i believe American frontier models like from Anthropic and OpenAI are still going to thrive, and coexist with Chinese models. They are just going to cater to different customers and different use cases.
Yeah, iPhones are just fashion statements in the US. Different than LLMs
They also just work better and aren't substantially priced different to the equivalent android option.
100x this it is why all of the AI giants are going to fail. They are too big and inefficient to scale properly. This is why Google is just casually taking its time in AI and not racing to a finish line. AI is essential but if it already does most things good enough then it can take longer to make it more efficient.
> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.
Phones are constrained by battery power and memory does not shrink as fast as CPU/GPU, so unless there's a battery breakthrough and/or memory breakthrough, you're not fitting 100Gb of RAM on your phone in 10 years.
Absolutely in a Mac Studio equivalent.
LLMs have emergent capabilities when they get smarter. So who knows how insanely big frontier models might be at that time, or what their capabilities may be.
Not just that memory shrinks slower, it has practically completely stalled. On chip cache seems stuck at 7nm and DRAM is stuck at 10nm. As transistors shrink, they hold less charge, creating weaker signals that are harder to read and prone to interference. Smaller nodes aren't a huge issue on CPU/GPU work load because they don't have to hold a static state.
I'm not saying we are at peak memory but future gains are going to come increasingly slower.
What's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for.
Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.
Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act.
These models might be smart but they're not close to being able to savor irony.
I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.
I asked Gemini to generate a picture of Peter Pan and Wendy (for a workbook I am putting together for youth summer reading) and it preceded to refuse due to copyright. Not everything about that work is owned by Disney. Thankfully the JM Barrie original artwork is public domain and available (and fantastic btw), so I used that instead.
I don’t know if you know this but Peter Pan’s copyright is weird in the UK. There is a legislated exception in the law that it never expires and the royalties will forever go to a specific children hospital. Here are the actual words: https://www.legislation.gov.uk/ukpga/1988/48/part/VII/crossh...
(Now i don’t think you are necessarily in the UK. Just wanted to explain that Disney is not the only reason an AI might be trained to thread carefully around copyright issues of Peter Pan.)
You can also try asking Gemini to create art that is "as close as possible without infringing" - I've had success with that.
I used Claude to build a complete data extraction pipeline for a popular current best seller book series: audiobook -> text (via whisper) -> local LLM (qwen) -> database. Not once did it seem to acknowledge or care about copyright. It even used knowledge it already had about the books to exclude certain ones before beginning since the character I was interested in did not appear in those. It definitely had context of what we were working on.
Why would you go from audiobook to text? Is there no epub available?
Not without DRM. It was easier to buy the audiobooks and use the analog loophole to get text. It's probably less accurate, but for what I'm doing it was fine. Names were the worst, but whisper at least made the same mistake each time so a simple search+replace handled most of the obvious edge cases.
If you are going to be illegal, you might as well use library genesis and get DRM free ebooks :)
On the other hand, it’s a beautiful example of the abilities LLMs have bestowed upon us, where it’s easier for a guy to transcribe audiobooks then to use a website to quickly download an epub