Why isn't the industry freaking out about DeepSeek 4.1 Flash?

dgt.is

200 points by jonotime 21 hours ago


vishvananda - 20 minutes ago

The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.

This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

giancarlostoro - 2 hours ago

Call me crazy but:

VRAM & Memory Requirements by Precision

• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

Even so why would anyone not sleep on a model they cannot run?

mlinsey - an hour ago

I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).

I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).

Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.

user43928 - 17 minutes ago

Because DeepSeek is not "a month or two" behind as claimed in the article.

These open models still did not beat February's Mythos / Fable 5.

DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.

Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.

It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.

lmf4lol - 2 hours ago

Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.

Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.

But as a main driver. I love flash. And it brought our bill down by A LOT :D

p1necone - an hour ago

I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).

I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.

However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.

serial_dev - 12 minutes ago

I've used frontier models for the better part of a year, because employer said "cost is not an issue". Who could have guessed, now we have a strict token budget since October, so I started experimenting with cheap models at work. I of course experimented with cheap models on personal projects, but they are not the same as figuring things out based on multiple million+ LOC codebases...

Unfortunately, it looks like Cursor doesn't support DeepSeek... but I can say that for lots of tasks, the cheapest models do just fine, too. Of course there are cases where they just keep spinning for 30 minutes without being able to figure things out, but in those cases, I just retry on a slightly more expensive, hopefully better model.

Some of the models are so good and cheap, not sure about the frontier labs and their multi-trillion dollar evaluation. Sometimes "good enough" is really good enough.

arush15june - 26 minutes ago

I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.

I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.

Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.

And it never says no for cyber tasks so that's a big win

hmontazeri - 2 hours ago

I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it

apitman - 30 minutes ago

> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited

My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).

I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.

This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.

simpaticoder - 2 hours ago

The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.

The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.

alex-moon - an hour ago

I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.

zug_zug - 44 minutes ago

I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.

That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.

swiftcoder - 2 hours ago

I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them

wg0 - 2 hours ago

While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.

I realized that mistake and guided DeepSeek where it should be.

Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.

RGS1811 - an hour ago

This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.

I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.

pants2 - 9 minutes ago

Probably because Luna is faster, cheaper, and approximately as smart

aguilaair - an hour ago

What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.

see https://artificialanalysis.ai/models/releases/comparisons?co...

elmer2 - an hour ago

DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.

f6v - 35 minutes ago

My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.

jbellis - an hour ago

I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/

And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking

robertlane0 - 13 minutes ago

Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.

browningstreet - 2 hours ago

What would freaking out look like, or is this just a stupid bloggish title flourish?

Is OpenAI coming in $20B under a sign of "freaking out"?

wildster - an hour ago

I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md

smallmancontrov - 2 hours ago

They might be. They would delay public admission as long as possible, because public admission would make stocks go down.

liuliu - an hour ago

DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.

aszen - an hour ago

Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out

booi - 2 hours ago

Because GLM 5.3 Flash is even cheaper?

xyzsparetimexyz - an hour ago

There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.

pianopatrick - 2 hours ago

I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.

Would be cool if they added it.

gsky - an hour ago

America bans Chinese models sooner or later just the China banned American big tech

tengbretson - 2 hours ago

I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.

pizza234 - an hour ago

People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.

I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).

Local models are also really slow, unless one spends insane amounts of money.

Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).

hypfer - 2 hours ago

Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?

- an hour ago
[deleted]
kristianp - an hour ago

> shrank the KV cache by roughly 437X

Can't you just say "shrank to 1/437th the size"? It's not that hard.

MisterMunchkin - an hour ago

I had it make 25 different things today and it cost $0.70

It’s disgustingly good value. I find it capable of doing anything I want.

Obviously can’t use it at work, but for home projects it’s awesome.

cactusplant7374 - an hour ago

Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.

anguralbanish2 - 2 hours ago

I would love to get them more better, it's good not a bad thing.

pessimizer - an hour ago

I'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.

If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.

I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"

I actually feel like 5.6 Luna seemed better.

sergiotapia - an hour ago

In my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.

m3kw9 - 42 minutes ago

i thought 6.1sol copied the caching architecture so this isn't such a big deal no more

doctorpangloss - an hour ago

because it doesn't work very well?

if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...

AIblemblio - 2 hours ago

No they can't.

And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.

But yes i'm glad that we have alternatives.

verdverm - 20 hours ago

Why would we freak out? The systems we use have always gotten better, faster, cheaper with time

kydanet - an hour ago

[flagged]

CurbStomper4 - 34 minutes ago

[dead]

distantsounds - an hour ago

because we've all figured out that AI is just a huge grift?

sroussey - 2 hours ago

Not comparing to gpt-6-luna which seems comparable and priced well.

wewewedxfgdf - an hour ago

You might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.

thefourthchime - 2 hours ago

For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.

Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5

Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...