DeepSeek-V4-Flash Update

api-docs.deepseek.com

637 points by dnhkng 15 hours ago


NitpickLawyer - 14 hours ago

This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.

DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.

Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.

f311a - 13 hours ago

I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast.

I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The rest 10% is to spot bugs, security problems and to investigate better architecture, which flash can also do pretty well, I just cross check it.

Faster iterations are way better for me, I hate waiting for 5-10 minutes on small changes. I tried to use recent versions of Kimi and GLM, but they use too much thinking for no reason and are pretty slow because of it. I also often feed a lot of data to it, without worrying about hitting the limits: dependencies (to find bottlenecks in them), logs, performance dumps and so on.

Also, it will never complain about security guards, I've been using it to reverse engineer binaries.

lionkor - 13 hours ago

I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days:

- Cost: $4.55USD

- API requests: 3,467

- Tokens: 323,183,886

And as an engineer who leads a small team, I have very high standards for quality, and these carry across to my personal projects where I use deepseek. It has not disappointed at all for coding or review tasks. For everything else, use another model.

kmarc - 12 hours ago

Essentially I'm running everything on flash now inside pi. With the correct set of MCP servers, context reducer tooling and skills it can implement any task I throw at it. Some sessions take 30+ turns, but it's fast and cheap; all this in an hour, with ~$0.5 cost.

(TBH though, in my multi-subagent workflow I do use other, more expensive models for planning, reviewing, oracle-ing)

I haven't used our slow opus subscription for weeks.

(Also set up an OpenWebUi self-hosted chat that works from my phone, has some mcp and skills. fully replaced perplexity. Monthly cost ~$18 for hosting and subscriptions)

wkcheng - 13 hours ago

If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the price decrease.

Crazy.

dannyw - 11 hours ago

Deepseek and moonshot are the only two providers I consent to training for.

ggcr - 14 hours ago

Woah, a 200B model competing with GLM-5.2 and getting close to Opus 4.8. Quite impressive.

If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.

heyalexej - an hour ago

Long story, I have humongous zai GLM 5.2 token budget that I'm using in a similar fashion as many comments explain here. GPT 5.6 or Fable 5 for planning, GLM for implementing, researching, extracting and many other tasks I consider grunt work. Very happy with the performance, speed isn't all that good though. I'd be curious to hear from someone who works with both, DeepSeek and GLM side by side.

Goranek - 13 hours ago

Kimi K3 (instead of Opus) for expensive stuff, DSV4 Flash for tasks (instead of Sonnet)?

Does this make sense?

thirtygeo - 13 hours ago

For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning falsified logic or information cooked in by the developers, rather than the training data speaking for itself

baalimago - 13 hours ago

Very promising. So it will both keep the speed and reduced price, yet exceed performance of the quite sufficient deepseek-v4-pro?

Should be extending the lead in intelligence/cost index, as deepseek-v4-flash already were the most price efficient model, which now becomes even better. Although, in the deepseek APIs, the cost is leaking all information about codebases to China.

wg0 - 3 hours ago

Don't know about the bench marks but I am getting Opus 4.7 level performance at fraction of cost with DeepSeek V4 Flash set to high. It is a reliable workhorse.

sim04ful - 12 hours ago

Why didn't they increment the version as atleast a patch update

namuol - 3 hours ago

Can someone please explain how these models aren’t just fine tuned for benchmarks? I’m not plugged in to this space much but it seems like such an obvious problem…

yewenjie - 10 hours ago

What was preventing them from calling it v4.1-Flash to distinguish it better?

throwa356262 - 8 hours ago

Gentlemen, start your DGX Sparks

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

troglodytetrain - 2 hours ago

This is very exciting, my own niche micro-saas has already been able to make heavy use of DeepSeek-V4-Flash for my use case, I am looking forward to seeing how performance improves.

alecsm - 8 hours ago

I've been using DeepSeek Pro for a while and Flash only for certain dumb tasks where I only need the speed of a LLM and not big brains.

I find the newest OpenAI and Anthropic models to be way better for big tasks that require many decisions but I don't like that anyway because I lose track of what's being done.

Knowing what I want for every prompt makes DeepSeek Pro the best LLM for me. It allows me to work relatively fast at a very low price.

sqemo - 11 hours ago

DeepSeek is great for tasks and software I already know well. Even if it gets something wrong, I can usually verify it myself. But when I'm working with a programming language I'm not familiar with, I prefer using Codex or Claude.

arjie - 13 hours ago

Oh my goodness what an update. I need these weights. It's an incredible model for the size. The improved tool calling etc. should be able to make my harness way simpler. This runs at mega-speed on prosumer hardware (2x RTX Pro 6000).

mordae - 12 hours ago

I was just using it when it landed. It started reasoning more extensively from nowhere and precision went up a lot. It also changed its prose style for the better. Looking forward to weights.

nickandbro - 14 hours ago

I wouldn't doubt GPT 4.6 Luna being in the top left quadrant's center on the Cost per Intelligence Index is not concerning for Liang Wenfeng. You have to remember DeepSeek v4 flash even though a bit cheaper, does not have vision abilities, which is a big draw for agentic tasks.

I admire DeepSeek's openness, but even they have been raising prices after their discounts.

amelius - 10 hours ago

Note: if you are having success with a model, then please post what you are using it for. Writing HTML/CSS is very different from writing Rust/C++ or doing maths.

Reubend - 14 hours ago

They're always very understated in their update descriptions. This is actually a HUGE improvement in the model's capabilities rather than just a small tweak.

f6v - 12 hours ago

I'm thinking of using ChatGPT for making plans and V4-Flash for execution. Does anyone have good advice on pairing different models?

ilaksh - 9 hours ago

I wonder when the antirez/ds4 group will have an update to their high accuracy 2 bit quant.

Although it's funny that I am thinking about that at all because I have a 2060 :P . My local inference is playing with Gemma 4 E2B and MiniCPM 5 1B.

Tepix - 10 hours ago

Sounds like a big improvement.

No mention of weights, just API. When will the updated weights be released?

miyuru - 13 hours ago

Judging by the openrouter leaderboard ranking for today, it looks like Dv4F us more popular than mimov2.5.

https://openrouter.ai/rankings?view=day#leaderboard-table

These days cost per task is more important, and SOTA models have become expensive.

KronisLV - 13 hours ago

Wonder how good the proper version of V4 Pro will be.

I'm still considering pulling the trigger on the annual subscription of Kimi for K3 but it's sometimes slower than I'd like (at least when compared to Anthropic) even on their Vivace plan, and the token limits on the GLM Coding subscription for GLM 5.2 were too easy to hit.

HyperL0gi - 9 hours ago

Is anyone using DSv4 for their agents that is not related to writing code? Curious about use cases specially for someone using gpt-5.4 mini for classification, categorization, etc

vladukha - 12 hours ago

Where do you guys get deepseek? I'm hearing a lot of good reviews and want to try it with my pi config. from the deeepseek themselves, openrouter, or anywhere else? does it make a difference? [edit]: whoa it is really fast. will take some time to evaluate quality thou

wolttam - 14 hours ago

Hooray! This model makes me very optimistic about the future of local inference. The CyberGym score stands out to me.

egeozcan - 13 hours ago

Every time I want to have fun coding something with natural language processing, I use deepseek flash. It's just incredible for the price. I have a fairly popular app with 400 users that uses DeepSeek in the background and it still didn't hit even 50 bucks of usage in a month.

PhilippGille - 13 hours ago

The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

Why not call it V4.1?

w2seraph - 2 hours ago

This made my day !

flysoft - 13 hours ago

Finally have a model with usable intelligence, at a reasonable price. Can't imagine what Pro GA would look like, considering pro preview has only 1.6t parameters.

throwaw12 - 12 hours ago

how different is their harness from Pi coding agent harness, is it possible to make an extension for Pi which can implement deepseek harness?

storywatch - 13 hours ago

How's their performance in English prose? We are currently searching for cost effective ways to keep story wikis up to date.

nathaah3 - 12 hours ago

DS v4 flash has been my goto model for tasks in work. its been unsurprisingly fast and cheap.

k__ - 11 hours ago

I'd take more throughput while everything else stays the same.

kamikazechaser - 13 hours ago

The flash variant is on par with Sonnet 5 on DeepSWE (54%). Big, if true.

znnajdla - 11 hours ago

The conspiracy theorist in me wants to think that the 80% drop in GPT 5.6 Luna prices today is correlated with this update from DeepSeek. Perhaps OpenAI has already hacked its competitors with it's Mythos-like models and is aware of what competitors are doing and is able to react in advance.

spwa4 - 14 hours ago

In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.

XCSme - 9 hours ago

Can't really use it now, without giving away your data:

> Trains: this provider may use prompts for training and may retain prompt data.

ra - 13 hours ago

What's the best way to run this on a 64GB M2 Pro?

sparse-Matrix - 10 hours ago

This may come as a surprise to a lot of AI concerns, but I have -zero- interest in paying for a model.

sreekanth850 - 8 hours ago

how this compare to luna high with reduced pricing.

mekky16 - 11 hours ago

if they were anthropic they wouldve just released it as a new model

- 14 hours ago
[deleted]
truth_seeker - 11 hours ago

The magic of post training with valuable dataset

sourcecodeplz - 12 hours ago

i've made a comparison between this and GPT Luna (recent %80 price drop)

https://x.com/SourceCodeplz/status/2083099712760987746

i prefer GPT-5.6 Luna honestly

dakolli - 11 hours ago

Chinese labs rushing to release models this week, because it's inevitable that Washington regulates Chinese models in the next 4 weeks. All the US AI leaders have been taking trips to Washington this week, what do you think they're there for..

Tepix - 10 hours ago

[dead]

tosh - 13 hours ago

[dead]

dnhkng - 14 hours ago

DeepSeek V4 Flash (Preview → 2026-07-31)

• Terminal Bench: 56.9 → 82.7 (+25.8)

• Toolathlon: 51.8 → 70.3 (+18.5)

Compared to GPT-5.6 Terra:

• Terminal Bench: Flash 82.7 vs Terra 78.4

• Toolathlon: Flash 70.3 vs Terra 53.1

• DeepSWE: Flash 54.4 vs Terra 69.6

• Agents' Last Exam: Flash 25.2 vs Terra 50.4

Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!

try-working - 13 hours ago

Let's see how the market reacts.

Havoc - 12 hours ago

>benchmark results far exceeding V4-Pro-Preview:

Wow that's crazy

Good times for those that don't need strict data protection