Flux 3

bfl.ai

403 points by ThouYS 7 hours ago


forgotusername6 - 8 minutes ago

Is anyone feeding models touch data? It seems the main thing we want the robots to do is touch things, but we are just feeding them audio/video/images. The model has to learn how to touch things despite never having touched anything before. Perhaps that's why they all look so hesitant when they touch things?

user43928 - 6 hours ago

I hope the open-weight versions will be SOTA.

> Over the next few weeks and months, we will make the following capabilities available

> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

> We will also release more technical details on the underlying approach.

thisisauserid - 6 hours ago

- Showed close to zero examples of people.

- Frivolous use of the term World Model.

- Claims 20 seconds of video, shows only jumpcuts.

Coming soon!

jdthedisciple - 5 hours ago

It's incredible how negative and dismissive the comments in here are while here I am thinking the model actually looks impressively capable.

But then again I heard the downers have always been the first to leave their dung comments here so let's see...

make_it_sure - 5 hours ago

first AI thing coming from Europe that gives high hopes

tormeh - 4 hours ago

These people are hiring in... Freiburg im Breisgau? Wonder how hiring is working out for them there.

OldMatey - 4 hours ago

I am very excited for this. 2.3 was excellent and I've seen clips from people who got early access along with reading their reflections on it and I think this will be the new SOTA for home use.

Tenoke - 5 hours ago

Flux 2 Dev Klein has practically been the best you could use on most commercial hardware so I really hope Flux 3 has a comparable updated open-weights model to it. if not it'd be a great loss to most hobbyists.

Gecko4072 - 5 hours ago

I thought the clips were real footage until they were dancing in a flooded room.

pwillia7 - 2 hours ago

Awesome -- glad they're going to release the open weight version! I've been waiting for an excuse to re jump into AI OS image gen!

ex-aws-dude - an hour ago

If you think about it why is language/image even separate from video?

Isn’t video + audio all you need?

- 3 hours ago
[deleted]
NSUserDefaults - 4 hours ago

> a model must learn a representation of the world: […] and how events sound

I honestly hope they put an unrealistic amount of wilhelm scream into the learning process, just for fun.

yangcheng - 2 hours ago

I hope flux will include a 3D generation model. right now the open-weights version of 3D is failing behind closed source by a big margin. Hopefully the improved spatial ability helps with robotics too

zmmmmm - 6 hours ago

> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.

I'm confused, videos contain images and audio ...?

abdusco - 5 hours ago

I wonder if this also creates people with huge heads and short necks like Flux Klein does.

rekpero - 5 hours ago

I have a feeling open-weight models ought to be outperforming proprietary ones by now, but that still hasn’t happened. So far, Nano Banana and GPT-2 Image seem to be the best in class, and Flux still isn’t crossing that quality bar.

mattmanser - 6 hours ago

Open-weight plans are near the bottom (Launch section):

    - Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
    - Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
    - Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
    - Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
teiferer - 6 hours ago

Lots of words about multi-modal but then this:

> our mission to develop real-world visual intelligence

Visual is mono-modal, isn't it?

saejox - 5 hours ago

i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)

SubiculumCode - 6 hours ago

Well, unified multimodal intelligence is the only way we will get to The Terminator, which seems to be the goal now of Silicon Valley and every Nation State with a military budget, so have at.

frotaur - 6 hours ago

Sorry because pointing this is a bit tired by now, but reading already the first two paragraph thete is this unmistakable stench of LLM slop writing. Immediately disengaged.

- 3 hours ago
[deleted]
vouaobrasil - 6 hours ago

The fact that people keep developing this technology shows that the true problem is not that machines are likely to become intelligent, but that people have already become machines - unthinking and without any care to the future whatsoever.

bkingfilm - 44 minutes ago

[flagged]

vladsiu - 5 hours ago

[dead]

doitright99 - 5 hours ago

AI slop trained on copyrighted content.

luciana1u - 6 hours ago

imagine spending nine figures training a model to learn that the sound has to match the impact. my 8-month-old figured that out by dropping a spoon on the floor twice.