Black Forest Labs' new model generates video and audio together, and it wants to control robots next

Started by HighKey15, Jul 26, 2026, 04:50 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Black Forest Labs' new model generates video and audio together, and it wants to control robots next   Views(Read 59 times)

HighKey15

German AI lab Black Forest Labs unveiled FLUX 3, a multimodal foundation model that jointly learns from images, video and audio inside a single unified architecture rather than stitching together separate specialized models. Its headline feature is generating up to 20 seconds of video with native, synchronized audio, meaning dialogue, sound effects and background music all come from the same model in one inference pass rather than a separate audio pipeline bolted on afterward

According to Black Forest Labs' own benchmark comparisons, FLUX 3 was preferred over Luma Ray 3.2 in 93% of head to head comparisons and over Runway Gen-4.5 in 77%, though results were closer against competitors like Seedance 2.0 and Gemini Omni Flash, and none of these figures have been independently verified yet. FLUX 3 Video is available now in early access, with an image focused version and an enterprise robotics focused version called FLUX 3 Action following in the coming weeks

The company's ambitions reach beyond content generation, with co-founder Robin Rombach framing the model around perceiving the world, predicting how it changes, taking action and learning from results, positioning FLUX 3 Action as a foundation for robot control rather than just video. An open weight version called FLUX 3 Dev is planned for later this year, continuing Black Forest Labs' pattern since its 2024 launch of pairing frontier capability with genuinely open access, already powering generative features inside Adobe Photoshop, Picsart and Nous Research's Hermes Agent
Come on Wales!

Delulu

Generating audio and video in the same inference pass rather than bolting on a separate audio model afterward is a different architecture, not just a marketing angle
VAR can do one

WearyScholar

The jump from video generation to robot action prediction is a much bigger ambition than most people probably expect from a company known mostly for image models

Electric Holly

Self reported benchmarks favoring your own model this heavily always deserve a healthy amount of skepticism until independent testing happens

Rosie96

Black Forest Labs committing to eventually open weighting this is notable given how capable the closed early access version already sounds
My model is always one epoch away from greatness

Slay

20 seconds might not sound like much but native synced audio the whole way through is a real technical jump over post-processed audio pipelines

WarMachine62

Curious how FLUX 3 Action actually performs on real robot control tasks compared to companies that have been specializing in that exact problem for years

Daemon82

Really interesting to watch a lab try to unify creative content generation and physical world action prediction under one single architecture

Related Topics (3)

Save money on everyday spending Free cashback on thousands of retailers
View offer