Black Forest Labs, the company behind the FLUX line of image-generation models, has unveiled FLUX 3, a new multimodal AI system that pushes well beyond the still pictures the firm is known for.
According to Crypto Briefing, the launch positions FLUX 3 as a multimodal model with both video generation and robotics capabilities rolled into a single system. Decrypt frames the shift bluntly, saying the company "ditches stills for video—and robot hands," signaling a move from generating individual images toward generating motion and controlling physical machines.
The company's own announcement, posted on its blog and shared on Hacker News, carries the name "Flux 3 X Mimic" and describes it as "the next generation of video-action models"—a phrase that ties together the two headline abilities: producing video and driving actions. On Hacker News, the post drew 123 points and 11 comments.
techi.com highlights the same core idea, reporting that FLUX 3 "folds robotics into one AI model," suggesting the appeal is consolidation—handling perception, video, and action within a single system rather than stitching together separate tools.
Beyond these points, the source items are largely headlines and do not detail benchmarks, pricing, availability, or the specific hardware the robotics features are meant to run on, so those aspects remain unclear from the material at hand.
Why it matters: a company best known for text-to-image AI reaching into video and robot control reflects a broader industry push to unify what were once separate AI specialties, and to connect generated content with machines that act in the physical world.