Black Forest Labs, the team behind the FLUX image models, just shipped FLUX 3 — and it’s no longer just an image generator. It’s one model that produces image, video, and audio from a single set of weights, jointly trained instead of three separate models bolted together behind one API.
What it actually does
Text-to-video, image-to-video, video-to-video, up to 20 seconds with native audio — including multilingual dialogue and keyframe transitions. In head-to-head preference tests it beat Runway Gen-4.5 (77%), Luma Ray 3.2 (93%), and Grok Imagine (69%). Same weights also predict robot actions, via a partnership with Mimic Robotics.
Why it matters
This is the open-camp’s first unified image-plus-video-plus-audio foundation model, and it topped HackerNews the day it dropped. Rollout is staged: video is in early access now, image lands in the coming weeks, and API access plus private and open weights (FLUX 3 Dev) arrive over the next few months. The API is where developers will plug generation into their own apps — think one call for a captioned, spoken-dialogue clip instead of stitching a video model to a separate TTS system.
You Might Also Like
- Runway gen 4 5 Just Took the top Spot in ai Video and its not Even Close
- Guide Labs Steerling 8b Finally Lets you see Inside the Black box
- Grok Imagine 1 on Designarena Across all Three Video Arenas Beating Sora 2 pro and veo 3 1
- Warp oz Just Dropped and its Exactly What dev Teams Have Been Missing
- Nessie Labs Turns Your Messy ai Chat History Into an Actual Second Brain

Leave a comment