The company that made FLUX the king of image generation just pointed the same architecture at robot arms.
What it actually is
FLUX 3 x mimic is a video-action model. Black Forest Labs took the FLUX 3 multimodal backbone and fused it with robotics startup mimic’s manipulation learning. The result is one model that watches a scene, predicts what happens if the arm moves a certain way, then acts — reacting in about 101 milliseconds, roughly human reflex speed. Training data: tens of millions of hours of general video plus hundreds of thousands of hours of human and robot manipulation footage.
The point is generalization. Conventional factory automation can’t touch soft, floppy stuff — seals, cables, tucking an ECU into a tight fixture. FLUX-mimic handles those, and some tasks fine-tune on just 30 minutes of robot data. Audi is running it on a real production line.
The API angle
FLUX 3 Video and FLUX 3 Action are in early access now. BFL says an API, private weights, and an open-source FLUX 3 Dev land before year-end. That means image, video, and physical action from one unified stack — the same team you already trust for pixels, now closing the loop into the real world.
You Might Also Like
- Flux 3 Black Forest Labs Beats Runway gen 4 5 in 77 of Video Preference Tests
- Guide Labs Steerling 8b Finally Lets you see Inside the Black box
- Arc agi 3 Turns ai Testing Into a Video Game and Every Frontier Model is Losing
- Meta Model api Muse Spark 1 1 Undercuts Openai and Anthropic at 1 25 4 25 per Million Tokens
- Sap Completes e1b Prior Labs Acquisition Tabular Foundation Model Tabpfn Enters the big Leagues

Leave a comment