The Video Model That Drives Robots: Black Forest Labs Open-Sources FLUX 3 Action
Black Forest Labs ships FLUX 3 Action, a 7B open-weights world action model that tops the RoboLab-120 leaderboard ahead of Nvidia's Cosmos 3 Nano — at 44% of its size.
Black Forest Labs (BFL) spent its first two years becoming the default name in open-weight image generation. This week it walked into a different room entirely. On September 22–23, the company quietly published the full FLUX 3 Action collection to Hugging Face: a 7-billion-parameter “world action model” that takes camera frames, a robot’s state, and a text instruction, and returns the next chunk of motor actions — denoised jointly with predicted future video frames. The weights, the fine-tuning recipes, and ready-to-run policies for two standard robot platforms are all public, and the model has already taken the top spot on the RoboLab-120 benchmark ahead of Nvidia’s much larger Cosmos 3 Nano.
What exactly shipped
The release is a collection of three repositories rather than a single checkpoint:
- flux-3-action-base — the adaptation backbone: 457 tensors covering the video and text streams, a shared action trunk, and two action heads (ee50 and gaming). BFL is explicit that this is not a complete robot policy; it is the component you adapt to a new embodiment, plus the frozen shared encoders — a video VAE and an unmodified Qwen3-VL-4B-Instruct text encoder carried over under its original Apache-2.0 terms.
- flux-3-action-so101 — a ready policy for the SO-101, the open-source community robot arm, trained on the SO-101 episodes of
lerobot/community_dataset_v3. Its observation/action contract is fully specified: two cameras (scene and wrist), six state dimensions, joint-delta actions with absolute gripper, 42 actions predicted with 32 executed at 30 Hz before replanning. - flux-3-action-droid — the policy fine-tuned on the DROID dataset for a Franka-style setup, with three RGB cameras, seven arm joints plus gripper fraction, and 32 absolute joint commands at 15 Hz.
Everything integrates with LeRobot, Hugging Face’s open robotics framework — a Flux3Policy.from_pretrained() call loads a policy, and the shipped lora.json recipe (rank-32 LoRA, BF16 mixed precision, 10,000 steps) lets anyone fine-tune onto a new robot, simulator, or even a game. For deployment flexibility, the DROID package ships six optimization variants: BF16, FP8, guidance-distilled, and step-distilled (single-step) versions, in every combination.
The numbers that matter
On RoboLab-120 — 120 tabletop manipulation tasks in Isaac Sim, ten trials each, on a DROID-style Franka setup — FLUX 3 Action scores 42.92% task success. The comparison table in the model card reads bluntly:
| Model | Type | Success | Parameters |
|---|---|---|---|
| FLUX 3 Action | WAM | 42.92% | 7B |
| Cosmos3-Nano-Policy | WAM | 36.8% | 16B |
| π0.5 | VLA | 28.0% | 3.3B |
That is a 6.1-point lead over Nvidia’s previous best open model using roughly 44% of its parameters, and BFL reports up to 3.95× faster runtime. Off-simulation, the company reports 93.3% single-item success on a real Franka arm and 28 of 30 successes on DROID real-world evaluation episodes.
The cost profile is equally notable for a robotics stack: BFL cites an H200 rollout example at $0.087, or about $0.0018 per successful sample when parallelized — an order-of-magnitude framing that positions robot-policy iteration closer to the economics of ordinary inference workloads than to traditional robotics research budgets.
Why “world action model” is the interesting part
The benchmark table distinguishes WAM (world action model) from VLA (vision-language-action). A VLA like π0.5 maps perception directly to actions. A WAM predicts future frames and actions together — the same denoising machinery that made FLUX famous for images and video is being asked to imagine what the scene will look like as the robot moves, and to output joint trajectories as part of that same generative pass.
This is the architecture BFL has been building toward since the FLUX 3 launch in July, when the company demonstrated the model family’s action-prediction head running manipulation tests on Audi production lines with partner mimic. That launch kept the Action component in gated early access. The significance of this week is the ungating: the base weights, two complete policies, the LoRA fine-tuning recipe, benchmark methodology, and hardware guidance are now downloadable by anyone, under the FLUX Kommunity License v1.0 (with the Qwen text encoder remaining Apache-2.0).
Safety framing — unusually concrete
The model cards carry an out-of-scope section that reads like it was written by people who expect the weights to actually control hardware: FLUX 3 Action outputs joint targets, and nothing in the model bounds joint velocity, force, or workspace — the application must enforce limits and keep a hardware stop within reach, with validation in a simulator or with safety limits engaged before operating near people. Prohibited uses include controlling machinery that endangers people without human oversight and a means of stopping it, and fully automated high-risk decision-making. BFL also points to its misuse-mitigation paper, “Capable, Open, and Safe: Combating AI Misuse,” and publishes a safety contact.
The competitive read
Three things make this release more than a niche robotics item.
First, it beat Nvidia at Nvidia’s own game. Nvidia has spent 2026 pushing Cosmos as the open Physical AI stack, tightly coupled to its simulation and GPU ecosystem. A 7B model from an image-generation lab outscoring a 16B Cosmos policy — while running faster — is exactly the kind of leaderboard result that reshapes assumptions about where robotics foundation models come from.
Second, it validates the “generation → action” thesis. If a video model’s learned world dynamics can be harvested directly as motor control, the trillions of tokens of visual pretraining behind generative media become transferable capital for embodied AI. BFL is now the most visible open-weights proof of that idea, and the fact that the same model family produces images, native-audio video, and robot actions keeps the training investment shared across revenue-generating and research lines.
Third, the openness is unusually operational. This is not a “weights dump” release. Complete observation contracts, normalization files, sampling parameters (Euler steps, shift, guidance values, seeds), six distilled optimization variants, a community dataset lineage, and a LoRA recipe mean a lab with a robot arm can realistically go from download to fine-tuned policy in days. The presence of a gaming action head in the base checkpoint also hints at where BFL thinks the next batch of embodiments will come from.
Caveats worth keeping
RoboLab-120 is a simulated benchmark, and sim-to-real gaps remain the field’s favorite caveat; BFL’s real-world numbers are self-reported on two platforms. The Kommunity License is not OSI-open in the strict sense, so commercial users should read its terms. And a 42.92% success rate, while first-place, still means the model fails the majority of attempts on the hardest board — this is a research frontier, not a deployable warehouse worker.
Even so, the direction is unmistakable. The same open-weights playbook that made FLUX the backbone of the image-generation ecosystem — release the base, ship the recipes, let the community build the long tail of applications — has now been pointed at physical intelligence. Nvidia’s Cosmos, Physical Intelligence’s π family, and now BFL are converging on the same idea from different starting points, and the open leaderboard just got a new leader.
Sources
- [1] https://huggingface.co/black-forest-labs/flux-3-action-base
- [2] https://huggingface.co/black-forest-labs/flux-3-action-droid
- [3] https://huggingface.co/black-forest-labs/flux-3-action-so101
- [4] https://venturebeat.com/infrastructure/black-forest-labs-debuts-flux-3-action-an-open-weights-ai-robotics-model-that-tops-the-leaderboard-at-half-the-size-of-its-competition
- [5] https://www.theneuron.ai/digest/everything-that-happened-in-ai-today-wednesday-september-23-2026/
- [6] https://bfl.ai/models/flux-3