A 7B model that returns two seconds of actions

FLUX 3 Action announcement graphic from Black Forest Labs on Hugging Face Caption: Preview image published with the 23 Sep 2026 FLUX 3 Action Hugging Face blog post · Source: Black Forest Labs via Hugging Face blog · link

On 23 September 2026 Black Forest Labs published FLUX 3 Action: a world action model you can fine-tune. FLUX 3 Action is an open-weights world action model (WAM). It is a 7B diffusion transformer. Each call takes one or more camera frames, a state vector and a text instruction. The model returns the next 32 actions, which the post describes as the next 2 seconds of actions. It can also return 32 decoded future frames. At control time, you skip the decode, execute the first few actions, observe again and replan.

A frozen video VAE encodes the frames. A frozen Qwen3-VL-4B model encodes the instruction. Actions form a second token stream in the same sequence. Each embodiment has its own input projection and output head. Video and action tokens share one noise level per sample, and the model denoises them jointly. The DROID and SO-101 checkpoints run as LeRobot policy classes.

42.92% on RoboLab-120 and a LeRobot training path

Black Forest Labs reports 42.92% overall success on RoboLab-120 for the DROID fine-tune. The table in the post lists 36.8% for the open 16B Cosmos 3 Nano policy. The DROID policy model card describes RoboLab-120 as 120 tabletop tasks in Isaac Sim. Each task runs 10 trials on a DROID-style Franka setup.

The post also shows SO-101 pick-and-place clips. The team adapted that policy on about 200 teleoperated episodes. It also trained policies for two games, GRUNT and VECTOR, on 800 recorded episodes per game. An indoor-drone policy learned from 800 simulated flights.

Four features make the release relevant to robotics and embodied-agent teams this week. The weights are open. The checkpoints run in LeRobot. The post cites a public simulation leaderboard number. Black Forest Labs built parameter-efficient fine-tuning recipes with NVIDIA.

FLUX Kommunity License, 94 stars and lab-ready code

The weights use the FLUX Kommunity License v1.0. The DROID model card lists license_name: flux-kommunity-license. Read LICENSE.md before commercial use, and check each intended use against its terms.

The code is in the flux-action repository on GitHub. The repository had 94 stars as of 28 Sep.

The FLUX 3 Action collection holds the weights, including flux-3-action-base and flux-3-action-droid.

The release is usable for lab and evaluation work. The post shows real SO-101 rollouts and simulation scores. It publishes no factory service-level agreement, safety case or certified collaborative-robot stack. Plan any production cell deployment as a separate engineering program.

Choosing between FLUX 3 Action, Cosmos 3 Nano and π0.5

  • Pick FLUX 3 Action if you want an open 7B WAM with a LeRobot training path. It fits best if you already run DROID-style Franka or SO-101 hardware, or if you plan to fine-tune with a few hundred episodes.
  • Compare it on your own tasks before you trust the 42.92% figure. RoboLab-120 runs tabletop tasks in Isaac Sim. Your lighting, cameras and objects will differ.
  • Stay on a smaller VLA if your edge VRAM or Jetson budget cannot hold a 7B diffusion policy. The table in the post lists π0.5 at 3.3B. Check the FP8 and distilled variants on the model card first.
  • Skip it if you only need text or browser agents. This release covers camera-to-action robotics, plus game and drone control demos.

Simulated scores, license terms and per-arm action heads

  1. RoboLab success measures simulated tabletop tasks. Run a real-robot evaluation before procurement documents cite 42.92%.
  2. The Kommunity License sets its own conditions on use. Ask counsel to read LICENSE.md before a customer pilot.
  3. Each embodiment has its own action head, so a new arm needs a fine-tune. A checkpoint swap alone will not retarget a new arm.
  4. The frozen Qwen3-VL-4B encoder, the VAE and the LeRobot version must stay pinned together. Version drift can break replay.
  5. The "places first" claim and the recovery-from-mistake clip come from Black Forest Labs. Reproduce them on your own seeds before you say that the system works in your cell.

Blog, weights, code and the RoboLab board