Revolute
Flow-matching visuomotor policies for a Seeed B601 RS arm on LeRobot — teleoperated demonstrations, a controlled U-Net vs. transformer comparison, and models sized for a Jetson Orin Nano.
Revolute is an imitation-learning stack for a Seeed B601 RS 7-DoF arm, built on LeRobot: teleoperated data collection, flow-matching policy training, and on-robot evaluation. The deployment target is a Jetson Orin Nano 8GB, and every default — frozen vision backbone, a single observation step, five ODE integration steps — is set by what fits a 533 ms re-plan window on that device, not by what trains best on a cluster GPU.
Why flow matching
Rectified flow learns a straight-path velocity field, so sampling an action chunk takes 5 Euler steps instead of a diffusion policy’s ~100 DDPM steps — the reason it was chosen for an edge deployment. (See also Flow Matching Policies for a benchmark against Diffusion Policy in robomimic.)
Policies
Four policies are registered as a LeRobot plugin:
| Policy | Velocity field / head |
|---|---|
flow_unet | FiLM-conditioned 1D conv U-Net (the Diffusion Policy backbone) |
flow_dit | DiT transformer with adaLN-Zero and cross-attention into the vision tokens |
act_dino | ACT-style transformer decoder with direct chunk regression — the non-generative control |
flow_lang | flow_dit plus a frozen text encoder — language-conditioned |
flow_unet and flow_dit share the same frozen DINOv2 encoder, objective, optimizer, and data pipeline — only the velocity field differs, so comparing them is a controlled experiment rather than a capacity comparison.
Sized for the edge
- A frozen
dinov2-smallbackbone with a small trainable transformer head per camera — about 40 demonstrations per task cannot fine-tune a ViT. - One observation step, a 48-step action horizon (1.6 s at 30 fps), and 16 executed actions per re-plan (533 ms).
- Measured off-board,
flow_dit_smallre-plans in 33 ms versus 83 ms forflow_dit_base; the larger model is more accurate on replayed episodes (0.200 vs. 0.284 joint units).
Data
Two teleoperated sessions at 30 fps, recording opposite directions of the same manipulation, plus a third lever-pulling task:
| Task | Demonstrations | Frames |
|---|---|---|
| Collect toast from the toaster and place it on the pan | 50 | 29,828 |
| Pick up a bread slice and put it in the toaster | 40 | 23,847 |
The repository also patches a uint8 overflow in LeRobot’s dataset statistics that collapsed the image standard deviations, and corrects a task annotation that described the reverse of the recorded demonstrations — harmless for the vision-only policies, but a language-conditioned policy would perform the opposite task.
Status
Trained flow_dit policies have been rolled out on the real arm; no success rates are reported yet. Deployment on the Jetson Orin Nano itself has not been run, and the Orin latency figures in the repository are desk estimates, not measurements.