ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm
The Armnet arm farm
Evaluation jobs run in parallel
across the arm farm.
v0.1 ran on two single-arm cells and one bimanual cell, co-located with a shared inference workstation. The clips show successful π0.5 rollouts on one task assigned to each cell.
Task suite
Eight single-arm tasks.
Four bimanual tasks.
For each task, all seven policies were trained or fine-tuned with the same budget of 50 demonstrations.
Single-arm
8 tasks
Bimanual
4 tasks
Evaluation protocol
30 rollouts per task–policy pair.
Each rollout was labelled
Leaderboard
Results under a shared 50-demo budget.
Strict success rate across all 12 tasks.
π0.5 led both embodiments.
45.4% single-arm and 52.1% bimanual.
Rankings changed by embodiment.
Diffusion Policy rose to second on bimanual tasks.
Cable clipping remained unsolved.
Every evaluated policy scored 0%.
View per-task results
| Task | ACT | Diffusion Policy | SmolVLA | π0 | π0.5 | GR00T | MolmoAct 2 |
|---|---|---|---|---|---|---|---|
| Single-arm | |||||||
| block_stack | 0% | 0% | 7% | 33% | 43% | 27% | 7% |
| ring_insert | 50% | 23% | 33% | 17% | 47% | 60% | 13% |
| tool_insert | 20% | 23% | 7% | 40% | 53% | 17% | 3% |
| tool_removal | 27% | 63% | 3% | 10% | 13% | 30% | 0% |
| cable_clip | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| cable_unclip | 60% | 7% | 37% | 47% | 70% | 33% | 23% |
| eye_drops_to_basket | 63% | 43% | 23% | 70% | 67% | 67% | 63% |
| eye_drops_to_shelf | 0% | 17% | 10% | 76% | 70% | 43% | 47% |
| Bimanual | |||||||
| fold_tea_towel | 10% | 50% | 17% | 50% | 70% | 13% | 20% |
| insert_candle | 0% | 7% | 7% | 27% | 30% | 3% | 3% |
| open_lamp_door | 0% | 73% | 30% | 7% | 23% | 13% | 33% |
| transfer_cube | 0% | 13% | 7% | 47% | 86% | 47% | 13% |
Released corpus
Every scored benchmark
rollout is released.
The main benchmark and leaderboard use 3,118 core episodes. The full release also includes 600 labelled episodes from extra runs.
Core benchmark labels · 3,118 episodes
Limitations and next steps
Priorities for the next release.
View limitations and planned improvements
- Automation and operator effort. Scene resets are still manual. Automating them would make initial states more repeatable and allow one operator to supervise more cells.
- Initial states. Operators randomised object placement but did not record the resulting positions. Future runs should log them to verify resets and measure sensitivity to starting conditions.
- Lighting. The cells used uncalibrated ambient office light. Per-cell lightboxes would make illumination more consistent.
- Scoring. On-site operators assigned every rollout label. Automatic scoring would reduce this manual workload.
- Coverage. The task suite should expand to include harder, longer-horizon interactions and tasks that require memory.