ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm

Praveen Selvaraj · Lorenzo Uttini · Ville Kuosmanen

Armnet

See how it ran

The Armnet arm farm

Evaluation jobs run in parallel
across the arm farm.

cell-1Single-arm · π0.5 · tool insert
cell-3Single-arm · π0.5 · cable unclip
cell-8Bimanual · π0.5 · transfer cube

v0.1 ran on two single-arm cells and one bimanual cell, co-located with a shared inference workstation. The clips show successful π0.5 rollouts on one task assigned to each cell.

7policies
12tasks
2,518core benchmark rollouts
600reference demonstrations

Task suite

Eight single-arm tasks.
Four bimanual tasks.

For each task, all seven policies were trained or fine-tuned with the same budget of 50 demonstrations.

Single-arm

8 tasks

01Block stack
02Cable clip
03Cable unclip
04Eye drops to basket
05Eye drops to shelf
06Ring insert
07Tool insert
08Tool removal

Bimanual

4 tasks

01Fold tea towel
02Insert candle
03Open lamp door
04Transfer cube

Evaluation protocol

30 rollouts per task–policy pair.

Each rollout was labelled

successful suboptimal failure

Leaderboard

Results under a shared 50-demo budget.

Strict success rate across all 12 tasks.

π0.5 led both embodiments.
45.4% single-arm and 52.1% bimanual.

Rankings changed by embodiment.
Diffusion Policy rose to second on bimanual tasks.

Cable clipping remained unsolved.
Every evaluated policy scored 0%.

View per-task results
Strict success rate by task and policy
TaskACTDiffusion PolicySmolVLAπ0π0.5GR00TMolmoAct 2
Single-arm
block_stack0%0%7%33%43%27%7%
ring_insert50%23%33%17%47%60%13%
tool_insert20%23%7%40%53%17%3%
tool_removal27%63%3%10%13%30%0%
cable_clip0%0%0%0%0%0%0%
cable_unclip60%7%37%47%70%33%23%
eye_drops_to_basket63%43%23%70%67%67%63%
eye_drops_to_shelf0%17%10%76%70%43%47%
Bimanual
fold_tea_towel10%50%17%50%70%13%20%
insert_candle0%7%7%27%30%3%3%
open_lamp_door0%73%30%7%23%13%33%
transfer_cube0%13%7%47%86%47%13%

Released corpus

Every scored benchmark
rollout is released.

The main benchmark and leaderboard use 3,118 core episodes. The full release also includes 600 labelled episodes from extra runs.

26 hours in full release
3 camera views per episode
3,718 released episodes

Core benchmark labels · 3,118 episodes

1,290 successful 89 suboptimal 1,739 failed

Limitations and next steps

Priorities for the next release.

View limitations and planned improvements
  • Automation and operator effort. Scene resets are still manual. Automating them would make initial states more repeatable and allow one operator to supervise more cells.
  • Initial states. Operators randomised object placement but did not record the resulting positions. Future runs should log them to verify resets and measure sensitivity to starting conditions.
  • Lighting. The cells used uncalibrated ambient office light. Per-cell lightboxes would make illumination more consistent.
  • Scoring. On-site operators assigned every rollout label. Automatic scoring would reduce this manual workload.
  • Coverage. The task suite should expand to include harder, longer-horizon interactions and tasks that require memory.