AirVLA V3-arm (pi0, step 25000, policy-controlled arm)

Ablation in which the policy commands the two arm joints directly instead of the deterministic task-phase controller. One-variable comparison against the v2 Base policy; it is worse on every axis.

Part of the AirVLA MSc dissertation campaign: adapting a pretrained pi0 vision-language-action policy to a quadcopter with a 2-DoF arm and parallel gripper, evaluated in MuJoCo 3.3.4.

Property Value
Base model lerobot/pi0
Training data hanapasta/airvla_v3 (v2 relabelled with arm-joint deltas)
Initialised from lerobot/pi0 base
Training steps 30,000; checkpoint 25,000 selected
Checkpoint selection lowest pinned-noise validation MSE over the saved ladder

Frozen-protocol result (n = 60 pick + 20 navigation, paired scenes)

0/60 grasped - 38/60 correct-target - 264.2 mm median (139.4 mm target-true) - 7/20 navigation, vs the Base policy's 1/60, 42/60, 170.8 mm, 9/20.

Usage

Evaluated with the frozen harness eval_v2.py from the code repository:

python eval_v2.py <this_checkpoint> 60 20 --torchseed 1000 --tag run --video

Code, reproduction guide and the full experimental ledger: https://github.com/robotics-hana/drone-version2

Evaluation uses MuJoCo 3.3.4; a different simulator version changes contact behaviour enough to invalidate comparison with the banked results.

Downloads last month
16
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading