molmoact2 · BusyBox "push the green button" · variant expert_only

Arm expert_only (group trainable_params) of the MolmoAct2 baseline pair in experimental/lerobot_policy_molmoact2/modal/ (alpha-robotics, branch policy_test_fanqi), run with the same protocol as the π0.5 grid in experimental/lerobot_policy_pi05/.

  • Description: b0_fft with the Molmo2 VLM (ViT + connector + LLM) frozen; only the ~578M flow-matching action expert trains (train_mode_vlm: freeze). Gradient checkpointing off: freeze mode fits 8 samples/rank without it.
  • Hypothesis: Counterpart of the pi0.5 grid's expert_only. This was the recipe behind every ArmnetBench MolmoAct2 number (18.9%); Ai2 calls action-expert-only "the clearest failure mode" on LIBERO (93.1 vs 97.2 full). If it matches b0_fft on the robot, MolmoAct2 sweeps can move to 1-2 cheap GPUs; if it is far behind, the ArmnetBench entry under-represented the model.
  • Base weights: allenai/MolmoAct2 @ e432d85f via policy.checkpoint_path (allenai/lerobot molmoact2-policy @ a4f15bf3)
  • Dataset: fanqi-robo/busybox_push_green_button — all 39 episodes, no hold-out (copy of armnet/busybox_push_green_button with exact q01/q99 quantile stats; see its meta/quantile_fix.json)
  • Trainable: train_mode_vlm=freeze (action expert always trains); gradient_checkpointing=False
  • Budget: global batch 32 (8/rank × 4 × H100) × 10000 steps; peak GPU memory 30263 MiB
  • Joint frame: LeRobot v3 (dataset frame); no joint_signs/joint_offsets anywhere — evaluate WITHOUT any v2.1→v3 joint fix
  • Published checkpoint: step 10000 (root of this repo, tag step-10000)
  • Every saved step: fanqi-robo/molmoact2_busybox_push_green_button_expert_only_step2000, fanqi-robo/molmoact2_busybox_push_green_button_expert_only_step4000, fanqi-robo/molmoact2_busybox_push_green_button_expert_only_step6000, fanqi-robo/molmoact2_busybox_push_green_button_expert_only_step8000, fanqi-robo/molmoact2_busybox_push_green_button_expert_only_step10000; full history (with optimizer state for the newest 2) on the Modal Volume molmoact2-busybox-ckpts
  • W&B: https://wandb.ai/fanqi-robo-saferobotics/molmoact2_busybox_push_green_button

Evaluate

The only test is the real robot: submit fanqi-robo/molmoact2_busybox_push_green_button_expert_only (or a _step<N> repo) on https://huggingface.co/spaces/armnet/armnet-eval (embodiment lerobot/so-101, task push_green_button, 20 rollouts, variation seed 42). Native horizon: execute the full 30-step chunk (n_action_steps=30).

Training config (lerobot-train --config_path=..., batch_size shown per rank)

policy:
  type: molmoact2
  checkpoint_path: allenai/MolmoAct2
  device: cuda
  action_mode: continuous
  train_mode_vlm: freeze
  chunk_size: 30
  n_action_steps: 30
  image_keys:
  - observation.images.top
  - observation.images.wrist
  - observation.images.front
  setup_type: single so100/so101 robotic arm in molmoact2
  control_mode: absolute joint pose
  model_dtype: bfloat16
  num_flow_timesteps: 8
  gradient_checkpointing: false
  normalize_gripper: true
  normalization_mapping:
    VISUAL: IDENTITY
    STATE: QUANTILES
    ACTION: QUANTILES
  optimizer_lr: 1.0e-05
  optimizer_vit_lr: 5.0e-06
  optimizer_connector_lr: 5.0e-06
  optimizer_action_expert_lr: 5.0e-05
  scheduler_warmup_steps: 200
  scheduler_decay_steps: 10000
  scheduler_decay_lr: 1.0e-06
  push_to_hub: false
  repo_id: fanqi-robo/molmoact2_busybox_push_green_button_expert_only
  checkpoint_revision: e432d85f6e039edca44afb93c262f3084ab72a9c
dataset:
  repo_id: fanqi-robo/busybox_push_green_button
  episodes:
  - 0
  - 1
  - 2
  - 3
  - 4
  - 5
  - 6
  - 7
  - 8
  - 9
  - 10
  - 11
  - 12
  - 13
  - 14
  - 15
  - 16
  - 17
  - 18
  - 19
  - 20
  - 21
  - 22
  - 23
  - 24
  - 25
  - 26
  - 27
  - 28
  - 29
  - 30
  - 31
  - 32
  - 33
  - 34
  - 35
  - 36
  - 37
  - 38
  image_transforms:
    enable: false
batch_size: 8
num_workers: 6
steps: 10000
save_freq: 2000
eval_freq: -1
log_freq: 20
seed: 1000
job_name: molmoact2_bb_green_expert_only
output_dir: /root/outputs/train/molmoact2_bb_green_expert_only
wandb:
  enable: true
  entity: fanqi-robo-saferobotics
  project: molmoact2_busybox_push_green_button
  mode: online
  disable_artifact: true
Downloads last month
20
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for fanqi-robo/molmoact2_busybox_push_green_button_expert_only

Finetuned
(52)
this model

Dataset used to train fanqi-robo/molmoact2_busybox_push_green_button_expert_only