Robotics
LeRobot
Safetensors
molmoact2
yam
clothes-folding
bf16

YAM folding — MolmoAct2, 70k steps, BF16 inference

Inference export of the latest complete 70,000-step YAM folding checkpoint at publication time. This model was full-finetuned on the selected folding episodes from MolmoAct2 BimanualYAM and ABC training data: global batch 512, joint discrete + continuous flow objective, 4 flow samples per training example, FP32 master weights and BF16 compute. The planned training length was 100k; this export is the saved 70k model.

This package is for continuous action inference using hq-fang/lerobot, branch main. Pin code commit 5f8be1af5380d89ba4f645bc69a9c10c632a8025. No additional source-code patches are required. The original checkpoint's machine-specific paths and unsupported gripper-range option were removed from the export metadata.

Precision and contents

dtype="bfloat16", model_params_fp32=false selects LeRobot's native BF16 storage/compute policy. Large vision/language weights are BF16; the action expert, normalization layers and other components designated by the implementation stay FP32. This is intentionally not an all-tensors-BF16 cast. Do not call policy.bfloat16() or policy.half() after loading, which would override those exceptions.

Contains model weights, policy configuration, saved pre/postprocessors, shared quantile statistics, this README, and a standalone inference example. No optimizer states, training RNG states, datasets or credentials. Normalization statistics remain at their original precision.

The loader also resolves the original base model's architecture, tokenizer and image processor from allenai/MolmoAct2-BimanualYAM, pinned at 8dcbed66f2380e4393189c303ea72488eb9e63c2. The current LeRobot constructor loads that base before applying the finetuned weights, so the first run also downloads/caches the base model. You can reuse an existing local copy with --base-model /path/to/MolmoAct2-BimanualYAM.

Training compilation, gradient checkpointing and inference CUDA graphs are disabled in the deployment config. They are unnecessary for loading the trained weights. Inference uses 10 flow integration steps, separate from the training flow-sample count of 4.

Install

Use Python 3.12 and a CUDA-enabled PyTorch installation supporting BF16 on your GPU.

git clone --branch main https://github.com/hq-fang/lerobot.git
cd lerobot
git checkout 5f8be1af5380d89ba4f645bc69a9c10c632a8025
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[molmoact2]'
hf download hqfang/yam_fold_70 --local-dir ./yam_fold_70

Custom camera and state inputs

Supply three RGB views with the corresponding trained roles: top, left, right. Map your hardware camera names to observation.images.top, observation.images.left, and observation.images.right. The example takes HWC uint8 arrays or RGB image files and converts them to CHW tensors. Training views were 640×360; let the saved image processor perform model resizing/normalization. Do not apply additional image normalization yourself. Arbitrary camera counts or swapping camera roles are not validated for this YAM model.

Provide a task instruction and the current 14-dimensional state in this order:

left_joint_0.pos, left_joint_1.pos, left_joint_2.pos,
left_joint_3.pos, left_joint_4.pos, left_joint_5.pos, left_gripper.pos,
right_joint_0.pos, right_joint_1.pos, right_joint_2.pos,
right_joint_3.pos, right_joint_4.pos, right_joint_5.pos, right_gripper.pos

Arm joint positions must use the training robot's joint convention and units (radians); grippers use the training convention in [0, 1]. The exported standard processor rejects out-of-range gripper inputs. Calibrate your sensor values to this convention rather than silently changing the normalization. State and action quantiles come from the saved mixture statistics; do not replace them with the original base-model statistics.

Save your measured state as state.npy, shape (14,), float32. Then:

python yam_fold_70/inference.py \
  --checkpoint ./yam_fold_70 \
  --top top.png --left left.png --right right.png \
  --state state.npy \
  --task 'Pick up the shirt, fold it neatly, and place it on the pile.' \
  --output actions.npy

The result is (30, 14): one 30-frame action chunk at the training rate of 30 Hz, in the same joint order. Outputs are absolute joint-position targets, not deltas. The saved postprocessor converts predictions back to physical units and preserves the gripper convention. The script only writes predictions; it does not command a robot.

Python / action queue

The provided inference.py exposes load_policy() and observation() for your application:

import sys
sys.path.insert(0, './yam_fold_70')
import torch
from inference import load_policy, observation

policy, pre, post = load_policy('./yam_fold_70')
policy.reset()  # once at each episode/task reset, not every control tick

# Inside your 30 Hz control loop, supply current RGB arrays and measured state:
raw = observation(top_rgb, left_rgb, right_rgb, state14, task_instruction)
with torch.inference_mode():
    action = policy.select_action(pre(raw), inference_action_mode='continuous')
    target14 = post(action)[0].cpu().numpy()

select_action queues 30 actions and generates the next chunk when its queue empties. If instead using predict_action_chunk, manage chunk execution yourself; that method generates a fresh chunk on every call. Do not pass ground-truth actions into the input batch. Reset the policy when starting a new episode or changing tasks.

Validation and limits

The merged main commit pinned above passed the same checkpoint-loading and inference smoke tests as the original molmoact2-compile export commit. Four continuous action chunks from the two training sources matched the original branch outputs exactly, for both native BF16 and FP32-master inference. The standalone custom-image CLI also matched exactly. Every loaded tensor, its dtype, and the published file checksums were verified. See main_validation.json. These checks apply to the pinned commit; future changes to main are not covered.

Tested runtime: PyTorch 2.11.0+cu128, Transformers 5.5.4, safetensors 0.8.0, huggingface_hub 1.30.0, on an NVIDIA GB200. See export_validation.json for the checks. Export validation checks strict weight loading, all exported tensors against the intended source casts, saved-normalization preservation, finite 30×14 continuous predictions on inputs from both training sources, and exact prediction agreement before/after BF16 serialization with matched random seeds.

BF16 conversion can change predictions relative to FP32 master-weight inference; that difference is measured in the validation report. These are small-sample open-loop checks, not a closed-loop success-rate evaluation or evidence of unchanged task performance. This package documents continuous inference only. Robot joint conventions, calibration and controller limits must match your deployment.

Downloads last month
26
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for hqfang/yam_fold_70

Finetuned
(23)
this model

Datasets used to train hqfang/yam_fold_70