pi0.5 joint-space policies β Unitree G1 dual arm, IKEA table assembly
Five Ο0.5 policies, one per subtask of an IKEA table assembly, fine-tuned from pi05_base
for the Unitree G1 (29-DoF, Dex1-1 grippers). Each policy ships in two forms:
<subtask>/β the original openpi JAX checkpoint: an Orbaxparams/store plus theassets/holding its normalization statistics.pytorch/<subtask>/β the same weights as bf16 safetensors for openpi's PyTorch implementation (model.safetensors+ the sameassets/), which is what runs on NVIDIA Jetson AGX Thor. Each conversion passed a same-noise JAX-vs-PyTorch equivalence check (numbers under PyTorch conversion below).
serving/ holds everything else needed to run them: the openpi overlay that defines this
action space, a multi-policy server, a Dockerfile for Thor, and the conversion + check
scripts. See serving/thor/README.md.
| subfolder | instruction | steps |
|---|---|---|
flip_table |
flip table | 30k |
pick_table_leg |
pick table leg | 30k |
insert_table_leg |
insert table leg to table base | 30k |
rotate_table_base |
rotate table base | 30k |
rotate_leg_tighten |
rotate leg to tighten | 30k |
The saved params/ are the EMA weights (decay 0.99, β100-step averaging window), not
the final raw optimizer iterate β openpi writes ema_params in place of params when EMA
is enabled.
Action space
State and action are 16-dimensional:
[0:7] left arm joints (shoulder p/r/y, elbow, wrist r/p/y) β G1_29_JointArmIndex order
[7:14] right arm joints
[14] left Dex1 gripper (motor rad, 0 closed .. 5.4 open)
[15] right Dex1 gripper
State comes from rt/lowstate; actions are rt/lowcmd (arm_sdk) targets, so the policy
commands what a teleoperator's controller commanded β no IK anywhere. The source
datasets store 19 dims (these 16 plus waist yaw/roll/pitch); the policies consume the
leading 16 and leave the waist to the robot's own balance controller.
Internally the 14 arm dims are trained as deltas against the observation's state while
both grippers stay absolute (make_bool_mask(14, -2)). openpi's AbsoluteActions
inverts this at inference, so what leaves the policy server is absolute joint targets
for all 16 dims β a client must not add the current state again.
Observations are three real cameras β head (binocular, left eye extracted upstream) in the
base_0_rgb slot plus two D405 wrist cameras β and the prompt is the dataset task string.
Action horizon is 50; the model pads 16 β action_dim 32 after normalization, so the norm
stats in assets/ are 16-dimensional.
Usage
Requires openpi with a matching
TrainConfig registered β one whose data config trims to 16 dims, uses the same delta
mask, and whose repo_id matches the assets/ path below. The checkpoint carries its norm
stats but not its transforms, so serving under a config with a different action space
will load successfully and produce garbage.
The overlay that teaches openpi this action space β the G1JointInputs/G1JointOutputs
transforms and the LeRobotG1JointDataConfig holding the make_bool_mask(14, -2) delta
mask β is in serving/openpi_overlay/. Install it into an openpi checkout
(commit 215abfb or compatible) with:
python serving/openpi_overlay/install_overlay.py --openpi ~/openpi \
--snippet config_g1_joint_snippet.py --policy-files g1_joint_policy.py --tag joint
This registers pi05_g1_ikea_joint_<subtask> for all subtasks; the repo_id in each is
chqzhu/g1-ikea-joint19-<subtask>, so the bundled norm stats resolve.
Then download and serve (JAX, one policy per process):
hf download chqzhu/pi05-g1-ikea-joint --include "flip_table/*" --local-dir ./ckpt
uv run scripts/serve_policy.py --port 8000 policy:checkpoint \
--policy.config=pi05_g1_ikea_joint_flip_table \
--policy.dir=./ckpt/flip_table
Norm stats are read from <checkpoint>/assets/<asset_id>/norm_stats.json, where asset_id
defaults to the training repo_id (chqzhu/g1-ikea-joint19-<subtask>). A config whose
repo_id differs will raise rather than silently substitute different statistics.
PyTorch checkpoints and the multi-policy server (Jetson Thor, offline)
The repo is self-sufficient for a Thor with no internet: download it elsewhere, copy it over, and either load the pre-built arm64 image or install from the bundled wheelhouse.
hf download chqzhu/pi05-g1-ikea-joint --include "pytorch/*" "serving/*" --local-dir ./g1 # 37 GB + 11 GB image + 1.5 GB wheels
# A. pre-built image (JetPack 7 has Docker + the NVIDIA runtime). The image is a runtime;
# the code it runs is serving/, mounted over /opt/pi05_vla β always pass that mount.
zstd -dc g1/serving/images/g1-policy-server-thor-arm64.tar.zst | docker load
docker run --rm --runtime nvidia --network host \
-v $PWD/g1/serving:/opt/pi05_vla -v $PWD/g1/pytorch:/ckpt g1-policy-server:thor-arm64
# B. native Python 3.12, no Docker: every aarch64 wheel is in serving/thor/wheelhouse-thor
bash g1/serving/thor/install_offline.sh ~/g1-serve && source ~/g1-serve/bin/activate
JAX_PLATFORMS=cpu python g1/serving/thor/serve_multi.py --ckpt-root g1/pytorch --check-contract
Limited bandwidth. Download only what you run: --include "pytorch/<subtask>/*" "serving/*"
is 7.5 GB per policy + 11 GB image + 1.5 GB wheelhouse (drop serving/thor/wheelhouse-thor/*
if you use the image, or serving/images/* if you use the wheelhouse). Code fixes are
shipped as files under serving/ and picked up through the mount, so after the first
download a fix costs a few MB; the image is re-published only when a dependency layer
changes. With internet, docker build -f g1/serving/thor/Dockerfile g1/serving rebuilds
the image for any architecture. Full details: serving/thor/README.md.
All five policies stay resident in one process on one port; the client picks one by URL
path (ws://thor:8000/flip_table) or per request (obs["policy"] = "flip_table"); a client
that sends neither is answered by the policy trained on its prompt. Every reply names the
policy that produced it. Protocol otherwise = openpi_client.
The PyTorch checkpoints load through the same create_trained_policy and the same
config name as the JAX ones, so the action-space contract above is unchanged.
PyTorch conversion
serving/thor/convert_ckpt.py (openpi's converter, memory-safe) produced pytorch/<subtask>/.
serving/thor/compare_backends.py then queried both backends with identical observations
from the subtask's own dataset and identical flow-matching noise; the JAX-vs-JAX gap
under fresh noise is the yardstick for "different sample"; a conversion is accepted below a ratio of 0.25. Six observations per subtask (episodes 0 and 1, stride 150); arm MAE is over the 50Γ14 chunk vs the recorded actions. The last column is the absolute-target contract: chunk[0] must sit near the input state.
| subfolder | JAX vs PyTorch, same noise (rad, mean / max) | JAX vs JAX, fresh noise (rad) | ratio | arm MAE JAX / PyTorch | PyTorch chunk[0] β state (rad) |
|---|---|---|---|---|---|
flip_table |
0.0014 (max 0.028) | 0.0158 | 0.09 | 0.0185 / 0.0185 | 0.020 |
pick_table_leg |
0.0017 (max 0.025) | 0.0359 | 0.05 | 0.0407 / 0.0404 | 0.015 |
insert_table_leg |
0.0015 (max 0.014) | 0.0309 | 0.05 | 0.0270 / 0.0271 | 0.023 |
rotate_table_base |
0.0017 (max 0.023) | 0.0247 | 0.07 | 0.0286 / 0.0282 | 0.013 |
rotate_leg_tighten |
0.0013 (max 0.024) | 0.0651 | 0.02 | 0.0759 / 0.0757 | 0.019 |
Evaluation
Open-loop against episodes of each subtask's own dataset: the predicted 50Γ16 chunk compared with ground truth, against a "hold the current pose" baseline.
| subfolder | 1-step MAE (rad) | hold-pose | k=49 MAE | hold-pose | ratio | gripper MAE |
|---|---|---|---|---|---|---|
flip_table |
0.0107 | 0.0199 | 0.0264 | 0.2649 | 0.10 | 0.043 |
pick_table_leg |
0.0119 | 0.0143 | 0.0436 | 0.2703 | 0.16 | 0.113 |
insert_table_leg |
0.0106 | 0.0271 | 0.0407 | 0.1312 | 0.31 | 0.070 |
rotate_table_base |
0.0096 | 0.0116 | 0.0205 | 0.2063 | 0.10 | 0.066 |
These are training-distribution episodes, so they show that the policies learned motion
rather than merely tracking the current pose; they are not a generalization measure. None of
these checkpoints has been validated on hardware. pick_table_leg is the weakest β its
1-step margin over hold-pose is thin and its gripper error is roughly 2.6Γ flip_table's.
Task-space output for the IROS decoupled lane
The policies are joint-space, but the IROS 2026 challenge's decoupled lane takes
task-space wrist poses and runs its own IK. serving/decoupled/ converts one into the
other, and the server emits either:
--action-space joint16 # raw policy output (default)
--action-space task25 # (T,25) task-space chunks; joint16 rides along as actions_joint16
--action-space both
Forward kinematics is pure numpy, replayed from a chain baked out of pinocchio
(export_chain.py), so nothing needs pinocchio at serve time. verify_fk.py checks it
against pinocchio to machine precision on real trajectories.
Conventions, all verified against the organizer's own modules β every one of them is accepted silently if you get it wrong:
| value | why it matters | |
|---|---|---|
| column layout | from boundary/actions.py TASKSPACE_SLICES |
navigate_cmd occupies [18:21]; base_height_cmd is [21], torso_rpy [22:25] |
| hands | 2 finger joints, [-1,+1], β1 open / +1 closed |
our Dex1 radians (0 closed β¦ +5.4 open) invert and rescale |
| quaternions | (w, x, y, z), unit norm |
the boundary checks the norm but not the ordering |
| tool offset | zero β raw wrist_yaw_link |
ik.py pins it to zero deliberately |
| waist | live, from body_q[12:15] |
ik.py locks non-arm joints at measured values; assuming zero costs 3.98 cm mean / 14.12 cm max at the wrist |
base_height_cmd |
0.74 m | forwarded to the WBC verbatim. 0.0 is not "no override" β it commands the pelvis to the floor |
Cameras on the bench
From reference/orin_bridge/real_orin_cameras.py in the organizer's Orin package:
- The head is a 3840x1080 side-by-side stereo USB camera. Each eye is 1920x1080, resized (not cropped) to 640x480 β 16:9 squashed into 4:3. Our training rig captured 4:3 per eye natively, so the bench head is a genuinely different image geometry. This, not mono-vs-stereo, is the real domain gap.
ego_viewis taken fromEGO_VIEW_EYE, which defaults toleft, andPUBLISH_STEREOalso defaults on β soego_viewandego_view_leftnormally carry identical pixels.EGO_VIEW_EYE=rightis the only way the mono key is a different eye.EGO_VIEW_RECTIFYdefaults off, which matches our raw unrectified training frames.- Wrist RealSense serials come from
LEFT_WRIST_SERIAL/RIGHT_WRIST_SERIALand are empty by default. Unset, those threads never start and the keys never publish, so the client sees black frames and loses both wrist views silently. HEAD_DEVICE_NAMEdefaults to"USB Camera"; if the v4l2 device is named anything else, noego_view*key publishes at all.
preflight.py also carries each subtask's training frame-0 pose. They are asymmetric and
none is near open/open, so starting from a default hand pose is wrong for every subtask
(rotate_leg_tighten starts with the left hand nearly closed, 0.71 rad).
Containers
| size | what | |
|---|---|---|
serving/images/g1-policy-server-thor-arm64.tar.zst |
11 GB | Thor policy server |
serving/images/g1-policy-client-orin-r35.tar.zst |
336 MB | PC2/Orin client, Python 3.8 / aarch64 |
The client is small because all inference is on Thor; it reads cameras and state, queries the server and publishes the chunk.
You do not need to rebuild either image to change this code. Both read it from a single path, so mount a newer copy over it:
docker run ... -v $PWD/serving:/opt/pi05_vla g1-policy-server:thor-arm64
docker run ... -v $PWD/serving:/app g1-policy-client:orin-r35
The same code ships as plain files under serving/, so a change is a file update rather
than an 11 GB upload. The published Thor image predates the task-space work β mount
serving/ over it, or rebuild from serving/thor/Dockerfile.
Provenance and limitations
Trained on teleoperated G1 demonstrations derived from the BitRobot
2026-humanoid-ikea-assembly-challenge recordings. Those recordings' true frame rate is
10.8β16.1 Hz despite fps=30 metadata, and it differs per subtask (median of the training
segments: flip_table 11.6, pick_table_leg 10.8, insert_table_leg 10.8, rotate_table_base 12.5,
rotate_leg_tighten 11.0 Hz), so play a chunk back at the subtask's own rate, not 30.
Commanding a real G1 with these weights moves a 29-DoF humanoid. Anything driving them should keep joint-limit clipping, a per-step delta cap, and an anomaly hold in the loop.
The demonstrations were recorded with gravity-compensating torque feedforward on the arms
(the teleop IK's rnea term), so the policies expect an arm that reaches its targets. A
position-only PD at the teleop gains (kp 80) sags ~0.15 rad below the shoulder-pitch and
elbow targets under load, and since every chunk restarts from the observed state, the
target then snaps back at each chunk boundary. serving/rollout/env_g1.py sends the
gravity torque and the chunk slope as feedforward and chunk_runner.py blends chunk
boundaries; a client of your own should do the same or raise its gains.
None of the serving stack has run on real hardware. The decoupled layer is verified
against the organizer's own validator and mocks β 40 chunks accepted, 0 rejected via
serving/decoupled/conformance_ours.py, and the Orin image verified the same way under
QEMU β but that establishes only that the output satisfies the contract. It says nothing
about whether the policies work on a robot. No Thor, no Orin, no G1 has run any of it.