RLinf PPO beta=0 on LIBERO-Spatial

RLinf PPO fine-tuning of the RLinf pi0.5 LIBERO SFT base, seed 42. Training ran for 100 optimizer steps and saved full and distributed-resume checkpoints every 10 steps.

Checkpoint selection uses the exact logged rollout reward at saved steps. Step 100 is the reward-best saved PPO checkpoint (reward 0.008400225).

Step-100 evaluation (generation horizon 10, denoising steps 3):

Suite Execute Success once Success at end
LIBERO-Spatial 5 0.874 0.652
LIBERO-Spatial 10 0.920 0.706
LIBERO-Plus Spatial 5 0.572856 0.389675
LIBERO-Plus Spatial 10 0.618651 0.482515

Each global_step_* directory contains portable full_weights.pt weights and DCP shards for resuming. Raw training and evaluation logs are included under training/ and evaluation/.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading