RLinf PPO beta=0 on LIBERO-Spatial
RLinf PPO fine-tuning of the RLinf pi0.5 LIBERO SFT base, seed 42. Training ran for 100 optimizer steps and saved full and distributed-resume checkpoints every 10 steps.
Checkpoint selection uses the exact logged rollout reward at saved steps. Step 100 is the reward-best saved PPO checkpoint (reward 0.008400225).
Step-100 evaluation (generation horizon 10, denoising steps 3):
| Suite | Execute | Success once | Success at end |
|---|---|---|---|
| LIBERO-Spatial | 5 | 0.874 | 0.652 |
| LIBERO-Spatial | 10 | 0.920 | 0.706 |
| LIBERO-Plus Spatial | 5 | 0.572856 | 0.389675 |
| LIBERO-Plus Spatial | 10 | 0.618651 | 0.482515 |
Each global_step_* directory contains portable full_weights.pt weights and DCP shards for resuming. Raw training and evaluation logs are included under training/ and evaluation/.