๐Ÿงฎ Reinforcement Learning for Math Reasoning

Model: Qwen/Qwen2.5-0.5B (494M params) ย |ย  Dataset: GSM8K ย |ย  Best accuracy: PPO โ†’ 37.30% (+1.44 pts over SFT)

GSM8K Math Reasoning โ€” Full Results Dashboard

All results were computed offline (no live training here). Charts are saved PNGs from the training notebooks.

Accuracy on GSM8K Test Set (1,319 examples)

Stage Correct Accuracy vs SFT
Base (zero-shot) 257 / 1319 19.48% -16.38 pts
SFT (LoRA, 3 epochs) 473 / 1319 35.86% +0.00 pts
GRPO (1 epoch, 1.5k) 477 / 1319 36.16% +0.30 pts
PPO (1 epoch, 1.5k) 492 / 1319 37.30% +1.44 pts โ† best
REINFORCE (1 epoch, 1.5k) 474 / 1319 35.94% +0.08 pts

Algorithm Comparison โ€” Base โ†’ SFT โ†’ GRPO โ†’ PPO โ†’ REINFORCE


Training Reward Curves

Each curve shows the mean batch reward during RL training. An upward trend means the model is learning to score higher rewards.


SFT Training Loss


Failure Taxonomy

Each bar shows how the model's 1,319 test answers are distributed across five outcome categories. The two green bars together = total correct answers.

Category Meaning
#### + Correct Model used #### format and got the right answer
#### + Wrong Model used #### but the number was wrong
No #### + Correct No #### but fallback extractor found the right number
No #### + Wrong No #### and fallback found a wrong number
No Answer No number extractable at all

_Project by Punam ย |ย  All training was done on Apple MPS (Mac M-series) ย |ย