๐งฎ Reinforcement Learning for Math Reasoning
Model: Qwen/Qwen2.5-0.5B (494M params) ย |ย Dataset: GSM8K ย |ย Best accuracy: PPO โ 37.30% (+1.44 pts over SFT)
GSM8K Math Reasoning โ Full Results Dashboard
All results were computed offline (no live training here). Charts are saved PNGs from the training notebooks.
Accuracy on GSM8K Test Set (1,319 examples)
| Stage | Correct | Accuracy | vs SFT |
|---|---|---|---|
| Base (zero-shot) | 257 / 1319 | 19.48% | -16.38 pts |
| SFT (LoRA, 3 epochs) | 473 / 1319 | 35.86% | +0.00 pts |
| GRPO (1 epoch, 1.5k) | 477 / 1319 | 36.16% | +0.30 pts |
| PPO (1 epoch, 1.5k) | 492 / 1319 | 37.30% | +1.44 pts โ best |
| REINFORCE (1 epoch, 1.5k) | 474 / 1319 | 35.94% | +0.08 pts |
Algorithm Comparison โ Base โ SFT โ GRPO โ PPO โ REINFORCE
Training Reward Curves
Each curve shows the mean batch reward during RL training. An upward trend means the model is learning to score higher rewards.
SFT Training Loss
Failure Taxonomy
Each bar shows how the model's 1,319 test answers are distributed across five outcome categories. The two green bars together = total correct answers.
| Category | Meaning |
|---|---|
#### + Correct |
Model used #### format and got the right answer |
#### + Wrong |
Model used #### but the number was wrong |
No #### + Correct |
No #### but fallback extractor found the right number |
No #### + Wrong |
No #### and fallback found a wrong number |
No Answer |
No number extractable at all |
Live Math Inference
Type any math word problem and see how the Base model and the PPO model (our best RL result) respond side by side.
โ ๏ธ First inference triggers model loading (~30โ60 s on CPU/MPS). Subsequent queries are fast.
Model Output
Extracted Answer
Quick examples โ click to load:
GSM8K Dataset Explorer
GSM8K contains 8,792 grade school math word problems with step-by-step solutions.
Training split: 7,473 examples | Test split: 1,319 examples
The dataset is downloaded from HuggingFace at runtime: openai/gsm8k
Raw Question
Formatted Training Prompt
This is exactly how the question was presented to the model during SFT training.
The Qwen chat template wraps the question; the assistant is expected to reason step-by-step and end with #### <number>.
Training Data Stats
| Stat | Value |
|---|---|
| Total training examples | 7,473 |
| Total test examples | 1,319 |
| SFT used | All 7,473 training examples |
| RL (GRPO / PPO / REINFORCE) used | First 1,500 training examples |
| Average question length | ~46 words |
| Average answer length | ~180 words (with steps) |
| Answer format | Multi-step reasoning + #### <number> |
_Project by Punam ย |ย All training was done on Apple MPS (Mac M-series) ย |ย