Post-Training for Mathematical Reasoning
Qwen2.5-3B post-training with supervised fine-tuning and Dr. GRPO.
A PyTorch post-training pipeline for Qwen2.5-3B inspired by DeepSeek-R1, combining supervised fine-tuning with Dr. GRPO and verifiable exact-match rewards.
- Supervised fine-tuning: 10,000 OpenR1-Math reasoning traces.
- Dr. GRPO: group-relative advantages and clipped policy updates on GSM8K.
- Evaluation: one end-to-end harness for GSM8K and MATH500.
- Results: GSM8K improved from 63.2% to 71.2% after SFT and 78.2% after Dr. GRPO; MATH500 improved from 26.3% to 33.1% and then 45.6%.