Post-Training for Mathematical Reasoning

Qwen2.5-3B post-training with supervised fine-tuning and Dr. GRPO.

A PyTorch post-training pipeline for Qwen2.5-3B inspired by DeepSeek-R1, combining supervised fine-tuning with Dr. GRPO and verifiable exact-match rewards.

  • Supervised fine-tuning: 10,000 OpenR1-Math reasoning traces.
  • Dr. GRPO: group-relative advantages and clipped policy updates on GSM8K.
  • Evaluation: one end-to-end harness for GSM8K and MATH500.
  • Results: GSM8K improved from 63.2% to 71.2% after SFT and 78.2% after Dr. GRPO; MATH500 improved from 26.3% to 33.1% and then 45.6%.

GitHub