EasyDeepLearn

GSM8K vs MATH vs AIME — what do they measure?

medium

Answer

  • GSM8K (Cobbe 2021): 8k grade-school word problems; 2-8 step arithmetic.
  • Frontier models are at ~95%.
  • MATH (Hendrycks 2021): 12.5k competition math problems (algebra, geometry, calculus, number theory) at high-school competition level.
  • Frontier ~85-90%.
  • AIME (American Invitational Math Exam): olympiad-level; the hardest widely-used math benchmark, ~15/30 for GPT-4, up to 90%+ for o1/o3/R1-style reasoning models.
  • Progression: GSM8K → MATH → AIME → IMO-level (still open).
  • Reasoning models trained with RL on verifiable rewards dominate these.
Check yourself — multiple choice
  • Same difficulty
  • GSM8K: grade-school word problems (~95% saturated); MATH: competition (~85-90%); AIME: olympiad (frontier of reasoning models)
  • MATH easier than GSM8K
  • AIME is trivial

GSM8K < MATH < AIME in difficulty; reasoning models trained via RL dominate.

#evaluation#benchmarks

Practise LLMs & GenAI

214 interview questions in this topic.

Related questions