GSM8K vs MATH vs AIME — what do they measure?
mediumAnswer
- GSM8K (Cobbe 2021): 8k grade-school word problems; 2-8 step arithmetic.
- Frontier models are at ~95%.
- MATH (Hendrycks 2021): 12.5k competition math problems (algebra, geometry, calculus, number theory) at high-school competition level.
- Frontier ~85-90%.
- AIME (American Invitational Math Exam): olympiad-level; the hardest widely-used math benchmark, ~15/30 for GPT-4, up to 90%+ for o1/o3/R1-style reasoning models.
- Progression: GSM8K → MATH → AIME → IMO-level (still open).
- Reasoning models trained with RL on verifiable rewards dominate these.
Check yourself — multiple choice
- Same difficulty
- GSM8K: grade-school word problems (~95% saturated); MATH: competition (~85-90%); AIME: olympiad (frontier of reasoning models)
- MATH easier than GSM8K
- AIME is trivial
GSM8K < MATH < AIME in difficulty; reasoning models trained via RL dominate.
#evaluation#benchmarks
Practise LLMs & GenAI
214 interview questions in this topic.