EasyDeepLearn

What does MMLU actually measure and its limitations?

medium

Answer

  • MMLU (Hendrycks 2020): 15,908 multiple-choice questions across 57 subjects (STEM, humanities, professional).
  • Format: question + 4 answer options, model picks A/B/C/D.
  • Measures general knowledge and reasoning across many domains.
  • Limitations: (1) MCQ format is easy to game via answer-key patterns; (2) many questions are contaminated in modern training sets; (3) frontier models are hitting 88-90% — saturation reduces signal; (4) doesn't test open-ended reasoning or tool use.
  • Replacements: MMLU-Pro (harder, 10 options), GPQA (grad-level, contamination-free).
Check yourself — multiple choice
  • MMLU is the best benchmark
  • 57-subject MCQ; saturated at ~90%, contaminated, gameable — replaced by MMLU-Pro / GPQA for frontier models
  • MMLU is open-ended
  • MMLU is generative

MMLU: 57-subject MCQ; saturated / contaminated; replaced by MMLU-Pro / GPQA.

#evaluation#benchmarks

Practise LLMs & GenAI

214 interview questions in this topic.

Related questions