What does MMLU actually measure and its limitations?
mediumAnswer
- MMLU (Hendrycks 2020): 15,908 multiple-choice questions across 57 subjects (STEM, humanities, professional).
- Format: question + 4 answer options, model picks A/B/C/D.
- Measures general knowledge and reasoning across many domains.
- Limitations: (1) MCQ format is easy to game via answer-key patterns; (2) many questions are contaminated in modern training sets; (3) frontier models are hitting 88-90% — saturation reduces signal; (4) doesn't test open-ended reasoning or tool use.
- Replacements: MMLU-Pro (harder, 10 options), GPQA (grad-level, contamination-free).
Check yourself — multiple choice
- MMLU is the best benchmark
- 57-subject MCQ; saturated at ~90%, contaminated, gameable — replaced by MMLU-Pro / GPQA for frontier models
- MMLU is open-ended
- MMLU is generative
MMLU: 57-subject MCQ; saturated / contaminated; replaced by MMLU-Pro / GPQA.
#evaluation#benchmarks
Practise LLMs & GenAI
214 interview questions in this topic.