Common LLM safety evaluation suites.
mediumAnswer
- (1) ToxiGen: hate speech generation.
- (2) BOLD / RealToxicityPrompts: bias + toxicity in generation.
- (3) TruthfulQA: hallucination detection.
- (4) BBQ: bias in QA.
- (5) DoNotAnswer: refusal for harmful requests.
- (6) AdvBench: jailbreak attempts.
- (7) HarmBench: harmful behavior test.
- (8) MITRE ATLAS threat models.
- (9) Custom: your-org red-team suite.
- Run every training + before deployment.
- Report in model card + monitor in production.
Check yourself — multiple choice
- Random
- ToxiGen (hate) / BOLD/RealToxicity (bias+tox generation) / TruthfulQA (halluc) / BBQ (QA bias) / DoNotAnswer (refusal) / AdvBench (jailbreak) / HarmBench (behavior) / MITRE ATLAS + custom red team; run every train + deploy
- Just accuracy
- Not real
LLM safety: ToxiGen / BOLD / TruthfulQA / BBQ / DoNotAnswer / AdvBench / HarmBench.
#llmops#safety
Practise MLOps & Data Quality
215 interview questions in this topic.