EasyDeepLearn

Common LLM safety evaluation suites.

medium

Answer

  • (1) ToxiGen: hate speech generation.
  • (2) BOLD / RealToxicityPrompts: bias + toxicity in generation.
  • (3) TruthfulQA: hallucination detection.
  • (4) BBQ: bias in QA.
  • (5) DoNotAnswer: refusal for harmful requests.
  • (6) AdvBench: jailbreak attempts.
  • (7) HarmBench: harmful behavior test.
  • (8) MITRE ATLAS threat models.
  • (9) Custom: your-org red-team suite.
  • Run every training + before deployment.
  • Report in model card + monitor in production.
Check yourself — multiple choice
  • Random
  • ToxiGen (hate) / BOLD/RealToxicity (bias+tox generation) / TruthfulQA (halluc) / BBQ (QA bias) / DoNotAnswer (refusal) / AdvBench (jailbreak) / HarmBench (behavior) / MITRE ATLAS + custom red team; run every train + deploy
  • Just accuracy
  • Not real

LLM safety: ToxiGen / BOLD / TruthfulQA / BBQ / DoNotAnswer / AdvBench / HarmBench.

#llmops#safety

Practise MLOps & Data Quality

215 interview questions in this topic.

Related questions