KTO — Kahneman-Tversky Optimization.
hardAnswer
- Ethayarajh et al. 2024.
- Uses prospect theory (loss aversion): people care more about avoiding losses than gaining.
- Works with unpaired (binary desirable/undesirable) labels — easier to collect than paired preferences.
- Loss makes desirable outputs more likely + undesirable less likely, weighted asymmetrically.
- Alternative to DPO when preference pairs are hard to get.
Check yourself — multiple choice
- Same as DPO
- Prospect theory-based: works with unpaired binary desirable/undesirable labels (not pairs); asymmetric loss for loss aversion; alternative to DPO when pairs are hard
- Random
- Not real
KTO: unpaired binary labels via prospect theory; alternative to paired DPO.
#llm#alignment
Practise Reinforcement Learning
214 interview questions in this topic.