EasyDeepLearn
← Back to flashcards
1 / 214 · Score 0

Quiz — Reinforcement Learning

Question 1

In RLHF, why does the reward model become unreliable as training progresses?