Flashcard

RLHF

RLHF aligns models using a reward model trained on human preferences and reinforcement learning.

Question

What are the stages of RLHF?

Click to reveal answer