Flashcard

DPO and Preference Alignment

DPO aligns models directly from preference pairs without training a separate reward model.

Question

How does DPO differ from RLHF?

Click to reveal answer