Explore Library
QuizIntermediate

Reinforcement Learning from Human Feedback

RLHF aligns model outputs with human preferences using feedback signals.

What is the purpose of RLHF in training LLMs?