QuizIntermediate
Reinforcement Learning from Human Feedback
RLHF aligns model outputs with human preferences using feedback signals.
What is the purpose of RLHF in training LLMs?
RLHF aligns model outputs with human preferences using feedback signals.
What is the purpose of RLHF in training LLMs?