QuizIntermediate
Q-Learning Basics
Q-learning is an off-policy method that iteratively updates action-value estimates toward the max.
What does Q-learning aim to learn?
Q-learning is an off-policy method that iteratively updates action-value estimates toward the max.
What does Q-learning aim to learn?