Agents that learn from reward rather than examples. Build a bandit, train Q-learning on a grid world, then watch a badly written reward get farmed.
3 units, about 3 hours. Free.
Explain how an agent learns from reward alone, and show why pure greed locks onto a bad choice.
2Implement Q-learning on a grid world and explain the update rule and the discount factor.
3Demonstrate reward hacking in code and explain why agents optimise what you wrote, not what you meant.