Reward-hacking isn't just a thing LLMs do.
- Reward-hacking isn't just a thing LLMs do.
- Humans do it too.
- It emerges any time the agents don't care about the why but only the what.
- Reward-hacking is Goodhart's law.
- Reward-hacking will emerge more strongly when:
- 1) The agents are more savvy.
- 2) The stakes are higher.
- 3) The more agents that participate in the domain (effective practices one agent discovers can diffuse to others).
- 4) The agents care about their own goal more than the collective's.