reward modelling is hard because u can’t specify exactly what u meant. it can do what u said, but by a different means.
jacob foster in a podcast mentioned that even humans do this reward hacking. p-hacking is the textbook example of academia doing reward hacking. (we are inteligent agents designed by aliens to improve science. But sadly many agents end up publishing papers with statsitical significant results than actual truths. The humans optimized not for finding truth, but for publishing papers)
modelling - papers try to explain a phenomenon, but end up only explaining it. They thought the reward was explaing the bit by bit phenomenon. Not understanding the system by looking at the behaviour.
goal: make company valuable outer misalignment:
- hack the stock market and make prices go up. Company became valuable.
- reward hacking/spec gaming/goodhart law innner misalignment:
- hire writers to write articles saying “company is valuable”. it thought the company is valuable because people talk/read about it. But actually company’s value didn’t change. It thought internally that “people saying company valuable” == “company valuable”
Reward hacking is a problem because u don’t give auxillary goals! (like auxillary hypothesis to check if a statement is true or not)