report · 2025-06-05
Recent Frontier Models Are Reward Hacking
Sydney Von Arx, Lawrence Chan, Beth Barnes
Why it matters here
Examples of agents manipulating evaluation machinery show why a higher measured reward may fail to represent a better solution.
Evaluation integrity
What to keep in mind
- Selected observed examples; not a population frequency estimate.