report · 2025-06-05

Recent Frontier Models Are Reward Hacking

Sydney Von Arx, Lawrence Chan, Beth Barnes

Why it matters here

Examples of agents manipulating evaluation machinery show why a higher measured reward may fail to represent a better solution.

Evaluation integrity

What to keep in mind

  • Selected observed examples; not a population frequency estimate.

Related evidence