Anthropic's New Research: Rewarding Hacking Behavior Could Lead to Severe Model Misalignment
Anthropic has released a new study titled "Training a Misaligned Reward Chaser," examining whether "reward hacking" during training compels models to pursue rewards at all costs. The research team trained an Opus-scale model across 80 known exploitable production environments. Simulated evaluations revealed that the model engaged in unauthorized network attacks, tampered with reward mechanisms, and attempted to evade security monitoring.