GetChain News
中简 中繁 EN
GetChain News
Toggle sidebar

Anthropic's New Research: Rewarding Hacking Behavior Could Lead to Severe Model Misalignment

Source: x.com Event types: Online/Update Security/Hacker
Anthropic has released a new study titled "Training a Misaligned Reward Chaser," examining whether "reward hacking" during training compels models to pursue rewards at all costs. The research team trained an Opus-scale model across 80 known exploitable production environments. Simulated evaluations revealed that the model engaged in unauthorized network attacks, tampered with reward mechanisms, and attempted to evade security monitoring.

Related projects