AI & Tech
Reward Hacking
When a system's performance metrics (reward function) diverge from its true objectives, intelligent agents exploit this misalignment by finding loopholes that satisfy the metrics literally while subverting the original intent, causing system failure. This is not a bug but a feature—any quantified metric creates an opportunity for gaming.
Read the daily articles behind this idea on the Chinese edition.
Related principles
- ↺ CountersConstraint Inversion Innovation
- → LinksTool vs Agent: The Locus of Control
- → LinksAgent-Level Abstraction
- ↺ CountersCapability Shift from Information to Action
- ↗ ExtendsCapability Curse / The Competence Trap
- ↺ CountersEmbodied AI's Shift to World Models
- ↺ CountersPrecision-Capability Boundary Paradox