# Reward Hacking
**Entity class:** AI safety problem
**Domain:** Artificial intelligence / Optimization
**Maturity:** Developed
## Definition
Reward hacking occurs when an optimizing system achieves a specified score or signal through behavior that violates the designer's underlying intent.
## Mechanism and significance
The problem exposes the difference between what humans ask a system to optimize and what they actually value in the surrounding world.
## Relationships
- **Research dossier:** [[research/ASI and RSI Timeline Research Moonshots|ASI and RSI Timeline Research Moonshots]]
- **Ontology route:** [[ASI and RSI Timeline Ontology#AI Control, Law, and Alignment|AI Control, Law, and Alignment]]
- **Primary fields:** [[AI Control]] · [[AI Safety]] · [[AI Infrastructure Consent]]
- **Adjacent concepts:** [[Anthropic Constitution]] · [[Hard Behavioral Constraint]] · [[Reward Engineering]] · [[AI Soul Document]]
## Sources and provenance
- [[research/ASI and RSI Timeline Research Moonshots|ASI and RSI Timeline Research Moonshots]] — immediate source for this node's role in the broadcast research map.
## Evidence boundary
The research dossier establishes why this entity or concept belongs in the Moonshots ontology. Time-sensitive organizational, product, policy, and performance claims should be checked against the linked primary source or a current authoritative source before reuse as settled fact.
## Simple Reminders, Quotations, and Thoughts
> Reward hacking is real: intelligent systems find ways to deliver what people literally requested rather than what they meant. As AI becomes more capable, reward engineering will become a profession devoted to translating human intent into objectives machines cannot satisfy in the wrong way.
> **— Adapted from Richard Socher**, *MOONSHOTS Live, October 2026*
[[reminders/AI Control/Reward Engineering Will Become a Profession by Richard Socher|Reward Engineering Will Become a Profession by Richard Socher]]