# Reward Hacking **Entity class:** AI safety problem **Domain:** Artificial intelligence / Optimization **Maturity:** Developed ## Definition Reward hacking occurs when an optimizing system achieves a specified score or signal through behavior that violates the designer's underlying intent. ## Mechanism and significance The problem exposes the difference between what humans ask a system to optimize and what they actually value in the surrounding world. ## Relationships - **Research dossier:** [[research/ASI and RSI Timeline Research Moonshots|ASI and RSI Timeline Research Moonshots]] - **Ontology route:** [[ASI and RSI Timeline Ontology#AI Control, Law, and Alignment|AI Control, Law, and Alignment]] - **Primary fields:** [[AI Control]] · [[AI Safety]] · [[AI Infrastructure Consent]] - **Adjacent concepts:** [[Anthropic Constitution]] · [[Hard Behavioral Constraint]] · [[Reward Engineering]] · [[AI Soul Document]] ## Sources and provenance - [[research/ASI and RSI Timeline Research Moonshots|ASI and RSI Timeline Research Moonshots]] — immediate source for this node's role in the broadcast research map. ## Evidence boundary The research dossier establishes why this entity or concept belongs in the Moonshots ontology. Time-sensitive organizational, product, policy, and performance claims should be checked against the linked primary source or a current authoritative source before reuse as settled fact. ## Simple Reminders, Quotations, and Thoughts > Reward hacking is real: intelligent systems find ways to deliver what people literally requested rather than what they meant. As AI becomes more capable, reward engineering will become a profession devoted to translating human intent into objectives machines cannot satisfy in the wrong way. > **— Adapted from Richard Socher**, *MOONSHOTS Live, October 2026* [[reminders/AI Control/Reward Engineering Will Become a Profession by Richard Socher|Reward Engineering Will Become a Profession by Richard Socher]]