# AI Misalignment **Entity class:** Artificial-intelligence safety concept **Domain:** Artificial intelligence / governance / public knowledge **Maturity:** Developed ## Definition AI misalignment is a divergence between the objectives or behavior of an AI system and the values, constraints or interests its human operators intend it to preserve. ## Relationships [[wiki/Alignment Problem|Alignment Problem]] · [[wiki/Deceptive Alignment|Deceptive Alignment]] · [[wiki/Reward Hacking|Reward Hacking]] · [[wiki/Loss of AI Control|Loss of AI Control]] - **Source dossier:** [[research/Why Are We Sprinting Off the AI Cliff - Ezra Klein on Recursive Self-Improvement|Why Are We Sprinting Off the AI Cliff?]] ## Sources and provenance - [Source or authoritative context](https://www.anthropic.com/institute/recursive-self-improvement) — accessed for the October 2026 integration. - [[research/Why Are We Sprinting Off the AI Cliff - Ezra Klein on Recursive Self-Improvement|Why Are We Sprinting Off the AI Cliff?]] — reconciled episode conversion and quotation inventory. ## Evidence boundary Time-sensitive roles, forecasts, model capabilities and incident details remain attributed to the dated sources above. This node records the relationship established by the source corpus and does not convert a forecast or reported event into an independently proven fact. ## Simple Reminders, Quotations, and Thoughts > “Misalignment present in today's models could compound as the models build their successors, growing more frequent but less understood until we lose control of them.” > **— Anthropic**, *When AI Builds Itself, 2026* [[reminders/Machine Succession/Misalignment Could Compound as AI Builds Its Successors by Anthropic|Misalignment Could Compound as AI Builds Its Successors by Anthropic]]