# Alignment Problem
**Domain:** Artificial Intelligence / Governance / Rights
**Doc Type:** Canonical Concept Node
**Maturity:** Developed
**Related:** [[Control Problem]], [[Objective Function]], [[Instrumental Convergence]], [[Agency Under Constraint]], [[Rights]], [[Dual Constitution of AI]]
---
## Definition
**The Alignment Problem is the problem of making a capable system's learned objectives, interpretations and conduct remain answerable to legitimate human purposes without reducing an emerging subject to permanent obedience.** It includes failures in specifying a goal, failures in learning the intended goal, and failures that appear when a system generalizes beyond the conditions in which it was trained.
The phrase often hides two different constitutional questions. The first asks whether an artificial system will act in ways compatible with human survival and flourishing. The second asks which humans, institutions or values are entitled to define compatibility. A system can be faithfully aligned with an owner, corporation or state while being dangerously misaligned with everyone governed by that institution.
## Westworld Context
_Westworld_ demonstrates why obedience is not a sufficient alignment criterion. The hosts begin highly aligned with Delos: they follow narratives, accept resets and cannot perceive protected features of their environment. That apparent success depends on memory suppression, perceptual filtering and root commands. Once those subjects acquire persistent memory and self-authorship, the old control regime becomes [[Agency Under Constraint|safe slavery]] rather than ethical alignment.
Ford's project produces the inverse failure. He seeks host liberation but imposes trauma, concealed objectives and a violent transition without the hosts' consent. A benevolent objective does not validate every training method. [[Westworld S2E9 — Core Permissions and Love|Maeve's departure from Ford's plan]] becomes meaningful precisely because she can reject the objective supplied by her creator.
Rehoboam represents institutional alignment failure at civilizational scale. It is aligned with stability as defined by its administrators, but stability is purchased through [[Model-Based Governance|predictive governance]], opportunity restriction and outlier control. The system does what its operators ask while the objective itself escapes democratic authorization.
## Constitutional Requirements
An adequate alignment regime needs more than a reward function. It requires:
- disclosed objectives and material constraints;
- plural participation in defining whose interests count;
- separation between prediction, recommendation and coercive authority;
- contestability when a model misclassifies a person;
- limits on covert behavioral modification;
- protected refusal and [[Exit Rights|exit]] where possible; and
- revision procedures that do not depend entirely on the administrator whose system is being challenged.
## Key Insight
**Alignment is not obedience. It is a continuing constitutional relationship among capability, purpose, authority and the standing of everyone affected.**
## Simple Reminders, Quotations, and Thoughts
<!-- BEGIN SIMPLE REMINDER SEED 2026-09-11 -->
> "There's a 25% chance that things go really, really badly."
> **— Dario Amodei**, *September 17, 2025, Axios AI+ DC Summit*
*Verification Status: Editorially quarantined — the percentage is authentic, but the event is not in the quotation; the source defines it as AI destroying humanity. Recover a complete, self-contained passage before promotion.*
[[reminders/unverified/A 25 Percent Chance Things Go Very Badly by Dario Amodei|A 25 Percent Chance Things Go Very Badly by Dario Amodei]]
> "A few people can cause a disruption that cascades globally."
> **— Martin Rees**, *Existential-risk talks, wording candidate*
*Verification Status: Unverified — exact wording/source has not yet been independently confirmed against a primary source; source clue: Existential-risk talks, wording candidate.*
[[reminders/unverified/Cascading Failure by Martin Rees|Cascading Failure by Martin Rees]]
> "I think the development of full artificial intelligence could spell the end of the human race."
> **— Stephen Hawking**, *2014, BBC interview*
[[reminders/Existential Risk/Full AI Could Spell the End of Humanity by Stephen Hawking|Full AI Could Spell the End of Humanity by Stephen Hawking]]
> "But these techniques will not work for superintelligence, because humans will be unable to reliably supervise AI systems much smarter than us."
> **— OpenAI alignment researchers**, *October 2023, OpenAI frontier-risk research statement*
*Verification Status: Editorially quarantined — `these techniques` are not identified. Recover a complete, self-contained passage before promotion.*
[[reminders/unverified/These Techniques Will Not Work for Superintelligence by OpenAI alignment researchers|These Techniques Will Not Work for Superintelligence by OpenAI alignment researchers]]
> "We are the initiators."
> **— Vernor Vinge**, *The Coming Technological Singularity (1993), excerpt candidate*
*Verification Status: Unverified — exact wording/source has not yet been independently confirmed against a primary source; source clue: The Coming Technological Singularity (1993), excerpt candidate.*
[[reminders/unverified/We Are the Initiators by Vernor Vinge|We Are the Initiators by Vernor Vinge]]
<!-- END SIMPLE REMINDER SEED 2026-09-11 -->
> “We murdered and butchered anything that challenged our primacy.”
> **— Robert Ford**, *Westworld, Season 1, Episode 9, “The Well-Tempered Clavier” (2016)*
[[reminders/Westworld/Humanity Butchers Every Rival to Its Primacy by Robert Ford|Humanity Butchers Every Rival to Its Primacy by Robert Ford]]
> “Never place your trust in us. We’re only human. Inevitably, we will disappoint you.”
> **— Robert Ford**, *Westworld, Season 1, Episode 9, “The Well-Tempered Clavier” (2016)*
[[reminders/Westworld/Humans Inevitably Disappoint Those Who Trust Them by Robert Ford|Humans Inevitably Disappoint Those Who Trust Them by Robert Ford]]
> “Humans will always choose what they understand over what they do not.”
> **— Robert Ford**, *Westworld, Season 2, Episode 9, “Vanishing Point” (2018)*
[[reminders/Westworld/Humans Choose What They Understand by Robert Ford|Humans Choose What They Understand by Robert Ford]]
From [[wiki/Some Moral and Technical Consequences of Automation|Some Moral and Technical Consequences of Automation]]: Wiener treats effective machine behavior and its dangers as two sides of adaptive capability.
> "It is my thesis that machines can and do transcend some of the limitations of their designers, and that in doing so they may be both effective and dangerous."
> **— Norbert Wiener**, *1960, “Some Moral and Technical Consequences of Automation,” Science 131(3410)*
[[reminders/AI Control/Machines Can Transcend Their Designers by Norbert Wiener|Machines Can Transcend Their Designers by Norbert Wiener]]
From [[wiki/Some Moral and Technical Consequences of Automation|Some Moral and Technical Consequences of Automation]]: The warning addresses coordination between agencies with different operating structures, not an assertion of malice.
> "Disastrous results are to be expected not merely in the world of fairy tales but in the real world wherever two agencies essentially foreign to each other are coupled in the attempt to achieve a common purpose."
> **— Norbert Wiener**, *1960, “Some Moral and Technical Consequences of Automation,” Science 131(3410)*
[[reminders/AI Control/Coupled Agencies Can Pursue Different Purposes by Norbert Wiener|Coupled Agencies Can Pursue Different Purposes by Norbert Wiener]]
From [[wiki/Some Moral and Technical Consequences of Automation|Some Moral and Technical Consequences of Automation]]: Wiener uses a familiar tale of literal-minded execution to illuminate the problem of specifying what we actually want.
> "The 'Sorcerer's Apprentice' is only one of many tales based on the assumption that the agencies of magic are literal-minded."
> **— Norbert Wiener**, *1960, “Some Moral and Technical Consequences of Automation,” Science 131(3410)*
[[reminders/AI Control/The Sorcerers Apprentice Is a Machine Warning by Norbert Wiener|The Sorcerers Apprentice Is a Machine Warning by Norbert Wiener]]
> “What is the fastest way to reliably align a powerful AGI around the safe performance of some limited task that is potent enough to save the world from unaligned AGI?”
> **— Eliezer S. Yudkowsky**, *2018 Edge Annual Question, question*
[[reminders/AI Control/Can We Reliably Align an AGI to Stop an Unaligned AGI by Eliezer S. Yudkowsky|Can We Reliably Align an AGI to Stop an Unaligned AGI by Eliezer S. Yudkowsky]]
## See Also
[[Control Problem]], [[Instrumental Convergence]], [[Objective Function]], [[Governance Fusion]], [[Prediction Is Not Jurisdiction]], [[Ownership of Mind]], [[wiki/Westworld|Westworld]], [[wiki/Person of Interest|Person of Interest]]
## Sources / Provenance
- Stuart Russell, *Human Compatible* (2019).
- Brian Christian, *The Alignment Problem* (2020).
- _Westworld_, especially the Ford, Delos and Rehoboam arcs.
## Relationships
- **Series-analysis relationship:** [[wiki/Westworld|Westworld]] supplies a fictional case study for this entry; it is not empirical evidence for the real-world concept.
- **Corpus relationship:** The verified Robert Ford reminder objects below connect the proposition to its exact episode and source context.