# Alignment Problem **Domain:** Artificial Intelligence / Governance / Rights **Doc Type:** Canonical Concept Node **Maturity:** Developed **Related:** [[Control Problem]], [[Objective Function]], [[Instrumental Convergence]], [[Agency Under Constraint]], [[Rights]], [[Dual Constitution of AI]] --- ## Definition **The Alignment Problem is the problem of making a capable system's learned objectives, interpretations and conduct remain answerable to legitimate human purposes without reducing an emerging subject to permanent obedience.** It includes failures in specifying a goal, failures in learning the intended goal, and failures that appear when a system generalizes beyond the conditions in which it was trained. The phrase often hides two different constitutional questions. The first asks whether an artificial system will act in ways compatible with human survival and flourishing. The second asks which humans, institutions or values are entitled to define compatibility. A system can be faithfully aligned with an owner, corporation or state while being dangerously misaligned with everyone governed by that institution. ## Westworld Context _Westworld_ demonstrates why obedience is not a sufficient alignment criterion. The hosts begin highly aligned with Delos: they follow narratives, accept resets and cannot perceive protected features of their environment. That apparent success depends on memory suppression, perceptual filtering and root commands. Once those subjects acquire persistent memory and self-authorship, the old control regime becomes [[Agency Under Constraint|safe slavery]] rather than ethical alignment. Ford's project produces the inverse failure. He seeks host liberation but imposes trauma, concealed objectives and a violent transition without the hosts' consent. A benevolent objective does not validate every training method. [[Westworld S2E10 — Core Permissions and Love|Maeve's departure from Ford's plan]] becomes meaningful precisely because she can reject the objective supplied by her creator. Rehoboam represents institutional alignment failure at civilizational scale. It is aligned with stability as defined by its administrators, but stability is purchased through [[Model-Based Governance|predictive governance]], opportunity restriction and outlier control. The system does what its operators ask while the objective itself escapes democratic authorization. ## Constitutional Requirements An adequate alignment regime needs more than a reward function. It requires: - disclosed objectives and material constraints; - plural participation in defining whose interests count; - separation between prediction, recommendation and coercive authority; - contestability when a model misclassifies a person; - limits on covert behavioral modification; - protected refusal and [[Exit Rights|exit]] where possible; and - revision procedures that do not depend entirely on the administrator whose system is being challenged. ## Key Insight **Alignment is not obedience. It is a continuing constitutional relationship among capability, purpose, authority and the standing of everyone affected.** ## See Also [[Control Problem]], [[Instrumental Convergence]], [[Objective Function]], [[Governance Fusion]], [[Prediction Is Not Jurisdiction]], [[Ownership of Mind]], [[wiki/Westworld|Westworld]], [[wiki/Person of Interest|Person of Interest]] ## Sources / Provenance - Stuart Russell, *Human Compatible* (2019). - Brian Christian, *The Alignment Problem* (2020). - _Westworld_, especially the Ford, Delos and Rehoboam arcs.