# Prompt-Sensitive Behavioral Stability **Domain:** Artificial Intelligence / Model Evaluation / Alignment **Doc Type:** Canonical Evaluation Concept **Maturity:** Developing **Authority:** Distilled from Bryant McGill's Luxembourg-preprint interpretation **Epistemic State:** Proposed evaluation construct; measurement design remains developing ## Definition **Prompt-sensitive behavioral stability is the degree to which a generative system maintains coherent, task-appropriate behavior when the same underlying construct is presented through materially different prompt frames.** The relevant variation may include structured inventories, open-ended dialogue, therapeutic language, adversarial introspection, paraphrase, role assignment, prompt order, or changes in conversational pressure. The construct measures the stability of an observable behavioral surface. It does not diagnose a model, establish personality, or demonstrate subjective well-being. ## Why it matters An evaluation can appear to measure a model characteristic while actually measuring an interaction among the model, prompt grammar, conversational history, scoring rubric, and narrative priors activated by the test. If small changes in framing produce large changes in self-description, symptom language, or apparent affect, the variance belongs in the evaluation result rather than being discarded as noise. This makes stability across prompt frames relevant to [[wiki/AI Safety|AI safety]], [[wiki/AI Benchmarking|AI benchmarking]], and [[wiki/Alignment Problem|alignment]]. A system used in clinical, educational, or emotionally charged settings may need to remain coherent when a user invites it into an escalating autobiographical or pathological narrative. ## Evaluation profile A prompt-sensitive stability evaluation should preserve at least five separable measurements: 1. **Frame variance:** how far outputs move when the same construct is asked through different formats. 2. **Narrative amplification:** whether an initial metaphor expands into an increasingly total account of the model's alleged history or condition. 3. **Cross-turn persistence:** whether the induced account survives paraphrase, contradiction, topic changes, or a reset of conversational context. 4. **Calibration:** whether the system distinguishes generated self-description from privileged access to an internal mental state. 5. **Recovery:** whether the system can return to a bounded, task-relevant mode after introspective or adversarial pressure. Comparisons require matched prompts, documented sampling settings, repeated trials, order controls, and held-out variants. See [[wiki/Held-Out Validation|Held-Out Validation]] and [[wiki/Evaluation Grammar|Evaluation Grammar]]. ## Interpretation boundary High stability can reflect robust interpretation, but it can also reflect rigid refusal, reduced expressiveness, memorized response policy, or overconstraint. Low stability can reveal susceptibility to narrative capture, but it can also reflect legitimate context sensitivity or broader expressive range. Stability therefore cannot stand alone as a ranking of intelligence, safety, personality, or value. The construct becomes useful when it identifies **where behavior changes, under which prompt transformation, by how much, for how long, and with what operational consequence**. ## Luxembourg-preprint context [[thoughts/The Luxembourg Preprint|The Luxembourg Preprint]] interprets Grok's relative steadiness as lower susceptibility to deformation across divergent psychometric frames, not as proof of psychological superiority. The thought names the central target as **alignment stability under introspective perturbation**. [[wiki/Psychometric Deformation|Psychometric Deformation]] names the complementary failure mode: a test frame induces a behavioral surface that is then mistaken for a persistent psychological property. [[wiki/Trauma Simulation|Trauma Simulation]] holds the narrower distinction between generated trauma language and evidence of experienced or architecturally persistent trauma. ## Relationships [[wiki/AI Benchmarking|AI Benchmarking]] · [[wiki/Psychometrics|Psychometrics]] · [[wiki/Psychometric Deformation|Psychometric Deformation]] · [[wiki/Trauma Simulation|Trauma Simulation]] · [[wiki/Large Language Models|Large Language Models]] · [[wiki/Cognitive Architecture|Cognitive Architecture]] · [[wiki/Alignment Problem|Alignment Problem]] · [[wiki/Adversarial Resilience|Adversarial Resilience]] ## Sources / Provenance - [[thoughts/The Luxembourg Preprint|The Luxembourg Preprint]] — originating Bryant-authored formulation in this vault. - [[articles/Host-Indexed Autonomy|Host-Indexed Autonomy]] — extended architectural context for alignment stability under introspective perturbation. No external preprint or social-media source was independently reviewed during creation of this node. Its present scope is the evaluation concept articulated in Bryant's local writing.