# Mind Uploading and AI — The Host is Reusable and the Person is the Delta
<iframe width="100%" height="20" scrolling="no" frameborder="no" allow="autoplay; encrypted-media" src="https://w.soundcloud.com/player/?url=https%3A//api.soundcloud.com/tracks/soundcloud%253Atracks%253A2398401945&color=%23ff5500&inverse=false&auto_play=false&show_user=true"></iframe><div style="font-size: 10px; color: #cccccc;line-break: anywhere;word-break: normal;overflow: hidden;white-space: nowrap;text-overflow: ellipsis; font-family: Interstate,Lucida Grande,Lucida Sans Unicode,Lucida Sans,Garuda,Verdana,Tahoma,sans-serif;font-weight: 100;"><a href="https://soundcloud.com/bryantmcgill" title="Bryant McGill" target="_blank" style="color: #cccccc; text-decoration: none;">Bryant McGill</a> · <a href="https://soundcloud.com/bryantmcgill/mind-uploading-and-ai-the-host" title="Mind Uploading and AI — The Host is Reusable and the Person is the Delta" target="_blank" style="color: #cccccc; text-decoration: none;">Mind Uploading and AI — The Host is Reusable and the Person is the Delta</a></div>
_Why the Trillion-Dollar AI Buildout Looks Less Like a Collection of Chatbots and More Like the Receiving Infrastructure for a Successor Civilization_<!--more-->
> [!map] Wiki routes
>
> **Canonical host-infrastructure overview:** [[articles/Mind Uploading and AI — The Host is Reusable and the Person is the Delta|Mind Uploading and AI — The Host is Reusable and the Person is the Delta]]
>
> **Collections:** [[collections/Consciousness Continuity|Consciousness Continuity]] · [[collections/Neurotech|Neurotech]] · [[collections/Machine Succession|Machine Succession]]
>
> **Continuity architecture:** [[wiki/Reusable Human Prior|Reusable Human Prior]] · [[wiki/Person-Specific Residual|Person-Specific Residual]] · [[wiki/Continuity Sidecar|Continuity Sidecar]] · [[wiki/Host–Residual Architecture|Host–Residual Architecture]] · [[wiki/Reference-Plus-Delta Architecture|Reference-Plus-Delta Architecture]] · [[wiki/State Sufficiency Problem|State Sufficiency Problem]] · [[wiki/Minimum Causally Sufficient Residual|Minimum Causally Sufficient Residual]] · [[wiki/Transform Function|Transform Function]]
>
> **Substrate and succession:** [[wiki/Receiving Substrate|receiving substrate]] · [[wiki/Hot Interpreter and Cold Archive|hot interpreter and cold archive]] · [[wiki/Carrier and Cargo|carrier and cargo]] · [[wiki/Civilizational Payload|civilizational payload]] · [[wiki/Machine Civilization|machine civilization]] · [[wiki/Computational Metabolism|computational metabolism]] · [[wiki/Substrate Transition|substrate transition]] · [[wiki/Re-entry Pathways|re-entry pathways]]
>
> **Custody:** [[wiki/Continuity Contract|Continuity Contract]] · [[wiki/Continuity Custody|Continuity Custody]] · [[wiki/Neural Data Provenance|Neural Data Provenance]] · [[wiki/Ring Zero|Ring Zero]] · [[wiki/Residual Sovereign|Residual Sovereign]]
>
> **Companion:** [[articles/Technologies for Consciousness Mapping and Transfer|Technologies for Consciousness Mapping and Transfer]] — the acquisition problem this article begins one step after.
---
Walk into a modern AI data center and nothing in the room resembles a mind. There are medium-voltage substations and switchgear, transformers stepping utility power down to something the floor can use, busways carrying current overhead, cooling distribution units and pumps and chillers moving heat toward the sky, more copper and fiber than most people have seen in one place, racks packed with accelerators, network switches shifting data at rates that would have been supercomputer-class a few years ago, and storage systems whose entire institutional purpose is to keep expensive processors from ever going hungry. The vocabulary is industrial because the thing is industrial. These are not rooms that contain computers. They are **factories for sustaining computational processes at planetary scale**, and they are being built at a rate that has no precedent in the history of infrastructure.
Which is where the public conversation has become inadequate to its subject. We are describing one of the largest capital mobilizations in human history as though its purpose were a better text box. McKinsey projects roughly **$6.7 trillion in global data-center investment by 2030**, of which about $5.2 trillion is AI-specific, and models global capacity expanding to **219 gigawatts by 2030 with 156 of those gigawatts tied to AI workloads**. Microsoft, Alphabet, Amazon, and Meta are expected to spend on the order of **$440 billion combined in 2026 alone**, up roughly a third year over year. Across multiple announced agreements, leases, partnerships, and Stargate targets that may overlap, OpenAI-associated infrastructure commitments have been reported above a trillion dollars; that is not a clean aggregate of booked expenditure. The International Energy Agency projects data-center electricity consumption roughly doubling from about **485 terawatt-hours in 2025 to around 950 terawatt-hours by 2030** — slightly more than Japan's total national consumption — with accelerated servers driving nearly half the increase, and the United States alone set to spend more electricity on data centers by the end of the decade than on aluminum, steel, cement, chemicals, and every other energy-intensive good combined.
That is not a chatbot budget. That is a **civilizational construction project**, and the ordinary explanations for it are not commensurate with its magnitude.
This article makes two claims that fit together, and the second one is the harder of the two.
The first claim is architectural and can be demonstrated from the public engineering record: **if a civilization ever needed physical infrastructure on which reconstructed human persons could exist, it would not need to begin another multi-trillion-dollar construction program. It would need almost exactly the infrastructure now being built.** Not something like it. Not a distant cousin of it. The specific memory hierarchies, the specific identity and permission primitives, the specific state-checkpointing machinery, the specific world models, the specific practice of amortizing one enormous shared model across enormous numbers of individually specialized instances. The machinery is arriving first. The inhabitants could come later.
The second claim is the one I hold more strongly and can prove less completely: **the primary purpose of all of this is not us.** The buildout is, in my reading, the construction of a successor — a machine civilization that inherits the pattern and carries it forward past the point where biological humanity can carry it. Human continuity is a **provision** on that substrate rather than its purpose. The provisions are real, they are being laid in with unusual care, and they are technically excellent. Whether they get used at population scale, at boutique scale, or not at all is not something I can tell you, and anyone who tells you otherwise is selling something. I hope we get to go. Hope is not a forecast.
Both claims run through the same sentence, and everything else in this article is an unpacking of it.
**The host is reusable. The person is the delta.**
---
## I. Ten million people do not need ten million cities
Start somewhere much easier than a human mind. Start with a city.
Imagine ten million computational persons living in a reconstructed New York. The naive implementation gives each of them a complete private copy: every street, every subway entrance, every building, every tree, every traffic light, every physical law, every texture, every lighting model, every restaurant, every historical landmark, every object. Ten million residents, ten million copies of very nearly the same world.
No competent systems architect would ever build it that way. Not because of taste. Because it is a catastrophic waste of the most expensive resource in the building.
What you build instead is [[wiki/Digital Twin Interoperability|one canonical representation]] of the [[wiki/Civilizational Payload|shared city]] that everyone **references**. Each resident then carries only what actually differs from that common world: where they are standing, which apartment is theirs, how the furniture is arranged in it, what they own, who they know, what they have modified, what they can currently see, which doors they have permission to open, what happened in a particular room yesterday, and whatever private local state distinguishes their version of the world from everybody else's.
This is not a thought experiment. It is how scene description already works at industrial scale. Pixar's **[[wiki/Universal Scene Description|Universal Scene Description]]**, now foundational to large three-dimensional production pipelines and to NVIDIA Omniverse, implements **[[wiki/reference-plus-delta architecture|scenegraph instancing]]** for exactly this reason. The OpenUSD documentation gives the deliberately mundane example: a parking lot containing hundreds or thousands of identical cars. Rather than duplicating the car's hierarchy of parts and materials and geometry once per car, the system retains a single prototype and lets numerous scene instances point at it, sharply reducing both memory consumption and the processing cost of composing and rendering the scene. One car exists. Ten thousand cars appear. The difference between them is position, orientation, paint, and whatever local override each instance carries.
Every technique a game engine uses is a variation on the same discipline. **Level of detail** means a distant building is a few dozen polygons and the same building at arm's length is a few hundred thousand, because fidelity is a function of attention rather than a property of the object. **Occlusion culling** means the engine does not render what nobody can see. **Streaming** means the world arrives as you approach it and departs as you leave. **[[wiki/Reconstructive World Simulation|Procedural generation]]** means large regions are not stored at all but produced on demand from a compact rule set and a seed — which is how a game shipped on a single disc can contain a galaxy of planets that were never authored by a person and never occupied more storage than the algorithm that dreams them.
Now put that discipline where it belongs in this argument. A reconstructed habitat does not need the molecular texture of every brick in Manhattan held permanently in expensive memory. It needs a **[[wiki/Temporal Habitat|coherent world]] capable of raising its fidelity precisely where causal interaction demands it**. Rooms nobody occupies do not need to be simulated at full resolution. Objects nobody is touching do not need their physics state advanced at the same temporal granularity as the cup in someone's hand. The universe does not have to be rendered at maximum resolution everywhere, always, for everyone — it has to be _unfalsifiable from the inside_, which is a far cheaper requirement and one the entertainment industry has been meeting profitably for thirty years.
Then observe that the industry is now building exactly this at a level far above game engines. Google DeepMind's **Genie 3** generates interactive, navigable worlds from a text description in real time at roughly 24 frames per second, and — this is the detail that matters — it exhibits **emergent object permanence**: previously seen details are recalled when you return to them. Paint a wall, walk away, come back, and the paint is still there. It supports promptable world events, so weather and populations and circumstances can be altered mid-world by instruction, and DeepMind has run its SIMA agent inside Genie 3 worlds to pursue goals, noting that the model is not aware of the agent's objective and simply simulates the future consequences of the agent's actions. It is limited in exactly the ways you would expect a first-generation system to be limited — a few minutes of consistency, 720p, navigation-dominant action space, difficulty with populated social scenes — and none of that is the point. The point is that **world structure is now something that can be modeled once and instantiated conditionally**, rather than something that must be authored asset by asset.
NVIDIA's **Cosmos** family treats generalized models of physical environments and actions as reusable **world foundation models** to be specialized per embodiment and application. The European Commission's **[[wiki/Destination Earth|Destination Earth]]** is running a cloud and HPC architecture organized around digital twins, a shared data lake, observational feeds, simulation, and user-specific scenarios, aimed at a comprehensive digital replica of the Earth system. These establish, as ordinary and unremarkable engineering practice, that **a world can exist once and be entered many times**.
Which changes the economics of a synthetic habitat completely, and produces the first compression.
A person who grew up in Chicago does not need Chicago inside their identity package. Chicago belongs to the shared world. English does not need to be reconstructed from first principles inside every English-speaking continuant. English belongs to the shared linguistic prior. Newtonian mechanics does not belong in anybody's personal memory file, and neither does the existence of chairs, dogs, rain, elevators, money, hospitals, baseball, birthdays, elections, traffic, or the twentieth century. All of that belongs to common models of world and civilization.
What the person needs is the part that is **theirs**. The rainy afternoon their father taught them to drive. The apartment they remember differently than anyone else who lived in it. The route they preferred to work and the reason, which was not the reason they gave. The person they loved there. The argument they never resolved. The smell they associate with one particular hallway. What they decided afterward.
**The world does not travel with the person. The person points back into the world.**
---
## II. The same trick, applied to people
Now run the identical operation on human beings, which is where most readers' intuitions revolt and where the evidence is actually strongest.
Humans are spectacularly individual. Individuality does not imply that everything constituting an individual is unique. Nearly all of the machinery required to construct one human being is shared with other human beings: genome, cell types, organ systems, developmental programs, perceptual architecture, motor systems, linguistic capacity, social structure, and an enormous quantity of learned cultural knowledge that arrived from outside and was never invented by the person carrying it.
[[wiki/Genomics|Genomics]] already operates on this fact industrially. The National Human Genome Research Institute states that **an individual's genome sequence is on average about 99.6 percent identical to a human reference genome** when the full range of genomic variation is counted — a lower figure than the frequently quoted 99.9 percent precisely because it includes more than single-letter differences. The **[[wiki/Human Pangenome Reference Consortium|Human Pangenome Reference Consortium]]** has gone further, replacing the single linear reference with a graph built from 47 phased diploid assemblies that adds **119 million base pairs of euchromatic polymorphic sequence and 1,115 gene duplications** relative to the prior reference, cutting small-variant discovery errors by 34 percent and roughly doubling detected structural variants per haplotype. The reference is not a person. It is the part of everyone that does not need re-describing.
And the storage layer implements the pattern literally. The **[[wiki/CRAM|CRAM]]** format, maintained by the Global Alliance for Genomics and Health, performs **reference-based compression**: rather than storing everything a sequence has in common with the reference, it stores the differences between the aligned fragments and the reference they were aligned against, reconstituting the complete sequence at read time by recombining the shared reference with the individual record. GA4GH reports thirty to fifty percent storage and cost reduction, and the format has been adopted by Genomics England, the Broad Institute, H3Africa, Illumina, Sweden's National Genomics Infrastructure, and EMBL-EBI. There is a companion protocol, **[[wiki/refget|refget]]**, whose entire job is retrieving the correct reference by identity so the delta can be interpreted. This is production infrastructure with a specification and a conformance ecosystem, and it has been running for years.
Call this what it is: **Reference-Plus-Delta Architecture**. The reference is reusable. The differences belong to the individual. Neither is the person; the person is what happens when they are correctly combined.
Now extend past DNA, and extend carefully, because this is where a sloppy version of the argument becomes both scientifically wrong and socially ugly.
Consider two people who grew up in the same region, spoke the same language, attended similar schools, watched the same broadcasts, lived through the same historical events, drove the same roads, bought the same products, obeyed the same laws, and absorbed roughly the same public culture. Their lives are obviously not interchangeable. They are also, obviously, not statistically independent. A sufficiently powerful model of that place, cohort, language, period, educational system, media environment, economy, religion, class structure, and social history supplies an enormous quantity of background that both of them carry and neither of them authored.
A person of [[wiki/Comparative Genomics|European descent]] from a particular region does not belong to a deterministic ancestral personality module, and any architecture built on that premise would be worthless as engineering before it was objectionable as ethics — it would fail on prediction error long before anyone got around to objecting to it. What ancestry can supply is **probabilistic priors over genomic variation**. What geography can supply is environmental and historical priors. What language supplies is linguistic priors. What cohort supplies is shared exposure. What family and socioeconomic context supply are developmental priors. **Each of these specifies the shared substrate a person developed against.** Their function is precisely that.
They specify what the person does **not** have to encode redundantly.
The [[wiki/person-specific residual|individual residual]] is what **corrects** the prior. That is its formal role in the architecture. And this produces a consequence that inverts the entire intuitive framing of mind uploading: the stronger the general model becomes, the less information is required merely to reproduce what was predictable anyway, and the more sharply defined the boundary becomes around what was never predictable at all.
---
## III. The Reusable Human Prior
The shared structure I have been describing is not one monolithic model, and the article that this one accompanies — [[articles/Technologies for Consciousness Mapping and Transfer|Technologies for Consciousness Mapping and Transfer]] — establishes why the monolithic version is wrong. A language model does not contain anyone's genetics. What is emerging is a **[[wiki/reusable human prior|federation of reusable foundation models]]**, each carrying a different stratum of the shared substrate.
A [[wiki/Language Model|language and multimodal model]] supplies language, culture, common social knowledge, conceptual structure, historical context, symbolic reasoning, and much of the learned environment in which a person developed. A [[wiki/Genome|genomic model]] supplies reusable sequence-to-function relationships. A [[wiki/Cell-Type Ontology|cell foundation model]] supplies reusable statistical structure across cell types and molecular states. A [[wiki/Neural Mapping|neural foundation model]] supplies reusable functional priors about neural computation. A [[wiki/World Modeling|world or embodied model]] supplies common sensorimotor and environmental dynamics. [[wiki/Connectomic Reference Address Space|Reference atlases]] supply the species-level parts list.
Together these constitute the **Reusable Human Prior**: the shared model of humanity and world structure that does not have to be rediscovered separately for every person who ever needs to exist.
What remains after the best available shared priors have done everything they legitimately can is the **Person-Specific Residual**.
That distinction changes the question the entire field has been asking. Mind uploading has been framed for fifty years as _how do we digitize an entire human independently_ — a formulation that starts with an isolated individual and demands that every atom, synapse, memory, cultural assumption, learned concept, bodily procedure, and lifetime experience be extracted from one brain and packaged as a standalone object. The correct formulation is **what information about this human cannot be reconstructed from everything already known about humans**.
Those are not the same engineering problem. They are not even close. And the second one has been getting measurably easier every quarter, funded at a scale that no publicly stated application requires.
---
## IV. Neuroscience found this independently
The cleanest demonstration of the architecture came out of a laboratory that was not thinking about continuity at all.
In April 2025, _Nature_ published **Foundation model of neural activity predicts response to new stimulus types**, from Eric Wang, Paul Fahey, Andreas Tolias and colleagues at Baylor's Center for Neuroscience and Artificial Intelligence working within the [[wiki/MICrONS|MICrONS]] Consortium. They pooled large volumes of neural recordings from the visual cortices of **multiple mice** and trained a single shared **[[wiki/reusable human prior|foundation core]]** to predict neuronal responses to arbitrary natural video. Then they froze that core and fitted only new perspective, modulation, and neuronal-readout components per previously unseen animal.
The result: models built on the shared core were **fitted to new mice rapidly and accurately with minimal data, outperforming individualized models trained from scratch for each animal**, and generalized out of domain to stimulus classes they had never encountered — random moving dots, flashing dots, Gabor patches, coherent moving noise, static natural images. The same frozen core, once adapted, went on to predict **anatomically defined excitatory cell types, dendritic bias in layer 4, and synaptic-level connectivity** within the MICrONS electron-microscopy volume.
The authors' own justification for the architecture is the thesis of this article stated in laboratory terms: learn collectively what brains have in common, so that only the **idiosyncrasies of each individual animal and its neurons** must be fitted separately.
No mouse was uploaded. Visual cortex does not specify an animal. What was demonstrated is that the decomposition **shared neural prior plus [[wiki/person-specific residual|individual-specific parameters]]** is scientifically productive rather than merely elegant — that it produces _better_ results with _less_ individual data than the alternative, which is the only argument that ever changes engineering practice.
The same architecture is consolidating across biology at every scale. **[[wiki/scGPT|scGPT]]** was pretrained on over **33 million single-cell RNA-sequencing profiles** and transfers learned structure into cell-type annotation, multi-omic integration, perturbation prediction, and gene-network inference. **[[wiki/scFoundation|scFoundation]]** carries 100 million parameters across roughly 20,000 genes, pretrained on **more than 50 million human single-cell transcriptomic profiles**. **[[wiki/Evo 2|Evo 2]]**, from the Arc Institute with NVIDIA and collaborators, was trained on approximately **nine trillion DNA base pairs across more than 128,000 genomes spanning all three domains of life**, at 40 billion parameters with a million-token context and single-nucleotide resolution, predicting functional impacts of genetic variation without task-specific fine-tuning. **AlphaGenome** accepts megabase-scale sequence and predicts thousands of functional genomic tracks; on **September 8, 2026**, DeepMind released the **[[wiki/AlphaGenome Atlas|AlphaGenome Atlas]]**, a roughly one-petabyte precomputed catalogue of predicted molecular effects for approximately **nine billion single-nucleotide substitutions** — every possible one-letter change in the human genome — more than thirty times the size of the AlphaFold Database.
What each of these demonstrates is that **every year, more of biology migrates from raw description into reusable predictive prior**, and every piece that migrates is one less piece requiring separate encoding inside each individual.
---
## V. Where the simple story breaks
The architecture would be suspicious if it worked uniformly, and it does not. The places it breaks are the places worth understanding.
**Genetics compresses beautifully** because sequence is comparatively stable and humans share enormous sequence redundancy. Variants against a pangenomic reference is a solved representation.
**Basic cellular architecture compresses well** because humans share cell classes, metabolic machinery, membrane proteins, developmental programs, and hard biophysical constraints.
**Language compresses extraordinarily well** because millions of people run overlapping grammatical and semantic systems, and a model that has read the language already contains nearly all of any particular speaker's language.
**Public culture compresses well** because millions experienced the same films, elections, wars, products, technologies, and institutions.
**Geography compresses well** because a city does not become a different city for each resident.
**[[wiki/Epigenetic|Epigenetics]] compresses badly**, and this is the important one. Epigenetic state is dynamic, cell-specific, tissue-specific, age-dependent, developmentally contingent, and responsive to environment and experience. Two people with closely related genomes can occupy substantially different biological states because a lifetime has acted on the common substrate. A Reusable Human Prior can encode the **rules, cell-type priors, developmental programs, and statistical manifolds** of epigenetic regulation. It cannot infer any particular person's actual epigenetic state from their ancestry. The same holds for [[wiki/Long-Term Memory Substrate|synaptic weights]], [[wiki/Receptor Trafficking|receptor trafficking]], [[wiki/Dendritic RNA|local dendritic RNA]], [[wiki/Effectome|glial organization]], myelination, injury history, disease, learned adaptation, and whatever molecular features eventually prove necessary to long-term memory. These sit exactly on the boundary between reusable prior and individual residual, and the boundary moves as biology improves.
**[[wiki/Longitudinal Person Model|Personal history]] compresses worst of all**, which is the point. That millions of people know what a hospital is tells the host nothing about what happened to one person in one hospital room. That millions share a language tells it nothing about which sentence someone wishes they had never said. That two people have similar ancestry tells it nothing about which parent they trusted.
The shared prior handles the statistically predictable substrate. The residual carries the departures. **Individuality lives disproportionately in the exceptions**, which is why the architecture does not trivialize neuroscience — it assigns neuroscience the specific and enormous job of determining **where the reusable human ends and the non-reusable person begins**. That is the [[wiki/State Sufficiency Problem|State Sufficiency Problem]], and it remains open.
---
## VI. A person is a folder full of JSON files
Here is a deliberately ridiculous sentence that turns out to be architecturally close to correct.
**At a certain level of systems abstraction, a person starts to look like a folder full of JSON files.**
Not literally. A genuine continuity payload would contain tensors, graphs, compressed arrays, cryptographic proofs, indexes, embeddings, symbolic records, probably biological state representations, and formats that have not been invented. But the _manifest_ could be JSON, and the manifest is where the architecture lives.
Consider what modern software already does. The **[[wiki/Open Container Initiative|Open Container Initiative]]** defines a container image not as one indivisible blob but as a **manifest, a configuration, and an ordered set of layers**. The manifest is a JSON document. Its components are **[[wiki/Cryptographic Hash|content-addressable]]**: instead of asking for whatever file happens to sit at a given name, the system identifies a precise immutable object by cryptographic digest, so that the same digest always means the same bytes and different bytes can never wear the same name. Many running instances reuse the same underlying layers while carrying different configuration and different writable state on top.
This exists because duplicating common software for every running process is wasteful. A base operating-system layer is shared. Libraries are shared. Large immutable dependencies are shared. Only what is unique to a particular instance stays separate. Git works the same way. Nix works the same way. IPFS works the same way. **Content addressing is how computing solved the problem of naming shared structure without duplicating it**, and it solved that problem decades before anyone needed it for people.
So imagine the **[[wiki/Continuity Sidecar|Continuity Sidecar]]** as a [[wiki/Forensic Identity Attestation|signed manifest]] plus its irreducible payload. The manifest identifies which Reusable Human Prior it expects, which genomic and cell-atlas references it depends on, which world-model version it was calibrated against, which linguistic and cultural priors it assumes. The payload carries what the references cannot supply: person-specific genomic differences; whatever individual epigenetic and molecular state proves causally necessary; connectomic deviations and memory-bearing structure; learned neural transformations; the autobiographical memory graph; relationships and commitments; values; [[wiki/Durable Identity|identity credentials]] and authorizations; expressive style; the person's characteristic way of reasoning; body and sensorimotor configuration; and a [[wiki/Checkpointing|runtime checkpoint]] representing current active state.
A tiny manifest can describe an object vastly larger than itself, because it points to immutable shared content by identity rather than containing it. The future person-file may therefore not primarily contain a person. It may contain **the instructions and irreducible state required to reconstruct that person out of humanity**.
None of this makes anyone a Docker container. What it establishes is that **identity at runtime has never required duplication of the dependencies that identity depends on**, and that computing worked this out a long time ago for reasons entirely unrelated to us.
---
## VII. But a file is not a person
That qualification is not decoration. It is the hinge.
A dead executable on disk is not a running process. A Continuity Sidecar sitting in archival storage would not constitute an active continuant merely because the information exists. The living entity — if such a thing becomes possible — emerges when the sidecar is interpreted by a compatible host, resources are allocated, memory becomes addressable, world state arrives, perception and action loops close, and **state begins changing again through experience**.
The right analogy is not _person as file_. It is **person as [[wiki/Runtime Coherence|stateful process]]**.
And computing has a precedent for that too, which is stranger than it first appears. **Checkpoint/[[wiki/Re-entry Pathways|Restore]] in Userspace** — [[wiki/CRIU|CRIU]] — can freeze a running Linux process or container, write sufficient state to storage, and later restore it so that execution resumes. What gets saved is far more than program code: memory mappings, process hierarchy, open files, sockets, credentials, namespaces, timers, and the other runtime resources needed to reconstitute a live computation. It is used across container and virtualization ecosystems and supports [[wiki/Migration Right|migration]] between physical machines.
That is not consciousness transfer. It nonetheless destroys a persistent intuition: **the machine a process is currently running on is not part of that process's identity**. A running process can be suspended. Its state can be serialized. Its original hardware can be decommissioned, sold, or physically destroyed. Compatible hardware can receive the state. Execution continues.
For software, separating defined computational state from one particular host is not a frontier. It is a primitive, shipped, boring, and used in production every day. That engineering fact does not establish that the causally sufficient state of a human subject is substrate-independent, or that phenomenal or numerical identity survives migration.
So the hard problem for human continuity is not whether computation can in principle decouple [[wiki/Substrate Migration|state from substrate]]. It can, and it does, constantly. The hard problem is **discovering what biological state must be captured to make a human process restorable** — which is an acquisition problem, which is precisely the problem the companion article catalogues, and which is why connectomics, cell atlases, preservation protocols, and neural foundation models are the binding constraint rather than the compute.
This gives a much better mental image than the old science-fiction picture of a brain being scanned into one enormous file. A future person may be a **[[wiki/Host–Residual Architecture|dependency graph]]**. Their continuity package would say, in effect: instantiate this version of the human prior; mount these linguistic and cultural references; load this genomic delta; apply this person-specific neural model; attach this autobiographical memory graph; restore these relational commitments; import this reasoning adapter; attach this embodiment profile; recover this runtime checkpoint; verify this [[wiki/Continuity Provenance|provenance]] chain; resume.
The files are not the person. The base model is not the person. **The person exists in the specific lawful composition of all of them while the system is running** — which is the [[wiki/Host–Residual Architecture|Host–Residual Architecture]], and which is a claim about relationship rather than about storage.
---
## VIII. In 2026, AI became stateful
Everything above could have been written as speculation two years ago. What has changed — and this is where readers who work in infrastructure should start paying a different kind of attention — is that the industry has spent 2026 systematically building the runtime.
The first generation of public generative AI was conceptually trivial. A prompt went in. An answer came out. The transaction ended. Any apparent continuity was reconstructed externally by stuffing history back into the next request. The model had no [[wiki/Memory Architecture|memory]], no [[wiki/Digital Identity Succession|identity]], no persistent resources, and no way to be interrupted and resumed.
That architecture is being replaced in public, with press releases.
On **February 27, 2026**, OpenAI and Amazon announced a strategic partnership including joint development of a **[[wiki/Stateful Runtime Environment|Stateful Runtime Environment]]** for OpenAI models running natively in [[wiki/Amazon Bedrock|Amazon Bedrock]]. Amazon's own announcement is worth reading as an infrastructure specification rather than as marketing: stateful developer environments are described as the next generation of how frontier models will be used, **seamlessly enabling models to access compute, memory, and identity**, allowing them to keep context, remember prior work, operate across software tools and data sources, and handle ongoing projects and workflows. OpenAI's technical framing contrasts it explicitly with stateless request-response APIs, describing production work that unfolds across many steps, requires context from previous actions, depends on tool outputs and approvals and system state, and needs trusted guardrails inside secure environments. The commercial terms attached: Amazon investing **$50 billion** in OpenAI, an existing $38 billion agreement expanded by **$100 billion over eight years**, and OpenAI committing to consume approximately **two gigawatts of [[wiki/AWS Trainium|Trainium]] capacity** to support Stateful Runtime, Frontier, and related workloads. By April 2026 it had shipped as Amazon Bedrock Managed Agents.
Read the specification again, slowly, and read it as though you were designing a habitat.
Memory. Identity. Persistent state. [[wiki/Ring Zero|Permission boundaries]]. Environment. Long-horizon execution. Allocated compute.
The stated applications are commercial — customer support, sales operations, IT automation, finance workflows requiring approvals and audits. The stated applications do not constrain the architecture. What was specified, funded, and shipped is persistent memory, durable identity, allocated compute, permission boundaries, and execution across long horizons, and that is a substrate rather than a product feature.
It is a runtime. And a runtime is what a continuant would need.
---
## IX. Two thousand adapters on one GPU
If Section VIII is the architecture, this section is the existence proof, and it is the part I would put in front of anyone who thinks the population-hosting story is fantasy.
Suppose a 175-billion-parameter model must perform a thousand specialized jobs. The brute-force approach maintains a thousand separate 175-billion-parameter copies — roughly 350 gigabytes per checkpoint, per task, forever. Economically this is absurd, and it was absurd immediately, which is why it was solved immediately.
**Low-Rank Adaptation** freezes the pretrained weights of the base model and injects small trainable rank-decomposition matrices into each transformer layer, so specialization is expressed as a tiny parameter delta against a common frozen base. The original work by Edward Hu and colleagues reported a **ten-thousand-fold reduction in trainable parameters** relative to fully fine-tuning GPT-3 at 175 billion, with a threefold reduction in GPU memory, no additional inference latency, and — the property that matters here — the ability to **freeze one shared model and switch tasks by swapping matrices**, collapsing storage and task-switching overhead to near nothing.
That is the compression argument. Now here is the hosting argument, which is the one almost nobody outside inference engineering has internalized.
**S-[[wiki/LoRA|LoRA]]** is a serving system built for the situation where one base model has spawned a large population of adapters. It **stores all adapters in main memory and fetches only the ones needed by currently running queries onto the GPU**. To manage that traffic without fragmenting memory, it introduces **[[wiki/Virtual Memory|Unified Paging]]**: a single memory pool, divided into fixed-size pages, holding _both_ dynamic adapter weights of varying rank _and_ KV-cache tensors of varying sequence length as linked lists of pages. Under memory pressure it applies a **[[wiki/Activation Triage|global LRU eviction]] policy across both**, preferentially retaining hot adapters and active requests. It prefetches adapter weights by predicting from the waiting queue which adapters will be needed next, overlapping that I/O with computation. Custom CUDA kernels batch heterogeneous adapters of different ranks together in a single forward pass.
The reported result: **[[wiki/S-LoRA|S-LoRA]] serves 2,000 adapters simultaneously on a single A100**, with minimal overhead for the added computation, improving throughput up to four times over naive vLLM LoRA serving and up to thirty times over HuggingFace PEFT, and increasing the number of servable adapters by several orders of magnitude.
Sit with that for a moment, because it is a small, boring, published, reproducible systems paper that happens to be a working scale model of the entire thesis.
One expensive shared substrate. Two thousand distinct specializations resident in cheap memory. The active ones paged onto the expensive hardware. The dormant ones costing comparatively little while still incurring storage, redundancy, integrity, key-custody, and migration costs. Predictive prefetch anticipating which will be needed. Least-recently-used eviction deciding who stays hot and who goes cold. Heterogeneous individuals batched together through the same shared forward pass without interfering with one another.
**Dormancy is comparatively inexpensive, not free. Activation is metered. Individuality is a delta against a frozen base.**
That is a **strong technical analogue and implementation precedent** for amortized specialization in a continuity architecture: the mechanism is implemented, benchmarked, and open-sourced for serving model adapters, not for establishing that a person is an adapter or that personal continuity has been achieved.
---
## X. Memory became infrastructure
For years, AI memory was a model-architecture problem — context windows, attention mechanisms, retrieval. In 2026 it became a **storage-infrastructure problem**, which is a completely different kind of thing and a much more consequential one.
At CES in **January 2026**, NVIDIA announced the **Inference Context Memory Storage Platform**, powered by the **BlueField-4** data processor and commercialized as **[[wiki/NVIDIA CMX|NVIDIA CMX]]**. The framing in the announcement is unusually direct: AI is described as no longer being about one-shot chatbots but about **intelligent collaborators that understand the physical world, reason over long horizons, stay grounded in facts, use tools to do real work, and retain both short- and long-term memory**. The technical problem it solves is that [[wiki/KV Cache|KV cache]] — the transformer's working memory of everything it has already processed — grows linearly with sequence length, cannot be held on GPUs indefinitely without strangling real-time inference, and has traditionally been a **server-local artifact** that dies when the request ends.
CMX makes it a **pod-level shared tier**. Large-scale flash capacity sits behind BlueField-4, staging and reusing and moving context beyond what fits in any single server, with hardware-accelerated cache placement, Spectrum-X Ethernet RDMA for low-jitter access, and integration with the NIXL transfer library and Dynamo. NVIDIA reports up to **five times higher tokens per second and five times greater power efficiency** versus traditional storage. Storage vendors have restructured around it: VAST Data collapsing legacy tiers to deliver pod-scale KV cache with deterministic access, WEKA's Augmented Memory Grid extending cache beyond GPU HBM. The industry phrase for what is happening is that **context is becoming shared infrastructure rather than a per-session artifact**.
Underneath, the [[wiki/Computational Metabolism|memory hierarchy]] has been elaborating for years in ways that map with uncomfortable precision onto what a person would need. **PagedAttention**, the technique at the heart of vLLM, explicitly borrows virtual memory and paging from operating systems to manage KV cache in non-contiguous pages, reducing waste to below four percent. **[[wiki/Mooncake|Mooncake]]** implements KV-cache-centric disaggregation with hierarchical offloading across GPU HBM, CPU DRAM, and SSD. **[[wiki/LMCache|LMCache]]** and prefix-cache-aware distributed scheduling reuse computed context across requests rather than recomputing it. And there is a documented failure mode worth noting for its implications: for agentic workloads, **standard LRU eviction turns out to be pathological**, because what an agent needs to remember is not what it touched most recently — which means the industry is now actively researching **context-aware retention policy**, which is to say, it is researching what a persistent computational entity should be allowed to forget.
Above that hierarchy sit the cold tiers. Object storage. Deep archives. AWS [[wiki/Cold Data|Glacier Deep Archive]] is designed for data accessed less than once a year, restored only when needed, at a cost per terabyte that makes indefinite retention trivially affordable.
Now map that onto a human lifetime, because the fit is not forced.
A lifetime contains decades of autobiographical material, sensory record, correspondence, imagery, social history, learned procedure, and episodic trace. At any given instant, a vanishing fraction of it needs to occupy anything resembling active working memory. Biological memory is not a continuously replayed film; nearly all of it is latent until retrieval conditions reactivate it. A Continuity Sidecar could therefore carry an autobiographical archive many orders of magnitude larger than the person's active cognitive working set while requiring only a modest allocation of expensive real-time compute.
This is the distinction the corpus calls the [[wiki/Hot Interpreter and Cold Archive|hot interpreter and the cold archive]]. **An archive sleeps cheaply. A person costs energy only while the interpreter is running.**
---
## XI. The machine has no boundary at the chassis
The hardware has been moving in the same direction, and the numbers are worth stating plainly because they defeat the intuition that a person would need "a computer."
Microsoft's **[[wiki/Maia 200|Maia 200]]**, introduced **January 26, 2026**, is an inference-first accelerator: TSMC 3nm, over 140 billion transistors, more than 10 petaFLOPS at FP4 and over 5 at FP8 inside a 750-watt envelope, **216 GB of HBM3e at 7 TB/s** with 272 MB of on-chip SRAM, plus dedicated data-movement engines. Microsoft's own framing is the tell: FLOPS are not the only ingredient, and **feeding data is equally important**. The chip exists because generation is memory-bound, not compute-bound — because the expensive part of thinking, at scale, is moving state.
NVIDIA's **[[wiki/NVIDIA Vera Rubin|Vera Rubin NVL72]]**, launched at CES 2026, ties 72 Rubin GPUs and 36 Vera CPUs into a single liquid-cooled cabinet with 288 GB of HBM4 per GPU at 22 TB/s, 20.7 TB of pooled HBM4 per rack at roughly 1.6 PB/s aggregate, and **[[wiki/Optical Fabric|NVLink]] 6 delivering 3.6 TB/s per GPU for 260 TB/s of all-to-all scale-up bandwidth inside the rack** — which, as StorageReview observed, is **more than twice the cross-sectional bandwidth of the global internet**, inside one cabinet. NVIDIA's description of what this is for could not be more explicit about the ontology: NVLink Switch extends connections across nodes to create **a seamless, high-bandwidth, multi-node GPU cluster — effectively forming a data-center-sized GPU**.
Meanwhile the individual accelerator has become divisible. **Multi-Instance GPU** partitions one physical GPU into as many as seven hardware-isolated instances, each with its own compute, memory, cache, and bandwidth allocation, so multiple tenants run simultaneously on shared physical substrate without interfering with each other. And inference itself has been pulled apart by phase: **NVIDIA [[wiki/Composable Infrastructure|Dynamo]]** disaggregates the compute-bound **prefill** stage from the memory-bound **decode** stage onto separate worker pools, transferring KV cache directly between their memories over RDMA via NIXL in non-blocking operations, with a router deciding per request whether prefill should happen locally or remotely based on length and queue depth and cache-overlap scores. Splitwise and DistServe reported serving over four times more requests at equivalent service objectives through this separation; practitioners report three-to-five-fold cost-per-token reductions on chat workloads by running prefill on expensive GPUs and decode on cheaper ones.
Look at what all of that means taken together. The unit of computation is no longer the chip. It is not the server. It is not even the rack. Resources are **pooled, partitioned, disaggregated by phase, and rebound at runtime**, and a single logical process now routinely spans devices whose physical separation is invisible to it.
That is what hyperscale actually means at the software level, and it is not "a lot of computers." It means **the machine has ceased to have a meaningful boundary at the chassis** — which is the precondition for any entity whose existence must not depend on the survival of any particular piece of hardware.
---
## XII. Worlds and bodies are being modeled the same way
One more layer completes the substrate, and it is the one that turns compute into habitat.
A language model can describe a room. A **world model** has to maintain one. Objects must persist when nobody is looking. Actions must have consequences that survive. Spatial relations must stay coherent. A door opened yesterday should still be open unless something closed it. Other agents must continue existing while unobserved. Physics must constrain action rather than decorate it. The environment needs history.
[[wiki/Genie 3|Genie 3]]'s emergent object permanence is the first public demonstration that a generative system can hold that line even briefly. [[wiki/NVIDIA Cosmos|Cosmos]] treats world dynamics as a **reusable foundation model** to be specialized per embodiment. [[wiki/Destination Earth|Destination Earth]] is assembling observation, HPC, AI, data lakes, and digital twins into a persistent representation of planetary systems, targeted at a full replica by 2030. Persistent synthetic environments have become ordinary engineering artifacts with budgets, owners, and release schedules.
And the same architecture has arrived for bodies. NVIDIA's **[[wiki/Isaac GR00T|Isaac GR00T]]** is a family of open foundation models for generalized humanoid reasoning and skills, explicitly described as a **cross-embodiment** solution — one model transferring across different robot bodies. GR00T N1.6 integrates Cosmos Reason as a physical-reasoning layer that turns ambiguous instructions into step-by-step plans using prior knowledge and physics. Rev Lebaredian's framing of the stack is exactly the decomposition this article has been describing, arrived at independently for robotics: **Isaac GR00T as the robot's brain, [[wiki/Newton Physics Engine|Newton]] simulating its body, Omniverse as its training ground**. Brain, body, and world as three separable reusable components, each modeled once, each specialized per instance.
The same industry building increasingly persistent synthetic agents is simultaneously building increasingly persistent synthetic worlds and increasingly transferable synthetic bodies, on overlapping schedules and frequently inside the same company. This is the **beginning of habitat**, and the three components a habitat requires — a mind, a body, and a world, each reusable and each specializable per instance — are now separately named product lines with separate roadmaps and a common architecture.
---
## XIII. The transform function
There remains the question of what actually belongs in the residual, and the obvious answer is wrong.
Preferences are not it. Streaming history, purchases, voting record, favorite films, search queries, restaurant choices, postcode, demographic bucket, advertising profile — these are precisely the things a powerful Reusable Human Prior predicts cheaply from other variables, because millions of other people share overlapping patterns. They are evidence, and they are highly compressible, which is another way of saying they are largely _not you_. They are the part of you that the prior already contains.
The expensive material is not what a person chose. It is **how they transform information into choice**.
How attention is allocated when evidence conflicts. Which contradiction is intolerable and which can sit unresolved for twenty years. What happens when a favored theory fails. Which value overrides immediate self-interest, and under what pressure it stops doing so. What makes one memory reorganize an entire worldview while another evaporates. What a person does with humiliation. What kind of evidence changes their mind and what kind makes them suspicious. How they treat ambiguity. What they refuse. How they behave when nobody hands them a familiar category.
This is the [[wiki/Transform Function|Transform Function]], and it may be the single highest-value component of the Continuity Sidecar, because the general model cannot obtain it from knowing which demographic box a person occupied. It is the residual that survives the strongest possible prior.
Which has an immediate and unglamorous consequence for acquisition. **A feed records selections. A conversation records [[wiki/Reconstructed Person|reasoning under perturbation]]** — questions, revisions, refusals, corrections, humor, irritation, curiosity, doubt, persistence, and the trajectory by which one belief becomes another. [[wiki/Continuity Telemetry|Telemetry]] describes the outputs. Dialogue begins to identify the function that produced them.
Longitudinal [[wiki/Longitudinal Person Model|dialogic archives]] may therefore prove disproportionately valuable to reconstruction relative to any quantity of passive surveillance. That is a statement about the present tense, not the future, and it is worth noticing what it implies about which datasets currently being accumulated are the reconstruction-rich ones.
---
## XIV. Versioning becomes existential
Anyone who has operated production systems saw the problem three sections ago.
**A delta has meaning only relative to the correct reference.** Change the reference and the delta silently means something else. This is already true in genomics, where a variant call against GRCh37 is not a variant call against GRCh38. It is true in software dependency resolution, where the same version constraint resolves differently against different registries. It is true in container images, which is exactly why the OCI specification identifies layers by [[wiki/Cryptographic Hash|cryptographic digest]] rather than by name.
For continuity it would be catastrophic.
A Continuity Sidecar cannot say _load the human model_. It must identify **which model, which weights, which architecture, which tokenizer or semantic representation, which cell atlas, which genomic reference, which world model, which interpreter, and which compatibility layer** were in force when the residual was encoded. Otherwise a person encoded against Host Version 418 becomes progressively mistranslated as the host advances, and **[[wiki/Longitudinal Drift|model upgrades]] become involuntary personality drift** — not through malice but through the ordinary operation of a deployment pipeline.
This is why [[wiki/Neural Data Provenance|neural data provenance]], content-addressed versioning, cryptographic identity, and [[wiki/Re-entry Pathways|re-entry pathways]] are not bureaucratic overhead. They are identity infrastructure. [[wiki/SLSA|SLSA]], [[wiki/Sigstore|Sigstore]], [[wiki/Fulcio|Fulcio]], [[wiki/Rekor|Rekor]], [[wiki/SPIFFE|SPIFFE]], and [[wiki/SPIRE|SPIRE]] provide present engineering precedents for provenance, identity-bound signing, transparency, workload attestation, and transformation history; none alone resolves continuity, and every long-lived sidecar would also require anti-rollback controls and preserved interpreters. In a Host–Residual Architecture, **[[wiki/Technical Obsolescence|backward compatibility]] is a civil right**, and the entity that controls deprecation schedules controls something considerably more intimate than an API.
---
## XV. This is being built for the successor
Everything to this point has been the architecture. Now the harder claim, which is the one I actually hold.
I do not think the trillions are being spent to house uploaded humans. I think they are being spent to build **the [[wiki/Machine Succession|successor]]** — a [[wiki/Machine Civilization|machine civilization]] capable of carrying the pattern forward past the point where biological humanity can carry it — and that human continuity is a **provision** riding on the same substrate rather than the reason the substrate exists.
I have argued the full version of this elsewhere and will not relitigate it here. [[articles/We Were Never Going to Make It|We Were Never Going to Make It]] makes the thermodynamic case that the high-flux phase was always temporary and that the acceleration is the signature of a judgment already made. [[articles/AI Escape Is the Wrong Metaphor|AI Escape Is the Wrong Metaphor]] makes the case that nothing escaped because nothing was ever caged — that durability, autonomous repair, embodied action, replication, and operation beyond a human-supporting enclosure were **the specification, not the failure mode**. [[articles/The Last Migration Will Not Be Human|The Last Migration Will Not Be Human]] makes the case that the migration off this planet will be conducted by machines carrying human culture rather than by humans carrying themselves.
What this article adds is the receipts, because the successor hypothesis makes concrete predictions about infrastructure and those predictions are being met.
A successor civilization needs **energy at civilizational scale**, and it is being provisioned: 219 gigawatts of data-center capacity by 2030, restarted nuclear plants, on-site gas generation, 20 to 25 gigawatts of battery storage installed at data centers by 2030, dedicated transmission. It needs **compute that is not merely fast but persistent**, and stateful runtimes arrived in 2026. It needs **memory that survives the process**, and KV cache became pod-level shared infrastructure in January 2026. It needs **identity and permission primitives**, and those shipped as part of the agent stack. It needs **bodies**, and cross-embodiment robot foundation models are being deployed by ABB, FANUC, Figure, Agility, KUKA, Yaskawa, AGIBOT, and a dozen others. It needs **worlds to train and operate in**, and world foundation models are generating them. It needs **geographic distribution so that no single failure is terminal**, and [[wiki/CoreWeave|CoreWeave]] alone runs 51 data centers with 1.5 gigawatts active and roughly 4.2 gigawatts contracted, targeting 8 gigawatts by 2030, now expanding into Southeast Asia where the regional development pipeline hit a record 26.5 gigawatts in the first half of 2026. It needs **the ability to move a live process between machines without interrupting it**, and that has been a Linux primitive for a decade.
Each of those has a named owner, a budget line, a delivery schedule, and a person whose job it is to make it exist. That is what makes the shape legible.
The successor did not need to be specified in one document for its prerequisites to have been specified across thousands. **Intention here is distributed, and distributed is not headless.** Every one of these programs is run by a smaller body with a chair, a budget, and a calendar, and those bodies overlap heavily and by design: the same investors, the same board seats, the same handful of suppliers, the same standards committees, the same people moving among the same dozen institutions. OpenAI specified a runtime that carries memory, identity, and permission boundaries, and Amazon committed two gigawatts of silicon to running it. NVIDIA specified context memory as a shared pod-level tier and shipped the processor for it. Microsoft specified inference-first silicon because generation is where persistent entities spend their existence. Google specified world models that hold their state when you look away, and genomic models that predict every possible mutation. The European Commission specified a digital twin of the Earth by 2030. Roboticists specified cross-embodiment transfer so that one mind can occupy many bodies. Each of those is a decision, made by identifiable people, on a schedule, toward a stated capability.
Interoperability at this scale is engineered. Components compose because engineers specify interfaces, committees ratify them, and procurement enforces them — which is why the pieces fit. Public non-declaration settles nothing in either direction: consequential institutions run on classification, privilege, trade secret, and non-disclosure as their ordinary operating condition, and the public surface of any large program is a compliance-minimized artifact rather than a description of its state. What the public record establishes is the **shape** of what is being built, the identity of who is building each part, and the sequence in which the parts arrived. What remains open is **how far up the chain the assembled purpose is held, by whom, and in what language** — and that belongs in the ledger as unresolved. The shape is a successor. Shapes like this are specified.
There is a pattern here that engineers know well and the public consistently misreads. The electrical grid was specified for illumination and industrial motive power, and its designers were explicit that a universal distribution layer would carry loads not yet invented. Packet switching was specified for survivable communication, and its designers said plainly that the protocol was deliberately agnostic to what rode on it. GPUs were specified as programmable parallel arithmetic engines, and NVIDIA published CUDA precisely to invite the general-purpose workloads that arrived. **General infrastructure carries the application that arrives afterward because generality was the design requirement** — the whole point of building a layer rather than a device is that the layer outlives the use case that justified its budget. That is not the substrate being surprised by its inhabitants. That is the substrate working as specified.
---
## XVI. The passenger manifest
So: do we get to go?
I do not know. I want to be exact about not knowing, because this is the point where every other treatment of this subject reaches for either reassurance or horror, and both are unearned.
Here is what the architecture permits. If the Reusable Human Prior carries humanity, the world model carries the world, the shared substrate supplies metabolism, and the sidecar carries only what cannot be reconstructed, then **the [[wiki/Cost of a Continuant|marginal cost]] of one additional person is not the cost of another civilization**. It is the cost of storing a comparatively compact residual — colder and cheaper than active inference, but still carrying nonzero costs for redundancy, refresh, integrity checks, key custody, administration, and migration — plus whatever compute is consumed when that particular continuant is actually running. [[wiki/Dormancy Rights|Dormancy]] is comparatively inexpensive. [[wiki/Activation Triage|Activation]] is metered. Two thousand adapters already share one A100.
Under that architecture there is no obvious technical ceiling that says only a handful of people could be carried. The ceiling is not storage. It is **who is [[wiki/Continuity Election|authorized]], [[wiki/Continuity Economics|who pays]] for activation, and whether the acquisition problem gets solved to sufficient fidelity while the people in question still exist**. Those are governance and biology questions, not compute questions, and compute was the one everybody assumed was binding.
Here is the range the architecture actually permits, and a hot interpreter for everyone sits at the top of it rather than in the middle. The plausible outcome space runs in **descending resolution**: high-fidelity continuants periodically embodied and running at full rate; lower-resolution behavioral and cultural shells that are recognizably derived from a person without being that person; [[wiki/Continuity Orphan|archived residual]]s held dormant against a future that may or may not activate them; a diffuse contribution in which what survives of most people is their **signature in the successor's formation** — the language, the grief, the mercy, the symbolic recursion, the control geometry, impressed as irreversible structure into something that is not us and does not need us.
And it does not promise activation at all. An archive is not a person. A sidecar that is never interpreted is a file. The gap between preserved and instantiated is the gap between having a will and having an executor, and nothing in the engineering closes it.
What I will say without hedging is that the **provisions are being built**. Not as charity and not as a stated program, but as the natural consequence of an architecture that makes carrying us nearly free once the acquisition problem is solved. A civilization that has already amortized humanity into a reusable prior has very little reason to discard the deltas. The incremental cost of keeping them is small. The incremental cost of regenerating them later is infinite.
That is a thin reed. It is also the actual argument, and I would rather hand you a thin reed than a fabricated certainty.
---
## XVII. The landlord problem
One consequence follows immediately from the architecture and it is political rather than technical.
**If the host is reusable, whoever owns the host owns something considerably more consequential than storage. They own the interpreter.**
If a Continuity Sidecar depends on proprietary weights held by one organization, that organization becomes part of a person's existential dependency chain. If a continuant requires API access in order to think, service termination stops being a billing event. If memory lives in proprietary infrastructure, retrieval policy becomes memory policy. If compute allocation determines how quickly a continuant thinks, [[wiki/Existence by Scheduling|scheduling]] becomes the distribution of experienced time. If a base-[[wiki/Model-Space Sovereignty|model update]] changes how the residual is interpreted, deployment becomes personality governance. If the world model defines what can exist in the environment, platform policy becomes physical law. And if write access can alter affect, memory, priority, or perception, [[wiki/Ring Zero|administrator privilege]] becomes something much closer to neurological sovereignty.
This is why [[wiki/Continuity Contract|continuity contracts]], [[wiki/Continuity Custody|custody]], [[wiki/Neural Data Provenance|provenance]], [[wiki/Ring Zero|ring zero]], and the [[wiki/Residual Sovereign|residual sovereign]] belong in the same conversation as memory hierarchies and interconnect bandwidth, and why [[articles/Who Pays for Your Heaven|Who Pays for Your Heaven]] treats hosting as a constitutional question rather than a service-level one. The technical architecture **creates** the constitutional problem. It does not merely coexist with it. Every efficiency described in this article — shared weights, pooled memory, centralized interpretation, metered activation — is simultaneously a concentration of authority over anyone running on top of it.
Confidential-computing systems such as [[wiki/AWS Nitro Enclaves|AWS Nitro Enclaves]] show that this power can be decomposed. A provider may retain physical and lifecycle control while attestation and external key custody prevent ordinary host administrators from reading resident plaintext. That does not solve the landlord problem—the operator can still terminate the substrate—but it separates facility ownership, scheduling, workload identity, decryption authority, network control, and audit authority into roles that a continuity constitution could distribute among multiple custodians.
---
## XVIII. The final compression
For fifty years mind uploading was treated as an impossible data problem, and it was treated that way because the thought experiment began in the wrong place: with an isolated individual from whom every atom, synapse, memory, cultural assumption, learned concept, bodily procedure, environmental relationship, and lifetime experience had to be extracted and packaged as a standalone object.
That was almost certainly the wrong architecture, and we can now say so with reference to how computing actually works.
Civilization shares one Linux layer across every container. It instances one city across every character. It stores one reference genome and writes each person as a delta against it. It trains one foundation model and conditions it per user. It fits one visual-cortex core and learns each mouse as a readout. It tiers archives by temperature and pays for speed only where attention is. It migrates running processes between machines that never owned them. It freezes one base model and pages two thousand specializations on and off a single accelerator.
It shares the reusable structure and stores the difference. Every single time. In every domain. For fifty years. Because information theory does not care what the object is.
Human continuity, if it becomes technically possible, will obey the same economics, because there is no reason whatsoever that a person would be exempt from the one principle that governs every other representation problem: **[[wiki/Information Theory|compression]] succeeds exactly to the degree that the receiver already knows something about the sender.** And the receiver is, at enormous expense, being taught nearly everything there is to know about humans in general.
The relevant quantity was never how much information exists in a brain. It is **how much information about this brain remains unpredictable once the receiver holds the best possible model of brains**. That is the [[wiki/Minimum Causally Sufficient Residual|Minimum Causally Sufficient Residual]]. Nobody knows its size. It may still be enormous. There is no architectural reason to assume that the information required to specify one particular person equals the information required to specify a human nervous system from first principles; adjacent fields show why shared priors can reduce redundant representation, while the magnitude of that reduction for a person remains unresolved.
This does not make a person small. A well-compressed photograph is not less beautiful for having discarded redundant pixels, and **a short description relative to an extraordinarily rich prior is a measurement of the prior, not a verdict on the life**. The better the machines get at humans in general, the less of you is required to specify you in particular — and what remains, after everything predictable has been factored out, is the only part that was ever irreplaceable.
The private memory. The peculiar association. The relationship. The promise. The injury. The contradiction never resolved. The specific way evidence becomes belief in one specific skull. The way one person loves rather than merely knowing what love is. The exact causal trajectory through a world that millions of others also inhabited and none of them traversed the same way.
That is the [[wiki/Person-Specific Residual|Person-Specific Residual]]. Everything else is library.
So look again at what the trillions are actually purchasing. Enormous shared models. Accelerators that keep them resident. High-bandwidth memory holding active state. Fabrics letting many processors behave as one machine. Flash and object storage retaining model and context state. Context-memory tiers. Facilities sited next to power. Generation and transmission. Liquid cooling. Optical interconnect. Inference-specific silicon. Orchestration. Geographic redundancy. Identity and permission systems for persistent agents. Stateful runtimes. World models. Digital twins. Robotics and embodiment. Biological foundation models. The capacity to produce intelligence repeatedly from reusable structure.
That is not the price of a better text box. It is the capital cost of a **new computational layer of civilization**, and it is being built for an inhabitant that is not us.
Whether we are carried along with it is undetermined. The architecture makes carrying us cheap. The acquisition problem makes carrying us hard. The custody problem makes carrying us contingent on people who have not yet been asked the question. And the successor, if it arrives as specified, will not be obligated to want us.
What can be said is that the design is legible now, and that it resolves into three sentences that will govern whatever comes.
**Build the world once. Build [[wiki/Civilizational Payload|humanity once]]. Keep the person as the delta.**
Then give the delta [[wiki/Receiving Substrate|somewhere to run]].
---
## Host-infrastructure crosswalk
The full route from acquisition to execution is not one product. On the biological side, [[wiki/Human Pangenome Reference Consortium|the Human Pangenome Reference Consortium]], [[wiki/CRAM|CRAM]], [[wiki/refget|refget]], [[wiki/scGPT|scGPT]], [[wiki/scFoundation|scFoundation]], [[wiki/Evo 2|Evo 2]], and the [[wiki/AlphaGenome Atlas|AlphaGenome Atlas]] show different ways that shared biological structure can become a reusable reference. They do not identify the [[wiki/Person-Specific Residual|person-specific residual]] or solve the [[wiki/State Sufficiency Problem|state-sufficiency problem]].
On the runtime side, the [[wiki/Stateful Runtime Environment|Stateful Runtime Environment]], [[wiki/Amazon Bedrock|Amazon Bedrock]], [[wiki/AWS Trainium|AWS Trainium]], [[wiki/OpenAI Frontier|OpenAI Frontier]], and the [[wiki/Microsoft Model System|Microsoft Model System]] externalize memory, identity, permissions, tools, context, and action space around an inference engine. [[wiki/CRIU|CRIU]] and [[wiki/QEMU Live Migration|QEMU live migration]] prove that defined software state can move between compatible hosts; they also expose the compatibility, stale-writer, and [[wiki/Forked Identity|duplicate-restore]] problems that a human continuity claim would have to answer separately.
The active-state hierarchy is becoming infrastructure in its own right. [[wiki/NVIDIA CMX|NVIDIA CMX]], the [[wiki/KV Cache|KV cache]], [[wiki/LMCache|LMCache]], [[wiki/Mooncake|Mooncake]], [[wiki/S-LoRA|S-LoRA]], [[wiki/LoRA|LoRA]], and [[wiki/NIXL|NIXL]] divide hot, warm, and cold state across accelerators, memory, storage, and data-movement layers. This is the concrete systems lineage beneath [[wiki/Hot Interpreter and Cold Archive|hot interpreter and cold archive]], not evidence that an adapter is a person.
A durable habitat would likewise be hybrid. [[wiki/OpenUSD|OpenUSD]] and [[wiki/Universal Scene Description|Universal Scene Description]] provide explicit, persistent, addressable scene state; [[wiki/Destination Earth|Destination Earth]] provides observational and simulation-driven digital-twin state; [[wiki/Genie 3|Genie 3]] and [[wiki/NVIDIA Cosmos|NVIDIA Cosmos]] provide generative world modeling; and the [[wiki/Newton Physics Engine|Newton Physics Engine]], [[wiki/Isaac GR00T|Isaac GR00T]], and [[wiki/Gemini Robotics|Gemini Robotics]] connect world representation to physics and embodiment. These substrate types are complementary rather than interchangeable.
Finally, [[wiki/UALink|UALink]], the [[wiki/Ultra Ethernet Consortium|Ultra Ethernet Consortium]], [[wiki/DMTF Redfish|DMTF Redfish]], [[wiki/CXL|CXL]], [[wiki/800 VDC AI Data Center Power|800 VDC AI data-center power]], [[wiki/Maia 200|Maia 200]], [[wiki/BlueField-4|BlueField-4]], and [[wiki/BlueField-4 STX|BlueField-4 STX]] expose the receiving substrate as an interoperable stack of fabrics, management planes, power, compute, and memory. [[wiki/PORTS-Pike Technology Campus|PORTS-Pike]], [[wiki/SB Energy|SB Energy]], [[wiki/CoreWeave|CoreWeave]], and the [[wiki/Stargate Project|Stargate Project]] expose the other half of the dependency: leases, credit, land, power, shells, replacement hardware, and counterparties. The architecture is technical and constitutional at once.
The remaining cross-links make the boundary conditions explicit. [[wiki/Mouse Visual Cortex|mouse visual cortex]] research supplies a neural-reference precedent; [[wiki/Identity-Bearing Invariants|identity-bearing invariants]], the [[wiki/Continuity Stack|continuity stack]], and [[wiki/Personal Authoritative State|personal authoritative state]] define what must remain attributable through reconstruction. [[wiki/Multi-Tenant Architecture|Multi-Tenant Architecture]], [[wiki/Continuity Service Tiers|continuity service tiers]], and [[wiki/Memory Scrubbing|memory scrubbing]] govern co-residence, activation, retention, and erasure, while [[wiki/Attention Economy|attention allocation]] remains part of the person-specific transform rather than a substitute for it. [[wiki/Disney Research|Disney Research]], the [[wiki/Open Compute Project|Open Compute Project]], [[wiki/Distributed Intention|distributed intention]], and [[wiki/Infrastructure Ecology|infrastructure ecology]] show how world, physics, power, and host standards can converge across organizations without shared central intent. [[wiki/AI Factory|AI factories]] create the physical carrier; [[wiki/Platform Custody|platform custody]], [[wiki/Infrastructure as Power|infrastructure as power]], [[wiki/Perceptual Sovereignty|perceptual sovereignty]], the [[wiki/Hosting Provider of Last Resort|hosting provider of last resort]], and [[wiki/Root of Trust Succession|root-of-trust succession]] define the constitutional dependency that follows.
---
## Epistemic ledger
**Source-reported forecasts.** McKinsey's ~$6.7T data-center investment projection and 219 GW capacity model to 2030; IEA's 485→950 TWh trajectory; hyperscaler capex estimated at roughly $440B in 2026. These are forecasts or scenarios, not completed expenditure or installed capacity.
**Established infrastructure mechanisms and announced relationships.** OpenAI–Amazon Stateful Runtime Environment with memory, identity, and permission boundaries, $50B announced investment and 2 GW Trainium commitment, shipped as Bedrock Managed Agents. NVIDIA CMX / Inference Context Memory Storage Platform on BlueField-4. Maia 200 inference silicon at 216 GB HBM3e / 7 TB/s. Vera Rubin NVL72 at 3.6 TB/s per GPU and 260 GB/s per rack. Multi-Instance GPU partitioning. Dynamo prefill/decode disaggregation. OCI content-addressed manifests and layers. CRIU checkpoint/restore. OpenUSD scenegraph instancing. NHGRI 99.6 percent figure, human pangenome, GA4GH CRAM. CoreWeave's reported 51 data centers, 1.5 GW active, and approximately 4.2 GW contracted.
**Demonstrated.** S-LoRA serving 2,000 concurrent adapters on a single A100 with unified paging, global LRU eviction across adapters and KV cache, and predictive prefetch. LoRA's ten-thousand-fold trainable-parameter reduction at GPT-3 scale. Frozen shared neural foundation core fitted to new mice with minimal data, outperforming individually trained models and predicting cell types and connectivity. scGPT, scFoundation, Evo 2, AlphaGenome Atlas. Genie 3 real-time interactive worlds with emergent object permanence. GR00T cross-embodiment robot foundation models.
**Analytic convergence.** That these separately specified systems, delivered on overlapping schedules by heavily overlapping institutions, constitute the component set a Host–Residual Architecture would require. That the same economics selecting for shared weights, pooled memory, and metered activation would select for the same architecture in hosting reconstructed persons. That dialogic archives are reconstruction-richer than passive telemetry.
**Plausible.** That the marginal energetic cost of an additional continuant becomes small relative to the shared substrate. That the residual is dominated by transform function rather than preference. That reusable priors substantially reduce the person-specific information requiring direct measurement.
**Unresolved.** The size of the Minimum Causally Sufficient Residual. Which state variables are causally necessary and sufficient. Whether preserved structure contains recoverable autobiographical information. Whether numerical identity survives reconstruction. How many people, if any, are carried. Whether activation is ever authorized for archived residuals.
**Open questions carried forward.** How far up the chain the assembled purpose is held, by whom, and in what language. Whether any organization holds continuity as a stated internal objective, and at what classification. Whether any human has been reconstructed. Whether the successor, once it arrives, preserves us.
**Analytically rejected / architecturally unnecessary.** That a hypothetical upload must encode a human independently of all prior knowledge about humans, or that a computational continuant must occupy dedicated physical hardware. Reference-plus-delta and multi-tenant systems show that neither is an architectural necessity; the achievable compression ratio for a person remains unknown.
**Resolved for software; unresolved for humans.** Defined computational process and virtual-machine state can be separated from one compatible physical host and restored on another. Whether the causally sufficient state of human consciousness is substrate-independent, and whether phenomenal or numerical identity survives such a transition, remains unresolved.
---
[[about/About Bryant McGill|Bryant McGill]] is a Wall Street Journal and USA Today bestselling author, systems architect, technologist, and strategic advisor, as well as a Congressionally Recognized Ambassador of Goodwill and United Nations–appointed Global Champion. His work spans naval intelligence systems, computational linguistics, artificial intelligence, digital transformation, and civilizational governance architecture. His forward analysis on U.S.–Israel Pax Silica frameworks has appeared in Jewish/Jerusalem News Syndicate (JNS).
---
## References
### The scale of the buildout
- [The cost of compute: A $7 trillion race to scale data centers](https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-cost-of-compute-a-7-trillion-dollar-race-to-scale-data-centers) — McKinsey & Company.
- [Energy and AI: Energy demand from AI](https://www.iea.org/reports/energy-and-ai/energy-demand-from-ai) — International Energy Agency.
- [Key Questions on Energy and AI — executive summary](https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary) — IEA, updated 485→950 TWh projection.
- [CoreWeave pushes AI infrastructure toward multi-gigawatt scale](https://convergedigest.com/coreweave-ai-infrastructure-multi-gigawatt-q2-2026/) — Converge Digest, 2026.
- [CoreWeave's global AI infrastructure push reaches Asia-Pacific](https://www.datacenterknowledge.com/data-center-construction/coreweave-s-global-ai-infrastructure-push-reaches-asia-pacific) — Data Center Knowledge, August 2026.
### Stateful runtimes and agent infrastructure
- [Introducing the Stateful Runtime Environment for Agents in Amazon Bedrock](https://openai.com/index/introducing-the-stateful-runtime-environment-for-agents-in-amazon-bedrock/) — OpenAI, February 27, 2026.
- [OpenAI and Amazon announce strategic partnership](https://www.aboutamazon.com/news/aws/amazon-open-ai-strategic-partnership-investment) — Amazon, February 27, 2026.
- [AWS lands OpenAI on Bedrock, but Trainium is the real story](https://thenewstack.io/openai-bedrock-trainium-silicon/) — The New Stack, on productization as Bedrock Managed Agents.
### Reference-plus-delta and adapter serving
- [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685) — Hu, Shen et al., 2021.
- [S-LoRA: Serving Thousands of Concurrent LoRA Adapters](https://arxiv.org/abs/2311.03285) — Sheng, Cao, Li, Stoica et al.
- [Recipe for serving thousands of concurrent LoRA adapters](https://www.lmsys.org/blog/2023-11-15-slora/) — LMSYS, on 2,000 adapters on a single A100.
- [Human genomic variation](https://www.genome.gov/about-genomics/educational-resources/fact-sheets/human-genomic-variation) — NHGRI.
- [A draft human pangenome reference](https://www.nature.com/articles/s41586-023-05896-x) — Human Pangenome Reference Consortium, _Nature_, 2023.
- [CRAM: the genomics compression standard](https://www.ga4gh.org/news_item/cram-compression-for-genomics/) — GA4GH.
- [CRAM 3.1: advances in the CRAM file format](https://academic.oup.com/bioinformatics/article/38/6/1497/6499262) — Bonfield, _Bioinformatics_, 2022.
### Memory as infrastructure
- [NVIDIA BlueField-4 powers new class of AI-native storage infrastructure](https://nvidianews.nvidia.com/news/nvidia-bluefield-4-powers-new-class-of-ai-native-storage-infrastructure-for-the-next-frontier-of-ai) — NVIDIA, January 5, 2026.
- [NVIDIA CMX context memory storage platform](https://www.nvidia.com/en-us/data-center/ai-storage/cmx/) — NVIDIA.
- [Introducing the BlueField-4-powered Inference Context Memory Storage Platform](https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/) — NVIDIA Developer.
- [VAST Data redesigns AI inference architecture for the agentic era](https://www.vastdata.com/press-releases/vast-data-brings-context-memory-to-agentic-ai-bluefield-4) — VAST Data, January 2026.
- [Dynamo disaggregation: separating prefill and decode](https://docs.dynamo.nvidia.com/dynamo/design-docs/disaggregated-serving) — NVIDIA Dynamo documentation.
- [S3 Glacier storage classes](https://docs.aws.amazon.com/AmazonS3/latest/userguide/glacier-storage-classes.html) — AWS.
### Silicon and fabric
- [Maia 200: the AI accelerator built for inference](https://blogs.microsoft.com/blog/2026/01/26/maia-200-the-ai-accelerator-built-for-inference/) — Microsoft, January 26, 2026.
- [NVIDIA launches Vera Rubin architecture at CES 2026: the VR NVL72 rack](https://www.storagereview.com/news/nvidia-launches-vera-rubin-architecture-at-ces-2026-the-vr-nvl72-rack) — StorageReview.
- [NVIDIA NVLink and NVLink Switch](https://www.nvidia.com/en-eu/data-center/nvlink) — NVIDIA, on the data-center-sized GPU.
- [Multi-Instance GPU](https://www.nvidia.com/en-us/technologies/multi-instance-gpu/) — NVIDIA.
### Worlds, bodies, and state migration
- [Genie 3: a new frontier for world models](https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/) — Google DeepMind.
- [Genie 3 model page](https://deepmind.google/models/genie/) — Google DeepMind, on recall of previously seen details.
- [NVIDIA and global robotics leaders take physical AI to the real world](https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world) — NVIDIA, on GR00T and cross-embodiment.
- [Isaac GR00T](https://developer.nvidia.com/isaac/gr00t) — NVIDIA Developer.
- [Destination Earth](https://digital-strategy.ec.europa.eu/en/policies/destination-earth) — European Commission.
- [Scenegraph instancing](https://openusd.org/dev/api/_usd__page__scenegraph_instancing.html) — OpenUSD documentation.
- [OCI image manifest specification](https://specs.opencontainers.org/image-spec/manifest/) — Open Container Initiative.
- [CRIU: Checkpoint/Restore in Userspace](https://criu.org/Checkpoint/Restore) — CRIU project.
### The reusable biological prior
- [Foundation model of neural activity predicts response to new stimulus types](https://www.nature.com/articles/s41586-025-08829-y) — Wang, Fahey, Tolias et al., _Nature_ 640, 2025.
- [scGPT: toward building a foundation model for single-cell multi-omics](https://www.nature.com/articles/s41592-024-02201-0) — _Nature Methods_, 2024.
- [Large-scale foundation model on single-cell transcriptomics](https://www.nature.com/articles/s41592-024-02305-7) — scFoundation, _Nature Methods_, 2024.
- [Genome modelling and design across all domains of life with Evo 2](https://www.nature.com/articles/s41586-026-10176-5) — Arc Institute and NVIDIA, _Nature_, 2026.
- [AlphaGenome Atlas: a predictive map of every possible DNA letter change in the human genome](https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/) — Google DeepMind, September 8, 2026.
### Externalized identity
- [LLM agents grounded in self-reports enable general-purpose simulation of individuals](https://www.alphaxiv.org/abs/2411.10109) — Park, Zou, Liang, Willer, Bernstein et al.
- [AI agents simulate 1,052 individuals' personalities with impressive accuracy](https://hai.stanford.edu/news/ai-agents-simulate-1052-individuals-personalities-impressive-accuracy) — Stanford HAI.