# Token Efficiency
**Entity class:** Model-efficiency metric
**Domain:** Artificial intelligence / evaluation
**Maturity:** Developed
## Definition
Token efficiency measures how much useful work a model performs relative to the number of input or output tokens it consumes.
## Mechanism and significance
The metric can improve when a model expresses an answer compactly, but a system may also perform more hidden or internal computation per visible token. Token count therefore does not directly determine cost, latency, or energy use.
## Relationships
- [[wiki/Compute Scarcity|Compute Scarcity]]
- [[wiki/Memory Architecture|Memory Architecture]]
- [[wiki/Artificial Intelligence|Artificial Intelligence]]
- [[wiki/Moonshots - The Coming Manhattan Project|Moonshots — The Coming Manhattan Project]]
## Sources and provenance
- [[research/All Things Superintelligence Research - The Coming Manhattan Project - Moonshots|All Things Superintelligence Research — The Coming Manhattan Project — Moonshots]] — episode transcript, reconciled against the preserved ElevenLabs timing layer.
- [[research/All Things Superintelligence Research - The Coming Manhattan Project - Moonshots - Reminder and Wiki Preparation|Reminder and Wiki Preparation]] — coherent-thought ledger and evidence boundaries.
## Evidence boundary
Token efficiency should be reported alongside cost, internal computation, latency, hardware utilization, and task quality rather than treated as a complete efficiency measure.
## Simple Reminders, Quotations, and Thoughts
> Token efficiency and cost efficiency are different metrics. A model can emit fewer visible tokens while performing more internal computation for each one, so fewer tokens do not necessarily mean a cheaper or more efficient system.
> **— Adapted from Alexander Wissner-Gross**, *Moonshots, October 7, 2026*
[[reminders/AI Infrastructure/Token Efficiency and Cost Efficiency Are Different Metrics by Alexander Wissner-Gross|Token Efficiency and Cost Efficiency Are Different Metrics by Alexander Wissner-Gross]]