# Token Efficiency **Entity class:** Model-efficiency metric **Domain:** Artificial intelligence / evaluation **Maturity:** Developed ## Definition Token efficiency measures how much useful work a model performs relative to the number of input or output tokens it consumes. ## Mechanism and significance The metric can improve when a model expresses an answer compactly, but a system may also perform more hidden or internal computation per visible token. Token count therefore does not directly determine cost, latency, or energy use. ## Relationships - [[wiki/Compute Scarcity|Compute Scarcity]] - [[wiki/Memory Architecture|Memory Architecture]] - [[wiki/Artificial Intelligence|Artificial Intelligence]] - [[wiki/Moonshots - The Coming Manhattan Project|Moonshots — The Coming Manhattan Project]] ## Sources and provenance - [[research/All Things Superintelligence Research - The Coming Manhattan Project - Moonshots|All Things Superintelligence Research — The Coming Manhattan Project — Moonshots]] — episode transcript, reconciled against the preserved ElevenLabs timing layer. - [[research/All Things Superintelligence Research - The Coming Manhattan Project - Moonshots - Reminder and Wiki Preparation|Reminder and Wiki Preparation]] — coherent-thought ledger and evidence boundaries. ## Evidence boundary Token efficiency should be reported alongside cost, internal computation, latency, hardware utilization, and task quality rather than treated as a complete efficiency measure. ## Simple Reminders, Quotations, and Thoughts > Token efficiency and cost efficiency are different metrics. A model can emit fewer visible tokens while performing more internal computation for each one, so fewer tokens do not necessarily mean a cheaper or more efficient system. > **— Adapted from Alexander Wissner-Gross**, *Moonshots, October 7, 2026* [[reminders/AI Infrastructure/Token Efficiency and Cost Efficiency Are Different Metrics by Alexander Wissner-Gross|Token Efficiency and Cost Efficiency Are Different Metrics by Alexander Wissner-Gross]]