# AI Benchmarking
**Domain:** Artificial Intelligence / Evaluation / Research Governance
**Doc Type:** Canonical Concept Node
**Maturity:** Developed
## Definition
**AI benchmarking** evaluates systems on shared tasks, datasets, metrics and protocols so that performance can be compared under specified conditions.
## Historical Context
The dispute surrounding [[wiki/Speech Understanding Research|Speech Understanding Research]] and [[wiki/HARPY|HARPY]] helped motivate a DARPA evaluation regime in which methods and benchmark tasks were prescribed in advance and systems were tested annually.
## Governance Context
Benchmarks make comparison possible, but they also decide what counts as performance. A system can optimize the test while leaving important capabilities, populations or deployment conditions outside the measured frame.
## Psychometric and introspective evaluation
[[thoughts/The Luxembourg Preprint|The Luxembourg Preprint]] shows why the prompt frame must be treated as part of the benchmark. A structured inventory and an open-ended therapeutic dialogue can appear to measure the same psychological construct while producing sharply different model behavior.
[[wiki/Prompt-Sensitive Behavioral Stability|Prompt-Sensitive Behavioral Stability]] names the comparative target across those frames. [[wiki/Psychometric Deformation|Psychometric Deformation]] names the failure mode in which test-induced output is mistaken for a persistent personality or mental state. A valid benchmark should preserve prompts, order, model version, sampling settings, complete interaction history, repeated trials, and the boundary between observed response and inferred interiority.
## Key Insight
**A benchmark does not merely measure success; it operationally defines the success that institutions can recognize.**
## See Also
[[wiki/Held-Out Validation|Held-Out Validation]], [[wiki/Objective Function|Objective Function]], [[wiki/Model-Based Governance|Model-Based Governance]], [[wiki/Psychometrics|Psychometrics]], [[wiki/Prompt-Sensitive Behavioral Stability|Prompt-Sensitive Behavioral Stability]], [[wiki/Psychometric Deformation|Psychometric Deformation]]
## Simple Reminders, Quotations, and Thoughts
> "More disturbing to me is the stubborn reluctance in many segments of society to allow computers to take over tasks that simple models perform demonstrably better than humans."
> **— Richard H. Thaler**, *2015, Edge annual question “What Do You Think About Machines That Think?”*
[[reminders/AI Control/Computers Already Make Better Routine Decisions by Richard H. Thaler|Computers Already Make Better Routine Decisions by Richard H. Thaler]]
> "A common theme in recent writings about machine intelligence is that the best new learning machines will constitute rather alien forms of intelligence."
> **— Andy Clark**, *2015, Edge annual question “What Do You Think About Machines That Think?”*
[[reminders/Information/Learning Machines May Develop Alien Intelligence by Andy Clark|Learning Machines May Develop Alien Intelligence by Andy Clark]]
> "Machine intelligence, while impressive in certain areas, is still narrow and inflexible."
> **— Timo Hannay**, *2015, Edge annual question “What Do You Think About Machines That Think?”*
[[reminders/Machine Succession/Machine Intelligence Is Still Narrow and Inflexible by Timo Hannay|Machine Intelligence Is Still Narrow and Inflexible by Timo Hannay]]
> "A human player can make generalizations and describe why certain types of moves are good, and use that to teach a human player."
> **— Rodney A. Brooks**, *2015, Edge annual question “What Do You Think About Machines That Think?”*
[[reminders/Risk Debate/Mistaking Performance For Competence Misleads Estimates Of AIs 21st Century Promise And Danger by Rodney A. Brooks|Mistaking Performance For Competence Misleads Estimates Of AIs 21st Century Promise And Danger by Rodney A. Brooks]]
> "Thought experiments about these matters are the source of practical insights into human and machine behavior and suggest how to build different and better kinds of machines."
> **— Robert Provine**, *2015, Edge annual question “What Do You Think About Machines That Think?”*
[[reminders/Machine Succession/Thought Experiments Can Build Better Machines by Robert Provine|Thought Experiments Can Build Better Machines by Robert Provine]]
## Sources / Provenance
- National Research Council, _Funding a Revolution_, 1999.
- [[articles/Modern Artificial Intelligence in the 1970s|Modern Artificial Intelligence in the 1970s]].