# Terminal-Bench
**Entity class:** AI benchmark
**Domain:** Artificial intelligence / Agents
**Maturity:** Developed
## Definition
Terminal-Bench evaluates AI agents on tasks performed through a computer terminal, emphasizing tool use, persistence, and execution across multiple steps.
## Mechanism and significance
It is most informative about the tested environment and harness; a large score change should not automatically be treated as general intelligence across unrelated domains.
## Relationships
- **Research dossier:** [[research/ASI and RSI Timeline Research Moonshots|ASI and RSI Timeline Research Moonshots]]
- **Ontology route:** [[ASI and RSI Timeline Ontology#Model Architecture and Compute Economics|Model Architecture and Compute Economics]]
- **Primary fields:** [[Machine Intelligence]] · [[AI Infrastructure]] · [[Compute Race]]
- **Adjacent concepts:** [[Externalized Model Knowledge]] · [[Knowledge Embedded in Model Weights]] · [[Long-Horizon Reasoning]] · [[Cost-Performance Frontier]]
## Sources and provenance
- [[research/ASI and RSI Timeline Research Moonshots|ASI and RSI Timeline Research Moonshots]] — immediate source for this node's role in the broadcast research map.
## Evidence boundary
The research dossier establishes why this entity or concept belongs in the Moonshots ontology. Time-sensitive organizational, product, policy, and performance claims should be checked against the linked primary source or a current authoritative source before reuse as settled fact.