TUA-Bench
A Benchmark for General-Purpose
Terminal-Use Agents
- 1 Meta AI
- 2 Duke University
- 3 Stanford University
∗Equal contribution
Abstract
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding. However, existing benchmarks do not adequately evaluate general-purpose terminal computer-use agents (TUAs): general computer-use benchmarks primarily target graphical user interfaces (GUIs), whereas terminal-based benchmarks largely emphasize technical and programming-centric workflows historically native to the shell.
We introduce TUA-Bench, a general-purpose benchmark
for terminal-use agents. TUA-Bench includes 120 real-world tasks across
five task families, covering routine digital activities—including
document editing, email management, and live-web information
seeking—as well as scientific and engineering workflows
co-designed with PhD-level domain experts that require specialized
software. This breadth distinguishes TUA-Bench from prior shell-focused
or domain-specific benchmarks. Each task is manually designed, runs in a
real terminal with a deterministic setup script, and is evaluated by an
execution-based scoring protocol. We find that the strongest frontier
agent, Claude Code with Claude Opus 4.8 max reasoning
effort, achieves 65.8% overall performance, with substantial gaps
across both tracks. By providing a broad and realistic evaluation of
terminal-use capabilities, TUA-Bench aims to accelerate the transition
from narrow, task-specific assistants to general-purpose agents capable
of operating reliably across diverse digital environments.
Leaderboard
Resolution rate across the TUA-Bench task suite. Results below are placeholders pending the official evaluation.
| # | Agent | Model | Reasoning Effort | Success Rate | Pass@1 | Pass@5 | All-5 | |
|---|---|---|---|---|---|---|---|---|
| Loading results… | ||||||||
Cost vs. Performance
Success rate against cost per run — the spend for one full pass over the 120-task suite. Hover a point for details, click it to open the per-task breakdown, and use the legend to isolate models or agents. The dotted line traces the cost–performance Pareto frontier.
Loading chart…
Citation
If you find this project useful, please use the following BibTeX entry.
@article{chen2026tua,
title={TUA-Bench: A Benchmark for Terminal-Use Agents},
author={Chen, Shoufa and Wang, Luyuan and Yang, Xuan and Liu, Zhiheng and Cong, Yuren and Ji, Yuanfeng and Zhou, Feiyan and Zhang, Xiaohui and Yang, Fanny and Zeng, Belinda},
journal={arXiv preprint arXiv:2606.28480},
year={2026}
}