TUA-Bench

TUA-Bench

A Benchmark for General-Purpose
Terminal-Use Agents

  1. Shoufa Chen1,∗
  2. Luyuan Wang1,∗
  3. Xuan Yang2
  4. Zhiheng Liu1
  5. Yuren Cong1
  6. Yuanfeng Ji3
  7. Feiyan Zhou1
  8. Xiaohui Zhang1
  9. Fanny Yang1
  10. Belinda Zeng1

Equal contribution

Abstract

As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding. However, existing benchmarks do not adequately evaluate general-purpose terminal computer-use agents (TUAs): general computer-use benchmarks primarily target graphical user interfaces (GUIs), whereas terminal-based benchmarks largely emphasize technical and programming-centric workflows historically native to the shell.

We introduce TUA-Bench, a general-purpose benchmark for terminal-use agents. TUA-Bench includes 120 real-world tasks across five task families, covering routine digital activities—including document editing, email management, and live-web information seeking—as well as scientific and engineering workflows co-designed with PhD-level domain experts that require specialized software. This breadth distinguishes TUA-Bench from prior shell-focused or domain-specific benchmarks. Each task is manually designed, runs in a real terminal with a deterministic setup script, and is evaluated by an execution-based scoring protocol. We find that the strongest frontier agent, Claude Code with Claude Opus 4.8 max reasoning effort, achieves 65.8% overall performance, with substantial gaps across both tracks. By providing a broad and realistic evaluation of terminal-use capabilities, TUA-Bench aims to accelerate the transition from narrow, task-specific assistants to general-purpose agents capable of operating reliably across diverse digital environments.

Leaderboard

Resolution rate across the TUA-Bench task suite. Results below are placeholders pending the official evaluation.

# Agent Model Reasoning Effort Success Rate Pass@1 Pass@5 All-5
Loading results…

Cost vs. Performance

Success rate against cost per run — the spend for one full pass over the 120-task suite. Hover a point for details, click it to open the per-task breakdown, and use the legend to isolate models or agents. The dotted line traces the cost–performance Pareto frontier.

Loading chart…

Citation

If you find this project useful, please use the following BibTeX entry.

@article{chen2026tua,
  title={TUA-Bench: A Benchmark for Terminal-Use Agents},
  author={Chen, Shoufa and Wang, Luyuan and Yang, Xuan and Liu, Zhiheng and Cong, Yuren and Ji, Yuanfeng and Zhou, Feiyan and Zhang, Xiaohui and Yang, Fanny and Zeng, Belinda},
  journal={arXiv preprint arXiv:2606.28480},
  year={2026}
}