TEAS Bench: innovative AI inference benchmarking
1 September 2026
TEAS is a new AI benchmark that maps the true cost, performance, and trade-offs of agentic AI across diverse hardware.
Researchers at the University of Edinburgh, EPCC, and Imperial College London have announced the release of TEAS, a new AI benchmark designed to measure realistic, agentic AI inference workloads, and the associated benchmarking suite, TEASBench.
Funded by the Advanced Research and Invention Agency (ARIA) under the “Scaling compute” programme, TEAS shifts the industry focus away from single-number leaderboards toward a comprehensive "hardware suitability map."
As AI models become increasingly integrated with external tools, traditional fixed-length benchmarking falls short. TEAS addresses this by running realistic, versioned workloads (including agentic and tool-using tasks) at their natural context lengths. Crucially, the benchmark spans both established data centre GPUs (NVIDIA, AMD) and emerging silicon (Tenstorrent, Cerebras) to map exactly what frontier AI inference costs and how fast it runs across the expanding hardware ecosystem.
“There is no single winner to crown. The cheapest chip and the fastest are rarely the same,” the research team notes. “TEAS shows the cost, performance, and accuracy trade-off, allowing organisations to choose the right hardware based on their specific priorities, model size, and budget.”
TEASBench provides analysis and metrics that can help AI users, people purchasing, owning, or operating hardware for AI inference, and hardware developers, to understand the performance characteristics of AI models and hardware.
What sets TEAS apart
- Real workloads, full system. Prompts run at variable, natural lengths rather than fixed stand-ins. Because agentic tool calls rely on the CPU, TEAS measures the whole system—splitting time between model compute and tool wait—to reveal the full operational time budget.
- CAP (cost, accuracy, performance) profiling. Every configuration is evaluated using a unified rent-and-buy cost model, tracking throughput, latency, energy efficiency, and cost simultaneously.
- Open and reproducible. Designed to be proven, not just reported. The results can be run end-to-end from open pipelines, ensuring the path from model to published number is transparent.
- Accuracy as a safety check, not a score. TEAS separates model capability from hardware execution. Accuracy is used solely to verify that a run finished cleanly, ensuring hardware is judged on how fast, cheaply, and efficiently it executes, rather than the intrinsic intelligence of the model.
- Hardware beyond the incumbents. TEAS validates and costs emerging accelerators next to established NVIDIA and AMD data centre GPUs without generation bias, tracking where the future of inference is heading.
TEAS is designed to complement, rather than replace, established benchmarks by supporting emerging accelerators, complex agentic workloads, and integrated cost modelling. The roadmap for TEAS includes expanding to measured energy metrics, increasing sample counts per configuration, and fostering independent reproduction built on its currently open pipelines.
Links
TEAS is built by collaborative research teams at the University of Edinburgh, EPCC, and Imperial College London, with funding provided by ARIA’s “Scaling compute” programme.