Real-time composite scores compared with Index v4.2 baseline.
| Rank | Model | Terminal v4 | AutomationBench | Reasoning | Index v4.3 Score | vs v4.2 Delta |
|---|
Benchmark Upgrade Changelog & Evidence
Terminal-Bench v4.0
Upgraded from v2.1. Addresses complaints on r/singularity that the earlier test suite failed to capture recent advances in command-line autonomous agents and multi-tool orchestration.
AutomationBench-AA
Replaces the previous 𝜏³-Banking evaluation. Features a private, contamination-resistant test set specifically geared towards agentic workflow automation in enterprise scenarios.
Bridging Release to v5
Community members observed that v5 has been in development for months; v4.3 was deployed to maintain trust and benchmark fidelity following the launch of high-performing frontier models.
Source: r/singularity: Artificial Analysis updates its Intelligence Index to version 4.3.