Artificial Analysis announced the release of v4.2, an updated version of its Intelligence Index. With this update, the index now covers complex and high-difficulty tasks that are closer to real-world use cases.
A new evaluation metric, "AA-Briefcase," has been added to measure agentic knowledge work. This evaluates a model's comprehensive agentic capabilities through complex projects constructed by industry experts. Additionally, "GDP.pdf," which involves synthesizing information from a massive volume of 4,592 PDF pages, has also been introduced.
Changes have also been implemented to increase the reliability of the evaluations. Private test sets now account for 40% of the index's weighting, which is double the amount used in v4.1. According to the company, this is intended to prevent developers from gaming the evaluations.
In the latest rankings, Anthropic's Claude Fable 5.1 took the top spot, followed by OpenAI's GPT-6 Astra. Meta ranked third, with SpaceXAI and Moonshot/Kimi following.
Regarding output token efficiency, it was reported that GPT-6 Astra demonstrates higher token efficiency compared to other models at the forefront of intelligence.
Source: Artificial Analysis Intelligence Index v4.2 (Hacker News Frontpage, 2026-09-05)