On September 17, 2026 (US time), Google released "Android Bench 2.0," the latest version of its benchmark for evaluating the Android development capabilities of Large Language Models (LLMs) and AI agents.
Google Releases "Android Bench 2.0" to Measure AI Agent Coding Capabilities
This article is a translation. Read the Japanese original
This update significantly overhauls the specification to assign "Long-Horizon Tasks (LHT)," which are tasks that would take an engineer several days to a week to complete. While previous benchmarks focused on limited tasks such as small changes or bug fixes in existing repositories—reaching a top success rate of approximately 91%—the top success rate for this new set of long-horizon tasks has plummeted to approximately 28%.
The introduced long-horizon tasks include upgrading dependencies such as libraries, adding new features, building applications from scratch, and porting cross-platform applications to Android.
Regarding evaluation metrics, the benchmark has moved away from traditional binary pass/fail judgments in favor of a continuous scoring method that calculates a "Completion Rate." This is calculated by combining factors such as functionality, UI fidelity, and prevention of regressions (where a fix breaks existing functionality), aiming to provide a more detailed understanding of a model's design capabilities.
Validation results revealed that AI models tend to show higher performance in generating new code rather than refactoring existing code. On the other hand, performance drops significantly in tasks requiring runtime verification, tasks involving breaking changes to frameworks, and tasks where models face "knowledge gaps" due to the use of undisclosed libraries.
Notably, the test subjects include the latest models such as OpenAI's "GPT-6 Astra," which recorded the top success rate of 28% in long-horizon tasks.
Sources
- 「1週間の開発タスク」でAIの限界を検証 Googleが「Android Bench 2.0」公開 (ITmedia AI+、2026-09-26)