androidcentral.com
Google just put the latest AI models through a brutal coding test — here’s how they did
Google unveiled Android Bench 2.0, a new benchmark that evaluates large language models on complex Android development tasks such as upgrading dependencies, adding major features, and building apps from scratch. The test introduces long‑horizon tasks that can take days to complete and shifts from a binary pass‑or‑fail to continuous scoring. Early results show GPT‑6 Astra leading with a 28% pass rate, while Gemini 3.8 Flash achieved only 8%. The benchmark aims to better gauge AI capabilities for real‑world software development. [...]