9to5google.com
Android Bench 2.0 focuses on long-horizon tasks, agent evaluations
Google unveiled Android Bench 2.0, a benchmark designed to assess large‑language models’ ability to perform software development tasks on Android. Unlike the first version, which measured incremental changes such as bug fixes, the new suite evaluates complex work like building new features, creating apps from scratch, and porting cross‑platform code to Android. Scoring shifts from binary pass/fail to a continuous metric that weighs functionality, visual fidelity, regression avoidance, and adherence to instructions, applying penalties for deviations. Tested [...]