VJOURNAL

AI • Global Desk • October 01, 2026

Android Bench 2.0 exposes the gap between impressive code and a working app

Android Bench results put full task success far below partial completion. Google’s new provider-agent tests show why a coding score needs a harness, a date and a clear cost unit.

AI-assisted conceptual editorial illustration of unbranded phones on a wooden test bench; not Google’s actual laboratory.

Answer in brief

Android Bench results put full task success far below partial completion. Google’s new provider-agent tests show why a coding score needs a harness, a date and a clear cost unit.

Evidence cutoff: 4 sources
Google announced Android Bench 2.0 on 17 September, adding long-horizon tasks and provider agents.
The table reports Google’s measurements; pass rate and partial completion answer different questions.
Benchmark-wide cost, overlapping uncertainty and harness choice limit direct procurement conclusions.

A tougher September test for app developers

Google’s 17 September announcement turns Android Bench results toward work that spans an app rather than a small patch. For development teams, the practical news is a changed question: can an agent carry a migration or new feature through to a functioning application? The release adds long-horizon tasks and evaluations using model providers’ own agents.

The distinction matters when a team delegates work over several sessions. A convincing initial diff may still leave a dependency unresolved or a screen unusable. Procurement therefore needs evidence about the whole workflow. A result from this suite should inform a shortlist for Android engineering, while a team’s acceptance requirements determine whether the resulting app is ready.

What Google measured, with the units intact

The official methodology describes thirty tasks with five independent attempts per task. The table preserves model-agent pairings, uncertainty ranges and partial completion. Google’s leaderboard labels dollars and hours as averages for a full benchmark run. These are published measurements observed on 1 October, not experiments conducted by VJOURNAL.

A buyer should retain those units in every copied spreadsheet. Dividing aggregate spending by an assumed task count can produce a planning estimate, but it cannot reveal the cost of a particular migration. Exact execution dates and reasoning effort are not identified in the summary, so those details remain unknown here rather than being inferred from product names.

Google Android Bench 2.0, long-horizon provider-agent results, checked 2026-10-01; release 2026-09-17. Thirty tasks, five runs per task. Cost and hours refer to a full benchmark run, not one task. Exact run dates and reasoning settings: Unknown in the summary.
Model / agentPass rate %Confidence range %Completion %Hours / full runUSD / full run
GPT-6 Astra / Codex28.013.3–42.082.27.9375.7
Claude Fable 5.1 / Claude Code22.710.7–36.082.422.2492.6
GPT-5.6 Sol / Codex19.37.3–32.074.38.6235.8
Claude Opus 5 / Claude Code16.75.3–29.377.827.0861.4
Gemini 3.8 Flash / Antigravity SDK8.03.3–13.347.412.134.5

An agent harness is part of the result

Codex, Claude Code and Antigravity SDK are execution environments around models. Their handling of files, tools and context can change what a model manages to finish. The comparison therefore evaluates deployed combinations. It does not isolate an underlying model by holding every surrounding component constant.

That is useful for teams buying an existing agent, but less decisive for teams building their own. Changing tools, permission boundaries or context management creates a different system. SWE-bench’s official documentation provides a separate reference for repository issue resolution; importing its scores into this Android table would erase meaningful differences in task design and verification.

Partial progress can hide a release blocker

Completion and pass rate should travel together. The former helps identify useful engineering work even when the final artifact fails; the latter records whether the requirements were satisfied. An almost complete app can still demand specialist repair before release. A team should ask which requirement failed and whether repair is predictable.

Google includes runtime and visual verification, which makes the benchmark relevant to interfaces as well as code structure. The uncertainty ranges also overlap across several listed combinations. That overlap discourages treating the displayed order as a precise estimate of how every agent will perform on a new company’s codebase.

The decision belongs to an acceptance workflow

A sensible adoption trial would use representative migrations, a reproducible starting repository and explicit build, runtime and accessibility checks. Record total spending, elapsed time and engineer repair time together. This is a suggested evaluation plan, not a report of tests we performed or a guarantee of a particular outcome.

The September release changes how ambitious coding delegation can be assessed. It gives managers a better reason to inspect completion, uncertainty and failure categories before choosing an agent. The next useful question is which combination reduces reviewed engineering effort on the team’s own tasks, with a working application as the acceptance condition.

Questions and answers

Does 82.2% completion mean the app passed?

No. Completion credits partially satisfied requirements. Google reports 28.0% full task success for Astra with Codex in this long-horizon suite; those measures have different meanings.

Are the costs charged for a single task?

No. The leaderboard defines these dollar values as average cost per full benchmark run. Treating $375.7 as a price for one task would misread the published unit.

Can these scores be compared directly with SWE-bench?

The suites use different tasks, environments and grading. Keep each result attached to its benchmark version and agent harness; the same percentage does not establish the same capability.