Answer in brief
Android Bench results put full task success far below partial completion. Google’s new provider-agent tests show why a coding score needs a harness, a date and a clear cost unit.
A tougher September test for app developers
Google’s 17 September announcement turns Android Bench results toward work that spans an app rather than a small patch. For development teams, the practical news is a changed question: can an agent carry a migration or new feature through to a functioning application? The release adds long-horizon tasks and evaluations using model providers’ own agents.
The distinction matters when a team delegates work over several sessions. A convincing initial diff may still leave a dependency unresolved or a screen unusable. Procurement therefore needs evidence about the whole workflow. A result from this suite should inform a shortlist for Android engineering, while a team’s acceptance requirements determine whether the resulting app is ready.
What Google measured, with the units intact
The official methodology describes thirty tasks with five independent attempts per task. The table preserves model-agent pairings, uncertainty ranges and partial completion. Google’s leaderboard labels dollars and hours as averages for a full benchmark run. These are published measurements observed on 1 October, not experiments conducted by VJOURNAL.
A buyer should retain those units in every copied spreadsheet. Dividing aggregate spending by an assumed task count can produce a planning estimate, but it cannot reveal the cost of a particular migration. Exact execution dates and reasoning effort are not identified in the summary, so those details remain unknown here rather than being inferred from product names.
| Model / agent | Pass rate % | Confidence range % | Completion % | Hours / full run | USD / full run |
|---|---|---|---|---|---|
| GPT-6 Astra / Codex | 28.0 | 13.3–42.0 | 82.2 | 7.9 | 375.7 |
| Claude Fable 5.1 / Claude Code | 22.7 | 10.7–36.0 | 82.4 | 22.2 | 492.6 |
| GPT-5.6 Sol / Codex | 19.3 | 7.3–32.0 | 74.3 | 8.6 | 235.8 |
| Claude Opus 5 / Claude Code | 16.7 | 5.3–29.3 | 77.8 | 27.0 | 861.4 |
| Gemini 3.8 Flash / Antigravity SDK | 8.0 | 3.3–13.3 | 47.4 | 12.1 | 34.5 |
An agent harness is part of the result
Codex, Claude Code and Antigravity SDK are execution environments around models. Their handling of files, tools and context can change what a model manages to finish. The comparison therefore evaluates deployed combinations. It does not isolate an underlying model by holding every surrounding component constant.
That is useful for teams buying an existing agent, but less decisive for teams building their own. Changing tools, permission boundaries or context management creates a different system. SWE-bench’s official documentation provides a separate reference for repository issue resolution; importing its scores into this Android table would erase meaningful differences in task design and verification.
Partial progress can hide a release blocker
Completion and pass rate should travel together. The former helps identify useful engineering work even when the final artifact fails; the latter records whether the requirements were satisfied. An almost complete app can still demand specialist repair before release. A team should ask which requirement failed and whether repair is predictable.
Google includes runtime and visual verification, which makes the benchmark relevant to interfaces as well as code structure. The uncertainty ranges also overlap across several listed combinations. That overlap discourages treating the displayed order as a precise estimate of how every agent will perform on a new company’s codebase.
The decision belongs to an acceptance workflow
A sensible adoption trial would use representative migrations, a reproducible starting repository and explicit build, runtime and accessibility checks. Record total spending, elapsed time and engineer repair time together. This is a suggested evaluation plan, not a report of tests we performed or a guarantee of a particular outcome.
The September release changes how ambitious coding delegation can be assessed. It gives managers a better reason to inspect completion, uncertainty and failure categories before choosing an agent. The next useful question is which combination reduces reviewed engineering effort on the team’s own tasks, with a working application as the acceptance condition.
Questions and answers
Does 82.2% completion mean the app passed?
No. Completion credits partially satisfied requirements. Google reports 28.0% full task success for Astra with Codex in this long-horizon suite; those measures have different meanings.
Are the costs charged for a single task?
No. The leaderboard defines these dollar values as average cost per full benchmark run. Treating $375.7 as a price for one task would misread the published unit.
Can these scores be compared directly with SWE-bench?
The suites use different tasks, environments and grading. Keep each result attached to its benchmark version and agent harness; the same percentage does not establish the same capability.
