Answer in brief
Google’s new Argon joins a crowded frontier. A shared benchmark harness and published API tariffs reveal tradeoffs that launch rankings alone cannot settle.
A new entrant changes the shortlist
Google announced Gemini 4 Argon on 30 September, bringing a new contender into the frontier model comparison with Claude Opus 5.5 and GPT-6 Astra. For developers and teams buying complex reasoning, the first distinction is operational: Google is starting with trusted cyber defenders, while the other two have documented API offerings.
Anthropic’s Opus announcement is dated 22 September. Astra’s official API card remains the reference for the demanding-work model compared here. These are specific products with specific access conditions. A newer family name, a promotional price or an impressive launch graph should trigger a closer inspection of what a team can actually deploy.
One harness provides a narrower comparison
Artificial Analysis’s pages, checked on 1 October, allow a more disciplined terminal comparison than mixing three launch charts. Its Terminal-Bench 4.0 evaluation uses 66 tasks, mini-swe-agent and pass@1 averaged over three repeats. The table preserves published rounding and reasoning settings. It describes these configurations, including Claude’s default fallback behavior, rather than every possible version of the models.
A shared harness reduces one source of variation; it does not equalize compute budgets or make a few percentage points decisive. Terminal success also leaves unanswered questions about readable explanations, visual interpretation and maintaining an existing application. Exact execution dates are not published on these comparison pages.
| Model | Reasoning setting | Terminal-Bench 4.0 (%) |
|---|---|---|
| Gemini 4 Argon | High (high) | 57% |
| Claude Opus 5.5 | Very high (xhigh) | 60% |
| GPT-6 Astra | High (high) | 54% |
Prices and output limits answer different questions
Provider documentation checked on 1 October supplies the tariff table. Google’s introductory rates later rise to the amounts after the arrows; the announcement does not establish an expiry date. Claude and Astra figures are ordinary API token prices. Google’s million-token headline concerns maximum output, so it should not be substituted for a measured document-retrieval result.
A cheap output token can still generate an expensive session if an agent repeatedly revisits the same problem. Conversely, a higher tariff may be worthwhile when an answer passes review sooner. Search, tool execution, cache behavior and Astra’s long-input surcharge need separate accounting before these figures become a project budget.
| Model | Input USD / 1M tokens | Output USD / 1M tokens | Maximum output tokens | Access at cutoff |
|---|---|---|---|---|
| Gemini 4 Argon | $2 → $4 | $10 → $20 | 1,000,000 | Trusted partners; broader access planned |
| Claude Opus 5.5 | $4 | $20 | 128,000 | Available through API |
| GPT-6 Astra | $10 | $50 | 128,000 | Available through API |
Capabilities need a task and an interface
All three comparisons require looking beyond the chat box. A model that identifies an image defect may still need a tool to edit the file, a browser to check its effect and a reviewer to accept the result. The tool environment shapes both cost and the boundary of what the model can complete.
This matters particularly for code migrations and research assembled from documents. A large output allowance does not guarantee factual citations or a patch that preserves behavior. A useful evaluation should inspect the deliverable: working code, traceable evidence and an explanation that lets another person understand the remaining uncertainty.
The buying decision starts with accepted work
The editorial conclusion is a routing decision, not an overall league table. Pick representative tasks, keep files and tool permissions consistent, record the chosen effort setting and count accepted deliverables. Include retries and human review in the cost. An unresolved answer belongs in the failure column even when its prose looks convincing.
Until Argon access broadens, teams can prepare this evaluation without assuming immediate deployment. The meaningful October question is which available configuration finishes their work within a tolerable budget and review process. Revisit that answer when access, prices or model revisions change; none of the published snapshots replaces the actual acceptance criteria.
Questions and answers
Which model wins this comparison?
The table describes one terminal benchmark at specified settings. It cannot establish a universal winner across writing, vision, research and your own production tasks.
Can everyone use Gemini 4 Argon now?
Google announced access for trusted cyber defenders first. Wider developer, enterprise and consumer access is planned, without a verified general opening date here.
Did VJOURNAL run these benchmarks?
No. We report Artificial Analysis’s independently published measurements and the providers’ documented prices, with the source snapshot dated 1 October 2026.
