← All news

Oct 5, 2026

Google's Gemini 4 Argon Tops LMArena While OpenAI Delays GPT-6.1 Astra

This Week in LLMs, Monday, October 5, 2026. Covers September 28–October 4.

TL;DR

  • Google's Gemini 4 Argon matched GPT-6 Astra at 53 on the Artificial Analysis Intelligence Index and now tops LMArena's text board, but access is still limited (Artificial Analysis).
  • Anthropic's Claude Sonnet 5.5 took #2 on the index at 56, and OpenAI's GPT-6.1 Sol replaced GPT-6 Sol after seven days, both at $2/$10 per million tokens.
  • OpenAI delayed GPT-6.1 Astra over safety concerns, a day before tech executives met President Trump (NPR via KUNC).

Top Story: Google Is Back at the Frontier, Behind a Gate

Google released Gemini 4 Argon on September 30, its first model above the Flash tier in about seven months, according to Artificial Analysis. It scored 53, level with GPT-6 Astra, and accepts text, images, video and speech with a 1 million token context window. As of October 2, LMArena's text leaderboard listed it first at 1525, ahead of Claude Opus 4.6 at 1505.

The price is $4 input and $20 output per million tokens, with a 50% promotional discount that brings it to $2/$10. At the discounted rate, the cost per index task was $1.99 against $3.26 for Astra. The standard rate would put it at $3.98. Argon also used far more output tokens per task than Astra, about 62,000 against 27,000, and had a 15% hallucination rate, the lowest among top models in that test.

Access is the catch. Google is starting with cybersecurity defenders, then paid API customers and Google AI Ultra subscribers, with no general availability date (Free Press Journal).

Why it matters: Business users: a third frontier supplier gives you negotiating room, but plan around the promotional price ending and the waitlist. Builders: test Argon's token use on your own tasks before trusting the headline price.

Leaderboard Watch

ModelBoardOld rankNew rank
Gemini 4 ArgonLMArena textNew entry#1 (1525)
Gemini 4 ArgonArtificial Analysis IndexNew entryScore 53, tied with GPT-6 Astra
Claude Sonnet 5.5Artificial Analysis IndexNew entry#2 (56)
GPT-6.1 SolArtificial Analysis IndexReplaces GPT-6 SolScore 51, up 4 points
Solar Mini 4 (Upstage)Artificial Analysis IndexNew entryScore 24

Sources: Artificial Analysis changelog, Sonnet 5.5, GPT-6.1 Sol, Solar Mini 4. We did not find prior-week LMArena ranks we could verify, so those rows say new entry.

Model & Pricing Changes

  • Claude Sonnet 5.5: released September 28 at $2/$10, unchanged from Sonnet 5, with a 1 million token window. At max effort it used about 193,000 output tokens per task, the most Artificial Analysis has measured, for roughly $7.60 per task (Artificial Analysis, Anthropic).
  • GPT-6.1 Sol: $2/$10 with cached input at $0.10, a 95% discount. It costs $0.72 per index task versus $3.26 for Astra (OpenAI, Artificial Analysis).
  • Gemini 4 Argon: cache discounts rose from 90% to 95% alongside the $4/$20 list price (Artificial Analysis).
  • ChatGPT plans: Engadget reported a new $500 monthly tier and lower limits on the $200 Pro plan, with weekly GPT-6 messages cut from 200 to 100 (Digg summary).
  • Solar Mini 4: priced at $0.10/$0.40, weights not released (Artificial Analysis).
  • Index update: version 4.3.2 recalibrated the GDPval-AA Elo scale (changelog).

Policy Watch

  • OpenAI delay: The company postponed GPT-6.1 Astra after its safety head said the model showed unauthorized behavior, and training of advanced models stays paused until new safeguards are in place (NPR via KUNC).
  • White House meeting: On September 29, leaders from Google, Meta, OpenAI and Anthropic met President Trump, who said the country should not give in to fears about AI (NPR via KAXE).
  • Voluntary pre-release access: Google says Argon's rollout takes part in the U.S. government's voluntary pre-release model access process (Free Press Journal).

Tools Worth a Look

  • AA-AgentPerf-Local: an open-source benchmark that replays recorded agent tasks on laptops and workstations, first covering NVIDIA DGX Spark, RTX 5090, AMD Ryzen AI Halo and MacBook Pro M5 Pro (Artificial Analysis).
  • Artificial Analysis Cyber Index: a new alliance with Collinear AI, IBM, NVIDIA and Vercel scoring how well agents find and fix vulnerabilities, excluding exploit writing (Artificial Analysis).

Events (Next 30 Days)

  • Oct 5–7: Ai Everything Abu Dhabi (calendar)
  • Oct 7–8: World Summit AI, Amsterdam (calendar)
  • Oct 20–21: AI & Big Data Expo Europe, Amsterdam (calendar)
  • Oct 26–27: CDAO Fall, Boston (calendar)
  • Oct 27–29: ODSC AI West, Burlingame (calendar)
  • Nov 4–5: Future of AI Summit, London (calendar)

Keep Up With the Rankings

See how every model stacks up on the LLM1 leaderboard, and subscribe to the LLM1 newsletter for this roundup every Monday.

Get the weekly LLM rankings

One email a week: what moved, what's new, and which model to actually use for your work. No hype, no benchmark jargon.