Model Watch
Every figure re-pulled from source on the date above.
Max / high reasoning where applicable. Numbers from Artificial Analysis, official benchmark boards, and vendor cards. Every charted number was re-pulled on the date shown.
- Index version All intelligence scores are Artificial Analysis Intelligence Index v4.3.2, re-pulled for every model in one sitting on 2026-09-26. AA rescored its whole catalog when it moved from v4.1.1, so every number dropped (Opus 5 went from 63 to 51) — compare bars with each other, not with older copies of this page.
- Pairs The intelligence chart pairs each lab's newest model with the one it replaced — pale bar = replaced, bright bar = current — so you see both the level and the rate of gain per month.
- One board Each capability benchmark charts one official board only (tbench.ai / deepswe.datacurve.ai) or the vendor's own reported score — never a mix of harnesses, and both bars in a pair share one harness.
- Google Google is charted on its Pro line only. Its Flash models are a different tier and are never swapped in to stand for a Pro — but Flash now outscores Pro, so 3.8 Flash rides the price/speed charts and the story is in the flags.
- Local models The last section covers openly-downloadable models a home user can run on their own PC, charted by the video-memory needed to run them and by open-weight intelligence.
- 2026-09-26 Full re-pull of every lab on Artificial Analysis Intelligence Index v4.3.2 (the index now uses Terminal-Bench 4.0 and dropped two older tests), plus four September releases: Opus 5.5, Grok 4.7, Muse Spark 1.3 and Gemini 3.8 Flash.
- Opus 5.5 Anthropic shipped Claude Opus 5.5 on 2026-09-22 and it now tops the whole index at 58, ahead of Fable 5.1 and GPT-6 Astra (both 53) — and it is cheaper than Opus 5 at $4 in / $20 out per million.
- Grok 4.7 xAI (now listed as SpaceXAI) shipped Grok 4.7 on 2026-09-21 at 46, up from Grok 4.6's 44, at the same $2 / $6 pricing, and it nearly doubled 4.6 on the Terminal-Bench 4.0 board (37.6 vs 20.3).
- Muse Spark 1.3 Meta's Muse Spark 1.3 (2026-09-02) scores 48, a jump of 8 points over 1.2 in four weeks, at the same $1.25 / $4.25 price and still the fastest frontier model measured (214 tokens/sec).
- Gemini 3.8 Flash Google's Gemini 3.8 Flash (2026-09-02) replaces 3.7 Flash in the price, speed and context charts: same $0.75 / $3.75 price, 41 on the index, 328 tokens/sec, and it ties for first on the DeepSWE coding board at 74.
- Terminal-Bench 4.0 The frozen Terminal-Bench 2.1 column is gone; the capability chart now uses the official tbench.ai Terminal-Bench 4.0 board, where GPT-6 Astra leads at 58.2.
- GPT-6 Astra speed Artificial Analysis has now measured GPT-6 Astra: 63 tokens/sec and 341.84 s to first token at max reasoning, so it joins the speed and latency charts.
- Google's flagship Gemini 3.5 Pro is still unreleased, so Google's newest Pro here remains February's 3.1 Pro (30) — and Google's own 3.8 Flash (41) now outscores it by eleven points.
- Too new for boards Opus 5.5, Grok 4.7 and Muse Spark 1.3 are not on the DeepSWE board yet (it last updated 2026-09-22), and Opus 5.5, Kimi K3, Gemini 3.1 Pro and Muse Spark are not on the Terminal-Bench 4.0 board — those bars are absent rather than guessed.
- GPQA gaps Anthropic published no GPQA Diamond score for Opus 5.5 and xAI none for Grok 4.6 or 4.7; a Muse Spark 1.3 figure circulates only in secondary blogs, not Meta's own materials, so none of those is charted.
- Opus 5.5 speed Artificial Analysis has not published speed or time-to-first-token for Opus 5.5 at max reasoning yet, so it is absent from those two charts — not slow, just unmeasured.
- OpenAI's lineup OpenAI also shipped GPT-6 Sol (2026-09-22, index 48), the direct successor to GPT-5.6 Sol; the OpenAI pair keeps charting its top model, GPT-6 Astra, so Sol is named here rather than charted.
- Gemini 3.0 Pro Artificial Analysis marks Gemini 3.0 Pro's v4.3.2 score (28) with an asterisk on its leaderboard; it is the replaced bar in Google's pair and is shown as published.
- Cost-per-task pulled The cost-per-task chart stays removed until Artificial Analysis's published cost-to-run-the-index figures are charted on one consistent basis.
- Auto-updater The weekly updater for this page is not firing (disabled on the bot, missed launches on its new host), and by design it only flags an index change rather than applying one. This refresh was done by hand.
- Local VRAM Video-memory figures in the local section are 4-bit estimates derived from verified parameter counts, not benchmarked measurements.
Grok Bot — xAI’s agent, not just a chatbot
A short note on the product sitting next to Grok 4.6 on this page. This is xAI’s Grok Bot, not a guest nickname on some other chat.
Grok Bot is xAI’s early-beta agent product. Each bot gets its own cloud computer — browser, files, terminal — so it can sign into tools and keep working after you close the laptop. That is a different product from the Grok chatbot on grok.com or X, and it is also different from the Grok 4.6 model scores in the charts below. Official page: x.ai/bot.
If you want a creator who actually set several bots up and said what stuck, watch Claire Vo on How I AI. That episode is already in the site’s AI video library. She covers Grok Bot setup, the virtual machine, and where she still will not switch.
Watch: How I AI — Grok Bot + Grok 4.6 + Cursor Origin How I AI channel Also in the library: Nate Herk’s week of Grok Bot lessons
Overall intelligence — each lab's newest model vs the one it replaced
Artificial Analysis Intelligence Index · v4.3.2 · every model re-pulled 2026-09-26
- Bars Pale = the model it replaced, bright = current. The block over each pair shows the gain and how fast it was earned (points per month).
- Anthropic Opus 5 to Opus 5.5 (2026-09-22) is +7 points to 58, the top score on the whole index. Anthropic's pricier Fable 5.1 (53) is a different line and is named here rather than swapped in.
- OpenAI GPT-6 Astra (53) over GPT-5.6 Sol (47). OpenAI's newer GPT-6 Sol (2026-09-22) scores 48 and is the mid-tier successor, so it is noted rather than charted.
- Google Google's bar is the oldest here because Gemini 3.5 Pro is still unreleased — February's 3.1 Pro is its newest Pro, and its own 3.8 Flash now outscores it.
- Meta Muse Spark 1.3 (2026-09-02) scores 48, up 8 points on 1.2 in four weeks — the fastest gain on the page.
Capability benchmarks (higher is better)
Score % · pale = previous model, bright = current · same test, same official board, same harness — or no bar at all
- Terminal-Bench 4.0 Official tbench.ai Terminal-Bench 4.0 board, read 2026-09-26: GPT-6 Astra 58.2 leads, Opus 5 53.9, Grok 4.7 37.6, GPT-5.6 Sol 37.3, Grok 4.6 20.3. Opus 5.5, Kimi, Gemini 3.x Pro and Muse Spark are not on this board, so they have no bar.
- DeepSWE Official deepswe.datacurve.ai board v1.1, updated 2026-09-22: GPT-6 Astra and Opus 5 tie at 74, GPT-5.6 Sol 73, Kimi K3 68.5, Grok 4.6 67, Muse Spark 1.2 55. Opus 5.5, Grok 4.7 and Muse Spark 1.3 are not on it yet.
- GPQA Diamond Vendor-reported only. GPT-6 Astra posts 96.0, the highest here. Anthropic published none for Opus 5.5, xAI none for Grok, and Meta none in its own materials for Muse Spark 1.2 or 1.3, so those bars are absent.
Token rates
USD per 1M tokens · Gemini 3.8 Flash is the cheapest here (cache hits $0.075); GPT-6 Astra is the priciest at $10 / $50
Output speed (higher is better)
Tokens per second · GPT-6 Astra now measured · Opus 5.5 (max) not yet measured by AA
Time to first token (lower is better)
Seconds · includes thinking time for reasoning models · Opus 5.5 (max) not yet measured by AA
Context window (higher is better)
Thousands of tokens
Local models for home users — will it run on your GPU?
Estimated video memory to run at 4-bit, in GB · green fits a 12-16GB card, blue needs ~24GB
- What this is Models you download and run on your own machine — private, no API bill, works offline. The real limit is video memory (VRAM), so this chart is what each one needs at 4-bit.
- 12-16GB card Best pick is Google's Gemma 4 12B (~8GB, Apache-2.0, ships ready-to-run GGUF); on a 16GB card gpt-oss-20b is the stronger reasoner. These are the green bars.
- 24GB card Qwen3.8-27B is the pick — the best all-round dense model at this size (~16GB, Apache-2.0). For more speed, the MoE options (Nemotron 3.5 Lightning, GLM-4.7-Flash) or Meta's agent-focused Muse Glimmer 30B.
- Big Mac / lots of RAM With 96-128GB of unified memory, gpt-oss-120b is the standout (~63GB, Apache-2.0); a 256GB Mac Studio can run DeepSeek V4 Flash (MIT) for strong open intelligence at home.
- Honesty These GB figures are 4-bit estimates from verified parameter counts, not measured. An AMD RX 6700 XT is a dead-end for local LLMs — assume NVIDIA, an Apple chip, or CPU+RAM.
The open-weight ceiling — AA Intelligence Index
Same v4.3.2 index as the charts above · the strongest models you can legally download
- Data-center only Xiaomi's MiMo-V2.6-Pro (46) is now the highest-scoring open-weight model, just ahead of GLM-5.3 (45) and Kimi K3 (44) — all free to download, but each needs data-center hardware, not a home PC.
- What you can run at home Qwen3.8-27B is now on the index at 34 — the best score you can run on a single 24GB consumer card, and close to DeepSeek V4 Pro (36), which needs a data center. gpt-oss-120b (12) is the big-Mac option.
Who leads each charted metric
- Fastest real iteration (same model line) Meta leads the window: Muse Spark 1.2 to 1.3 was +8 points (40 to 48) in four weeks. Anthropic's Opus 5 to 5.5 added +7 in two months, OpenAI's GPT-5.6 Sol to GPT-6 Astra +6, xAI's Grok 4.6 to 4.7 +2 in about six weeks, Google's Gemini 3.0 to 3.1 Pro +2, and Moonshot's Kimi K2.6 to K3 +17 over three months. The chart derives every rate from the release dates, so the points-per-month figures cannot drift from the bars.
- Opus 5.5 takes the top spot — and costs less than Opus 5 Claude Opus 5.5 (2026-09-22) scores 58 on the new v4.3.2 index, five points clear of Fable 5.1 and GPT-6 Astra (both 53). List price fell to $4 in / $20 out per million from Opus 5's $5 / $25. It is too new for the DeepSWE and Terminal-Bench boards, and Anthropic did not publish a GPQA score, so those bars are absent rather than borrowed.
- Muse Spark 1.3 is the fastest and one of the cheapest at the frontier Meta's Muse Spark 1.3 (48) sits level with GPT-6 Sol and ahead of Grok 4.7, while running at 214 tokens/sec for $1.25 in / $4.25 out — roughly a fifth of GPT-6 Astra's input price.
- Google's Flash now beats Google's Pro — by eleven points Gemini 3.8 Flash scores 41 on the index against 30 for Google's newest Pro, the February 3.1 Pro, and it ties GPT-6 Astra and Opus 5 for first on the DeepSWE coding board (74). Gemini 3.5 Pro is still unreleased. Flash is also the cheapest ($0.75 / $3.75) and fastest (328 tokens/sec) model here. The intelligence chart still holds Google to its Pro line so the cross-lab comparison stays Pro-vs-Pro.
- Current-gen leaders (AA Index v4.3.2) Opus 5.5 58 · Fable 5.1 53 · GPT-6 Astra 53 · Muse Spark 1.3 48 · GPT-6 Sol 48 · Grok 4.7 46 · MiMo-V2.6-Pro 46 · GLM-5.3 45 · Kimi K3 44 · Gemini 3.1 Pro 30.
- The open-weight ceiling moved Kimi K3 is no longer the top downloadable model: Xiaomi's MiMo-V2.6-Pro (46) and Z AI's GLM-5.3 (45) now edge it (44). All three need data-center hardware; the best you can run on one consumer card is Qwen3.8-27B at 34.
- A benchmark score is not a verdict on whether a model can do YOUR job Worth saying plainly, because these charts invite the opposite read. Gemini 3.1 Pro scores 11.8 on DeepSWE — dead last here by a wide margin — yet it does real day-to-day coding perfectly well inside an agentic IDE. A benchmark measures one harness, one prompt style, one task shape; change any of them and the ranking moves. Read these bars as evidence about a specific test, never as a ceiling on what a model can do for you.