Model Watch

Frontier model benchmarks, pricing, and speed — charted. Short, current, no hype.
Updated 2026-07-27 · window May–Jul 2026
Auto-refreshed weekly and when a new model ships.

Max / high reasoning where applicable. Numbers from Artificial Analysis, lab cards, and published lab tables.

54 Grok 4.5 · AA Index ▲ 16.4 vs prev gen
57 Kimi K3 · AA Index ▲ 13 vs prev gen
61 Opus 5 · AA Index ▲ 5 vs prev gen
59 GPT-5.6 Sol · AA Index ▲ 4 vs prev gen
46.5 Gemini 3.1 Pro · AA Index ▲ 7 vs prev gen
51 Muse Spark 1.1 · AA Index ▲ 8 vs prev gen
How to read these charts
  • Pairs The intelligence chart pairs each lab's newest model with the one it replaced — pale bar = replaced, bright bar = current.
  • Index All scores are AA Intelligence Index v4.1. The widely-quoted 64.9 for Fable 5 comes from a different index version and is not comparable to these numbers.
  • One board Terminal-Bench 2.1 and DeepSWE chart only their official boards (tbench.ai / deepswe.datacurve.ai) — one board per category, never vendor-harness numbers mixed in, and both bars in a pair share the same agent harness.
  • Anthropic Price, speed, latency and context bars chart Opus 5 — the mid tier most people actually use. Fable 5 is Anthropic’s top creative tier above it ($10 in / $50 out) and now sits just below Opus 5 on the intelligence index at twice the price.
  • Google Google is charted on its Pro line only. Its Flash models are a different tier, so they are never swapped in to stand for a Pro — even when they outscore one.
Verified this update
  • 2026-07-27 All six labs re-pulled from Artificial Analysis in one sitting. Every intelligence-index score came back unchanged, and every token price came back unchanged — the roster held completely steady on both.
  • Speed & latency Speed and latency moved on all six models and are all newly measured. The big ones: Kimi K3 fell to 33.0 tok/s with a 158.9s wait for its first token, and Opus 5 now reads 52.8 tok/s at 66.4s.
  • Opus 5 latency Opus 5’s time-to-first-token is a real Opus 5 reading for the first time (66.36s) instead of a figure carried over from Opus 4.8. It is far slower to start than the 21.51s the page showed last week.
  • Grok GPQA Grok 4.5’s GPQA Diamond bar has been removed. Two sources confirm xAI never published one: the 93.1 previously shown is Artificial Analysis’s own harness, and this axis charts vendor-reported scores only.
  • Opus 5 boards Re-checked and still absent: Opus 5 has no Terminal-Bench 2.1 entry and no DeepSWE entry on either board mirror. GPT-5.6 Sol and Kimi K3 are likewise still missing from Terminal-Bench.
  • Kimi K3 Moonshot published Kimi K3’s weights for free public download on 2026-07-26 — the strongest openly downloadable model on this page by a wide margin. No score changed.
Open flags & carried-forward numbers
  • Google’s flagship Gemini 3.5 Pro has now missed three announced dates and is still unreleased, so Google’s newest Pro model here is still February’s 3.1 Pro. Google shipped three Flash models instead, two of which outscore it — whether this page should switch Google to the Flash line is Phil’s call.
  • Gemini 3 Pro Gemini 3 Pro’s Terminal-Bench bar is pulled this week. Three reads of the official board returned three different numbers (65.8, 74.4 and 83.8 across different harnesses), none of them the 73.9 the page had been showing, so no bar is honest until one read wins.
  • Cost per task The whole cost-per-task row is still carried from earlier weeks and is the last unverified thing on this page. Opus 5’s $1.80 is very likely too low — it was inherited from Opus 4.8, and Opus 5 is markedly more verbose, which is what actually drives the bill.
  • Better cost basis Artificial Analysis now publishes the full cost of running its index per model, which is directly readable and would cover all six labs including Muse Spark. Switching the row to that basis changes what the chart measures, so it waits for Phil.
  • Roster watch Two outsiders are now scoring inside our six-lab range — Z.AI’s GLM-5.2 on the intelligence index and Sakana’s Fugu on GPQA. Neither is charted; adding a seventh lab is Phil’s decision, not an automatic one.

Overall intelligence — each lab's newest model vs the one it replaced

Artificial Analysis Intelligence Index · v4.1 · single snapshot

  • Bars Pale = the model it replaced, bright = current. Rate-of-gain shown only where both are the same model line.
  • Meta Meta joined as a sixth lab on 2026-07-21 — Muse Spark (43) to Muse Spark 1.1 (51).
  • Anthropic Opus 5 (24 Jul) replaces Opus 4.8 as Anthropic’s newest — a real same-line step (56 → 61), so the rate is shown. Fable 5 (60) is the top creative tier above Opus and sits out the 2-model window.
  • Google Google’s bar is the oldest on the chart for a reason: Gemini 3.5 Pro is still unreleased after three missed dates, so February’s 3.1 Pro remains its newest Pro model.

Capability benchmarks (higher is better)

Score % · pale = previous model, bright = current · same test, same official board, same agent harness — or no bar at all

  • Terminal-Bench 2.1 Official tbench.ai board only, re-checked 2026-07-27. Opus 5 is still not scored (its 89.1 is the AA harness, excluded); Kimi K3/K2.6, Grok 4.3, Muse Spark 1.0 and GPT-5.6 Sol (only Terra/Luna variants) remain absent — so Opus 4.8 and GPT-5.5 show without a bright partner.
  • Gemini 3 Pro Pulled this week: three reads of the board gave three different Gemini 3 Pro scores across different harnesses, so its bar is withheld rather than guessed. Gemini 3.1 Pro’s 65.6 (Terminus 2) confirmed twice and stands.
  • Terminal-Bench 2.1 Muse Spark 1.1 charts its board entry via the mini-SWE-agent harness (76.2); Spark 1.0 has no board entry, so the pale bar is absent, not zero.
  • DeepSWE Official deepswe.datacurve.ai board v1.1, cross-checked 2026-07-27 against the benchlm.ai and llm-stats.com mirrors. Opus 5 is still held out — absent from both mirrors; a third-party blog claims 74.0 but that is not the board. Grok 4.3, Kimi K2.6 and Muse Spark 1.0 remain absent from both.
  • GPQA Diamond Vendor-reported via aggregator. Grok 4.5’s bar was removed 2026-07-27: xAI published no GPQA score, and the 93.1 once shown here is Artificial Analysis’s harness, which does not belong on a vendor-reported axis.
Price · USD · lower is better

Token rates

USD per 1M tokens · Meta cache hits $0.15 (88% off)

Cost per AA task

Weighted USD per Intelligence Index task · Muse Spark and Gemini omitted — no per-task figure published as of 2026-07-27 · all four values carried, see flags

Speed & latency

Output speed (higher is better)

Tokens per second

Time to first token (lower is better)

Seconds · includes thinking time for reasoning models

Context

Context window (higher is better)

Thousands of tokens

Who leads each charted metric

  • Fastest real iteration (same model line) xAI 6.1 pts/mo (Grok 4.3 -> 4.5, +43.6% in 2.7 mo) · Moonshot 4.5 pts/mo (K2.6 -> K3, +29.5% in 2.9 mo) · Anthropic 2.7 pts/mo (Opus 4.8 -> Opus 5, +8.9% in 1.9 mo) · Meta 2.6 pts/mo (Muse Spark -> 1.1, +18.6% in 3.0 mo) · Google 2.3 pts/mo (Gemini 3.0 Pro -> 3.1 Pro, +17.7% in 3.1 mo) · OpenAI 1.6 pts/mo (GPT-5.5 -> 5.6 Sol, +7.3% in 2.5 mo). All six labs, every same-line pair on the chart — not a top-five.
  • Opus 5 lands — and it’s the value story Opus 5 (24 Jul) tops Anthropic’s own board at AA Index 61, a hair above Fable 5’s 60 — at $5 in / $25 out it is half Fable’s price ($10 / $50). It replaces Opus 4.8 in the 2-model window, and because it is a true same-line successor (not a new tier like Fable was), Anthropic finally posts a real iteration rate here. Fable 5 stays the top creative/writing tier above it; its capability-board scores are archived in git history.
  • Current-gen leaders Opus 5 61 · GPT-5.6 Sol 59 · Kimi K3 57 · Grok 4.5 54 · Muse Spark 1.1 51 · Gemini 3.1 Pro 46.5
  • Google is the one lab that did not ship a flagship this quarter Google's bar is the oldest on this chart, and that is the story rather than an oversight. Gemini 3.5 Pro was previewed on stage in May and has now missed three announced dates, with Bloomberg reporting the base model was scrapped over coding weaknesses — Alphabet shed roughly $200B in market value on the July news. What Google did ship is Flash: 3.6 Flash, 3.5 Flash-Lite, and a restricted-access 3.5 Flash Cyber built for finding and patching vulnerabilities. Two of those Flash models now score above Gemini 3.1 Pro on the intelligence index. This page still charts Google's Pro line, because a Flash standing in for a Pro would quietly change what the comparison means — so Google shows February's model until its flagship actually lands.
  • Who is Meta and what is Muse Spark 1.1 good for New to this page. Meta's first frontier model behind a PAID developer API (9 Jul 2026), built under Alexandr Wang's Superintelligence Labs — the end of Meta's open-weights-only era. It is an AGENT model: computer use, zero-shot MCP tool generalization, parallel subagent orchestration, and a managed 1M-token context. It is the cheapest thing on this board by a wide margin ($1.25 in / $4.25 out, cache hits $0.15) and among the fastest (127.6 tok/s, 1.41s to first token). Strong at tool orchestration (MCP Atlas 88.1) and computer use (OSWorld-Verified 80.8); middling at sustained coding-agent work (DeepSWE 53.3 — near Grok 4.5's 53.8, below Opus 4.8's 59.0 and well below GPT-5.6 Sol's 72.7).
  • A benchmark score is not a verdict on whether a model can do YOUR job Worth saying plainly, because these charts invite the opposite conclusion. Gemini 3.1 Pro scores 11.8 on DeepSWE — dead last here by a wide margin — yet it does real day-to-day coding work perfectly well inside an agentic IDE. A benchmark measures one harness, one prompt style, one task shape. Change the harness or the prompt and the ranking moves. Read these bars as evidence about a specific test, never as a ceiling on what a model can do for you.
  • Climbing hardest from the lowest base Grok 4.3 (37.6) -> 4.5 (54) is the biggest single jump on the board; Kimi K2.6 (44) -> K3 (57) is second. Both started well behind the US frontier and closed most of the gap in under three months.
Sources: Artificial Analysis model pages — all six re-pulled together 2026-07-27 (Opus 5 max, Grok 4.5 high, Kimi K3 high, GPT-5.6 Sol max, Muse Spark 1.1 xhigh, Gemini 3.1 Pro Preview) for index, speed, latency, price and context · tbench.ai Terminal-Bench 2.1 official board · deepswe.datacurve.ai board v1.1 cross-checked against the benchlm.ai and llm-stats.com mirrors · benchlm.ai GPQA Diamond aggregator · vendor pricing pages. Every charted number on this page was re-read on 2026-07-27 except the cost-per-task row, which is flagged below.
Data lives in src/data/model-watch.json — charts regenerate from it automatically.