A leaderboard that makes models draw pelicans has unexpectedly become the best window for watching three years of reasoning-model evolution.

When OpenAI shipped o1 in September 2024, there was one detail that got quoted again and again: it would "think" for 10 to 20 seconds before answering, and users at the time complained it felt "like a top student who types very slowly." Two years later, when GPT-6 Astra was handed the prompt "draw a lighthouse at night," it could "think" for a full 971 seconds — sixteen minutes. What happened in between wasn't simply models getting stronger; it was that the whole industry changed how it defines "intelligence": from "who answers fast" to "who is willing to think longer, and whether thinking longer is worth it."

This article strings together a single slice of real data — the OpenRouter Sketch leaderboard (a drawing board that has text models produce an image in one shot, captured on 2026-10-04 via the "AI LLM Leaderboard" plugin on our NAS) — to lay out the full development path of ChatGPT, Claude, and DeepSeek's reasoning models. You'll find: every leap in reasoning capability is written in plain sight in two numbers — "cost" and "time."


Chapter Zero: First, Understand — What Exactly Do Reasoning Models Do "More"?

Before the timeline, you need three minutes to understand one concept: the difference between a reasoning model and an ordinary model is, at its core, "turning an answer into a computation process."

An ordinary LLM takes a question and generates the answer directly, spitting out one token after another. A reasoning model does two extra things:

  1. Chain of Thought (CoT): before giving the final answer, it first generates a large stretch of "thinking process" in its head — decomposing the problem, trying solutions, self-checking, correcting course. o1's "thinks for 10 seconds," in engineering terms, is generating thousands of extra thinking tokens.
  2. Reinforcement Learning with Verifiable Rewards (RLVR): during training, instead of scoring only on human preference, it rewards the model's "thinking process" using tasks with a clear right-or-wrong (math problems, code tests, competition problems), forcing it to learn to spend compute where it counts.

This combination spawned a new term: test-time compute. o1's technical report has a key chart: on reasoning tasks, the model's accuracy (pass@1) rises monotonically as "thinking time / thinking-token budget" increases — which is the equivalent of turning "reasoning" itself into a new scaling law. The pretraining scaling law says "more data makes the model smarter"; the test-time scaling law says "think longer, answer more accurately."

So the whole industry gained three new tunable parameters, and they are the through-line of every story that follows:

  • Thinking effort (effort): low / medium / high / max, or a multi-level numeric value, controlling "how deeply to think";
  • Thinking budget (budget_tokens): a token cap on the chain of thought;
  • Model routing (routing): the system automatically judges "should this question be answered in a flash or thought through deeply."

"Thinking" went from an invisible engineering detail to a product dimension that can be priced, tuned, and compared. That is why every number that follows revolves around "cost" and "time."

Chapter One: Autumn 2024 — o1 Presses the "Think" Key

The true birth year of reasoning models was 2024. On September 12, OpenAI released o1-preview and o1-mini, internal codename "Strawberry." Its fundamental difference from every earlier large model: it runs the answer through its head first, then speaks.

The official numbers at the time: on the IMO International Mathematical Olympiad qualifier it scored 83% (GPT-4o had only 13%); physics, chemistry, and biology PhD-level benchmarks beat human experts. The cost was equally visible — slower, more expensive, and in the first week after launch you could only send 30 messages, with the chat experience mocked as "like a top student who types very slowly."

o1's technical significance was that it validated two things: RLVR can scale (replacing human feedback with verifiable rewards, the model learned self-correction); test-time compute can scale (give the model more thinking tokens and accuracy goes up). On December 6, the multimodal-capable full o1 launched.

On December 20, o3 appeared on the closing night of OpenAI's "12-day event": 87.5% on ARC-AGI visual reasoning (human baseline 85%), and records across the board on math and GPQA. What makes the ARC-AGI board special is that it tests abstract reasoning — it doesn't give the model problems it has seen before, so "memorizing questions" is useless; only real reasoning works.

But the story turned here. On February 13, 2025, Altman announced: o3 would not ship as a standalone product. Reasoning capability would be merged directly into GPT-5, because standalone reasoning models like o1/o3 were to be folded into a single system that can automatically toggle thinking on and off. o3-mini became the last standalone member of the o series (launched 2025-01-31, the first time ChatGPT opened a reasoning model to free users).

Looking back, the significance of the o1→o3 year: reasoning became a product dimension that can be priced and quantified for the first time, and "how long to think" went from an engineering detail to a product parameter; and the route disagreement over "should reasoning be a standalone product" planted the seeds for the next two years of industry structure.

Chapter Two: January 2025 — DeepSeek Pries Open the Price Gap with Open Source

If o1 was "the iPhone of reasoning models," then DeepSeek-R1 was "Android" — open source, deployable, paper-level transparency.

First, the foundation. On December 26, 2024, DeepSeek-V3 shipped: a 671B-total / 37B-active MoE architecture, paired with multi-head latent attention (MLA), FP8 low-precision training, and multi-token prediction (MTP); the official paper disclosed a training cost of about $5.6 million — a number that made all of Silicon Valley recompute "how much does it actually cost to train a large model."

On January 20, 2025, R1 and R1-Zero launched, punching straight through the pricing structure for reasoning models:

  • R1-Zero: no human annotation at all, reasoning "grown" purely by reinforcement learning (GRPO — Group Relative Policy Optimization, which even skips the critic value model). The training logs contain the famous "Aha Moment" — the model stops on its own to reflect: "Wait, wait, this is an epiphany worth marking... let me re-evaluate this solution."
  • R1: o1-level on math and code; at the launch-time API price, output cost was roughly 1/27 to 1/30 of o1's. Weights released under MIT, with distilled small models for Qwen and Llama as a bonus.
  • The diffusion effect: it hit #1 on the US App Store free chart 7 days after launch, writing "OpenAI should cut prices" into the whole industry's calendar, and global AI chip stocks swung violently for a while.

After R1, DeepSeek's iteration almost followed OpenAI's roadmap step by step:

When Version Key move
2025-03 V3-0324 Reasoning boost, AIME 39.6→59.4
2025-05 R1-0528 AIME 70→87.5, GPQA 71.5→81
2025-08 V3.1 One model with both thinking and non-thinking modes — head-on collision with GPT-5's "unified system" idea
2025-09 V3.2-Exp In-house sparse attention (DSA), API price cut 50%+
2025-12 V3.2 official World-leading reasoning performance, gold medals at IMO, CMO, ICPC, and IOI; first to fold "thinking" into tool calling (agents)
2026-04 V4 series Pro (1.6T total / 49B active) and Flash (285B / 13B), 1M-token context as standard, fully open source under MIT
2026-08 V4-Pro official Major agent-capability boost (DeepSWE 62.7), native OpenAI Responses API support, three thinking-effort levels low/high/max; official evals claim a better experience than Sonnet 4.5 and delivery quality close to Opus 4.6 non-thinking mode
2026-09 V4.1-Flash Brand-new architecture, see below

V4.1-Flash (2026-09-10) is a milestone on DeepSeek's efficiency route:

  • Architecture: a brand-new asymmetric Causal Encoder-Decoder structure, 552B backbone parameters (plus a 196B sparse memory module), but prefill activates only 8B per token and decode activates 16B — "reading" and "writing" use differently sized compute, purpose-built for the agent work pattern of "massive input, little output";
  • KV Cache compression: the global cache squeezed to 890 bytes/token (about 1/4 of V4-Flash), a 437x shrinkage versus the first-generation model — a smaller cache means cliff-edge cost drops for long context and multi-turn agent tasks;
  • Native multimodality: vision understanding merged into the main model (no more bolted-on experimental Vision edition), pretrained on 45T tokens of multimodal corpus;
  • Scores: GPQA Diamond 90.9, Codeforces rating 3471, Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2, CyberGym 88.1, AutomationBench 54.8;
  • Pricing: the API was renamed deepseek-flash; off-peak cache-hit input as low as $0.003 per million tokens, cache miss $0.15, output $0.6 — another 60% below the same tier of V4-Flash.

One industry aside worth remembering: DeepSeek announced that after 2026-09-14, V4 Pro is retired and requests are automatically routed to V4.1-Flash, billed at Flash prices; the official reason is "continuing to offer the weaker V4 Pro at a higher price and slower speed is no longer appropriate." Community reaction split in two: some said "a free upgrade that also cuts the price," others said "my production workflow I had validated got force-switched on 48 hours' notice" (compare OpenAI's Sora API shutdown, which gave 6 months' notice). The lesson: when "cheap" becomes company strategy, even existing customers' migration costs get folded into the efficiency ledger.

The keyword on DeepSeek's line has always been efficiency: others buy intelligence with compute, it buys intelligence with architecture — MoE sparse activation, attention compression, asymmetric input/output, KV cache slimming. It drove the unit price of "top-tier reasoning" down to a fraction of closed-source models, and turned "open weights + paper transparency" into the pricing anchor for the whole industry.

Chapter Three: Claude's Third Road — Hybrid Reasoning → Adaptive Thinking

Anthropic was the only one of the three that didn't ship a reasoning model in 2024. CEO Dario Amodei publicly opposed a "fast model / reasoning model" product split, arguing that making users manually choose "think or not" is a failure of interaction design — that position directly determined Claude's route for the next three years.

On February 24, 2025, Claude 3.7 Sonnet arrived with Extended Thinking, positioned as "the market's first hybrid reasoning model": the same weights can both answer in a flash and think deeply; whether to think, and for how long, is controlled by the API parameter budget_tokens. It was a frontal rebuttal of the o1 route — don't split the product; hand the choice to the caller.

Everything that followed was polishing that "thinking switch":

  • 2025-05 Claude 4 (Opus/Sonnet): can call tools while thinking, plus a "thinking summary" mechanism (a small model compresses the chain of thought, cutting cost and latency);
  • 2025-08 Opus 4.1: coding and long-context improvements;
  • 2025-09/10 Sonnet 4.5 / Haiku 4.5: mid-tier and entry-tier fill-ins;
  • 2025-11 Opus 4.5: new coarse-grained effort control in three levels — low/medium/high — while cutting API prices by 66% ($15/$75 → $5/$25) — the "expensive" label on reasoning models got its first voluntary discount;
  • 2026-02 Opus 4.6: adaptive thinking — the model itself decides whether to think and how deeply; manual mode begins to be deprecated;
  • 2026-04 Opus 4.7: manual mode removed entirely (returns 400); only adaptive + the xhigh tier remain;
  • 2026-05 Opus 4.8: Dynamic Workflows (orchestration of parallel sub-agents).

In the summer of 2026, Anthropic entered its fifth generation, and the product line shifted from "version numbers" to "role division":

  • 2026-06-09 Fable 5 / Mythos 5: the first Fable-tier model, aimed at long-horizon agent work (briefly suspended on June 12 over US export controls, restored July 1);
  • 2026-06-30 Sonnet 5: $2/$10, the speed-versus-intelligence balance tier;
  • 2026-07-23 Opus 5: $5/$25, the workhorse for everyday professional work;
  • 2026-09-01 Fable 5.1: adaptive thinking is always on and cannot be disabled (manually setting disabled returns 400 directly), five effort levels (low/medium/high/xhigh/max, default high), 1M context, 128K output, $10/$50, cache reads dropped to $0.25 (75% cheaper than Fable 5, a big cut for long agent sessions); the same-weights restricted Mythos 5.1 is open only to vetted institutions (capability tiering in cybersecurity/bio, similar to OpenAI's restricted access);
  • 2026-09-22 Opus 5.5 ($4/$20) and Sonnet 5.5 ($2/$10) round out the line.

By October 2026, Claude's price ladder is: Haiku 4.5 ($1/$5) → Sonnet 5.5 ($2/$10) → Opus 5.5 ($4/$20) → Fable 5.1 ($10/$50). Each tier corresponds to a "thinking intensity": from "fastest instant answer" to "adaptive always-on, default high."

The essence of the Claude route: reasoning is not a switch, it is a runtime resource. From "developers turning the token-budget knob in code," to "the model allocates compute by task difficulty," to "a classifier downgrades high-risk requests to lower-capability models" — the threshold for reasoning keeps sinking, while safety tiering runs in parallel as a theme of its own.

Chapter Four: After GPT-5 — Reasoning Becomes "Automatic Transmission"

OpenAI's answer was GPT-5 on August 7, 2025: the o series was formally merged into the main GPT series, ChatGPT no longer lets you pick the model — it automatically decides "instant answer or deep thought" — Auto mode. GPT-5's capability layer covers coding, math, writing, and vision, and is regarded as the finished form of the "unified system."

After that, OpenAI entered a high-frequency, small-steps mode:

When Version Key points
2025-08 GPT-5 / mini / nano Unified system, dual Auto/Thinking modes
2025-09 GPT-5-Codex Dedicated edition for agentic coding
2025-11 GPT-5.1 Stronger coding and instruction following
2025-12 GPT-5.2 Better quality and tool calling
2026-01~02 GPT-5.2/5.3-Codex Iterating specifically on coding agents
2026-03 GPT-5.4 / 5.4-mini/nano Performance and cost tiering
2026-04 GPT-5.5 Low-cost tier approaches flagship performance
2026-07 GPT-5.6 (Sol/Terra/Luna) Capability tiers formally priced separately: Sol (flagship) > Terra > Luna (entry), routed by task complexity
2026-09-03 GPT-6 Astra Officially "the smartest, most aligned model in the world," see below
2026-09-22 GPT-6 Sol / Luna Handle everyday work at different capability×cost combos
2026-09-29 GPT-6.1 Sol Near-Astra capability, API price less than half of Astra ($2/$10), plus multi-agent beta support

GPT-6 Astra is the current endpoint of this history (released 2026-09-03, rolled out to full availability before 2026-09-22):

  • Specs: 1.05M-token context window, 128K max output, knowledge cutoff 2026-04-30, API pricing $10 per million input / $50 per million output (cache hit $1);
  • Benchmarks (OpenAI's own evals): FrontierMath Tier 4 at 98% (already used to help solve several long-open math problems), ARC-AGI-3 99.9%, ExploitBench 100%, Terminal-Bench 4.0 57.9% (vs. GPT-5.6 Sol 37.3% and Claude Fable 5.1 55.8%, and officially about 9% and 63% cheaper per task respectively), GPQA Diamond 96.0%;
  • Safety tiering: the first model to reach the "Critical" cyber-capability level under the OpenAI Preparedness Framework — able to autonomously discover and exploit unknown vulnerabilities — so a stricter safety stack, full chain-of-thought monitoring, and "misalignment monitoring" shipped alongside it;
  • Alignment: across 54,000 internal Codex task simulations, high-severity misalignment signals were about half of GPT-5.6 Sol's;
  • Interaction: supports async tool calls, mid-turn steering (changing instructions mid-task), adjusting effort mid-conversation; the lowest effort tier is no longer supported (Astra does not allow "no thinking at all").

Note two trends in OpenAI's stretch: model naming shifted from "version numbers" to "capability tiers" (Sol/Terra/Luna, GPT-6 Astra/Sol/Luna/6.1 Sol), and reasoning intensity shifted from "user picks" to "system routing + dynamic adjustment during the session". Reasoning capability itself is becoming infrastructure rather than a selling point.

Chapter Five: Using the Sketch Board to Take a Group Photo of Three Generations

The above is "official narrative"; below is "measured data." OpenRouter's Sketch board does something clever: it gives dozens to hundreds of text models the same prompts ("pelican on a bike," "lighthouse at night," "penguin with an umbrella," ...) and makes them produce images directly. Text models aren't as strong as dedicated image models, but this task's combined demand on "understanding, planning, and detail execution" is exactly a touchstone for reasoning ability — and it leaves viewable outputs for every model.

Data spec (captured 2026-10-04, via the Halo "AI LLM Leaderboard" plugin cache on the NAS)

  • Across the board's 4 prompt columns there are 640 "model × sample" records in total, deduplicated to 204 models; the default column "Pelican on a bike" has 143 models (the lighthouse column has 173, the most site-wide);
  • The upstream for that column gives no quality score; the main metrics are cost per run (USD, lower is better) and generation time (seconds); this article does not weight-synthesize any "overall quality score" itself;
  • The board-wide cost median is about $0.009~0.014, the time median about 52~79 seconds — in other words, the "market average price" of a one-shot image is very cheap; the money you pay extra all goes to "thinking."

"Cost × Time" comparison of the three companies' flagship generations (default column)

Vendor Generation (release date) Cost per run Generation time
OpenAI GPT-4 (2023-03) $0.043 5 s
OpenAI GPT-4o mini (2024-07) $0.00035 8 s
OpenAI GPT-5 (2025-08) $0.099 106 s
OpenAI GPT-5.5 (2026-04) $0.15 69 s
OpenAI GPT-6 Astra (2026-09) $0.87 391 s
Anthropic Claude 3 Haiku (2024-03) $0.002 13 s
Anthropic Claude 4 Sonnet (2025-05) $0.029 16 s
Anthropic Claude 4.5 Sonnet (2025-09) $0.034 25 s
Anthropic Claude Opus 5 (2026-07) $0.5 275 s
Anthropic Claude Fable 5.1 (2026-09) $2.13 519 s
DeepSeek V3-0324 (2025-03) $0.001 26 s
DeepSeek R1-0528 (2025-05) $0.004 94 s
DeepSeek V3.2 (2025-12) $0.001 25 s
DeepSeek V4 Pro (2026-08) $0.055 214 s
DeepSeek V4.1 Flash (2026-09) $0.019 62 s

Source: OpenRouter Sketch board (openrouter.ai/benchmarks/media/sketch, default prompt column), captured 2026-10-04; OpenRouter data is restated under CC BY 4.0. The full lists are in the next section.

The three families' complete lists in the default column

Laying out all the models each company submitted to the default column makes the evolution clearer (sorted by cost, ascending):

OpenAI (20): gpt-4.1-nano ($0.00027/7s) → gpt-4o-mini ($0.00035/8s) → gpt-oss-120b ($0.00065/40s) → gpt-5-nano ($0.01/144s) → gpt-4.1 ($0.017/24s) → gpt-5.1-codex ($0.018/14s) → gpt-5.1-codex-mini ($0.019/171s) → gpt-5-mini ($0.025/100s) → gpt-5.6-luna ($0.032/263s) → gpt-4 ($0.043/5s) → gpt-5 ($0.099/106s) → gpt-5.1 ($0.11/101s) → gpt-5.5 ($0.15/69s) → gpt-5.3-codex ($0.21/156s) → gpt-5.6-terra ($0.29/308s) → gpt-5.2-codex ($0.33/346s) → gpt-5.4 ($0.38/429s) → gpt-5.2 ($0.4/415s) → gpt-5.6-sol ($0.57/364s) → gpt-6-astra ($0.87/391s).

Anthropic (11): claude-3-haiku ($0.002/13s) → claude-4.5-haiku ($0.012/12s) → claude-4-sonnet ($0.029/16s) → claude-4.5-sonnet ($0.034/25s) → claude-4.5-opus ($0.055/24s) → claude-4-opus ($0.088/21s) → claude-4.1-opus ($0.11/71s) → claude-4.8-opus ($0.46/222s) → claude-4.6-opus ($0.49/234s) → claude-opus-5 ($0.5/275s) → claude-fable-5.1 ($2.13/519s).

DeepSeek (9): v3-0324 ($0.001/26s) → v3.2 ($0.001/25s) → v3.1-terminus ($0.002/77s) → r1-distill-llama-70b ($0.002/107s) → r1-0528 ($0.004/94s) → v4-flash ($0.007/450s) → v4.1-flash ($0.019/62s) → v4-flash-vision-exp ($0.021/51s) → v4-pro ($0.055/214s).

Flagships across prompt columns: stability is no accident

The same model's cost/time under different prompts (captured data):

Model Pelican on a bike Lighthouse at night Penguin with umbrella Pelican riding a bicycle
GPT-6 Astra $0.87 / 391s $0.74 / 971s $0.91 / 947s —
Claude Opus 5 $0.5 / 275s $0.57 / 277s $0.83 / 379s —
DeepSeek V4 Pro $0.055 / 214s $0.12 / 358s — $0.093 / 484s

Two ways to read this: first, cost is stable across columns (a model's price is set by how deeply it reasons, not by the difficulty of the prompt); second, GPT-6 Astra thought for 971 seconds on "Lighthouse at night" — sixteen minutes, pushing the phrase "the cost of thinking" to its limit.

Four images to understand the price of "reasoning depth"

① The OpenAI line: from "stick figure" to "red scarf." The pelican GPT-4 drew is minimal line work, done in a few strokes; GPT-4o mini's picture is incomplete with a blacked-out background; GPT-5 has a complete body and a clear bicycle structure; GPT-5.5 gains a sun, clouds, and grass; by GPT-6 Astra the pelican wears a red scarf, and the background has a halo and vegetation — the finish looks like two generations apart. Cost rose in step from $0.043 to $0.87 (about 20x), and time from 5 seconds to 391 seconds (about 78x). Every extra cent and second was spent on "running the picture through its head a few more times."

The OpenAI line: GPT-4 to GPT-6 Astra, same-prompt image evolution (prompt 'Pelican on a bike')

② The Anthropic line: from "no image produced" to "the priciest on the board." Claude 3 Haiku produced no valid image in the default column (the sample is a placeholder icon); Claude 4 Sonnet began producing a decent composition; Opus 5 reached scene completeness; by Fable 5.1, a single run costs $2.13 and takes 519 seconds — the priciest order on the whole board. That is fully consistent with its product design of "adaptive thinking always on, default high effort": expensive, but it hands the choice of "how deep to think" to the runtime.

The Anthropic line: Claude 3 Haiku to Fable 5.1, same-prompt image evolution

③ The DeepSeek line: hugging the value curve in the lower-left. V3-0324 costs only $0.001 and takes 26 seconds; R1-0528 drew a scene-based image for $0.004; V4 Pro's image completeness closes in on top closed-source models at $0.055 — only one sixteenth of GPT-6 Astra and one thirty-ninth of Fable 5.1. By V4.1 Flash, $0.019 and 62 seconds, it is the most balanced point of "cheap + fast + image not bad."

The DeepSeek line: V3-0324 to V4.1 Flash, same-prompt image evolution

④ The four flagships on "Lighthouse": change the exam, the gap remains. On "Lighthouse at night": GPT-5.5's picture is simple with sparse detail ($0.55/370s); GPT-6 Astra spent 16 minutes thinking and delivered the most complete picture ($0.74/971s); Claude Opus 5 handed in a near-complete picture in under 5 minutes ($0.57/277s); DeepSeek V4 Pro spent just $0.12 and 6 minutes, with completeness in the second tier — **on this chart, the optimum of efficiency versus quality is "the further lower-left, the better"**.

Same prompt 'Lighthouse at night': GPT-5.5 / GPT-6 Astra / Claude Opus 5 / DeepSeek V4 Pro, four flagships compared

Three counterintuitive findings

  1. The cheapest models are not the strongest, but they are not the weakest either. The five cheapest models on the whole board are all small open-source models (Gemma 3 12B $0.00015, Llama 4 Scout 17B, Ministral 3B/8B, Tencent hy-mt2 1.8B); their image completeness is visibly low — but they prove that "reasoning" and "generation cost" are different things: a small model can complete a one-shot image fast and cheap, it just can't draw the details.
  2. The competition among reasoning models was never "who is cheaper." The flagships get more and more expensive (GPT-6 Astra $0.87, Fable 5.1 $2.13), because buyers are not buying "one image" — they are buying "the deep thinking required to make a good image." The real competition is "for the same budget, whose intelligence density is higher."
  3. The board has "absentees," and the absence itself is information. As of the snapshot (2026-10-04), the newest generation — Claude Opus 5.5 / Sonnet 5.5, GPT-6 Sol/Luna, DeepSeek V4.1 Pro — is not on the board yet; OpenRouter's coverage lags official releases. In other words, the board's "ceiling" keeps being refreshed, but the long-term trend "reasoning cost only rises with time, never falls" is already locked in.

Conclusion: Three Routes, One Convergence Direction

Lay out the three years from 2024-09 to 2026-09 and the three companies' route differences are clear enough to draw as three arrows:

  • OpenAI: standalone o series → unified GPT system with automatic transmission → capability-tier pricing (Sol/Terra/Luna, Astra/6.1 Sol). Direction: capability integration and routing;
  • Anthropic: single-model hybrid reasoning → adaptive thinking always on. Never accepted the "dual-model split"; direction: runtime autonomy and safety tiering;
  • DeepSeek: pure-RL open source → sparse attention for cost cuts → asymmetric architecture + KV cache compression. Direction: democratizing efficiency.

But the three lines converge at the endpoint, forming five industry-wide consensuses:

  1. Reasoning = test-time compute, and cost and time become hard metrics on par with capability;
  2. Tunable thinking effort becomes the universal interface (OpenAI's effort, Claude's five effort levels, DeepSeek's low/high/max plus a continuous 1–100 scale);
  3. Model routing and dynamic adjustment emerge: the system/session automatically decides "how long to think" (Astra supports adjusting effort mid-conversation, Fable 5.1 supports per-message effort);
  4. The efficiency race escalates into architecture-level competition: from price-cut promotions (R1's 1/30, Opus 4.5's -66%, V3.2's -50%) to architecture rebuilds (CED, CSA2, FP4 caches, $0.003 cache-hit pricing on 1M context);
  5. Capability tiering and safety governance become a parallel second battlefield: OpenAI's "Critical" cyber-capability tier with chain-of-thought monitoring, Anthropic's safety classifiers and restricted access (Mythos/Project Glasswing), DeepSeek's MIT open source and self-hosting bar (roughly 2,000 GPUs).

For people who make content, tools, or games, the practical meaning of this is: choosing a model is no longer about who is #1 on the board, but about "how long does your task need it to think." For art concept work, coding agents, and long-document research, pick the "expensive but thinks deep" flagships; for high-frequency, low-cost, deterministic tasks, pick the "cheap and fast" tiers and save the budget for scenarios that really need reasoning. What the Sketch board mirrors is exactly this shift in progress: the industry has moved from "who is smarter" to "whose compute is better spent."

Appendix

1. Data and methods

  • Board: the OpenRouter Sketch board (openrouter.ai/benchmarks/media/sketch), one of six multimodal boards; text models produce a one-shot image and results are scored on the image. This site captured it via the cache of the Halo "AI LLM Leaderboard" plugin on the NAS, capture time 2026-10-04.
  • Spec: the default prompt column "Pelican on a bike" has 143 models; the upstream for that column gives no quality score, and the main metrics are cost per run (lower is better) and generation time (seconds); no "overall quality score" was weight-synthesized by us.
  • The flagship-generation samples from the three companies are hand-picked representative versions; the full 640-row dataset is in the asset file or-sketch-live-full.json, and each company's complete list is in the main text.
  • OpenRouter data is restated under CC BY 4.0; official release dates per vendor are from the official release records and model documentation of OpenAI / Anthropic / DeepSeek respectively.

2. The three companies' complete iteration timeline

When OpenAI Anthropic DeepSeek
2024-09 o1-preview / o1-mini (Strawberry) — —
2024-12 o1 official; o3 teased (ARC-AGI 87.5%) — DeepSeek-V3 (MoE/MLA/FP8)
2025-01 o3-mini launched (first free reasoning model) — R1 / R1-Zero (pure-RL open source)
2025-02 — Claude 3.7 Sonnet (first hybrid reasoning) —
2025-03 — — V3-0324
2025-04 o3 official + o4-mini — —
2025-05 — Claude Opus 4 / Sonnet 4 (thinking + tools) R1-0528
2025-08 GPT-5 (unified system Auto/Thinking) Opus 4.1 V3.1 (hybrid architecture)
2025-09 GPT-5-Codex Sonnet 4.5 V3.1-Terminus; V3.2-Exp (sparse attention)
2025-10 — Haiku 4.5 —
2025-11 GPT-5.1 Opus 4.5 (three effort levels, 66% price cut) —
2025-12 GPT-5.2 — V3.2 (four golds at IMO/IOI, thinking folded into tools)
2026-01/02 GPT-5.2/5.3-Codex Opus 4.6 (adaptive thinking) —
2026-03 GPT-5.4 / 5.4-mini/nano — —
2026-04 GPT-5.5 Opus 4.7 (manual mode removed) V4 series (1M context, open source)
2026-05 — Opus 4.8 (Dynamic Workflows) —
2026-06 — Fable 5 / Mythos 5 (6-09); Sonnet 5 (6-30) —
2026-07 GPT-5.6 (Sol/Terra/Luna) Opus 5 (7-23) V4-Flash official
2026-08 — — V4-Pro official; V4-Flash-Vision-Exp
2026-09 GPT-6 Astra (9-03); Sol/Luna (9-22); GPT-6.1 Sol (9-29) Fable 5.1 / Mythos 5.1 (9-01); Opus 5.5 (9-22) V4.1-Flash (9-10, new architecture); V4 Pro retired (9-14)

3. Key API pricing comparison (per million tokens, official public prices)

Model Input Output Notes
GPT-6 Astra $10 (cache hit $1) $50 1.05M context, 128K output
GPT-6.1 Sol $2 $10 Near Astra, multi-agent beta
Claude Fable 5.1 $10 (cache read $0.25) $50 Adaptive thinking always on, default high
Claude Opus 5.5 $4 $20 Workhorse for everyday professional work
Claude Sonnet 5.5 $2 $10 Speed/intelligence balance
Claude Haiku 4.5 $1 $5 Fastest tier
DeepSeek V4.1 Flash (off-peak) $0.15 (cache hit $0.003) $0.6 Peak/off-peak pricing, doubles at peak; 1M context
DeepSeek V4 Pro (before retirement, off-peak) $0.66 $1.98 Routed to V4.1-Flash after 2026-09-14

4. Glossary

  • Chain of Thought (CoT): the internal reasoning steps a model generates before answering; the core mechanism of reasoning models.
  • Test-time compute: extra compute spent during inference (thinking tokens); the third scaling dimension alongside "pretraining compute" and "post-training compute."
  • RLVR: reinforcement learning with verifiable rewards — training reasoning on objective right/wrong (math/code/competition) rather than human preference.
  • GRPO: Group Relative Policy Optimization, DeepSeek's RL algorithm; skips the critic model and lowers training cost.
  • MoE (Mixture of Experts): a large model that activates only part of its parameters per step; the mainstream architecture of the DeepSeek/open-source route.
  • MLA / DSA / CSA2: multi-head latent attention / sparse attention / compressed sparse attention — technologies that make attention caches smaller and cheaper.
  • effort / budget_tokens: thinking intensity / thinking budget — the two universal knobs for "how long to think."
  • Adaptive thinking: the model decides its own thinking depth (from Claude 4.6; always on from Fable 5.1).
  • Model routing: the system automatically assigns model tiers by task difficulty/cost (GPT-5.6's Sol/Terra/Luna, the GPT-6 family).