📊 Model Timeline
When did which model arrive — and when does something disappear? The chronicle of major AI models: releases, sunsets, and price changes, each with a source.
Claude Sonnet 5 — end of introductory pricing
The introductory price of $2/$10 per million tokens ends; after that the regular price of $3/$15 applies.
DeepSeek: old model names get migrated
From this day, the old API names deepseek-chat and deepseek-reasoner point to V4-Flash — anyone using the old models should check their integration.
Claude Opus 5
Anthropic released Claude Opus 5: the same price as Opus 4.8 ($5/$25 per million tokens), a 1M-token context, thinking enabled by default, a five-level effort control, and a fast mode at double the price ($10/$50). Since Claude Code v2.1.219 (July 24, 2026), Opus 5 is the new standard Opus there; Opus 4.8 remains in use as an automatic fallback for safety-related refusals.
Claude Fable 5 — permanent in Max & Team Premium
From 20 July 2026 included in all Max and Team Premium plans (at 50% of limits). Pro and Team Standard keep access via usage credits and receive a one-time $100 credit. The previously announced sunset did not happen.
Kimi K3
2.8 trillion-parameter MoE with 1M context — according to Moonshot the largest open model in the world; the weights follow on July 27, until then API-only.
Grok 4.5
Positioned for agentic coding, 500K context, $2/$6 per million tokens — the cheapest model in xAI's own comparison chart.
Inkling (Thinking Machines)
First open-weights model from Mira Murati's lab: an MoE with 975B total / 41B active parameters, 1M context, 77.6% SWE-bench Verified — billed as the strongest open US model. A smaller preview variant, Inkling-Small, has 12B active parameters.
Muse Spark 1.1 + Meta Model API
Meta's first closed-weight frontier model, accessible via the new OpenAI-SDK-compatible Meta Model API (public preview).
GPT-5.6 family (Sol, Terra, Luna)
Three tiers: Sol as the flagship ($5/$30), Terra as the mid-range option ($2.50/$15), and Luna as the cheapest variant ($1/$6).
Gemini CLI + Code Assist for GitHub discontinued
Google is discontinuing Gemini CLI for free users; the successor is the Antigravity CLI. Code Assist for GitHub: deprecation starts June 18, full shutdown July 17.
Qwen3.7-Max
Alibaba's agentic flagship ("designed for the agent era") — proprietary for the first time instead of open-weight, with 1M context.
Gemini 3.5 Flash (GA)
According to Google, 4x faster output and ahead of its own 3.1 Pro on coding benchmarks — Flash is no longer a budget tier ($1.50/$9).
GLM-4.7
358B MoE under the MIT license — an open-weights workhorse for agentic coding, with a free Flash variant.
Mistral Large 3
Europe's open-weights flagship: 675B MoE under Apache 2.0, with vision and an EU hosting option.