promptgarten 🌱

What do I use when? β€” Models & tools compared

Seven common situations, one click: see the tool, the model, and a copy-paste example. Below that: every current model from the big providers, and what each tool uses by default.

Stand / as of: 2026-07-28 Β· Benchmark overview β†’ Β· πŸ“Š Model timeline β†’

🧭 WHAT DO YOU WANT TO DO? β€” CLICK YOUR SCENARIO

Intelligence per dollar: all major models at a glance

Each dot is a model. Further right means more expensive (price per 1M tokens, blended 3:1 input:output, log scale), further up means smarter (Intelligence Index from Artificial Analysis, v4.1). Top left is the sweet spot.

0.5 $1 $2 $5 $10 $20 $25354555Price per 1M tokens (blended 3:1, log scale)Intelligence Index (Artificial Analysis v4.1)Opus 5Fable 5GPT-5.6 SolKimi K3Opus 4.8GPT-5.6 TerraGrok 4.5Sonnet 5GPT-5.6 LunaMuse Spark 1.1Gemini 3.5 FlashGemini 3.1 ProQwen3.7-MaxDeepSeek-V4-Proβ˜… Ratio KingKimi K2.5GLM-4.7Mistral Medium 3.5Haiku 4.5
FrontierMaximum capability β€” when the outcome matters more than the price.
Price-performanceLots of intelligence for the money β€” the everyday zone for most tasks.
BudgetCheap for scale, simple tasks, and experiments β€” often open weights.
Open weights (self-hostable)

Prices = official API list prices (excluding batch/cache), Index = Artificial Analysis Intelligence Index v4.1 in the configuration tested there (usually highest reasoning). As of: 07/17/2026 β€” numbers change fast.

🎬 AS A SHORT VIDEO

The chart in 10 seconds: pricier to the right, smarter towards the top β€” and where the sweet spot sits.

The ranking: intelligence per dollar

Intelligence Index divided by the blended price (3 parts input + 1 part output, per 1M tokens). Higher = more thinking power per dollar. This says nothing about the absolute top spot β€” that's what the index itself is for.

#MODELINTELLIGENCE/$INDEXPRICE (BLENDED)
1DeepSeek-V4-Pro ✳ · DeepSeek81.444.27$0.54
2GLM-4.7 ✳ · Zhipu/Z.ai33.733.70$1.00
3Kimi K2.5 ✳ · Moonshot29.535.37$1.20
4Muse Spark 1.1 Β· Meta25.350.62$2.00
5GPT-5.6 Luna Β· OpenAI22.851.24$2.25
6Grok 4.5 Β· xAI17.953.83$3.00
7Gemini 3.5 Flash Β· Google14.950.20$3.38
8Sonnet 5 (Intro-Preis) Β· Anthropic13.353.35$4.00
9Qwen3.7-Max Β· Alibaba12.345.99$3.75
10Haiku 4.5 Β· Anthropic11.923.71$2.00
11Gemini 3.1 Pro Β· Google10.346.46$4.50
12Mistral Medium 3.5 ✳ · Mistral10.029.95$3.00
13GPT-5.6 Terra Β· OpenAI9.854.95$5.63
14Kimi K3 Β· Moonshot9.557.11$6.00
15Opus 5 Β· Anthropic6.160.6910,00 $
16Opus 4.8 Β· Anthropic5.655.69$10.00
17GPT-5.6 Sol Β· OpenAI5.258.89$11.25
18Fable 5 Β· Anthropic3.059.86$20.00
  • Grok 4.5: price applies up to 200K context β€” above that it doubles (ratio then ~9.0).
  • Gemini 3.1 Pro: price applies up to a 200K prompt β€” above that $4/$18 (ratio then ~6.2).
  • Qwen3.7-Max: currently a time-limited 50% discount (ratio then ~24.5) β€” the list price counts here.
  • Sonnet 5: intro price until 08/31/2026, then $3/$15 (ratio 8.9).
  • Mistral Large 3 ($0.50/$1.50) and Llama 4 Maverick (from ~$0.27/$0.85 at hosters) are missing from the ranking because Artificial Analysis doesn't list an index for them.
  • Sources: official provider pricing pages + artificialanalysis.ai, all accessed on 07/17/2026.

Models from the major providers

The current models from Anthropic, OpenAI, Google, xAI β€” and now also Moonshot (Kimi), DeepSeek, Zhipu (GLM), Meta, Alibaba (Qwen), and Mistral. With price, context, strengths, and what they're meant for.

ANTHROPIC

Fable 5

Anthropic's most capable widely released model β€” built for the most demanding reasoning and long-horizon agentic work (per Anthropic).

🧠 1M tokensπŸ’Ά $10 / $50 per MTok (as of July 2026)
  • βœ… Highest capability tier per Anthropic, for long autonomous agent runs
  • βœ… 1M token context by default
  • βœ… Verifies its own work more often than smaller models
  • βœ… From 20 July 2026 permanently included in Max and Team Premium (at 50% of the regular limits); Pro and Team Standard get a one-time $100 credit and pay API rates after that.
  • ⚠️ Most expensive model of the current Claude generation ($10/$50, as of July 2026)
  • ⚠️ Not the default in Claude Code β€” you must select it explicitly (`/model fable`)

β†’ Use it for very large, multi-hour tasks where a lot is at stake.

Opus 5

For complex agentic coding and enterprise work β€” comes close to the frontier intelligence of Fable 5, but at half the price (according to Anthropic).

🧠 1M tokensπŸ’Ά $5 / $25 per MTok (as of July 2026)
  • βœ… According to Anthropic, state-of-the-art on coding benchmarks (Frontier-Bench, GDPval-AA) and stronger at research/knowledge work than Opus 4.8
  • βœ… Five-level effort control (low to maximum) for balancing intelligence and token usage
  • βœ… Intelligence index 60.69 (Artificial Analysis v4.1, max-effort configuration) β€” putting it numerically slightly ahead of Fable 5 (59.86)
  • βœ… Since Claude Code v2.1.219 (July 24, 2026), the new standard Opus in Claude Code and the new default model on Claude Max
  • βœ… 1M-token context by default, 128K max output
  • ⚠️ Fast mode costs twice as much ($10 / $50 per MTok) and only runs via the Anthropic API or subscription credit β€” not on Amazon Bedrock, Google Clouds Agent Platform, Microsoft Foundry, or Claude Platform on AWS
  • ⚠️ More expensive than Sonnet 5 for everyday tasks

β†’ Use it for complex agentic coding tasks and enterprise knowledge work when quality close to Fable 5 is needed without paying its price.

Opus 4.8

For complex agentic coding and enterprise work β€” a premium model for serious coding and knowledge work (per Anthropic).

🧠 1M tokensπŸ’Ά $5 / $25 per MTok (as of July 2026)
  • βœ… Recommended for complex architectural decisions and multi-step reasoning
  • βœ… 1M token context
  • ⚠️ More expensive than Sonnet 5 for similar everyday tasks
  • ⚠️ Superseded by Opus 5: same price, higher intelligence index. Anthropic no longer lists Opus 4.8 in its main comparison table and offers a migration guide β€” however, it is not deprecated, and there is no retirement date so far.

β†’ For new projects, Opus 5 is the better choice: same price, more recent generation. Opus 4.8 remains useful when existing workflows should not be switched over.

Sonnet 5

The best combination of speed and intelligence β€” the most agentic Sonnet generation yet (per Anthropic).

🧠 1M tokensπŸ’Ά $2 / $10 per MTok until Aug 31, 2026, then $3 / $15 (as of July 2026)
  • βœ… Handles most coding tasks well
  • βœ… Cheaper than Opus
  • βœ… Default model in Claude Code for Pro/Team Standard accounts
  • ⚠️ Weaker than Opus/Fable on very complex architectural questions

β†’ Use it as your everyday model for most coding tasks.

Haiku 4.5

The fastest model with near-frontier intelligence β€” a lightweight version of the strongest Anthropic model at a lower price (per Anthropic).

🧠 200K tokensπŸ’Ά $1 / $5 per MTok (as of July 2026)
  • βœ… Cheapest Claude model
  • βœ… Very fast
  • βœ… 73.3% SWE-bench Verified per Anthropic's own figure
  • ⚠️ Smaller context window (200K instead of 1M)
  • ⚠️ Weaker than Sonnet/Opus on very complex tasks

β†’ Use it for cheap experimentation and simple, clearly scoped tasks.

OPENAI

GPT-5.6 Sol

Frontier model for complex professional work β€” flagship for complex reasoning and coding (per OpenAI).

🧠 ~1.05M tokensπŸ’Ά $5 / $30 per MTok, cached input $0.50 (as of July 2026)
  • βœ… Strongest model in the GPT-5.6 family
  • βœ… Default in Codex CLI (medium reasoning effort)
  • ⚠️ Most expensive model in the 5.6 family

β†’ Use it for tasks where detail and polish matter.

GPT-5.6 Terra

Balances intelligence and cost β€” roughly corresponds to the mini model tier of earlier GPT-5 families (per OpenAI).

🧠 ~1.05M tokensπŸ’Ά $2.50 / $15 per MTok, cached input $0.25 (as of July 2026)
  • βœ… Good middle ground of cost and capability
  • βœ… Recommended in Codex CLI as the "everyday workhorse"
  • ⚠️ Weaker than Sol on the hardest tasks

β†’ Use it as your everyday model when Sol is too expensive.

GPT-5.6 Luna

Optimized for cost-sensitive workloads β€” roughly corresponds to the nano model tier (per OpenAI).

🧠 ~1.05M tokensπŸ’Ά $1 / $6 per MTok, cached input $0.10 (as of July 2026)
  • βœ… Cheapest model in the 5.6 family
  • βœ… Recommended in Codex CLI for "clear, repeatable work"
  • ⚠️ Positioned by OpenAI only for clear, repeatable work

β†’ Use it for simple, repetitive coding tasks at scale.

GPT-5.3-Codex

The most capable agentic coding model to date β€” optimized for agentic coding tasks in Codex or similar environments (per OpenAI).

🧠 400K tokensπŸ’Ά $1.75 / $14 per MTok, cached input $0.175 (as of July 2026)
  • βœ… Specifically optimized for agentic coding
  • βœ… Cheaper than Sol with strong coding performance
  • ⚠️ Marked deprecated in Codex under ChatGPT sign-in β€” availability inconsistent (as of July 2026)

β†’ Use it when you want quality specifically tuned for agentic coding instead of a general-purpose model.

GOOGLE

Gemini 3.5 Flash

Most intelligent model for sustained frontier performance on agentic and coding tasks (per Google).

🧠 ~1M tokens (output up to 65K)πŸ’Ά $1.50 / $9 per MTok, batch/flex $0.75 / $4.50 (as of July 2026)
  • βœ… Positioned by Google as its "most intelligent" model β€” despite the Flash name
  • βœ… Free tier available
  • βœ… Cheaper than Google's own Pro model
  • ⚠️ Knowledge cutoff January 2025

β†’ Use it as Google's top recommendation for agentic coding tasks.

Gemini 3.1 Flash-Lite

Frontier-class performance rivaling larger models at a fraction of the cost β€” Google's most cost-efficient model for high-volume agentic tasks (per Google).

🧠 ~1M tokens (output up to 65K)πŸ’Ά $0.25 (text/image/video) or $0.50 (audio) / $1.50 per MTok (as of July 2026)
  • βœ… Cheapest model in this whole comparison
  • βœ… Free tier available
  • ⚠️ Not positioned for complex tasks

β†’ Use it for translation, simple data processing, and high-volume, simple agent runs.

Gemini 3.1 Pro

Advanced intelligence, complex problem-solving skills, and powerful agentic and vibe-coding capabilities (per Google) β€” still in preview.

🧠 ~1M tokens (output up to 65K)πŸ’Ά $2–4 / $12–18 per MTok, tiered by prompt length (as of July 2026)
  • βœ… Google's most capable model for complex multi-step tasks
  • βœ… Strong vibe-coding capabilities per Google
  • ⚠️ Still in preview, no free tier
  • ⚠️ More expensive than Flash (with an extra tier above 200K prompts)

β†’ Use it for complex, multi-step problems where Flash isn't enough.

XAI

Grok 4.5

xAI: "For everything else, including code, use Grok 4.5. It is the most intelligent and fastest model we’ve built." Cursor co-trains it and calls it its flagship model.

🧠 500K tokensπŸ’Ά $2 / $6 per MTok, cached input $0.50 (as of July 2026)
  • βœ… Co-trained by Cursor as its flagship partner model
  • βœ… Cheaper than most flagship models from other vendors
  • ⚠️ Smaller context window than competitors (500K instead of 1M)

β†’ Use it in Cursor for the hardest tasks β€” per xAI their most intelligent and fastest model.

Grok 4.3

No official positioning statement found β€” listed on the same xAI models page as Grok 4.5, as a cheaper sibling model with a larger context window.

🧠 1M tokensπŸ’Ά $1.25 / $2.50 per MTok (as of July 2026)
  • βœ… Larger context than Grok 4.5 (1M instead of 500K)
  • βœ… Cheaper than Grok 4.5
  • ⚠️ Barely documented β€” no dedicated positioning or model page found

β†’ Use it as a cheaper xAI alternative with more context, when Grok 4.5 is too expensive.

MOONSHOT AI

Kimi K3

Brand new (07/16/2026): according to Moonshot, the largest open-weight model in the world (2.8 trillion parameters, MoE) β€” rank 3 in the Intelligence Index, right behind Fable 5 and GPT-5.6 Sol.

🧠 1M tokens (output up to 1M possible)πŸ’Ά $3 / $15 per MTok, cache-hit input $0.30 (as of: July 2026)
  • βœ… Intelligence Index 57.11 β€” rank 3 of all models (Artificial Analysis, 07/17/2026)
  • βœ… Agentic web research: BrowseComp 91.2 (Moonshot's own figure)
  • βœ… Flat price across the entire 1M context, automatic caching
  • ⚠️ As of 07/17/2026, weights are NOT yet downloadable β€” open release announced only for 07/27/2026
  • ⚠️ Thinking is permanently on β€” fast, cheap answers aren't its thing

β†’ Use it for demanding agent and research tasks when you want frontier performance cheaper than Fable/Sol.

Kimi K2.5

The predecessor: 1-trillion-parameter MoE, open weights (Modified MIT), still available and significantly cheaper.

🧠 256K tokensπŸ’Ά $0.60 / $3 per MTok, cache hit $0.10 (as of: July 2026)
  • βœ… SWE-bench Verified 76.8% (Moonshot's own figure) at a budget price
  • βœ… Open weights β€” self-hostable
  • βœ… Multimodal + agentic search
  • ⚠️ Index 35.37 β€” noticeably below K3 and the frontier models

β†’ Use it for cheap agentic coding or self-hosting with decent performance.

DEEPSEEK

DeepSeek-V4-Pro

The price-performance king of this overview: intelligence per dollar unmatched (ratio 81) β€” open weights under MIT license.

🧠 1M tokens (output up to 384K)πŸ’Ά $0.435 / $0.87 per MTok, cache-hit input ~99% cheaper (as of: July 2026)
  • βœ… By far the best intelligence-per-dollar ratio (index 44.27 at $0.54 blended)
  • βœ… SWE-bench Verified 80.6% (DeepSeek's own figure, Pro-Max mode)
  • βœ… OpenAI- AND Anthropic-compatible API, open weights (MIT)
  • ⚠️ Absolute top-end performance is below the frontier models
  • ⚠️ Benchmark numbers are vendor-reported, not independently validated

β†’ Use it if you want to process a lot of volume cheaply β€” or want to self-host a strong model.

ZHIPU / Z.AI

GLM-4.7

China's open-source workhorse for coding: 358B MoE under MIT license, strong at agentic coding and frontend generation.

🧠 200K tokens (output up to 128K)πŸ’Ά $0.60 / $2.20 per MTok; Flash variant free (as of: July 2026)
  • βœ… SWE-bench Verified 73.8% (Zhipu's own figure) at a budget price
  • βœ… Open weights (MIT) β€” self-hosting via vLLM/SGLang, 44 quantized variants
  • βœ… Free Flash variant to try out
  • ⚠️ Index 33.70 β€” others take the crown for complex reasoning
  • ⚠️ Zhipu already has a newer GLM-5 series β€” 4.7 is no longer the house's flagship model

β†’ Use it for cheap agentic coding, web/frontend generation, or as a self-hosting base.

META

Muse Spark 1.1

Meta's first closed-weight frontier model (07/09/2026) β€” via the new Meta Model API, built for agentic orchestration across apps, MCP servers, and skills.

🧠 1M tokens (with built-in context compaction)πŸ’Ά $1.25 / $4.25 per MTok, cached input $0.15 (as of: July 2026)
  • βœ… Index 50.62 at just $2 blended β€” fourth-best intelligence-per-dollar ratio
  • βœ… Agentic orchestration (apps/MCP/skills, zero-shot) as core positioning
  • βœ… OpenAI-SDK-compatible API, $20 starting credit
  • ⚠️ NOT open-weights β€” a break with Meta's Llama tradition
  • ⚠️ Public benchmarks partly published only as chart images

β†’ Use it for multi-app agents and tool orchestration at a good price.

Llama 4 Maverick

Meta's current open-weights flagship (400B MoE, natively multimodal) β€” runs extremely cheap at hosters like Together, Groq, or Fireworks.

🧠 ~1M tokensπŸ’Ά from ~$0.27 / $0.85 per MTok depending on hoster (Together AI, as of: July 2026)
  • βœ… Cheapest entry in this overview
  • βœ… Open weights + self-hosting
  • βœ… Multimodal, 12 languages
  • ⚠️ No Artificial Analysis index listed β€” limited comparability
  • ⚠️ The Llama 4 Community License is not OSI-compliant (restrictions apply above 700M users)

β†’ Use it for mass processing at minimal cost or your own hosting.

ALIBABA (QWEN)

Qwen3.7-Max

Alibaba's agent flagship ("designed for the agent era") β€” proprietary for the first time instead of open-weight, with hybrid reasoning at the same price.

🧠 1M tokens (output up to 65K)πŸ’Ά $2.50 / $7.50 per MTok list price; currently temporarily βˆ’50% (as of: July 2026)
  • βœ… SWE-bench Verified 80.4% and GPQA Diamond 92.4 (Alibaba's own figures)
  • βœ… Long-horizon agentic work across hundreds of steps as core focus
  • βœ… Very strong multilingual performance (WMT24++ 85.8)
  • ⚠️ No self-hosting β€” a break with the Qwen open-weight tradition
  • ⚠️ Index 45.99 is below the Western frontier models

β†’ Use it for long agent chains and multilingual tasks β€” especially during the discount promotion.

MISTRAL

Mistral Large 3

Europe's open-weights flagship (675B MoE, Apache 2.0) β€” vision, function calling, and an EU hosting option.

🧠 256K tokensπŸ’Ά $0.50 / $1.50 per MTok (model page; an older FAQ still lists $2/$6) (as of: July 2026)
  • βœ… Apache 2.0 open weights β€” the freest license among the big players
  • βœ… Very cheap for its size
  • βœ… EU hosting/data sovereignty
  • ⚠️ No Artificial Analysis index listed
  • ⚠️ Benchmark numbers mostly published only as chart images

β†’ Use it if open source, EU hosting, or the Apache license are decisive for you.

Mistral Medium 3.5

Mistral's coding specialist: 128B dense, Modified MIT, with a strong SWE-bench score for its size.

🧠 256K tokensπŸ’Ά $1.50 / $7.50 per MTok (as of: July 2026)
  • βœ… SWE-bench Verified 77.6% (Mistral's own figure)
  • βœ… Open weights (Modified MIT)
  • βœ… Compact enough for your own hosting
  • ⚠️ Index 29.95 β€” weaker than the MoE giants for broad reasoning

β†’ Use it for coding agents on your own infrastructure.

Claude Docs: Models Overview β†— Β· Claude Docs: Pricing β†— Β· Anthropic: Claude Sonnet 5 Launch β†— Β· OpenAI Docs: Pricing β†— Β· OpenAI Docs: GPT-5.6 Sol β†— Β· Google AI: Gemini Models β†— Β· Google AI: Gemini Pricing β†— Β· xAI Docs: Models β†— Β· Moonshot AI: Kimi K3 Pricing β†— Β· DeepSeek API Docs: Pricing β†— Β· Z.ai Docs: Pricing β†— Β· Meta Model API (Developer) β†— Β· Alibaba Cloud Model Studio: Pricing β†— Β· Mistral Docs: Models β†— Β· Artificial Analysis: Intelligence Index β†— Β· Claude Docs: Migrating to Claude Opus 5 β†—

Model Duels

Not β€œwhich model is best” but β€œwhich model for this one task”. Four head-to-head comparisons, derived from pricing and vendors' own statements β€” including the answer to the question that matters most: when the pricier model is NOT worth it. Prices and figures: as of July 2026.

βš”οΈ Haiku 4.5 vs. Sonnet 5

Is the task mechanical and unambiguous β€” or do you need a model that actually thinks it through?

Claude Haiku 4.5

for mechanical, clearly-scoped tasks with no room for interpretation β€” the instruction is unambiguous, there's nothing to decide, only to execute.

  • Generate commit messages from a diff
  • Filter 100 log lines for a known error pattern
  • Rename a JSON field consistently across 40 files

Claude Sonnet 5

when you're building a feature, debugging across multiple steps, or need to understand how the code works before you change it.

  • Build a new feature that touches several files and modules
  • Chase a bug whose root cause isn't in the error message
  • Understand unfamiliar code, then extend it cleanly

πŸ’Έ When it is NOT worth it: Rule of thumb: using Sonnet 5 for simple, unambiguous edits is money wasted β€” Sonnet costs $2/$10 per MTok (in/out, intro pricing through August 31, 2026) or $3/$15 after that, while Haiku 4.5 is just $1/$5 (as of July 2026) β€” 2x to 3x more for a task Haiku solves just as well. The flip side: Haiku fails once requirements are implicit β€” β€œmake this cleaner” without defining β€œcleaner”, an architecture call between two valid approaches, or a task where you first need to understand the code context before you even know what to do.

πŸ“Š SWE-bench Verified: 73.3% β€” Anthropic states Haiku 4.5 itself scores 73.3% on SWE-bench Verified β€” close to models that cost far more. That's Anthropic's own figure, not an independent test run, but it shows Haiku is strong on clearly-specified coding tasks, not just β€œfast and cheap”. (as of July 2026)

Anthropic: Claude Haiku β†— Β· Anthropic platform: Pricing β†— Β· Claude Code docs: Costs β†—

βš”οΈ Sonnet 5 vs. Opus 4.8 / Fable 5

Does the standard workhorse cover it β€” or do you need the flagship model?

Claude Sonnet 5

as your default for everyday coding work. Anthropic's own Claude Code guidance: β€œSonnet handles most coding tasks well and costs less than Opus.”

  • Everyday features and bug fixes in an active project
  • Code reviews and routine refactors
  • Debugging sessions that don't need hours of autonomous self-correction

Claude Opus 4.8 / Fable 5

for complex architectural decisions, hard bugs spanning many files and subsystems, or long autonomous agent runs with heavy self-correction.

  • Plan an architecture redesign touching several modules at once
  • Chase a bug that spans 10+ files and multiple subsystems
  • An hours-long autonomous run where the model researches, acts, and verifies its own results β€” per Anthropic, Fable 5's strength: β€œsustains long autonomous sessions, investigates before acting, and verifies its work more often than smaller models.”

πŸ’Έ When it is NOT worth it: Rule of thumb: using the flagship model for everyday edits burns money. Sonnet 5 costs $2/$10 per MTok (in/out, intro pricing through August 31, 2026) or $3/$15 after that. Opus 4.8 costs $5/$25 β€” 1.7x to 2.5x Sonnet, depending on which price tier you compare against. Fable 5 costs $10/$50 β€” 3.3x to 5x (all figures as of July 2026). Anthropic's own guidance reserves Opus for β€œcomplex architectural decisions or multi-step reasoning”; Fable 5 isn't the default on any account type and has to be selected explicitly with `/model fable` β€” it isn't positioned for routine work.

Anthropic: Claude Opus β†— Β· Anthropic platform: Introducing Fable 5 β†— Β· Anthropic platform: Pricing β†— Β· Claude Code docs: Model config β†— Β· Claude Code docs: Costs β†—

βš”οΈ Gemini 3.5 Flash vs. Gemini 3.1 Pro Preview

The fast model Google itself calls β€œmost intelligent” β€” or the pricier preview flagship?

Gemini 3.5 Flash

as your default for agentic and coding tasks. Google itself positions it as the β€œmost intelligent model for sustained frontier performance on agentic and coding tasks” β€” not the Pro model.

  • Agentic coding workflows with many chained tool calls
  • Tasks needing search and grounding (Google calls it β€œsuperior search and grounding”)
  • The first thing to try for pretty much anything you'd have reached for β€œPro” for by default

Gemini 3.1 Pro Preview

when you specifically want to test Google's Pro-tier positioning for complex problem-solving β€” and can live with the preview status. Note: Flash and Pro share the same 1M context.

  • Complex problem-solving sessions Google explicitly positions Pro for β€” knowing that prompts beyond 200,000 tokens push Pro into the pricier tier
  • Exploratory sessions where Google describes Pro as β€œthe best model family … for multimodal understanding, agentic capabilities, and vibe-coding”

πŸ’Έ When it is NOT worth it: Curious fact (as of July 2026): Google's own model docs list the fast Flash model β€” not Pro β€” as the β€œmost intelligent model”, while the Pro model is still in preview and costs noticeably more per token: Flash is $1.50/$9 per MTok (in/out) versus Pro at $2–4/$12–18 (tiered by prompt length). Rule of thumb: for most coding tasks, the premium for Pro Preview is hard to justify right now β€” you're paying more for a model Google itself doesn't rank as the smartest. Both models share the same ~1M context β€” long context is not a Pro argument; beyond 200k prompts Pro even gets pricier while Flash stays flat. Pro only pays off when you specifically need its problem-solving positioning.

Google AI: Gemini models overview β†— Β· Google AI: Pricing β†— Β· Google AI: Gemini 3.5 Flash β†— Β· Google AI: Gemini 3.1 Pro Preview β†—

βš”οΈ Subscription vs. API pricing (pay-per-token)

Fixed cost with a cap β€” or usage-based with no ceiling?

Subscription (e.g. Claude Pro/Max, ChatGPT Plus/Pro, Cursor subscription)

when you want predictable, capped costs and use it regularly but not at extreme volume β€” a flat monthly price instead of a fluctuating bill.

  • Daily personal use of Claude Code or Cursor as a solo developer
  • Team seats with a predictable monthly budget instead of a variable API bill

API / pay-per-token

when your usage swings a lot, you run automated batch jobs, or you need more volume than a subscription cap allows.

  • Large overnight batch processing β€” the Batch API gives 50% off input and output on every Claude model (as of July 2026)
  • A CI step that calls a cheap model like Haiku 4.5 automatically a few times a day
  • Repeated large contexts via prompt caching β€” a cache hit costs just 0.1x the normal input price, i.e. 90% less (as of July 2026)

πŸ’Έ When it is NOT worth it: Rule of thumb: pay-per-token API pricing rarely pays off for occasional users β€” with irregular, low usage, a flat-rate subscription is usually cheaper and you don't have to track tokens. The reverse is also true: a subscription is the wrong choice for mass batch processing β€” subscription caps are built for interactive session use, not thousands of automated calls. That's where the API, with batch discounts (50% at Claude) and prompt caching (up to 90% off repeated contexts), is clearly cheaper (as of July 2026).

Anthropic platform: Pricing (batch, caching) β†— Β· Claude: Pricing (Abos) β†—

πŸ’Ά PRICE CALCULATOR β€” WHAT DOES YOUR SETUP COST PER MONTH?

β‰ˆ MONTHLY COST (30 DAYS, LIST PRICES)

DeepSeek-V4-Pro
$15.66
GLM-4.7
$24.60
Kimi K2.5
$27.00
Haiku 4.5
$45.00
GPT-5.6 Luna
$48.00
Muse Spark 1.1
$50.25
Gemini 3.5 Flash
$72.00
Grok 4.5
$78.00
Gemini 3.1 Pro
$96.00
Qwen3.7-Max
$97.50
GPT-5.6 Terra
$120
Sonnet 5
$135
Kimi K3
$135
Opus 4.8
$225
GPT-5.6 Sol
$240
Fable 5
$450

Sonnet 5: intro pricing $2/$10 until Aug 31, 2026 β€” calculator uses the regular $3/$15. Gemini 3.1 Pro: price applies up to 200k prompts; above that $4/$18.
List prices from vendor pricing pages (as of July 2026), excluding caching/discounts/subscriptions β€” real costs may differ. Sources: Anthropic Pricing Β· OpenAI Models Β· Gemini API Pricing Β· xAI Models Β· Kimi Pricing Β· DeepSeek Pricing Β· Z.ai Pricing Β· Meta Model API Β· Alibaba Model Studio

Inside the tool: which model for what?

Every tool ships with a default model and its own recommendations for specific tasks. Here's how to switch models β€” and when it's worth it.

Claude Code

The default depends on the account (mostly Opus 5 or Sonnet 5 since Claude Code v2.1.219) β€” switch with `/model <name>` or the picker via `/model`.

Everyday coding tasks

sonnet handles most tasks well and costs less than Opus (per Anthropic).

Complex architectural decisions

opus recommended for complex reasoning tasks (per Anthropic).

Simple subagent tasks

haiku fast and efficient for simple tasks (per Anthropic).

Claude Code Docs: Model Config β†— Β· Claude Code Docs: Costs β†—

Codex CLI

The default is GPT-5.6 Sol with "medium" reasoning effort β€” switch with `/model` or `codex --model <id>`.

Detail and polish work

5.6 Sol "for detail and polish" β€” start with Sol if unsure (per OpenAI).

Everyday workhorse

5.6 Terra competitive with GPT-5.5 at lower cost (per OpenAI).

Clear, repeatable work

5.6 Luna lowest cost in the family (per OpenAI).

OpenAI: Codex models β†— Β· OpenAI: Codex CLI quickstart β†—

Cursor

Pick a model from the picker; "Auto" automatically balances intelligence, cost, and reliability.

Everyday tasks

Auto "good for everyday tasks" (per Cursor).

Complex, multi-step tasks

Premium / Grok 4.5 Premium selects the most capable models; Grok 4.5 is Cursor's own flagship model (per Cursor).

Fast interactive coding

Composer Cursor's fast, cost-efficient model, built for interactive coding (per Cursor).

Cursor Docs: Models β†— Β· Cursor Help: Available Models β†—

Aider

No single documented default model anymore β€” you supply your own model via `--model` or an API key environment variable.

Choosing a model

--model <name> Aider is deliberately model-agnostic β€” works with almost any LLM (per Aider docs).

Listing a provider's models

--list-models anthropic/ shows available models for a provider before you pick one.

Fine-tuning (e.g. thinking budget)

.aider.model.settings.yml persistent, advanced per-project model settings.

Aider Docs: LLMs β†— Β· Aider Docs: Anthropic β†—

Antigravity CLI

All plans include both Gemini core models (3.5 Flash, 3.1 Pro) β€” switch with `/model`; third-party models (Claude, GPT-OSS) are available on every plan except Enterprise.

Fast everyday tasks

Gemini 3.5 Flash available on every plan, positioned for agentic and coding tasks (per Google).

More complex problems

Gemini 3.1 Pro more advanced intelligence, also available on every plan (per Google).

Using third-party models

Claude Opus/Sonnet 4.6 (not on Enterprise) available on Free, Google AI Plus, Pro, and Ultra; only the Enterprise plan doesn't offer them.

Antigravity Docs: Models β†— Β· Antigravity Docs: Plans β†—

The tools, side by side

Claude Code

CLI + IDE extensions + webOpen Source: noClaude subscription or API costs

Agentic work in the terminal: long tasks, automation, loops, subagents.

  • βœ… Subagents, hooks and MCP support built in
  • βœ… Artifacts: shareable live pages straight from a session
  • βœ… Default model depends on your plan (Opus 4.8 or Sonnet 5, 1M token context on Sonnet 5) β€” switch via /model
  • ⚠️ No graphical editor β€” terminal-first is a matter of taste
  • ⚠️ Tied to the Claude ecosystem

Claude Code docs: What's new β†— Β· Claude Code docs: Subagents β†—

Cursor

IDE (VS Code base)Open Source: nosubscription

IDE-first: see, click, edit β€” with agents in the background.

  • βœ… Background and cloud agents keep working while you edit
  • βœ… Composer for multi-file changes
  • βœ… Automations: agents start on repo events or timers
  • ⚠️ Closed source, subscription model
  • ⚠️ Acquisition by SpaceX announced (June 16, 2026) β€” watch this space

Cursor changelog β†— Β· TechCrunch: SpaceX to acquire Cursor β†—

OpenAI Codex CLI

CLIOpen Source: yesChatGPT subscription or API

If you live in the OpenAI ecosystem and want an open CLI.

  • βœ… Subagents: up to 6 in parallel
  • βœ… Fast roundtrips over a persistent WebSocket
  • βœ… Bundled with the ChatGPT desktop app
  • ⚠️ Strengths depend heavily on the OpenAI model lineup
  • ⚠️ Younger than the competition, ecosystem still growing

OpenAI Codex changelog β†—

Aider

CLIOpen Source: yesfree β€” you only pay your model API

Full control: your model, your keys, git-native workflow.

  • βœ… Model-agnostic β€” works with almost any LLM
  • βœ… Git-native: clean auto-commits per change
  • βœ… Fully open source, no vendor lock-in
  • ⚠️ Less automation comfort than the big suites
  • ⚠️ Setup and model choice are on you

Aider β€” official site β†—

Antigravity CLI

CLIOpen Source: noGoogle account β€” Free/Pro/Ultra quota (no bring-your-own API key)

Google's successor to Gemini CLI for individual users: a terminal agent in the Google ecosystem with MCP, plugins, and skills.

  • βœ… Officially named by Google as the successor to Gemini CLI for individual users
  • βœ… Full MCP support (stdio + remote) plus its own plugin/skill system (`agy plugin`, `/skills`)
  • βœ… Non-interactive mode (`agy -p`) for scripts, just like the other CLIs
  • ⚠️ Very young product (repo created May 2026) β€” no open-source license file in the repo (`license: null`)
  • ⚠️ Free quota only described qualitatively ("meaningful"/"generous"), no concrete numbers

Google Blog: Transitioning Gemini CLI to Antigravity CLI β†— Β· Antigravity CLI Docs: Command Reference β†—

Sources: All claims come from the linked official sources; the loop fixes reported errors. πŸ› button bottom right.