
There is a question I get asked more than almost any other.
Which AI model should I use?
It is the wrong question. Not because it is a bad question to ask. Because it assumes there is one right answer. There is not. The right model depends entirely on what you are trying to do, how much it costs to be wrong, and how much you are willing to pay for the difference.
This week, we map the whole landscape. Every major model family, what each one is good at, what it costs, and how to build a simple decision framework so you stop guessing and start choosing deliberately.
We also cover the new GPT-5.6 Terra and Luna pricing, because the gap between the cheapest and most expensive options is now large enough that getting this decision wrong is an expensive habit.
The mistake almost everyone makes
Most people pick one model and use it for everything. Not because they have decided it is the best choice for every task. Because switching feels like effort, and the interface they already have open is the path of least resistance.
This is the equivalent of using a sports car to move furniture because it is the car parked in your driveway. It works. It is also wildly inefficient, and eventually something breaks that a more suitable vehicle would have handled without issue.
The professionals getting the most out of AI right now are not the ones with access to the most powerful model. They are the ones who understand what each model is for and route their tasks accordingly.
The frontier model landscape in one place
Here is where things stand as of this week. Three major closed-model families, each now offering a tiered structure similar to what we covered with GPT-5.6 last week.
OpenAI: GPT-5.6 family
Sol, Terra, and Luna. Sol is the flagship for complex reasoning. Terra is the balanced everyday model. Luna is the fast, low-cost tier for volume and speed.

Terra at $2.50 input and $15 output is roughly half the price of Sol for work that does not require the deepest reasoning tier. Luna at $1 and $6 is a fifth of the cost of Sol. For a high-volume product feature or an internal tool processing thousands of requests a day, that difference compounds into real money very quickly.
Anthropic: the Claude family
Claude currently ships four active models, each with a distinct role. Fable 5 sits above Opus 5 as the most capable and most expensive tier, with Opus, Sonnet, and Haiku forming a familiar capability ladder beneath it.

Sonnet 5’s introductory pricing of $2 input and $10 output runs through the end of August 2026, making it, for the moment, one of the strongest value propositions across every frontier provider. Claude also offers prompt caching that cuts input costs by up to 90 percent on repeated context, and a Batch API with a flat 50 percent discount for work that does not need to happen in real time.
Google: the Gemini family
Gemini’s pricing structure has become notably aggressive on the low end, and the Flash tier in particular has become one of the strongest price-to-performance options across the entire market.

Two things are worth knowing before you budget around Gemini. First, thinking tokens, the tokens Gemini spends reasoning before it answers, are billed at output rates, so your actual bill can run higher than the sticker price suggests for complex tasks. Second, Gemini 3.1 Pro’s 2 million token context window is the largest available from any major provider, which makes it the strongest choice specifically when you need to load an entire codebase, a long legal document, or a large research corpus into a single call.
The open-source alternative
Alongside the three closed frontier families, a genuinely competitive open-source ecosystem now exists, and for certain workloads it is the most economically sensible choice available.
DeepSeek V4. The current value leader. MIT-licensed, meaning fully open with no restrictions on commercial use. Splits into V4-Pro (1.6 trillion parameters, 49 billion active) and the lighter V4-Flash. Both run a 1 million token context window. Scores 83.7% on SWE-bench Verified, competitive with frontier closed models on coding. Hosted API pricing runs as low as $0.27 per million tokens depending on the provider, though the hosted API itself runs from China, so regulated workloads should self-host the open weights to keep data in their own environment.
Llama 4 (Meta). Ships under the Llama 4 Community License, not a strict open-source license, but with weights freely available for most commercial use. The Scout variant packs 109 billion total parameters with only 17 billion active, enabling deployment on a single H100 GPU while supporting a 10 million token context window, the largest of any model in this comparison. Maverick scales further and outperforms GPT-4o on MMLU-Pro.
Qwen 3.6 (Alibaba). Punches above its parameter class and is frequently benchmarked as matching Claude Sonnet on several reasoning tasks at a fraction of the API cost through hosting providers.
Where to run open-weight models. Groq, Together AI, Fireworks AI, and DeepInfra all host these models at 50 to 90 percent lower cost than frontier APIs, without you needing to manage any infrastructure. Groq in particular runs on custom hardware built specifically for inference speed rather than training, which makes it the fastest option available for real-time applications built on open models.
The self-hosting question
If you are processing enough volume, self-hosting an open-weight model on your own infrastructure can beat API pricing entirely. The commonly cited break-even point is around 2 million tokens processed daily. Below that volume, API access through a hosting provider is almost always cheaper once you account for GPU costs, engineering time, and maintenance. Above that volume, self-hosting starts to win, particularly for output-heavy workloads.
A side-by-side view: the whole landscape

A pattern is visible immediately. Google’s budget tier, Gemini 3.1 Flash-Lite at ten cents per million input tokens, is dramatically cheaper than any other provider’s equivalent tier. If your workload is genuinely high-volume and low-complexity, Gemini’s cheapest tier deserves serious consideration purely on cost grounds.
The decision framework: matching task to model
Here is the practical system. Ask three questions before you open any model.
How complex is the reasoning required? If the task involves multiple steps, ambiguous instructions, competing considerations, or a decision that would be expensive to get wrong, you need a flagship-tier model: Sol, Opus, Fable, or Gemini 3.1 Pro. If the task is well-scoped with a clear expected output, a mid-tier model handles it well: Terra, Sonnet, or Gemini 3.6 Flash.
How much volume are you processing? A single important document deserves a flagship model. Ten thousand customer support tickets per day do not. For high-volume, repetitive, or latency-sensitive work, the budget tiers, Luna, Haiku, or Gemini Flash-Lite, deliver strong results at a fraction of the cost.
How much context do you need to load? If you are working with an entire codebase, a lengthy legal document, or a large research corpus in a single call, context window size becomes the deciding factor. Gemini 3.1 Pro’s 2 million token window and Llama 4 Scout’s 10 million token window are the two standout options here, ahead of anything Anthropic or OpenAI currently offers.
Practical routing examples
Drafting a client proposal. Mid-tier is sufficient. Claude Sonnet 5 or GPT-5.6 Terra. The task is well-defined and the cost of using a flagship model adds no meaningful quality improvement.
Reviewing a 200-page contract for risk. Flagship tier, and consider the context window carefully. Gemini 3.1 Pro if the document plus supporting material exceeds what fits comfortably elsewhere. Claude Opus 5 or Fable 5 if the reasoning depth matters more than raw context size.
Classifying 5,000 support tickets by category. Budget tier. GPT-5.6 Luna, Claude Haiku, or Gemini Flash-Lite. This is a high-volume, low-complexity task where speed and cost matter far more than reasoning depth.
Building an internal tool that queries a document set thousands of times a day. Consider open-weight models via a hosting provider, or self-hosting if you cross the volume threshold. DeepSeek V4-Flash or Llama 4 Scout via Together AI or Fireworks will handle this at a fraction of frontier API cost.
A complex, multi-step coding task spanning several files. Flagship tier. Claude Opus 5 currently leads the Agentic Index benchmark at 55.3, ahead of GPT-5.6 Sol at 54.0. For genuinely hard, long-running coding work, this is where the gap between tiers is most visible.
The strategy most teams are missing: model routing
The most sophisticated approach, and the one increasingly used by teams who take their AI costs seriously, is not choosing one model. It is building a system that routes each request to the appropriate tier automatically.
A simple version of this: default every request to the cheapest capable tier. If the output fails a quality check, or if the task is flagged as complex during intake, escalate to the next tier up. Anthropic explicitly recommends this pattern: route to Sonnet first, escalate to Opus or Fable only when needed.
This single practice, routing by default rather than defaulting to the most expensive option out of habit, is consistently cited as the highest-leverage cost optimisation available to any team running AI at scale.
Reflection for the week
The AI model landscape has matured to the point where the question is no longer which model is best. Every major provider now ships a tiered family designed around the same insight: different tasks need different amounts of intelligence, and paying frontier prices for every request is not a sign of quality. It is a sign of not having thought about the decision.
The professionals building genuine advantage right now are not the ones with access to the most powerful model. They are the ones who understand the full landscape well enough to route deliberately, task by task, and treat every model choice as a decision rather than a default.
The question worth sitting with this week:
Look at the AI tasks you ran this week. How many of them used your most expensive available model by default, and how many of them actually needed that level of capability?
Help shape future Learn with Tochii articles
I’m putting together future guides on practical ways professionals can use AI at work, in school and in everyday life.
What is one AI question you have always wanted answered?
Your response will help decide what I write about next.
Want to build the skill of choosing and directing AI strategically?
Model selection, cost management, and building AI workflows that scale are core parts of what we teach in Amakora Group’s AI programmes.
Explore our courses: amakoragroup.com/programs
Apply for the fellowship: amakoragroup.com/apply
Until next week,
Tochii
Founder, Learn with Tochii | Amakora Group
Inspire · Educate · Empower
📧 contact@tochukwuachebe.com
🌐 https://www.tochukwuachebe.com