Skip to content
Learn Kiro.

Core knowledge · Chapter 10 of 63

Choosing a Model

Every model Kiro offers with context windows and credit multipliers, plus what Auto model selection does, how to change the model, and reasoning effort.

All levels 12 min read last reviewed 2026-09-24

◎ Learning objective

Pick the right model for a given task by weighing capability, context window, and credit cost, and explain what the Auto router and reasoning effort do.

Kiro’s model picker holds Anthropic’s Claude line, OpenAI’s GPT-5.6 tiers, several open-weight models, and an Auto router that chooses for you. Each carries a credit multiplier, and the spread between the cheapest and the most expensive is 88 times for identical work, which is why this dropdown deserves more thought than it usually gets.

There is a small dropdown in Kiro’s chat panel that most people set once and never think about again. It deserves more attention than that, because it is one of the few settings that changes both the quality of your results and what they cost you.

Every model in Kiro has a credit multiplier: how many credits a unit of work consumes compared to the default. Auto sits at 1.0x. GPT-5.6 Sol, the most expensive entry, is 4.4x. Claude Opus 5 is 2.2x. Qwen3 Coder Next is 0.05x. That is an 88x spread between the top and bottom of the menu for the exact same request, and it widens to 176x on the largest GPT-5.6 requests. Model choice is not a preference setting; it is a cost lever with a dial on it.

Credit multiplier A number showing how many credits a model consumes relative to Kiro's Auto setting, which is defined as 1.0x. A 2.2x model spends 2.2 credits where Auto would spend 1. and context window The maximum amount of text (code, conversation, files) a model can consider at once, measured in tokens. 1M means roughly a large codebase; 200K means a substantial slice of one. are the two numbers on this page worth memorizing. Everything else follows from them.

What models does Kiro use?

Verified 2026-09-24 against Kiro’s model documentation. Every model below is available on all paid tiers unless noted; multipliers are relative to Auto at 1.0x.

ModelContextMultiplierWhere it fits
Auto (router)1M1.0xThe default, and the baseline every other number is measured against
Claude Fable 5.11M6.0xEnterprise-only preview (Sep 2026, US East); traces root causes across complex tasks
Claude Opus 51M2.2xNewest (July 2026); long agentic runs, multi-agent coordination, code review
Claude Opus 4.8 / 4.7 / 4.61M2.2xPrevious Opus generations, same price as the newest
Claude Opus 4.5200K2.2xOldest Opus still listed; smaller window at the same cost
Claude Sonnet 51M1.3x”Anthropic’s most agentic Sonnet model yet”; the natural daily driver
Claude Sonnet 4.61M1.3xThe previous Sonnet with a full 1M window
Claude Sonnet 4.5 / 4.0200K1.3xSonnet 4.5 is the model the Free tier gets
Claude Haiku 4.5200K0.4xFast and cheap; built for volume, not for subtlety
OpenAI GPT-5.6 Sol1M4.4x / 8.8xFlagship of the GPT-5.6 tier, and the priciest model in Kiro
OpenAI GPT-5.6 Terra1M2.2x / 4.4xThe balanced middle tier, now priced level with Opus 5
OpenAI GPT-5.6 Luna1M1.1x / 2.2xThe fastest of the three; cheap next to Sol, not cheap outright
GLM-5200K0.5xOpen-weight; a capable budget option
DeepSeek 3.2128K0.25xOpen-weight; smallest window on the list
MiniMax M2.5200K0.25xOpen-weight
MiniMax M2.1200K0.15xOpen-weight; older and cheaper than M2.5
Qwen3 Coder Next256K0.05xOpen-weight; the cheapest option by a wide margin

The GPT-5.6 rows carry two numbers because they are the only models in Kiro priced in two tiers. A request up to 272K tokens bills at the first rate; a request above 272K tokens bills at double it. That threshold is invisible in the picker — you do not choose it, your session size does — so a long GPT-5.6 session can cost twice what the table appears to promise without anything announcing the change. Every other model on the list has one flat multiplier whatever the request size.

Two things stand out. First, the Anthropic families price by family, not by version: every Opus costs 2.2x whether it is 4.5 or 5, and every Sonnet costs 1.3x. So within a family, there is rarely a reason to pick an older version — you pay the same either way. Second, the open-weight tier isn’t slightly cheaper, it is cheaper by an order of magnitude. That gap is what makes it worth thinking about at all.

The GPT-5.6 arrival matters historically too: Sol, Terra, and Luna landed on July 14, 2026 as the first non-Anthropic frontier models offered in Kiro. Roughly, Sol competes with Opus, Terra with Sonnet, and Luna is the speed option.

And their prices have moved twice. On July 31, 2026 Kiro passed an OpenAI price cut through: Terra dropped from 1.2x to 1.0x and Luna from 0.6x to 0.1x, which briefly made Luna the cheapest non-open-weight model in Kiro. On September 14, 2026 that reversed. The 1M context window arrived and the whole family was re-priced onto the two-tier scheme: Sol 2.4x to 4.4x, Terra 1.0x to 2.2x, Luna 0.1x to 1.1x. Terra no longer matches Auto, it matches Opus 5. Luna is no longer a budget option at all — at 1.1x it costs more than Haiku 4.5, more than every open-weight model, and more than Auto. If you built a high-volume habit on Luna in August, that habit now costs eleven times what it did. Treat this as the standing reminder that multipliers are pricing, not physics: re-read the models page before you build a habit on a number.

What is Kiro’s default model?

Auto. If you have never opened the model picker, that is what has been answering you, and it is a perfectly defensible place to stay. Auto is also the reference point for the whole pricing table: it is defined as 1.0x, so every other multiplier on this page is a comparison against it.

What does Auto model selection do?

Auto is not a model. It is a router: for each request, it decides which underlying model should answer, combining frontier models with smaller specialized ones, reading intent from how you phrase things, and reusing cached work where it can. You get a 1M context window and, by definition, a 1.0x multiplier.

Three details from Kiro’s documentation are worth having:

  • It routes to frontier models with automatic fallback, so a model being busy or unavailable does not stop your request.
  • The class depends on your plan: Sonnet-class or better on Free, Opus-class or better on paid tiers.
  • It never routes to experimental models. The experimental entries in the lineup are opt-in only.

Auto being the baseline is the part people miss. When Kiro says a model is 2.2x, it means 2.2x Auto, so choosing Opus 5 for everything is a standing decision to spend a bit over double what the default would have spent.

How do I change the model in Kiro?

Three places, in increasing order of how permanent the choice is.

  • In the IDE, click the model name in the chat input. The picker opens there, and the same control is where the effort panel lives.
  • In the CLI, choose from the session’s model selector. The CLI has remembered your model preference across sessions since v2.6.0 (June 2026), so setting it once sticks.
  • In a custom agent, model is a configuration field: model in the CLI’s agent JSON, and model in the YAML front matter of an IDE Markdown agent. This is the underused one. A “quick-fixer” agent pinned to Haiku and a “reviewer” agent pinned to Opus 5 means the right model is selected by the job, not by whatever you last clicked.

That third option is the closest thing to a set-and-forget answer. If you find yourself switching models manually several times a day, that is a signal to encode the pattern into agents instead. See Custom agents for the file formats.

What is reasoning effort?

Reasoning effort is the second dial, and most people never find it. It controls how much thinking the model does per turn, independently of which model you picked.

Five levels: low, medium, high, xhigh, max. Support varies by model. Claude Opus 5, Opus 4.8, and Sonnet 5 offer all five. Opus 4.7, Opus 4.6, and Sonnet 4.6 offer everything except xhigh. The GPT-5.6 tiers (Sol, Terra, Luna) offer all levels and default to high.

Where to set it:

WhereHow
IDEClick the model name in the chat input, then the Effort panel
CLI, one session/model (includes effort since v2.23.0) or /effort
CLI, at launch--effort
CLI, as a defaultchat.modelDefaults in ~/.kiro/settings/cli.json or .kiro/settings/cli.json

Precedence runs session > workspace > user > built-in, so a /model or /effort change during a session beats anything configured underneath it.

The trade-off is exactly what you would guess: lower effort is faster, produces shorter answers, and consumes fewer credits. That makes effort a genuine cost lever alongside the multiplier, and a better one for routine work than dropping to a weaker model. A high-effort Sonnet 5 and a low-effort Sonnet 5 are the same model with different budgets for thinking.

A practical decision guide

The facts above are Kiro’s. The advice below is recommended practice, not official guidance — a starting point to adjust once you have your own sense of how each model behaves on your codebase.

  • Daily driver: Auto, or Sonnet 5. Auto if you want the cheapest sensible default and don’t mind that the underlying model varies. Sonnet 5 (1M, 1.3x) if you want consistent frontier behavior you can build habits around. The 30% premium buys predictability.
  • The genuinely hard work: Opus 5. Large refactors, long autonomous runs, debugging that has already defeated one other model, code review where a missed issue is expensive. Accept the 2.2x; you are buying it for a reason.
  • Quick, mechanical, high-volume tasks: Haiku 4.5 or an open-weight model. Renaming things, writing boilerplate tests, formatting, drafting commit messages, first-pass summaries. Here the multiplier is the point: at 0.05x, Qwen3 Coder Next lets you do twenty tasks for the price of one on Auto.
  • The GPT-5.6 tiers when you want a second opinion. A different model family sometimes sees a problem your usual one keeps walking past. Since the September 14 re-pricing this is a capability argument and no longer a cost one: Sol at 4.4x is the most expensive choice available, Terra at 2.2x costs the same as Opus 5, and Luna at 1.1x is dearer than Auto. All three now hold 1M context. Reach for them when a second family is worth paying for, not to save credits — and watch the 272K line, past which each of those rates doubles.

Worked example

Say a mid-sized spec task — requirements, design, tasks, and a first implementation pass — costs about 10 credits on Auto. Same task, different picker settings:

  • GPT-5.6 Sol: ~44 credits, and ~88 if the session runs past 272K tokens
  • Opus 5 or any Opus: ~22 credits
  • GPT-5.6 Terra: ~22 credits (the same as Opus 5)
  • Sonnet 5: ~13 credits
  • GPT-5.6 Luna: ~11 credits
  • Haiku 4.5: ~4 credits
  • Qwen3 Coder Next: ~0.5 credits

Nobody should run a real spec on Qwen3 and expect Opus results. But look at the middle of that list: Sonnet 5 costs three credits more than Auto and a full nine credits less than Opus. Most of the time, that middle row is the right row. And note the top of the list — the same task on Sol costs four times what it costs on Auto, or eight times on a long session.

Context windows, in practice

The context window is a ceiling, not a target. A 1M window doesn’t make a model smarter on a small task; it makes it possible to work on a large one.

  • 1M models (Opus 5 and 4.8/4.7/4.6, Sonnet 5 and 4.6, the GPT-5.6 tiers since September 14, and Auto) earn their window when you point the agent at a whole codebase, run a long multi-file refactor, or hold a session open across many turns. On the GPT-5.6 tiers, remember that using the top of that window is what doubles the rate.
  • 200K–256K models are plenty for focused work: one module, one bug, one feature. Most of what you do in a day fits comfortably.
  • DeepSeek 3.2’s 128K is the tightest on the list. Fine for a file or two; it will run out on a sprawling session.

Worth remembering either way: filling a large window with irrelevant material makes results worse, not better, and costs more doing it. A big window is permission to include what matters, not an invitation to include everything.

A note on tiers

The Free tier gets only Sonnet 4.5 and the open-weight models. Everything else on the table — every Opus, Sonnet 5, the GPT-5.6 tiers, Haiku, and Auto itself — needs a paid plan. If you are following a tutorial that assumes Opus and you are on Free, that is why the option is not in your dropdown. (How plans and credits work is covered in Credits & Pricing; kiro.dev/pricing is the authoritative source for current numbers.)

Opus 5 also arrived through a gradual rollout to paid customers, so a brand-new model may take a little while to appear for you specifically.

This page will go stale

Kiro’s model lineup changes roughly monthly: new models arrive, multipliers get adjusted, older versions retire. The GPT-5.6 family was re-priced twice in seven weeks, in opposite directions, which is the clearest evidence available that this table has a shelf life. The numbers here were checked on September 17, 2026 and are honest as of that date.

The authorities are kiro.dev/docs/models for the current lineup and kiro.dev/changelog/models/ for what changed and when. When this page and those pages disagree, they are right.

In short: Auto is 1.0x and the sane default. Sonnet 5 is the consistent daily driver at 1.3x. Opus 5 at 2.2x is for work that earns it. Haiku and the open-weight models exist so mechanical tasks cost almost nothing. Check the multiplier before you switch, and check the docs before you trust this table.

Frequently asked questions

What models does Kiro use?

Kiro offers Auto (a router), Anthropic's Claude Opus 5 and 4.8, 4.7, 4.6 and 4.5, Claude Sonnet 5, 4.6, 4.5 and 4.0, Claude Haiku 4.5, OpenAI's GPT-5.6 Sol, Terra and Luna, and the open-weight models GLM-5, DeepSeek 3.2, MiniMax M2.5 and M2.1, and Qwen3 Coder Next. Each carries a credit multiplier relative to Auto at 1.0x.

What is Kiro's default model?

Auto. It is a router rather than a model: for each request it picks which underlying model should answer. It has a 1M context window and is defined as the 1.0x baseline every other multiplier is measured against, so if you have never touched the picker you have been on Auto.

How do I change the model in Kiro?

In the IDE, click the model name in the chat input and choose from the picker. In the CLI, select a model from the session's model selector; the CLI has remembered your model preference across sessions since v2.6.0. You can also pin a model per custom agent using the model field in the agent's configuration.

What does Auto model selection do?

Auto routes each request to a frontier model with automatic fallback, combining frontier models with smaller specialized ones, reading intent from your phrasing, and reusing cached work. On the Free plan it routes to Sonnet-class models or better, and on paid plans to Opus-class or better. It never routes to experimental models.

What is reasoning effort in Kiro?

Reasoning effort controls how much thinking a model does per turn. The five levels are low, medium, high, xhigh, and max. Opus 5, Opus 4.8, and Sonnet 5 support all five; Opus 4.7, Opus 4.6, and Sonnet 4.6 have no xhigh; the GPT-5.6 tiers support all levels and default to high. Lower effort is faster, shorter, and consumes fewer credits.

Which models does the Kiro Free plan include?

The Free plan lists Claude Sonnet 4.5 plus the open-weight models. Every Opus, Sonnet 5, the GPT-5.6 tiers, Haiku, and Auto itself need a paid plan, which is why a tutorial that assumes Opus will not match your dropdown on Free.

☰ Chapter summary

  • Every model carries a credit multiplier measured against Auto, which is 1.0x and the default.
  • The spread is enormous: GPT-5.6 Sol is 4.4x and Opus 5 is 2.2x, while Qwen3 Coder Next is 0.05x.
  • Auto is a router: it picks a model per request instead of making you decide every time.
  • Sonnet 5 (1M context, 1.3x) is the sensible daily driver; Opus 5 is for the genuinely hard work.
  • Haiku 4.5 (0.4x) and the open-weight models exist for mechanical, high-volume tasks.
  • The Free tier gets only Sonnet 4.5 plus open-weight models, and the lineup changes monthly.
  • Reasoning effort (low, medium, high, xhigh, max) is a second dial: lower effort is faster, shorter, and cheaper.

All chapter summaries are collected on the revision page.

Was this chapter helpful?