docs: update kilo-auto model routing documentation

- Add claw mode to frontier tier's Opus routing table
- Update balanced tier: routes by API interface type (Completions→Qwen3.6+, Responses→GPT-5.3-Codex, Messages→Claude Haiku 4.5)
- Update free tier: remove hardcoded 80/20 split, describe dynamic server-side routing
This commit is contained in:
kiloconnect[bot]
2026-04-22 21:33:09 +00:00
parent 6804c71a9b
commit b2c9467a88
3 changed files with 16 additions and 19 deletions
@@ -30,7 +30,7 @@ The underlying models behind each Auto Model tier are updated server-side as bet
## Tiers
- **Frontier** — Routes to the latest and most capable paid models. Uses different models for reasoning-heavy tasks (planning, architecture, debugging) versus implementation tasks (coding, building, exploring), pairing the right capability to each type of work.
- **Balanced** — Uses a single cost-effective model across all modes. A good default for most developers who want strong AI assistance without paying frontier prices.
- **Balanced** — Routes to a cost-effective model for all modes. The specific model is selected based on the API interface in use, but does not vary by mode. A good default for most developers who want strong AI assistance without paying frontier prices.
- **Free** — Routes to the best available free models on OpenRouter, splitting traffic across them. Because free model availability shifts over time as providers change promotional periods, the mapping is updated server-side — you always get the best free option without having to track what's currently available. Quality will be lower than paid tiers, and the models may change over time.
## Benefits
@@ -42,7 +42,7 @@ Free models on OpenRouter appear and disappear based on promotional periods. A m
**Who it's for**: Users who want the best available models and are willing to pay for them.
**What it does**: Routes between the best paid models based on the task — stronger reasoning models for planning and architecture, faster models for code generation and editing. Optimizes for the best balance of capability, speed, and token efficiency.
**What it does**: Routes between the best paid models based on the task — stronger reasoning models for planning and architecture, faster models for code generation and editing. This includes the `claw` mode for agentic coding tasks. Optimizes for the best balance of capability, speed, and token efficiency.
**Pricing**: Paid. Uses credits.
@@ -52,7 +52,7 @@ For the current mode-to-model mappings, see the [Auto Model user docs](/docs/cod
**Who it's for**: Cost-conscious developers who want better results than free models at a fraction of frontier cost.
**What it does**: Uses GPT 5.3 Codex — a cost-effective model with strong reasoning and coding capabilities — for every mode. Unlike Frontier, Balanced does not vary its underlying model by mode.
**What it does**: Routes to a cost-effective model based on the API interface used by the client. Requests using the Completions API (default) route to `qwen/qwen3.6-plus`; Responses API requests route to `openai/gpt-5.3-codex`; Messages API requests route to `anthropic/claude-haiku-4.5`. Unlike Frontier, Balanced does not vary its underlying model by mode.
**Pricing**: Paid, but significantly cheaper than Frontier.
@@ -62,7 +62,7 @@ For the current mode-to-model mappings, see the [Auto Model user docs](/docs/cod
**Who it's for**: Users who want to try Kilo without a credit card, students, hobbyists, and anyone exploring AI-assisted coding.
**What it does**: Splits requests across the best available free models, weighted by a deterministic per-session hash so a given session sticks with one model. As free model availability changes due to promotional periods, the split and the underlying models are updated transparently server-side. Users always get the best free option without having to track which models are currently available.
**What it does**: Routes each session to one of the best available free models, selected deterministically based on the session (or user/IP) so a given session sticks with one model. The full candidate pool is determined server-side from curated preferred free models, and updated transparently as availability changes due to promotional periods. Users always get the best free option without having to track which models are currently available.
**Pricing**: Free. No credits required.
@@ -84,28 +84,25 @@ The mappings below reflect the current routing. The underlying models behind eac
Highest performance and capability for any task. Frontier requests are sent with medium reasoning effort and medium verbosity.
| Mode | Resolved Model |
| -------------------------------------------------------------- | ----------------------------- |
| `plan`, `general`, `architect`, `orchestrator`, `ask`, `debug` | `anthropic/claude-opus-4.7` |
| `build`, `explore`, `code` | `anthropic/claude-sonnet-4.6` |
| Default (no / unknown mode) | `anthropic/claude-sonnet-4.6` |
| Mode | Resolved Model |
| ---------------------------------------------------------------------- | ----------------------------- |
| `claw`, `plan`, `general`, `architect`, `orchestrator`, `ask`, `debug` | `anthropic/claude-opus-4.7` |
| `build`, `explore`, `code` | `anthropic/claude-sonnet-4.6` |
| Default (no / unknown mode) | `anthropic/claude-sonnet-4.6` |
### `kilo-auto/balanced`
Great balance of price and capability. Balanced routes to the same model regardless of mode, with low reasoning effort.
Great balance of price and capability. The resolved model depends on the API interface used by the client.
| Mode | Resolved Model |
| --------- | ---------------------- |
| All modes | `openai/gpt-5.3-codex` |
| API interface | Resolved Model | Reasoning effort |
| --------------------- | ---------------------------- | ---------------- |
| Completions (default) | `qwen/qwen3.6-plus` | enabled |
| Responses API | `openai/gpt-5.3-codex` | low |
| Messages API | `anthropic/claude-haiku-4.5` | medium |
### `kilo-auto/free`
Free with limited capability. No credits required. Requests are split across the available free models; the mapping updates server-side as free model availability shifts.
| Routing | Resolved Model |
| ------- | ----------------------------- |
| 80% | `minimax/minimax-m2.5:free` |
| 20% | `stepfun/step-3.5-flash:free` |
Free with limited capability. No credits required. The resolved model is selected dynamically per session from a curated set of available free models; the mapping updates server-side as free model availability shifts.
### `kilo-auto/small`