From 89baf26dd4d28df30ce030dcd238242dbdddb76f Mon Sep 17 00:00:00 2001 From: Josh Lambert Date: Mon, 23 Feb 2026 09:00:40 -0500 Subject: [PATCH 1/9] First draft --- packages/kilo-docs/lib/nav/contributing.ts | 4 +++ .../contributing/architecture/features.md | 25 ++++++++++--------- 2 files changed, 17 insertions(+), 12 deletions(-) diff --git a/packages/kilo-docs/lib/nav/contributing.ts b/packages/kilo-docs/lib/nav/contributing.ts index 0ef87db4d37..0381811b513 100644 --- a/packages/kilo-docs/lib/nav/contributing.ts +++ b/packages/kilo-docs/lib/nav/contributing.ts @@ -34,6 +34,10 @@ export const ContributingNav: NavSection[] = [ href: "/contributing/architecture/agent-observability", children: "Agent Observability", }, + { + href: "/contributing/architecture/auto-model-tiers", + children: "Auto Model Tiers", + }, { href: "/contributing/architecture/benchmarking", children: "Benchmarking", diff --git a/packages/kilo-docs/pages/contributing/architecture/features.md b/packages/kilo-docs/pages/contributing/architecture/features.md index 667fcf8ffbd..d73518f5596 100644 --- a/packages/kilo-docs/pages/contributing/architecture/features.md +++ b/packages/kilo-docs/pages/contributing/architecture/features.md @@ -7,17 +7,18 @@ description: "Overview of current and planned features in Kilo Code" These pages document the architecture and design of current or planned features, as well as any unique development patterns. -| Feature | Description | -| ---------------------------------------------------------------------------------------- | ------------------------------------------------ | -| [Agent Observability](/docs/contributing/architecture/agent-observability) | Observability and monitoring for agentic systems | -| [Benchmarking](/docs/contributing/architecture/benchmarking) | Benchmarking Kilo Code across models and agents | -| [Enterprise MCP Controls](/docs/contributing/architecture/enterprise-mcp-controls) | Admin controls for MCP server allowlists | -| [MCP OAuth Authorization](/docs/contributing/architecture/mcp-oauth-authorization) | OAuth 2.1-based authorization for MCP servers | -| [Onboarding Improvements](/docs/contributing/architecture/onboarding-improvements) | User onboarding and engagement features | -| [Organization Modes Library](/docs/contributing/architecture/organization-modes-library) | Shared modes for teams and enterprise | -| [Agentic Security Reviews](/docs/deploy-secure/security-reviews) | AI-powered security vulnerability analysis | -| [Track Repo URL](/docs/contributing/architecture/track-repo-url) | Usage tracking by repository/project | -| [Vercel AI Gateway](/docs/contributing/architecture/vercel-ai-gateway) | Vercel AI Gateway integration | -| [Voice Transcription](/docs/contributing/architecture/voice-transcription) | Live voice input for chat | +| Feature | Description | +| ---------------------------------------------------------------------------------------- | ---------------------------------------------------- | +| [Agent Observability](/docs/contributing/architecture/agent-observability) | Observability and monitoring for agentic systems | +| [Auto Model Tiers](/docs/contributing/architecture/auto-model-tiers) | Multi-tier auto model routing (Frontier, Free, Open) | +| [Benchmarking](/docs/contributing/architecture/benchmarking) | Benchmarking Kilo Code across models and agents | +| [Enterprise MCP Controls](/docs/contributing/architecture/enterprise-mcp-controls) | Admin controls for MCP server allowlists | +| [MCP OAuth Authorization](/docs/contributing/architecture/mcp-oauth-authorization) | OAuth 2.1-based authorization for MCP servers | +| [Onboarding Improvements](/docs/contributing/architecture/onboarding-improvements) | User onboarding and engagement features | +| [Organization Modes Library](/docs/contributing/architecture/organization-modes-library) | Shared modes for teams and enterprise | +| [Agentic Security Reviews](/docs/deploy-secure/security-reviews) | AI-powered security vulnerability analysis | +| [Track Repo URL](/docs/contributing/architecture/track-repo-url) | Usage tracking by repository/project | +| [Vercel AI Gateway](/docs/contributing/architecture/vercel-ai-gateway) | Vercel AI Gateway integration | +| [Voice Transcription](/docs/contributing/architecture/voice-transcription) | Live voice input for chat | To propose a new feature design, consider using the [Spec Template](/docs/contributing/architecture/feature-template). From 0f557832f4bab715c3acf60f50be14e1086422ac Mon Sep 17 00:00:00 2001 From: Josh Lambert Date: Mon, 23 Feb 2026 16:16:20 -0500 Subject: [PATCH 2/9] Actually add the proposal --- .../architecture/auto-model-tiers.md | 158 ++++++++++++++++++ 1 file changed, 158 insertions(+) create mode 100644 packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md new file mode 100644 index 00000000000..6ebf79bb891 --- /dev/null +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -0,0 +1,158 @@ +--- +title: "Auto Model Tiers" +description: "Extending Kilo Auto to a family of smart model tiers that match users to the right models without requiring AI expertise" +--- + +# Auto Model Tiers + +## Overview + +Today, Kilo Auto (`kilo/auto`) solves a real problem: users shouldn't have to know which AI model is best for planning versus coding. It picks the right model for the task automatically. + +But it only solves that problem for one audience — users willing to pay for frontier models. Everyone else is left navigating a long, intimidating model list where the "best" free or open-weight option changes monthly. This is a siyrce of friction for new free users and cost-conscious teams. + +This spec proposes extending Kilo Auto into a family of tiers so that every user — regardless of budget, preference, or expertise — gets the same "just works" auto experience. + +## Problem + +### Users shouldn't need to be AI model experts + +The AI model landscape is overwhelming. There are hundreds of models across dozens of providers, with different pricing, capabilities, context windows, and availability. Most developers just want to write code — they don't want to research which model is best for their task, budget, and workflow. + +Today, Kilo Auto handles this for users with a willingness to pay for frontier models. But we leave three underserved groups to fend for themselves: + +1. **Free users** — They see a list of free models that changes on promotional periods and shifting models. Which one is the best? Which is good for a particular task? They have no way to know without trial and error. + +2. **Cost-conscious users** — They want something better than free but cheaper than frontier. Open-weight models like Kimi K2.5 and Deepseek are useful and significantly cheaper, but which one? Which version? The answer changes every few weeks. + +3. **Background tasks** — Kilo uses small models for things like generating session titles and commit messages. Today this is hardcoded to a specific model. If a user runs out of credits, these background tasks fall back to the current large model. This can cause Kilo unnecessary costs and slow responsiveness for the user. + +### Free model churn creates a moving target + +Free models on OpenRouter appear and disappear based on promotional periods. A model that works well today may be gone next week. Users who manually selected a free model discover it's unavailable. Kilo Auto tiers would absorb this churn — when the best free model changes, we update the mapping and users can just keep working. + +## Solution + +Extend Kilo Auto into four tiers. + +### Auto: Frontier + +**Who it's for**: Users who want the best available models and are willing to pay for them. + +**What problem it solves**: Professional developers and teams don't want to spend time evaluating which paid model is best for which task — and the answer keeps changing as providers release new versions. Frontier eliminates that research burden. Users get the best paid models automatically, matched to their task, without ever having to compare benchmarks or read release notes. + +**What it does**: This is what `kilo/auto` is today. It routes between the best paid models based on the task — stronger reasoning models for planning and architecture, faster models for code generation and editing. It optimizes for the best balance of capability, speed, and token efficiency. + +**Why it matters**: Frontier models offer the highest code quality, best instruction following, and most reliable tool use. For professional developers and teams where output quality directly impacts productivity, this is the right choice. + +**Pricing**: Paid. Uses credits. + +**Backward compatibility**: The existing `kilo/auto` model ID becomes an alias for `kilo/auto-frontier`. No behavior change for existing users. + +### Auto: Free + +**Who it's for**: Users who want to try Kilo without a credit card, students, hobbyists, and anyone exploring AI-assisted coding. + +**What problem it solves**: Free users today face a confusing model list where the "best" option is a moving target — models appear and disappear based on promotional periods, quality varies, and there's no guidance on which one to use for which mode. Most users don't have the context to make this choice, and may make a poor choice or feel intimidated. Auto: Free makes this decision for them so the onboarding experience is just "start coding." + +**What it does**: Automatically maps to the best available free model(s) for each mode. As free model availability changes due to promotional periods, the mapping updates transparently. Users always get the best free option without having to track which models are currently available. + +**Why it matters**: This solves removes a potentially intimidating choice for free users, and sets them up for a good experience. "Auto: Free" is selected by default (or it's selected for them by default when unauthenticated) and start coding immediately. When a free model promotion ends, they don't get an prompt to pick a new model — the routing silently falls back to the next best option. + +**Pricing**: Free. No credits required. + +**Constraints**: Free models may not provide sufficient breadth to justify different models based on modes. In that case, a single model may be used for all modes. Quality will be lower than Frontier or Open tiers — this is a tradeoff users accept by choosing free. + +### Auto: Open + +**Who it's for**: Cost-conscious developers who want better results than free models, users who prefer open-source/open-weight models for philosophical or compliance reasons, and teams that want model transparency. + +**What problem it solves**: The open-weight model landscape moves faster than any other segment — new releases from DeepSeek, Minimax, Moonshot AI, Z.AI, Qwen, and others land every few weeks, leapfrogging each other on benchmarks and real-world coding tasks. Users who want to use open-weight models face frequent research overhead and real-world time to stay on the best one. Auto: Open solves this for them, always routing to the current best open-weight option so they get the cost and transparency benefits without the cognitive burden. + +**What it does**: Routes to the best open-weight models (DeepSeek, Moonshot, Minimax etc.) for each task type. May use a blend of paid and free models depending on what's available and the current state of the art. Like Frontier, it can switch models based on task type when the available have sufficient difference in capabilities and cost. + +**Why it matters**: Open-weight models have gotten remarkably capable and are significantly cheaper than frontier closed models. Auto: Open absorbs model churn and always routes to the current best option. It also appeals to users and organizations that value model transparency and auditability. + +**Pricing**: Generally cheaper than Frontier. May use a mix of paid and free models depending on availability. + +### Auto: Small (internal) + +**Who it's for**: Not user-facing. Used internally by Kilo cleints for lightweight background tasks. + +**What problem it solves**: Kilo uses small models behind the scenes for tasks like generating session titles, commit messages, and conversation summaries. Today, the small model is configurable by the end user by defaults to a paid frontier model. If a free user doesn't have access to that model, the task falls back to the user's main (large) model — wasting credits and adding latency for a task that should be instant and cheap. Auto: Small makes background tasks reliably work for all users regardless of their payment status. + +**What it does**: Automatically selects the right small model for lightweight tasks. When credits are available, it uses a fast paid small model (Haiku, GPT Nano, etc.). When no credits are available, it falls back to a capable free small model. Since we can use a free model for this, we can support it for all users and providers. + +**Why it matters**: Users never think about background tasks, and they shouldn't have to. Auto: Small ensures these tasks always work, always feel fast, and never waste credits on an expensive model when a cheap one will do. + +## User experience + +### Model picker + +The three user-facing tiers appear in the model selector: + +| Display Name | Description shown to user | +| -------------- | ---------------------------------------------------- | +| Auto: Frontier | Best paid models, automatically matched to your task | +| Auto: Free | Best free models, no credits required | +| Auto: Open | Open-weight models, strong capability at lower cost | + +Auto: Small does not appear in the model picker. It is selected automatically for background tasks. + +### Defaults + +- **Authenticated users** (have credits): Default to Auto: Frontier +- **Unauthenticated users** (no credits): Default to Auto: Free + +This means a brand-new user who hasn't signed in gets a working experience immediately — no model selection required. + +### What users see + +The UI shows the tier name (e.g., "Auto: Frontier"), not the underlying model. Users don't need to know or care that their planning request went to Opus and their coding request went to Sonnet. The abstraction is the product. + +## Requirements + +- `kilo/auto` remains as a backward-compatible alias for `kilo/auto-frontier` +- Unauthenticated users default to `kilo/auto-free` with no configuration required +- Free tier must handle model availability changes gracefully — fallback to next-best free model, never surface a "model unavailable" error if any free model exists +- Open tier must use open-weight models as the primary routing targets +- Auto: Small must detect credit availability and select paid or free small models accordingly +- All tiers use mode-based routing where the underlying models support it +- When a tier routes to different model families across turns in a conversation, thinking/reasoning blocks from the previous model must be stripped to prevent compatibility errors + +### Non-requirements + +- Per-agent tier overrides (e.g., Frontier for code, Free for explore) — future work +- Showing the resolved underlying model name in the UI — future work +- User-configurable tier preferences — future work +- Custom user-defined tiers — out of scope + +## Risks + +| Risk | User impact | Mitigation | +| -------------------------------------------------------- | ---------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | +| Free model disappears mid-session | User's next message fails | Fallback chain: primary → secondary → tertiary free model. Graceful error only if all options exhausted. | +| Model quality variance across free/open tiers | Inconsistent experience compared to Frontier | Set clear expectations in UI ("Free" and "Open" imply tradeoffs). Curate model lists, don't just pick the cheapest. | +| Cross-family model switching breaks conversation context | Thinking blocks from Model A are incompatible with Model B | Strip thinking blocks when the underlying model family changes between turns. Frontier stays within one family so this only affects Free and Open. | +| Users don't understand the tier differences | Wrong tier selected, poor experience | Clear descriptions in the model picker. Good defaults (Frontier for paid, Free for unpaid) so most users never need to actively choose. | + +## Data and compliance + +- **Frontier**: Same compliance posture as today. Uses Anthropic models with no training on user data. +- **Free and Open**: The underlying models may have different data handling policies depending on the provider. This must be documented per-tier so enterprise users can make informed choices. +- **Small**: Same concern as Free/Open — the model selected depends on credit status, which may route to providers with different policies. + +## Success criteria + +- New unauthenticated users can start coding without selecting a model (Free tier auto-selected) +- Free model churn is invisible to users — no "model unavailable" errors when alternatives exist +- Conversion from Free to Frontier increases as users experience the product and want better quality +- Background tasks (titles, summaries) never fail due to model availability regardless of credit status + +## Features for the future + +- **Resolved model transparency**: Show the actual model being used on hover/click for users who want to know +- **Per-agent tier overrides**: Let users pick Frontier for their code agent but Free for explore +- **Auto model changelog**: A status page or in-product notification when tier mappings change +- **Tier analytics**: Dashboard showing which models each tier resolves to, latency, error rates, quality metrics +- **Enterprise open-weight preference**: Organizations that require open-weight models for auditability could enforce the Open tier across their team From 41fc9dd52ab411c76c07458b383f59a755d1aa40 Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Mon, 23 Feb 2026 17:22:34 -0500 Subject: [PATCH 3/9] Update packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com> --- .../pages/contributing/architecture/auto-model-tiers.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index 6ebf79bb891..f38e4e3c881 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -9,7 +9,7 @@ description: "Extending Kilo Auto to a family of smart model tiers that match us Today, Kilo Auto (`kilo/auto`) solves a real problem: users shouldn't have to know which AI model is best for planning versus coding. It picks the right model for the task automatically. -But it only solves that problem for one audience — users willing to pay for frontier models. Everyone else is left navigating a long, intimidating model list where the "best" free or open-weight option changes monthly. This is a siyrce of friction for new free users and cost-conscious teams. +But it only solves that problem for one audience — users willing to pay for frontier models. Everyone else is left navigating a long, intimidating model list where the "best" free or open-weight option changes monthly. This is a source of friction for new free users and cost-conscious teams. This spec proposes extending Kilo Auto into a family of tiers so that every user — regardless of budget, preference, or expertise — gets the same "just works" auto experience. From 62ce52d8d2642575310726ec425acbcd2fcc84dc Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Mon, 23 Feb 2026 17:22:44 -0500 Subject: [PATCH 4/9] Update packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com> --- .../pages/contributing/architecture/auto-model-tiers.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index f38e4e3c881..215d053509d 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -57,7 +57,7 @@ Extend Kilo Auto into four tiers. **What it does**: Automatically maps to the best available free model(s) for each mode. As free model availability changes due to promotional periods, the mapping updates transparently. Users always get the best free option without having to track which models are currently available. -**Why it matters**: This solves removes a potentially intimidating choice for free users, and sets them up for a good experience. "Auto: Free" is selected by default (or it's selected for them by default when unauthenticated) and start coding immediately. When a free model promotion ends, they don't get an prompt to pick a new model — the routing silently falls back to the next best option. +**Why it matters**: This removes a potentially intimidating choice for free users, and sets them up for a good experience. "Auto: Free" is selected by default (or it's selected for them by default when unauthenticated) and start coding immediately. When a free model promotion ends, they don't get a prompt to pick a new model — the routing silently falls back to the next best option. **Pricing**: Free. No credits required. From a81ca55386571430e93c7bc8fdbb0057f580f745 Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Mon, 23 Feb 2026 17:22:53 -0500 Subject: [PATCH 5/9] Update packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com> --- .../pages/contributing/architecture/auto-model-tiers.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index 215d053509d..456de2cba6d 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -69,7 +69,7 @@ Extend Kilo Auto into four tiers. **What problem it solves**: The open-weight model landscape moves faster than any other segment — new releases from DeepSeek, Minimax, Moonshot AI, Z.AI, Qwen, and others land every few weeks, leapfrogging each other on benchmarks and real-world coding tasks. Users who want to use open-weight models face frequent research overhead and real-world time to stay on the best one. Auto: Open solves this for them, always routing to the current best open-weight option so they get the cost and transparency benefits without the cognitive burden. -**What it does**: Routes to the best open-weight models (DeepSeek, Moonshot, Minimax etc.) for each task type. May use a blend of paid and free models depending on what's available and the current state of the art. Like Frontier, it can switch models based on task type when the available have sufficient difference in capabilities and cost. +**What it does**: Routes to the best open-weight models (DeepSeek, Moonshot, Minimax etc.) for each task type. May use a blend of paid and free models depending on what's available and the current state of the art. Like Frontier, it can switch models based on task type when the available models have sufficient difference in capabilities and cost. **Why it matters**: Open-weight models have gotten remarkably capable and are significantly cheaper than frontier closed models. Auto: Open absorbs model churn and always routes to the current best option. It also appeals to users and organizations that value model transparency and auditability. From 550eb319636a510c78f928cbd14ec7bdc5c27c11 Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Mon, 23 Feb 2026 17:22:58 -0500 Subject: [PATCH 6/9] Update packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com> --- .../pages/contributing/architecture/auto-model-tiers.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index 456de2cba6d..82d27cb7be9 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -77,7 +77,7 @@ Extend Kilo Auto into four tiers. ### Auto: Small (internal) -**Who it's for**: Not user-facing. Used internally by Kilo cleints for lightweight background tasks. +**Who it's for**: Not user-facing. Used internally by Kilo clients for lightweight background tasks. **What problem it solves**: Kilo uses small models behind the scenes for tasks like generating session titles, commit messages, and conversation summaries. Today, the small model is configurable by the end user by defaults to a paid frontier model. If a free user doesn't have access to that model, the task falls back to the user's main (large) model — wasting credits and adding latency for a task that should be instant and cheap. Auto: Small makes background tasks reliably work for all users regardless of their payment status. From 1b03c3de80489f78b5bfabc6c11aaf13df65b24d Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Mon, 23 Feb 2026 17:23:06 -0500 Subject: [PATCH 7/9] Update packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com> --- .../pages/contributing/architecture/auto-model-tiers.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index 82d27cb7be9..6ebc02b3169 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -79,7 +79,7 @@ Extend Kilo Auto into four tiers. **Who it's for**: Not user-facing. Used internally by Kilo clients for lightweight background tasks. -**What problem it solves**: Kilo uses small models behind the scenes for tasks like generating session titles, commit messages, and conversation summaries. Today, the small model is configurable by the end user by defaults to a paid frontier model. If a free user doesn't have access to that model, the task falls back to the user's main (large) model — wasting credits and adding latency for a task that should be instant and cheap. Auto: Small makes background tasks reliably work for all users regardless of their payment status. +**What problem it solves**: Kilo uses small models behind the scenes for tasks like generating session titles, commit messages, and conversation summaries. Today, the small model is configurable by the end user but defaults to a paid frontier model. If a free user doesn't have access to that model, the task falls back to the user's main (large) model — wasting credits and adding latency for a task that should be instant and cheap. Auto: Small makes background tasks reliably work for all users regardless of their payment status. **What it does**: Automatically selects the right small model for lightweight tasks. When credits are available, it uses a fast paid small model (Haiku, GPT Nano, etc.). When no credits are available, it falls back to a capable free small model. Since we can use a free model for this, we can support it for all users and providers. From 17ec96a1107a45b4457735a5e55c222932c1321b Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Mon, 23 Feb 2026 19:20:40 -0500 Subject: [PATCH 8/9] Apply suggestions from code review Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com> --- .../pages/contributing/architecture/auto-model-tiers.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index 6ebc02b3169..e174f3cea90 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -23,7 +23,7 @@ Today, Kilo Auto handles this for users with a willingness to pay for frontier m 1. **Free users** — They see a list of free models that changes on promotional periods and shifting models. Which one is the best? Which is good for a particular task? They have no way to know without trial and error. -2. **Cost-conscious users** — They want something better than free but cheaper than frontier. Open-weight models like Kimi K2.5 and Deepseek are useful and significantly cheaper, but which one? Which version? The answer changes every few weeks. +2. **Cost-conscious users** — They want something better than free but cheaper than frontier. Open-weight models like Kimi K2.5 and DeepSeek are useful and significantly cheaper, but which one? Which version? The answer changes every few weeks. 3. **Background tasks** — Kilo uses small models for things like generating session titles and commit messages. Today this is hardcoded to a specific model. If a user runs out of credits, these background tasks fall back to the current large model. This can cause Kilo unnecessary costs and slow responsiveness for the user. @@ -57,7 +57,7 @@ Extend Kilo Auto into four tiers. **What it does**: Automatically maps to the best available free model(s) for each mode. As free model availability changes due to promotional periods, the mapping updates transparently. Users always get the best free option without having to track which models are currently available. -**Why it matters**: This removes a potentially intimidating choice for free users, and sets them up for a good experience. "Auto: Free" is selected by default (or it's selected for them by default when unauthenticated) and start coding immediately. When a free model promotion ends, they don't get a prompt to pick a new model — the routing silently falls back to the next best option. +**Why it matters**: This removes a potentially intimidating choice for free users, and sets them up for a good experience. "Auto: Free" is selected by default (or it's selected for them by default when unauthenticated) and they start coding immediately. When a free model promotion ends, they don't get a prompt to pick a new model — the routing silently falls back to the next best option. **Pricing**: Free. No credits required. From 177a93e2f0968054f3b6706ebd190a0b18c6f7cc Mon Sep 17 00:00:00 2001 From: "kiloconnect[bot]" <240665456+kiloconnect[bot]@users.noreply.github.com> Date: Tue, 24 Feb 2026 13:38:40 +0000 Subject: [PATCH 9/9] Add gpt-oss-20b and GPT-5 Nano as options for auto:small tier --- .../pages/contributing/architecture/auto-model-tiers.md | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md index e174f3cea90..578fbdd4f63 100644 --- a/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md +++ b/packages/kilo-docs/pages/contributing/architecture/auto-model-tiers.md @@ -83,6 +83,15 @@ Extend Kilo Auto into four tiers. **What it does**: Automatically selects the right small model for lightweight tasks. When credits are available, it uses a fast paid small model (Haiku, GPT Nano, etc.). When no credits are available, it falls back to a capable free small model. Since we can use a free model for this, we can support it for all users and providers. +**Model options for Auto: Small**: + +| Model | Cost | Capability | Notes | +| ----- | ---- | ---------- | ----- | +| gpt-oss-20b | ~50% cheaper than GPT-5 Nano | Lower — suitable for simple background tasks like titles and summaries | Open-weight, cost-optimized option for high-volume lightweight tasks | +| GPT-5 Nano | Higher than gpt-oss-20b | Higher — better instruction following and output quality | Preferred when credits are available and task quality matters | + +Auto: Small should prefer **gpt-oss-20b** when minimizing cost (e.g., free users, high-volume background tasks) and **GPT-5 Nano** when credits are available and higher output quality is desired. + **Why it matters**: Users never think about background tasks, and they shouldn't have to. Auto: Small ensures these tasks always work, always feel fast, and never waste credits on an expensive model when a cheap one will do. ## User experience