From 82254696ec351ccaf1ebee2c3595ae5a353f13de Mon Sep 17 00:00:00 2001 From: Joshua Lambert <25085430+lambertjosh@users.noreply.github.com> Date: Thu, 30 Apr 2026 04:02:58 -0400 Subject: [PATCH] docs(kilo-docs): document codebase indexing for CLI and new VS Code extension (#9714) * docs(kilo-docs): document codebase indexing for CLI and new VS Code extension The existing Codebase Indexing page was flagged as legacy-only. Rewrite it to cover the new shared implementation in the CLI and VS Code extension, including the `/indexing` command, the `indexing` config section, LanceDB, and the full list of supported embedding providers. * docs(kilo-docs): address review feedback on codebase indexing - Add a top-level warning that indexing is experimental in the CLI and new VS Code extension - Add a new 'Enabling the feature' section that documents the `experimental.semantic_indexing` flag and where to toggle it (CLI `kilo.jsonc`, VS Code Experimental settings, or legacy direct config) - Split Embedding Providers and Vector Stores into new-platform vs VSCode (Legacy) tabs so the legacy section only lists providers it actually supports (OpenAI, Gemini, Ollama; Qdrant only) - Fix the Vercel AI Gateway link to point at the product docs (https://vercel.com/docs/ai-gateway) instead of the raw API endpoint - Clarify that the Mistral indexing provider uses a La Plateforme API key, not the Codestral key from the autocomplete setup guide - Call out the experimental flag prerequisite at each configuration entry point * docs(kilo-docs): rename combined tab to 'VSCode & CLI' * chore: revert unrelated package-lock changes * docs(kilo-docs): order tabs VSCode, CLI, VSCode (Legacy) * docs(kilo-docs): consolidate platform-specific indexing docs into one tabs block --- packages/kilo-docs/lib/nav/customize.ts | 1 - .../customize/context/codebase-indexing.md | 351 +++++++++++++----- 2 files changed, 256 insertions(+), 96 deletions(-) diff --git a/packages/kilo-docs/lib/nav/customize.ts b/packages/kilo-docs/lib/nav/customize.ts index 7687bad119f..83006b9e1bd 100644 --- a/packages/kilo-docs/lib/nav/customize.ts +++ b/packages/kilo-docs/lib/nav/customize.ts @@ -37,7 +37,6 @@ export const CustomizeNav: NavSection[] = [ { href: "/customize/context/codebase-indexing", children: "Codebase Indexing", - platform: "legacy", }, { href: "/customize/context/context-condensing", diff --git a/packages/kilo-docs/pages/customize/context/codebase-indexing.md b/packages/kilo-docs/pages/customize/context/codebase-indexing.md index 163afeaa00c..f577678670b 100644 --- a/packages/kilo-docs/pages/customize/context/codebase-indexing.md +++ b/packages/kilo-docs/pages/customize/context/codebase-indexing.md @@ -1,14 +1,15 @@ --- title: "Codebase Indexing" description: "Index your codebase for improved AI understanding" -platform: legacy --- # Codebase Indexing Codebase Indexing enables semantic code search across your entire project using AI embeddings. Instead of searching for exact text matches, it understands the _meaning_ of your queries, helping Kilo Code find relevant code even when you don't know specific function names or file locations. -{% image src="/docs/img/codebase-indexing/codebase-indexing.png" alt="Codebase Indexing Settings" width="800" caption="Codebase Indexing Settings" /%} +{% callout type="warning" title="Experimental" %} +Codebase Indexing is currently **experimental** in the CLI and the new VS Code extension. You must explicitly opt in before the feature becomes available — see the **Setup** section below. Behavior, configuration, and defaults may change in future releases. +{% /callout %} ## What It Does @@ -16,7 +17,7 @@ When enabled, the indexing system: 1. **Parses your code** using Tree-sitter to identify semantic blocks (functions, classes, methods) 2. **Creates embeddings** of each code block using AI models -3. **Stores vectors** in a Qdrant database for fast similarity search +3. **Stores vectors** in a vector database for fast similarity search 4. **Provides the [`semantic_search`](/docs/automate/tools/semantic-search) tool** to Kilo Code for intelligent code discovery This enables natural language queries like "user authentication logic" or "database connection handling" to find relevant code across your entire project. @@ -28,43 +29,217 @@ This enables natural language queries like "user authentication logic" or "datab - **Cross-Project Discovery**: Search across all files, not just what's open - **Pattern Recognition**: Locate similar implementations and code patterns -## Setup Requirements +## Setup -### Embedding Provider +{% tabs %} +{% tab label="VSCode" %} -Choose one of these options for generating embeddings: +### 1. Enable the experimental flag -**OpenAI (Recommended)** +Codebase Indexing is gated behind an experimental flag. Until the flag is on, the Indexing UI is hidden and `semantic_search` is unavailable. -- Requires OpenAI API key -- Supports all OpenAI embedding models -- Default: `text-embedding-3-small` -- Processes up to 100,000 tokens per batch +1. Open Kilo Code **Settings** → **Experimental**. +2. Toggle **Semantic Indexing** on. +3. The **Indexing** tab will appear in Settings and the indexing status indicator will appear at the bottom of the prompt input panel. -**Gemini** +Alternatively, set `experimental.semantic_indexing` to `true` in your `kilo.jsonc`: -- Requires Google AI API key -- Supports Gemini embedding models including `gemini-embedding-001` -- Cost-effective alternative to OpenAI -- High-quality embeddings for code understanding +```json +{ + "experimental": { + "semantic_indexing": true + } +} +``` -**Ollama (Local)** +### 2. Configure indexing -- Requires local Ollama installation -- No API costs or internet dependency -- Supports any Ollama-compatible embedding model -- Requires Ollama base URL configuration +1. Open Kilo Code **Settings** → **Indexing**, or click the indexing indicator at the bottom of the prompt input panel. +2. Toggle **Enable Indexing** on. +3. Pick an **Embedding Provider** and fill in its required fields. +4. Pick a **Vector Store** (`Qdrant` or `LanceDB`) and configure it. +5. Optionally adjust **Tuning Parameters** (search score, batch size, retries, max results). +6. Save to start the initial scan. -### Vector Database +You can also edit the `indexing` section in `kilo.jsonc` directly: -**Qdrant** is required for storing and searching embeddings: +```json +{ + "indexing": { + "enabled": true, + "provider": "openai", + "model": "text-embedding-3-small", + "vectorStore": "lancedb", + "openai": { "apiKey": "sk-..." }, + "lancedb": {} + } +} +``` -- **Local**: `http://localhost:6333` (recommended for testing) -- **Cloud**: Qdrant Cloud or self-hosted instance -- **Authentication**: Optional API key for secured deployments +### Embedding providers + +| Provider | How to use | Notes | +|---|---|---| +| **OpenAI** | API key | Default model: `text-embedding-3-small`. `text-embedding-3-large` for higher accuracy. | +| **Ollama** | Local base URL | No API costs. Runs fully offline. | +| **OpenAI-Compatible** | Base URL + API key | For self-hosted or third-party OpenAI-compatible endpoints. | +| **Gemini** | Google AI API key | Supports `gemini-embedding-001` and other Gemini embedding models. | +| **Mistral** | API key from [La Plateforme](https://console.mistral.ai/api-keys/) | Use a standard Mistral API key. The Codestral-specific keys from the [Mistral autocomplete setup guide](/docs/code-with-ai/features/autocomplete/mistral-setup) are **not** interchangeable — those only work for completion. | +| **Vercel AI Gateway** | API key | Routes requests through [Vercel AI Gateway](https://vercel.com/docs/ai-gateway). | +| **AWS Bedrock** | AWS region + profile | Uses the AWS SDK credential chain. | +| **OpenRouter** | API key (optional specific provider) | Routes through [OpenRouter](https://openrouter.ai/). | +| **Voyage** | API key | Voyage `voyage-code-3` is tuned for code. | + +### Vector stores + +- **Qdrant** (default) — external server. Recommended for team deployments and larger codebases. See [Setting Up Qdrant](#setting-up-qdrant). +- **LanceDB** — embedded, file-based. No server to run. Stores data under your Kilo data directory by default. + +{% callout type="tip" %} +For a fully local, zero-cost setup, combine **Ollama** (embeddings) with **LanceDB** (vector store — no separate server needed). +{% /callout %} + +### Status indicator + +The prompt input panel shows a compact indexing status indicator that reflects the current state (Standby / In Progress / Complete / Error) along with progress when scanning or embedding. + +{% /tab %} +{% tab label="CLI" %} + +### 1. Enable the experimental flag + +Codebase Indexing is gated behind an experimental flag. Until the flag is on, the `/indexing` command is hidden and `semantic_search` is unavailable. + +Set the flag in your `kilo.jsonc`: + +```json +{ + "experimental": { + "semantic_indexing": true + } +} +``` + +Restart the CLI for the change to take effect. The `/indexing` command (and aliases `/index`, `/embedding`) will appear in the command palette once the flag is active. + +### 2. Configure indexing + +Open a Kilo TUI session and run: + +```text +/indexing +``` + +(aliases: `/index`, `/embedding`) + +This opens an interactive configuration dialog where you can: + +- **Toggle** indexing on/off +- Choose an **Embedding Provider** and fill in provider settings (API key, base URL, AWS region, etc.) +- Set the **Embedding Model** (blank = provider default) +- Set the **Vector Dimension** (blank = auto-detect from the model) +- Choose a **Vector Store** (`Qdrant` or `LanceDB`) and configure its connection +- Adjust **Tuning Parameters** (search threshold, batch size, retries, max results) + +All changes are written to your `kilo.jsonc` config and take effect immediately. + +You can also edit the `indexing` section directly. This is the full shape of the section: + +```json +{ + "indexing": { + "enabled": true, + "provider": "voyage", + "model": "voyage-code-3", + "dimension": 1024, + "vectorStore": "qdrant", + "voyage": { + "apiKey": "pa-..." + }, + "qdrant": { + "url": "http://localhost:6333", + "apiKey": "" + }, + "searchMinScore": 0.4, + "searchMaxResults": 50, + "embeddingBatchSize": 60, + "scannerMaxBatchRetries": 3 + } +} +``` + +### Embedding providers + +| Provider | Config key | Settings | Notes | +|---|---|---|---| +| **OpenAI** | `openai` | `{ apiKey }` | Default: `text-embedding-3-small`. | +| **Ollama** | `ollama` | `{ baseUrl }` | No API costs. Runs fully offline. | +| **OpenAI-Compatible** | `openai-compatible` | `{ baseUrl, apiKey }` | For self-hosted or third-party endpoints. | +| **Gemini** | `gemini` | `{ apiKey }` | Supports `gemini-embedding-001`. | +| **Mistral** | `mistral` | `{ apiKey }` | Use a [La Plateforme](https://console.mistral.ai/api-keys/) key — the Codestral-specific keys from the [autocomplete setup guide](/docs/code-with-ai/features/autocomplete/mistral-setup) don't work for embeddings. | +| **Vercel AI Gateway** | `vercel-ai-gateway` | `{ apiKey }` | Routes through [Vercel AI Gateway](https://vercel.com/docs/ai-gateway). | +| **AWS Bedrock** | `bedrock` | `{ region, profile }` | Uses AWS SDK credential chain. | +| **OpenRouter** | `openrouter` | `{ apiKey, specificProvider? }` | Routes through [OpenRouter](https://openrouter.ai/). | +| **Voyage** | `voyage` | `{ apiKey }` | `voyage-code-3` is tuned for code. | + +### Vector stores + +- `qdrant` — `{ url?, apiKey? }` (default). See [Setting Up Qdrant](#setting-up-qdrant). +- `lancedb` — `{ directory? }` — embedded, file-based. No server to run. Uses a default Kilo data directory when omitted. + +{% callout type="tip" %} +For a fully local, zero-cost setup, combine **Ollama** (embeddings) with **LanceDB** (vector store — no separate server needed). +{% /callout %} + +### Status indicator + +When indexing is enabled, the CLI shows an indexing status badge at the bottom of the TUI in the form `IDX ` (for example `IDX In Progress 40% 120/300`, `IDX Complete`, `IDX Standby`, or `IDX Error `). + +{% /tab %} +{% tab label="VSCode (Legacy)" %} + +The legacy extension does not require an experimental flag. + +### Open Codebase Indexing Settings + +1. In the chat header, click the database icon (indexing status). +2. The Codebase Indexing settings panel opens. +3. If you don't see the icon, open Kilo Code settings ({% codicon name="gear" /%}) and search for **Codebase Indexing**. + +{% image src="/docs/img/codebase-indexing/codebase-indexing.png" alt="Codebase Indexing Settings" width="800" caption="Codebase Indexing Settings (legacy)" /%} + +### Configure Settings + +1. Enable **"Enable Codebase Indexing"** using the toggle switch. +2. Configure your embedding provider: + - **OpenAI**: Enter API key and select model + - **Gemini**: Enter Google AI API key and select embedding model + - **Ollama**: Enter base URL and select model +3. Set Qdrant URL and optional API key. +4. Configure **Max Search Results** (default: 20, range: 1-100). +5. Click **Save** to start initial indexing. + +### Embedding providers + +The legacy extension supports a smaller set of providers: + +| Provider | How to use | Notes | +|---|---|---| +| **OpenAI** | API key | Default: `text-embedding-3-small`. | +| **Gemini** | Google AI API key | Supports Gemini embedding models including `gemini-embedding-001`. | +| **Ollama (local)** | Local base URL | No API costs. | + +### Vector store + +The legacy extension only supports **Qdrant**. See [Setting Up Qdrant](#setting-up-qdrant). + +{% /tab %} +{% /tabs %} ## Setting Up Qdrant +If you choose **Qdrant** as your vector store, you need a running Qdrant server. + ### Quick Local Setup **Using Docker:** @@ -92,55 +267,27 @@ volumes: For team or production use: -- [Qdrant Cloud](https://cloud.qdrant.io/) - Managed service +- [Qdrant Cloud](https://cloud.qdrant.io/) — managed service - Self-hosted on AWS, GCP, or Azure - Local server with network access for team sharing -## Configuration - -### Open Codebase Indexing Settings - -1. In the chat header, click the database icon (indexing status) -2. The Codebase Indexing settings panel opens -3. If you don't see the icon, open Kilo Code settings () and search for **Codebase Indexing** - -### Configure Settings - -1. Enable **"Enable Codebase Indexing"** using the toggle switch -2. Configure your embedding provider: - - **OpenAI**: Enter API key and select model - - **Gemini**: Enter Google AI API key and select embedding model - - **Ollama**: Enter base URL and select model -3. Set Qdrant URL and optional API key -4. Configure **Max Search Results** (default: 20, range: 1-100) -5. Click **Save** to start initial indexing - -### Enable/Disable Toggle - -The codebase indexing feature includes a convenient toggle switch that allows you to: - -- **Enable**: Start indexing your codebase and make the search tool available -- **Disable**: Stop indexing, pause file watching, and disable the search functionality -- **Preserve Settings**: Your configuration remains saved when toggling off - -This toggle is useful for temporarily disabling indexing during intensive development work or when working with sensitive codebases. - ## Understanding Index Status -The interface shows real-time status with color indicators: +The interface shows real-time status: -- **Standby** (Gray): Not running, awaiting configuration -- **Indexing** (Yellow): Currently processing files -- **Indexed** (Green): Up-to-date and ready for searches -- **Error** (Red): Failed state requiring attention +- **Standby**: Not running, awaiting configuration or paused +- **In Progress**: Currently processing files (with a progress percentage and `processed/total` count) +- **Complete**: Up-to-date and ready for searches +- **Error**: Failed state, with an error message +- **Disabled**: Indexing is turned off or not yet configured ## How Files Are Processed ### Smart Code Parsing - **Tree-sitter Integration**: Uses AST parsing to identify semantic code blocks -- **Language Support**: All languages supported by Tree-sitter -- **Markdown Support**: Full support for markdown files and documentation +- **Language Support**: Broad language coverage via Tree-sitter — C, C#, C++, CSS, Elisp, Elixir, Go, HTML, Java, JavaScript, Kotlin, Lua, OCaml, PHP, Python, Ruby, Rust, Scala, Solidity, Swift, SystemRDL, TLA+, TOML, TSX, TypeScript, Vue, Zig, and more +- **Markdown Support**: Dedicated parser for markdown and documentation - **Fallback**: Line-based chunking for unsupported file types - **Block Sizing**: - Minimum: 100 characters @@ -159,40 +306,56 @@ The indexer automatically excludes: ### Incremental Updates -- **File Watching**: Monitors workspace for changes +- **File Watching**: Monitors the workspace for changes and re-indexes in the background - **Smart Updates**: Only reprocesses modified files - **Hash-based Caching**: Avoids reprocessing unchanged content - **Branch Switching**: Automatically handles Git branch changes +## Tuning Parameters + +These advanced settings live under the `indexing` key and are exposed in the CLI's `/indexing → Tuning Parameters` menu and the VS Code extension's Indexing settings: + +| Setting | Default | Description | +|---|---|---| +| `searchMinScore` | `0.4` | Minimum cosine similarity (0-1) for a result to be returned. | +| `searchMaxResults` | `50` | Maximum number of results returned per search. | +| `embeddingBatchSize` | `60` | Number of code segments per embedding batch. Lower this if your embedding endpoint has strict rate limits. | +| `scannerMaxBatchRetries` | `3` | Maximum retry attempts for a failed embedding batch. | + ## Best Practices ### Model Selection -**For OpenAI:** +**OpenAI:** - **`text-embedding-3-small`**: Best balance of performance and cost - **`text-embedding-3-large`**: Higher accuracy, 5x more expensive - **`text-embedding-ada-002`**: Legacy model, lower cost -**For Ollama:** +**Ollama:** -- **`mxbai-embed-large`**: The largest and highest-quality embedding model. -- **`nomic-embed-text`**: Best balance of performance and embedding quality. -- **`all-minilm`**: Compact model with lower quality but faster performance. +- **`mxbai-embed-large`**: The largest and highest-quality embedding model +- **`nomic-embed-text`**: Best balance of performance and embedding quality +- **`all-minilm`**: Compact model with lower quality but faster performance + +**Voyage:** + +- **`voyage-code-3`**: Code-tuned embeddings; strong default for source-heavy repos ### Security Considerations -- **API Keys**: Stored securely in VS Code's encrypted storage -- **Code Privacy**: Only small code snippets sent for embedding (not full files) -- **Local Processing**: All parsing happens locally -- **Qdrant Security**: Use authentication for production deployments +- **API Keys**: Stored in your `kilo.jsonc` config. Treat that file as a secret in shared environments. +- **Code Privacy**: Only small code snippets are sent for embedding — never whole files. +- **Local Processing**: All parsing (Tree-sitter) happens locally. +- **Fully Local Option**: Pair **Ollama** (embeddings) with **LanceDB** (local vector store) for a setup that never leaves your machine. +- **Qdrant Security**: Use authentication for production deployments. ## Current Limitations - **File Size**: 1MB maximum per file - **Single Workspace**: One workspace at a time -- **Dependencies**: Requires external services (embedding provider + Qdrant) -- **Language Coverage**: Limited to Tree-sitter supported languages for optimal parsing +- **Dependencies**: Requires an embedding provider, and — for Qdrant — a running Qdrant instance +- **Language Coverage**: Optimal parsing is limited to Tree-sitter supported languages ## Troubleshooting @@ -200,6 +363,16 @@ The indexer automatically excludes: If your local embedding server is based on llama.cpp (including Ollama), indexing can fail with errors about `n_ubatch` or `GGML_ASSERT`. Ensure both batch size (`-b`) and micro-batch size (`-ub`) are set to the same value for embedding models, then restart the server. For Ollama, configure `num_batch` in your Modelfile or request options to match the same effective value. +### Indexing status stays on "Disabled" + +- Check that `indexing.enabled` is `true` in your `kilo.jsonc` +- Verify that the selected provider has all required credentials set +- If using Qdrant, make sure the Qdrant server is reachable at the configured URL + +### Rate-limit or batch errors with a hosted provider + +Lower `embeddingBatchSize` under `indexing` (default `60`). Smaller batches send fewer segments per request and are less likely to hit per-request or per-minute rate limits. + ## Using the Search Feature Once indexed, Kilo Code can use the [`semantic_search`](/docs/automate/tools/semantic-search) tool to find relevant code: @@ -213,34 +386,22 @@ Once indexed, Kilo Code can use the [`semantic_search`](/docs/automate/tools/sem The tool provides Kilo Code with: -- Relevant code snippets (up to your configured max results limit) +- Relevant code snippets (up to your configured `searchMaxResults`) - File paths and line numbers - Similarity scores - Contextual information ### Search Results Configuration -You can control the number of search results returned by adjusting the **Max Search Results** setting: +Tune result volume and quality via: -- **Default**: 20 results -- **Range**: 1-100 results -- **Performance**: Lower values improve response speed -- **Comprehensiveness**: Higher values provide more context but may slow responses +- **`searchMaxResults`** — default `50`. Lower for faster, more focused responses; higher for more context. +- **`searchMinScore`** — default `0.4`. Raise to require closer matches; lower to include more tangentially related code. ## Privacy & Security -- **Code stays local**: Only small code snippets sent for embedding +- **Code stays local**: Only small code snippets are sent for embedding - **Embeddings are numeric**: Not human-readable representations -- **Secure storage**: API keys encrypted in VS Code storage -- **Local option**: Use Ollama for completely local processing -- **Access control**: Respects existing file permissions - -## Future Enhancements - -Planned improvements: - -- Additional embedding providers -- Multi-workspace indexing -- Enhanced filtering and configuration options -- Team sharing capabilities -- Integration with VS Code's native search +- **Secure storage**: API keys are stored in your local `kilo.jsonc` configuration +- **Fully local option**: Use **Ollama + LanceDB** for completely local processing +- **Access control**: Respects existing file permissions and `.kilocodeignore`