mirror of
https://github.com/coder/coder.git
synced 2026-09-24 15:04:27 +08:00
docs: add Models page and restructure agents docs into directory (#22643)
Adds a Models page documenting LLM provider and model configuration for Coder Agents. Moves agents pages into `docs/ai-coder/agents/` directory. URLs are unchanged. <img width="1343" height="633" alt="image" src="https://github.com/user-attachments/assets/e870340b-9ae5-4904-9936-49f51ab0e0c4" />
This commit is contained in:
@@ -0,0 +1,255 @@
|
||||
# Architecture
|
||||
|
||||
Coder's AI agent interacts with workspaces over the same
|
||||
connection path as a developer's IDE, web terminal, and SSH session already
|
||||
use. There is no sidecar process and no new network paths. If your developers
|
||||
can already connect to their workspaces, the agent can too.
|
||||
|
||||
## Architecture at a glance
|
||||
|
||||
Three components are involved in every agent interaction:
|
||||
|
||||
1. **The control plane** runs the agent loop. It receives prompts, streams them
|
||||
to the LLM provider, interprets tool calls, and dispatches them to
|
||||
workspaces.
|
||||
1. **The LLM provider** (Anthropic, OpenAI, Google, Azure, AWS Bedrock, or any
|
||||
OpenAI-compatible endpoint) performs model inference. It never communicates
|
||||
with the workspace directly.
|
||||
1. **The workspace** is standard compute infrastructure. It runs shell commands,
|
||||
reads and writes files, and executes processes — exactly what occurs when a
|
||||
developer connects via their IDE.
|
||||
|
||||

|
||||
|
||||
## The same connection your IDE uses
|
||||
|
||||
This is the key architectural insight: the agent reaches into a workspace
|
||||
over the same Tailnet tunnel that a developer's tools already use.
|
||||
|
||||
When a developer opens a web terminal in the Coder dashboard, connects via
|
||||
VS Code Remote, or runs `coder ssh`, the traffic follows this path:
|
||||
|
||||
1. The client connects to the control plane.
|
||||
1. The control plane routes the connection through its internal Tailnet node.
|
||||
1. The connection reaches the workspace daemon over a DERP relay or
|
||||
direct peer-to-peer link.
|
||||
1. The workspace daemon handles the request — spawning a shell,
|
||||
forwarding a port, or serving a file.
|
||||
|
||||
When the agent executes a tool call — reading a file, running a command,
|
||||
writing code — it follows the same tunnel:
|
||||
|
||||
1. The agent loop in the control plane issues a tool call.
|
||||
1. The control plane routes the call through its internal Tailnet node.
|
||||
1. The call reaches the workspace daemon over the same DERP relay or
|
||||
peer-to-peer link.
|
||||
1. The workspace daemon handles the request via its HTTP API — reading a file,
|
||||
starting a process, or writing content.
|
||||
|
||||
The underlying tunnel is identical. IDE connections use SSH, web terminals use
|
||||
a WebSocket protocol, and the agent uses the workspace daemon's HTTP API — but
|
||||
all three traverse the same Tailnet connection and rely on the same security
|
||||
boundary. No additional ports or network paths are introduced.
|
||||
|
||||
### No inbound ports
|
||||
|
||||
The workspace daemon always dials _out_ to the control plane — never
|
||||
the reverse. The control plane then uses that established tunnel to reach back
|
||||
in. This means:
|
||||
|
||||
- The workspace needs no inbound ports or exposed services.
|
||||
- You can block all inbound traffic to the workspace.
|
||||
- The only required outbound connection from the workspace is to the control
|
||||
plane itself.
|
||||
|
||||
This is unchanged from how workspaces already operate in Coder. Enabling
|
||||
Coder Agents does not change your workspace network requirements.
|
||||
|
||||
## The agent loop
|
||||
|
||||
When a user submits a prompt, the control plane processes it as a background
|
||||
job:
|
||||
|
||||
1. The prompt is saved to the database and the chat is marked `pending`.
|
||||
1. The control plane picks up the chat and marks it `running`.
|
||||
1. The control plane streams the conversation to the configured LLM provider.
|
||||
1. The model responds with text, reasoning, or tool calls.
|
||||
1. If the response includes tool calls, the control plane executes them
|
||||
(connecting to the workspace as needed) and returns the results to the model.
|
||||
1. Steps 3–5 repeat until the model produces a final response with no further
|
||||
tool calls.
|
||||
1. The chat is marked `waiting` for the next user message.
|
||||
|
||||
This loop runs inside the control plane process. There is no separate service
|
||||
to deploy — it is part of the same binary that serves the dashboard and API.
|
||||
|
||||
### Context compaction
|
||||
|
||||
As conversations grow, the agent automatically summarizes older context to stay
|
||||
within the model's context window. When token usage exceeds a threshold, the
|
||||
agent generates a compressed summary and inserts it as a new message. Earlier
|
||||
messages remain in the database and are still visible to users, but are excluded
|
||||
from the model's context window. This happens transparently and keeps
|
||||
long-running sessions productive.
|
||||
|
||||
### Message queuing
|
||||
|
||||
Users can send follow-up messages while the agent is actively working. Messages
|
||||
are queued in the database and delivered when the agent completes its current
|
||||
turn — the full sequence of steps until the model stops calling tools. There is
|
||||
no need to wait for a response before providing additional context or
|
||||
redirecting the agent.
|
||||
|
||||
## Tool execution
|
||||
|
||||
Tools are how the agent takes action. Each tool call from the LLM translates to
|
||||
a concrete operation — either inside a workspace or within the control plane
|
||||
itself.
|
||||
|
||||
### Workspace connection lifecycle
|
||||
|
||||
The connection to a workspace is **lazy**. It is not established when a chat
|
||||
starts — only when something needs to reach the workspace. This is typically
|
||||
triggered by the first tool call that requires workspace access. Once
|
||||
established, the connection is cached and reused for the duration of that chat
|
||||
session.
|
||||
|
||||
Chats that don't need workspace access (answering questions, planning an
|
||||
approach, discussing architecture) never provision or connect to a workspace.
|
||||
|
||||
### Workspace tools
|
||||
|
||||
These tools execute inside the workspace via the workspace daemon's HTTP API.
|
||||
They traverse the same Tailnet tunnel used by web terminals and IDE connections.
|
||||
|
||||
| Tool | What it does |
|
||||
|------------------|--------------------------------------------------------------------|
|
||||
| `read_file` | Reads file contents with line-number pagination. |
|
||||
| `write_file` | Writes content to a file. |
|
||||
| `edit_files` | Performs atomic search-and-replace edits across one or more files. |
|
||||
| `execute` | Runs a shell command (foreground or background). |
|
||||
| `process_output` | Retrieves output from a background process. |
|
||||
| `process_list` | Lists all tracked processes in the workspace. |
|
||||
| `process_signal` | Sends a signal (SIGTERM or SIGKILL) to a background process. |
|
||||
|
||||
### Platform tools
|
||||
|
||||
These tools run entirely within the control plane. They do not require a
|
||||
workspace connection.
|
||||
|
||||
| Tool | What it does |
|
||||
|--------------------|-------------------------------------------------------------------|
|
||||
| `list_templates` | Browses available workspace templates, sorted by popularity. |
|
||||
| `read_template` | Gets template details and configurable parameters. |
|
||||
| `create_workspace` | Creates a workspace from a template and waits for it to be ready. |
|
||||
|
||||
### Orchestration tools
|
||||
|
||||
These tools manage sub-agents — child chats that work on independent tasks in
|
||||
parallel.
|
||||
|
||||
| Tool | What it does |
|
||||
|-----------------|--------------------------------------------------------------|
|
||||
| `spawn_agent` | Delegates a task to a sub-agent with its own context window. |
|
||||
| `wait_agent` | Waits for a sub-agent to finish and collects its result. |
|
||||
| `message_agent` | Sends a follow-up message to a running sub-agent. |
|
||||
| `close_agent` | Stops a running sub-agent. |
|
||||
|
||||
## What runs where
|
||||
|
||||
Understanding the split between the control plane and the workspace is central
|
||||
to the security model.
|
||||
|
||||
| Responsibility | Where it runs | Details |
|
||||
|---------------------|---------------|---------------------------------------------------------------------------|
|
||||
| Agent loop | Control plane | Prompt processing, tool dispatch, step iteration. |
|
||||
| LLM inference | LLM provider | The control plane streams requests to the external provider. |
|
||||
| Chat state | Control plane | All messages, token usage, and status stored in the database. |
|
||||
| Git authentication | Control plane | Uses existing Coder external auth (GitHub, GitLab, Bitbucket). |
|
||||
| User identity | Control plane | Every action is tied to the user who submitted the prompt. |
|
||||
| Model/prompt config | Control plane | Administrators configure providers, models, and system prompts centrally. |
|
||||
| File read/write | Workspace | The workspace file system is the source of truth for code. |
|
||||
| Shell execution | Workspace | Commands run in the workspace's environment with its packages and tools. |
|
||||
| Git operations | Workspace | Commits, pushes, and branch management happen inside the workspace. |
|
||||
| Build and test | Workspace | Compilation, test suites, and dev servers run on workspace compute. |
|
||||
|
||||
The workspace has **zero AI awareness**. There are no LLM API keys, no agent
|
||||
processes, and no AI-specific software installed. If you inspect a workspace
|
||||
created by the agent, it looks identical to one a developer created
|
||||
manually.
|
||||
|
||||
## Chat state and persistence
|
||||
|
||||
All chat data is stored in the control plane database, not in the workspace.
|
||||
|
||||
- **Chat metadata** — status, owner, associated workspace, timestamps, and
|
||||
parent/child relationships for sub-agents.
|
||||
- **Messages** — every message (user, assistant, tool calls, tool results) is
|
||||
stored as a separate record with role, content, and token usage.
|
||||
- **Compressed context** — when the agent compacts the conversation, summaries
|
||||
are stored with a compression flag so the original context budget is
|
||||
preserved.
|
||||
- **Queued messages** — follow-up messages sent while the agent is working are
|
||||
held in a queue and delivered in order.
|
||||
|
||||
Because state lives in the database:
|
||||
|
||||
- Chat history survives workspace stops, rebuilds, and deletions.
|
||||
- An administrator can inspect any chat for audit or debugging.
|
||||
- The agent can resume work by targeting a new workspace and continuing from the
|
||||
last git branch or checkpoint.
|
||||
|
||||
## Security implications
|
||||
|
||||
The control plane architecture has direct consequences for how you secure AI
|
||||
coding workflows.
|
||||
|
||||
### No API keys in workspaces
|
||||
|
||||
LLM provider credentials exist only in the control plane. The workspace never
|
||||
sees them. There is nothing for a developer, a compromised dependency, or a
|
||||
rogue process to exfiltrate.
|
||||
|
||||
### Workspaces can be fully network-isolated
|
||||
|
||||
Because the workspace does not need to reach any LLM provider, you can restrict
|
||||
its network access to only:
|
||||
|
||||
- The control plane (required for the workspace daemon to function).
|
||||
- Your git provider (for push/pull operations).
|
||||
|
||||
Everything else can be blocked. The AI functionality comes from the control
|
||||
plane, not from the workspace's network.
|
||||
|
||||
> [!TIP]
|
||||
> For sensitive environments, create dedicated templates for agent workloads
|
||||
> with stricter egress rules than your standard developer templates. Because
|
||||
> the AI comes from the control plane, these templates do not need any
|
||||
> outbound access to LLM providers.
|
||||
|
||||
### Centralized enforcement
|
||||
|
||||
Administrators control which models are available, the system prompt, and tool
|
||||
configuration from the control plane. Developers can select from the set of
|
||||
admin-enabled models when starting or continuing a chat, but cannot add their
|
||||
own providers or override system prompts or tool permissions. When an
|
||||
administrator removes a model or modifies the system prompt, the change applies
|
||||
to all agent sessions immediately.
|
||||
|
||||
### User identity on every action
|
||||
|
||||
Every action the agent takes — PRs opened, code committed, commands executed —
|
||||
is tied to the user who submitted the prompt. There is no shared bot account or
|
||||
anonymous identity. If a developer submits a prompt that results in a pull
|
||||
request, that pull request is attributed to them via the git authentication
|
||||
already configured in your Coder deployment.
|
||||
|
||||
## Scaling and resource impact
|
||||
|
||||
The control plane overhead for Coder Agents is minimal. The heavy computation
|
||||
happens elsewhere:
|
||||
|
||||
- **LLM inference** runs on the external provider's infrastructure.
|
||||
- **File I/O, builds, and tests** run on workspace compute.
|
||||
- **The control plane** primarily proxies streaming responses and dispatches
|
||||
tool calls over existing network connections.
|
||||
@@ -0,0 +1,246 @@
|
||||
# Coder Agents
|
||||
|
||||
> [!NOTE]
|
||||
> Coder Agents is currently in internal preview. We are actively developing
|
||||
> the feature and demoing it with customers for feedback.
|
||||
|
||||
Coder Agents is a chat interface and API for delegating development work and research to coding agents in your Coder deployment. Developers describe the work they want done, and Coder Agents handles selecting a template, provisioning a workspace, and executing the task.
|
||||
|
||||
Coder Agents includes its own self-hosted AI coding
|
||||
agent that runs the agent loop directly within the Coder control plane.
|
||||
|
||||
No specialized software, API keys, or network access is required inside your workspace. The only requirement is network access between the control plane and external LLM providers.
|
||||
|
||||
<video autoplay playsinline loop>
|
||||
<source src="https://github.com/coder/coder/blob/main/docs/images/guides/ai-agents/coder-agents-ui.mp4?raw=true" type="video/mp4">
|
||||
Your browser does not support the video tag.
|
||||
</video>
|
||||
|
||||
## What Coder Agents is and isn't
|
||||
|
||||
It is a standalone agent written in Go that implements standard
|
||||
agentic patterns — sub-agent delegation, context compaction, file editing, and
|
||||
shell execution — and works with any LLM provider you configure.
|
||||
|
||||
It is not a wrapper around third-party agent tools like Claude Code
|
||||
or Codex.
|
||||
|
||||
## Who Coder Agents is for
|
||||
|
||||
Coder Agents is designed for organizations that need to self-host their AI
|
||||
coding workflows and maintain full control over how agents operate. It is a
|
||||
strong fit for:
|
||||
|
||||
- **Regulated industries** such as financial services, healthcare, and
|
||||
government, where AI tools must run on controlled infrastructure with
|
||||
auditable access and strict network boundaries.
|
||||
- **Platform engineering teams** that want to provide developers with a
|
||||
high-quality AI coding experience without managing per-workspace agent
|
||||
installations, API key distribution, or third-party agent licensing.
|
||||
- **Organizations with existing Coder deployments** that want to add agentic
|
||||
capabilities using their current templates, workspaces, and identity
|
||||
providers rather than adopting a separate SaaS product.
|
||||
|
||||
Coder Agents runs entirely self-hosted. There is no SaaS or managed component — the agent
|
||||
loop, chat history, and all tool execution happen within your Coder deployment.
|
||||
|
||||
Coder Agents is not a replacement for your text editor or IDE. It is the
|
||||
primary interface where developers work with and orchestrate coding agents.
|
||||
Developers still connect to workspaces via VS Code, Cursor, JetBrains, or any
|
||||
other editor to review, refine, and complete work that the agent produces.
|
||||
|
||||
## How it works
|
||||
|
||||
The agent loop runs inside [the control plane](./architecture.md). When a user
|
||||
submits a prompt, the control plane:
|
||||
|
||||
1. Sends the prompt to the configured LLM provider (Anthropic, OpenAI, Google,
|
||||
Azure, AWS Bedrock, or any OpenAI-compatible endpoint).
|
||||
1. Receives the model's response, which may include tool calls such as reading
|
||||
files, writing code, or running shell commands.
|
||||
1. Executes tool calls by connecting to a Coder workspace over the existing
|
||||
workspace connection — the same path used for web terminals, port
|
||||
forwarding, and IDE access.
|
||||
1. Returns tool results to the model and continues the loop until the task is
|
||||
complete.
|
||||
|
||||
The workspace itself has no knowledge of AI. It is standard compute
|
||||
infrastructure — there are no LLM API keys, no agent harnesses, and no special
|
||||
software installed. All intelligence lives in the control plane.
|
||||
|
||||

|
||||
|
||||
<small>The agent loop runs in the control plane. It makes outbound requests to LLM
|
||||
providers and connects to workspaces only when tool execution is needed.</small>
|
||||
|
||||
### Automatic workspace provisioning
|
||||
|
||||
Not every chat requires a workspace. The agent runs in the control plane and can
|
||||
answer questions, discuss architecture, or plan an approach without any
|
||||
infrastructure. Workspaces are only provisioned when the agent needs to take
|
||||
action — reading code, running commands, or editing files.
|
||||
|
||||
This means:
|
||||
|
||||
- **Faster responses** — conversations that don't require workspace access
|
||||
start immediately with no provisioning delay.
|
||||
- **Lower infrastructure cost** — workspaces are only created when the agent
|
||||
needs to do real development work.
|
||||
|
||||
When a workspace _is_ needed, the agent reads the available templates —
|
||||
including their descriptions and parameters — selects the appropriate one, and
|
||||
creates a workspace automatically. Users can also manually choose which workspace is used when starting a new chat.
|
||||
|
||||
Platform teams control template routing by writing clear template descriptions.
|
||||
For example, a description like "Use this template for Python backend services
|
||||
in the payments repo" helps the agent select the correct infrastructure.
|
||||
|
||||
**Examples of what triggers workspace creation:**
|
||||
|
||||
| No workspace needed | Workspace provisioned |
|
||||
|------------------------------------------------------|----------------------------------------------------------|
|
||||
| "What are the tradeoffs between REST and gRPC?" | "Find and fix the nil pointer crash in the auth service" |
|
||||
| "Help me draft an RFC for adding a caching layer" | "Run the test suite and fix any failures" |
|
||||
| "What's the best way to handle retry logic in Go?" | "Refactor the handler to use the new SDK types" |
|
||||
| "Compare connection pooling strategies for Postgres" | "Read the config file and add the new feature flag" |
|
||||
|
||||
### Sub-agents
|
||||
|
||||
Coder Agents supports sub-agent delegation. The root agent can spawn child
|
||||
agents to work on independent tasks in parallel. Each sub-agent gets its own
|
||||
context window, which keeps individual conversations focused and avoids the
|
||||
quality degradation that occurs as context windows grow large.
|
||||
|
||||
For example, an agent tasked with "explore this repository and document its
|
||||
structure" might spawn separate sub-agents to analyze the backend, frontend,
|
||||
and infrastructure directories simultaneously.
|
||||
|
||||
### Chat persistence
|
||||
|
||||
All chat state is stored in the Coder database, not in the workspace. If a
|
||||
workspace is stopped, deleted, or rebuilt, the full conversation history
|
||||
survives. The agent can resume work by creating a new workspace with the same
|
||||
template and continuing from the last known state, such as a git branch.
|
||||
|
||||
Users can also fork a chat at any point to explore a different direction while
|
||||
preserving the original conversation.
|
||||
|
||||
### Message queuing
|
||||
|
||||
Users can send follow-up messages while the agent is actively working. Messages
|
||||
are queued and delivered when the agent completes its current step, so there is
|
||||
no need to wait for a response before providing additional context or changing
|
||||
direction.
|
||||
|
||||
## Security benefits of the control plane architecture
|
||||
|
||||
Running the agent loop in the control plane rather than inside the developer
|
||||
workspace is an architectural decision that directly addresses the primary
|
||||
concerns regulated organizations have with AI coding tools: how do you give
|
||||
developers access to coding agents without introducing unnecessary risk?
|
||||
|
||||
Traditionally, agents run inside the same compute where code
|
||||
lives. This means the agent needs LLM API keys in the workspace, outbound
|
||||
network access to model providers, and often elevated permissions. In a
|
||||
regulated environment, this creates a surface area that is difficult to lock
|
||||
down.
|
||||
|
||||
Coder Agents eliminates this by moving the agent loop out of the workspace
|
||||
entirely:
|
||||
|
||||
- **No API keys in workspaces.** LLM provider credentials never enter the
|
||||
workspace. The control plane makes all outbound requests to model providers
|
||||
directly, so there is nothing for a developer or a compromised process to
|
||||
exfiltrate.
|
||||
- **No agent software to manage.** Workspaces don't need Claude Code, Codex,
|
||||
or any agent harness installed. This eliminates a class of supply chain risk
|
||||
and removes the need to keep agent software up to date across all workspaces.
|
||||
- **Network boundaries are simpler.** Because the workspace doesn't need access
|
||||
to LLM APIs, you can apply strict egress rules. An agent-only template might
|
||||
permit access to only your git provider (e.g., `github.com`) and nothing
|
||||
else. The workspace never needs to reach the internet for AI functionality.
|
||||
- **Centralized, enforced control.** Platform teams configure models, system
|
||||
prompts, and tool permissions from the control plane. These settings are
|
||||
enforced server-side — they are not user preferences that developers can
|
||||
override.
|
||||
- **User identity is always attached.** Every action the agent takes — PRs
|
||||
opened, code pushed, commands run — is tied to the user who submitted the
|
||||
prompt. There is no shared bot identity or anonymous execution.
|
||||
|
||||
> [!TIP]
|
||||
> For highly sensitive environments, create a dedicated set of templates for
|
||||
> agent workloads with stricter network policies than your standard developer
|
||||
> templates. Because the AI comes from the control plane, these templates don't
|
||||
> need any outbound access to LLM providers.
|
||||
|
||||
## LLM provider support
|
||||
|
||||
Coder Agents works with any LLM provider. Administrators configure providers
|
||||
and models from the Coder dashboard or API. Supported providers include:
|
||||
|
||||
| Provider | Description |
|
||||
|-------------------|------------------------------------------|
|
||||
| Anthropic | Claude models via Anthropic API |
|
||||
| OpenAI | GPT and Codex models via OpenAI API |
|
||||
| Google | Gemini models via Google AI API |
|
||||
| Azure OpenAI | OpenAI models hosted on Azure |
|
||||
| AWS Bedrock | Models available through AWS Bedrock |
|
||||
| OpenAI Compatible | Any endpoint implementing the OpenAI API |
|
||||
| OpenRouter | Multi-model routing via OpenRouter |
|
||||
| Vercel AI Gateway | Models via Vercel AI SDK |
|
||||
|
||||
Most providers support custom base URLs, which allows integration with
|
||||
enterprise LLM proxies, self-hosted model endpoints, and internal gateways.
|
||||
|
||||
Administrators can configure multiple providers simultaneously and set a default
|
||||
model. Developers select from enabled models when starting a chat.
|
||||
|
||||

|
||||
|
||||
<small>The model configuration panel in the Coder dashboard.</small>
|
||||
|
||||
## Built-in tools
|
||||
|
||||
The agent has access to a set of workspace tools that it uses to accomplish
|
||||
tasks:
|
||||
|
||||
| Tool | Description |
|
||||
|--------------------|---------------------------------------------------------|
|
||||
| `list_templates` | Browse available workspace templates |
|
||||
| `read_template` | Get template details and configurable parameters |
|
||||
| `create_workspace` | Create a workspace from a template |
|
||||
| `read_file` | Read file contents from the workspace |
|
||||
| `write_file` | Write a file to the workspace |
|
||||
| `edit_files` | Perform search-and-replace edits across files |
|
||||
| `execute` | Run shell commands in the workspace |
|
||||
| `spawn_agent` | Delegate a task to a sub-agent running in parallel |
|
||||
| `wait_agent` | Wait for a sub-agent to complete and collect its result |
|
||||
| `message_agent` | Send a follow-up message to a running sub-agent |
|
||||
| `close_agent` | Stop a running sub-agent |
|
||||
|
||||
These tools connect to the workspace over the same secure connection used for
|
||||
web terminals and IDE access. No additional ports or services are required in
|
||||
the workspace.
|
||||
|
||||
## Comparison to Coder Tasks
|
||||
|
||||
Coder Agents is a new approach that differs from
|
||||
[Coder Tasks](../tasks.md) in several ways:
|
||||
|
||||
| Aspect | Coder Agents | Coder Tasks |
|
||||
|---------------------|--------------------------------------|----------------------------------------------------------------|
|
||||
| Agent execution | Runs in the control plane | Runs inside the workspace |
|
||||
| Agent harness | Built-in, no installation needed | Requires Claude Code, Codex, or similar installed in workspace |
|
||||
| API keys | Stored in control plane only | Injected into workspace environment |
|
||||
| Chat state | Persisted in database | Stored in workspace |
|
||||
| Workspace selection | Automatic, based on task description | Manual, user selects template |
|
||||
| Sub-agents | Built-in parallel delegation | Not supported |
|
||||
| Modern chat UI | Native chat with diffs, queuing | Terminal-based interface |
|
||||
|
||||
## Product status
|
||||
|
||||
Coder Agents is currently in internal preview. We are actively developing the
|
||||
feature and demoing it with customers for feedback.
|
||||
|
||||
Our next step is to offer an early access program for interested customers. If
|
||||
you would like to participate, [contact us](https://coder.com/contact).
|
||||
@@ -0,0 +1,199 @@
|
||||
# Models
|
||||
|
||||
Administrators configure LLM providers and models from the Coder dashboard.
|
||||
These are deployment-wide settings — developers do not manage API keys or
|
||||
provider configuration. They select from the set of models that an administrator
|
||||
has enabled.
|
||||
|
||||
## Providers
|
||||
|
||||
Each LLM provider has a type, an API key, and an optional base URL override.
|
||||
|
||||
Coder supports the following provider types:
|
||||
|
||||
| Provider | Description |
|
||||
|-------------------|------------------------------------------|
|
||||
| Anthropic | Claude models via Anthropic API |
|
||||
| OpenAI | GPT and o-series models via OpenAI API |
|
||||
| Google | Gemini models via Google AI API |
|
||||
| Azure OpenAI | OpenAI models hosted on Azure |
|
||||
| AWS Bedrock | Models available through AWS Bedrock |
|
||||
| OpenAI Compatible | Any endpoint implementing the OpenAI API |
|
||||
| OpenRouter | Multi-model routing via OpenRouter |
|
||||
| Vercel AI Gateway | Models via Vercel AI SDK |
|
||||
|
||||
The **OpenAI Compatible** type is a catch-all for any service that exposes an
|
||||
OpenAI-compatible chat completions endpoint. Use it to connect to self-hosted
|
||||
models, internal gateways, or third-party proxies like LiteLLM.
|
||||
|
||||
### Add a provider
|
||||
|
||||
1. Navigate to the **Agents** page in the Coder dashboard.
|
||||
1. Click **Admin** in the top bar to open the configuration dialog.
|
||||
1. Select the **Providers** tab.
|
||||
1. Click the provider you want to configure.
|
||||
1. Enter the **API key** for the provider.
|
||||
1. Optionally set a **Base URL** to override the default endpoint. This is
|
||||
useful for enterprise proxies, regional endpoints, or self-hosted models.
|
||||
1. Click **Save**.
|
||||
|
||||

|
||||
|
||||
<small>The providers list shows all supported providers and their configuration
|
||||
status.</small>
|
||||
|
||||

|
||||
|
||||
<small>Adding a provider requires an API key. The base URL is optional.</small>
|
||||
|
||||
### Provider API keys and security
|
||||
|
||||
Provider API keys are stored encrypted in the Coder database. They are never
|
||||
exposed to workspaces, developers, or the browser after initial entry. The
|
||||
dashboard shows only whether a key is set, not the key itself.
|
||||
|
||||
Because the agent loop runs in the control plane, workspaces never need direct
|
||||
access to LLM providers. See
|
||||
[Architecture](./architecture.md#no-api-keys-in-workspaces) for details
|
||||
on this security model.
|
||||
|
||||
## Models
|
||||
|
||||
Each model belongs to a provider and has its own configuration for context limits,
|
||||
generation parameters, and provider-specific options.
|
||||
|
||||
### Add a model
|
||||
|
||||
1. Open the **Admin** dialog and select the **Models** tab.
|
||||
1. Click **Add** and select the provider for the new model.
|
||||
1. Enter the **Model Identifier** — the exact model string your provider
|
||||
expects (e.g., `claude-opus-4-6`, `gpt-5.3-codex`).
|
||||
1. Set a **Display Name** so developers see a human-readable label in the model
|
||||
selector.
|
||||
1. Set the **Context Limit** — the maximum number of tokens in the model's
|
||||
context window (e.g., `200000` for Claude Sonnet).
|
||||
1. Configure any provider-specific options (see below).
|
||||
1. Click **Save**.
|
||||
|
||||

|
||||
|
||||
<small>The models list shows all configured models grouped by provider.</small>
|
||||
|
||||

|
||||
|
||||
<small>Adding a model requires a model identifier, display name, and context
|
||||
limit. Provider-specific options appear dynamically based on the selected
|
||||
provider.</small>
|
||||
|
||||
### Set a default model
|
||||
|
||||
Click the **star icon** next to a model in the models list to make it the
|
||||
default. The default model is pre-selected when developers start a new chat.
|
||||
Only one model can be the default at a time.
|
||||
|
||||
## Model options
|
||||
|
||||
Every model has a set of general options and provider-specific options.
|
||||
The admin UI generates these fields automatically from the provider's
|
||||
configuration schema, so the available options always match the provider type.
|
||||
|
||||
### General options
|
||||
|
||||
These options apply to all providers:
|
||||
|
||||
| Option | Description |
|
||||
|-----------------------|--------------------------------------------------------------------------------------------------|
|
||||
| Model Identifier | The API model string sent to the provider (e.g., `claude-opus-4-6`). |
|
||||
| Display Name | The label shown to developers in the model selector. |
|
||||
| Context Limit | Maximum tokens in the context window. Used to determine when context compaction triggers. |
|
||||
| Compression Threshold | Percentage (0–100) of context usage at which the agent compresses older messages into a summary. |
|
||||
| Max Output Tokens | Maximum tokens generated per model response. |
|
||||
| Temperature | Controls randomness. Lower values produce more deterministic output. |
|
||||
| Top P | Nucleus sampling threshold. |
|
||||
| Top K | Limits token selection to the top K candidates. |
|
||||
| Presence Penalty | Penalizes tokens that have already appeared in the conversation. |
|
||||
| Frequency Penalty | Penalizes tokens proportional to how often they have appeared. |
|
||||
|
||||
### Provider-specific options
|
||||
|
||||
Each provider type exposes additional options relevant to its models. These
|
||||
fields appear dynamically in the admin UI when you select a provider.
|
||||
|
||||
#### Anthropic
|
||||
|
||||
| Option | Description |
|
||||
|------------------------|---------------------------------------------------------|
|
||||
| Thinking Budget Tokens | Maximum tokens allocated for extended thinking. |
|
||||
| Effort | Thinking effort level (`low`, `medium`, `high`, `max`). |
|
||||
|
||||
#### OpenAI
|
||||
|
||||
| Option | Description |
|
||||
|-----------------------|-----------------------------------------------------------------------|
|
||||
| Reasoning Effort | How much effort the model spends reasoning (`low`, `medium`, `high`). |
|
||||
| Max Completion Tokens | Cap on completion tokens for reasoning models. |
|
||||
| Parallel Tool Calls | Whether the model can call multiple tools at once. |
|
||||
|
||||
#### Google
|
||||
|
||||
| Option | Description |
|
||||
|------------------|-----------------------------------------------------|
|
||||
| Thinking Budget | Maximum tokens for the model's internal reasoning. |
|
||||
| Include Thoughts | Whether to include thinking traces in the response. |
|
||||
| Safety Settings | Content safety thresholds by category. |
|
||||
|
||||
#### OpenRouter
|
||||
|
||||
| Option | Description |
|
||||
|-------------------|---------------------------------------------------|
|
||||
| Reasoning Enabled | Enable extended reasoning mode. |
|
||||
| Reasoning Effort | Reasoning effort level (`low`, `medium`, `high`). |
|
||||
| Provider Order | Preferred provider routing order. |
|
||||
| Allow Fallbacks | Whether to fall back to alternative providers. |
|
||||
|
||||
#### Vercel AI Gateway
|
||||
|
||||
| Option | Description |
|
||||
|-------------------|-----------------------------------------------|
|
||||
| Reasoning Enabled | Enable extended reasoning mode. |
|
||||
| Reasoning Effort | Reasoning effort level. |
|
||||
| Provider Options | Routing preferences for underlying providers. |
|
||||
|
||||
> [!NOTE]
|
||||
> Azure OpenAI uses the same options as OpenAI. AWS Bedrock uses the same
|
||||
> options as Anthropic.
|
||||
|
||||
## How developers select models
|
||||
|
||||
Developers see a model selector dropdown when starting or continuing a chat on
|
||||
the Agents page. The selector shows only models from providers that have valid
|
||||
API keys configured. Models are grouped by provider if multiple providers are
|
||||
active.
|
||||
|
||||
The model selector uses the following precedence to pre-select a model:
|
||||
|
||||
1. **Last used model** — stored in the browser's local storage.
|
||||
1. **Admin-designated default** — the model marked with the star icon.
|
||||
1. **First available model** — if no default is set and no history exists.
|
||||
|
||||
Developers cannot add their own providers, models, or API keys. If no models
|
||||
are configured, the chat interface displays a message directing developers to
|
||||
contact an administrator.
|
||||
|
||||
## Using an LLM proxy
|
||||
|
||||
Organizations that route LLM traffic through a centralized proxy — such as
|
||||
Coder's AI Bridge or third parties like LiteLLM — can point any provider's **Base URL** at their proxy endpoint.
|
||||
|
||||
For example, to route all OpenAI traffic through Coder's AI Bridge:
|
||||
|
||||
1. Add or edit the **OpenAI** provider.
|
||||
1. Set the **Base URL** to your AI Bridge endpoint
|
||||
(e.g., `https://example.coder.com/api/v2/aibridge/openai/v1`).
|
||||
1. Enter the API key your proxy expects.
|
||||
|
||||
Alternatively, use the **OpenAI Compatible** provider type if your proxy serves
|
||||
multiple model families through a single OpenAI-compatible endpoint.
|
||||
|
||||
This lets you keep existing proxy-level features like per-user budgets, rate
|
||||
limiting, and audit logging while using Coder Agents as the developer interface.
|
||||
Reference in New Issue
Block a user