From 18128b7b52038084425c53b7de75e51f9bc3ec18 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Pawe=C5=82=20Banaszewski?= Date: Wed, 29 Jul 2026 19:38:22 +0200 Subject: [PATCH] docs: add standalone AI Gateway docs (#27592) Documents standalone AI Gateway deployment, Gateway key authentication, monitoring, and the updated embedded vs standalone topology in the AI Gateway docs. --------- Co-authored-by: Cian Johnston --- docs/admin/infrastructure/architecture.md | 10 +- .../validated-architectures/index.md | 10 + docs/admin/integrations/prometheus.md | 11 +- docs/admin/users/sessions-tokens.md | 2 + .../ai-gateway/ai-gateway-proxy/setup.md | 3 +- docs/ai-coder/ai-gateway/auth.md | 165 ++++++--- docs/ai-coder/ai-gateway/clients/index.md | 5 + docs/ai-coder/ai-gateway/index.md | 4 +- docs/ai-coder/ai-gateway/mcp.md | 4 +- docs/ai-coder/ai-gateway/monitoring.md | 233 +++++++++---- docs/ai-coder/ai-gateway/providers.md | 10 +- docs/ai-coder/ai-gateway/reference.md | 90 ++++- docs/ai-coder/ai-gateway/setup.md | 21 +- docs/ai-coder/ai-gateway/standalone.md | 319 ++++++++++++++++++ docs/install/kubernetes.md | 8 + docs/manifest.json | 10 +- helm/ai-gateway/README.md | 2 +- scripts/metricsdocgen/metrics | 2 +- 18 files changed, 773 insertions(+), 136 deletions(-) create mode 100644 docs/ai-coder/ai-gateway/standalone.md diff --git a/docs/admin/infrastructure/architecture.md b/docs/admin/infrastructure/architecture.md index 4d3c85dc21..79b709a6e9 100644 --- a/docs/admin/infrastructure/architecture.md +++ b/docs/admin/infrastructure/architecture.md @@ -138,7 +138,15 @@ as OpenAI and Anthropic. Users authenticate through Coder instead of managing se provider API keys. All prompts, token usage, and tool invocations are recorded for compliance and cost tracking. -Learn more: [AI Gateway](../../ai-coder/ai-gateway/index.md) +AI Gateway supports 2 deployment topologies: + +- **Embedded:** `coderd` runs the AI Gateway data plane in the same process. +- **Standalone:** AI Gateway runs outside `coderd`, as replicas that serve AI traffic and send requests directly to upstream providers. + +Standalone replicas hold no durable state. `coderd` is the source of truth and the only component that writes AI Gateway state to the database. +Each replica maintains a control connection to `coderd` for Coder API key validation, provider configuration, and AI session recording, and becomes unready when that connection is unavailable. + +Refer to [AI Gateway](../../ai-coder/ai-gateway/index.md) and [standalone deployment](../../ai-coder/ai-gateway/standalone.md) for configuration and operational guidance. ### Agent Firewall diff --git a/docs/admin/infrastructure/validated-architectures/index.md b/docs/admin/infrastructure/validated-architectures/index.md index 9fa4968bf8..8c037a6433 100644 --- a/docs/admin/infrastructure/validated-architectures/index.md +++ b/docs/admin/infrastructure/validated-architectures/index.md @@ -123,6 +123,16 @@ or helper scripts. Please note that the Registry is a hosted service and isn't available for offline use. +### AI Gateway + +[AI Gateway](../../../ai-coder/ai-gateway/index.md) proxies AI provider traffic +and records each AI session. It runs inside `coderd` by default, and can also +run as a [standalone deployment](../../../ai-coder/ai-gateway/standalone.md) +that scales independently of the control plane. Size replicas from your own AI +request volume and `CODER_AI_GATEWAY_MAX_CONCURRENCY`. For the chart's resource +requests and autoscaling defaults, refer to the +[AI Gateway Helm chart README](https://github.com/coder/coder/blob/main/helm/ai-gateway/README.md). + ## Kubernetes Infrastructure Kubernetes is the recommended, and supported platform for deploying Coder in the diff --git a/docs/admin/integrations/prometheus.md b/docs/admin/integrations/prometheus.md index fc913277ae..fe96bf992a 100644 --- a/docs/admin/integrations/prometheus.md +++ b/docs/admin/integrations/prometheus.md @@ -72,6 +72,13 @@ scrape_configs: apps: "coder" ``` +If you run a [standalone AI Gateway](../../ai-coder/ai-gateway/standalone.md), +each replica exports its own metrics on its own listener. Its Helm chart uses the +same `0.0.0.0:2112` default as the `coder` chart, but sets up no scrape +discovery. Refer to +[AI Gateway monitoring](../../ai-coder/ai-gateway/monitoring.md#kubernetes-discovery) +for more details. + To use the Kubernetes Prometheus operator to scrape metrics, you will need to create a `ServiceMonitor` in your Coder deployment namespace. The following is an example `ServiceMonitor`. @@ -102,6 +109,8 @@ You must first enable `coderd_agentstats_*` with the flag `CODER_PROMETHEUS_COLLECT_AGENT_STATS` before they can be retrieved from the deployment. They will always be available from the agent. +The `coder_ai_gateway_cost_control_*` metrics are exported only by `coderd`. + | Name | Type | Description | Labels | @@ -124,7 +133,7 @@ deployment. They will always be available from the agent. | `coder_ai_gateway_key_pool_exhaustions_total` | counter | The number of times the key pool was exhausted with no usable key (outcome: rate_limited, auth_failed). | `outcome` `provider` | | `coder_ai_gateway_key_pool_failover_attempts` | histogram | The number of keys attempted before success or exhaustion, per interception for bridged requests and per request for passthrough requests. | `provider` | | `coder_ai_gateway_key_pool_state` | gauge | The number of keys currently in each state (state: valid, temporary, permanent). | `provider` `state` | -| `coder_ai_gateway_key_pool_state_transitions_total` | counter | The number of API key state transitions during failover (reason: rate_limited, unauthorized, forbidden). | `provider` `reason` | +| `coder_ai_gateway_key_pool_state_transitions_total` | counter | The number of API key state transitions during failover (reason: rate_limited, unauthorized). | `provider` `reason` | | `coder_ai_gateway_non_injected_tool_selections_total` | counter | The number of times an AI model selected a tool to be invoked by the client. | `model` `name` `provider` | | `coder_ai_gateway_passthrough_total` | counter | The count of requests which were not intercepted but passed through to the upstream. | `method` `provider` `route` | | `coder_ai_gateway_prompts_total` | counter | The number of prompts issued by users (initiators). | `client` `initiator_id` `model` `provider` | diff --git a/docs/admin/users/sessions-tokens.md b/docs/admin/users/sessions-tokens.md index 8d31426694..f07e4e4474 100644 --- a/docs/admin/users/sessions-tokens.md +++ b/docs/admin/users/sessions-tokens.md @@ -119,6 +119,8 @@ To hard-delete a token, use the `--delete` flag: coder tokens remove --delete ``` +Deleting the user that owns a token revokes every token that user holds at the same time. + ## API Key Scopes API key scopes allow you to limit the permissions of a token to specific operations. By default, tokens are created with the `all` scope, granting full access to all actions the user can perform. For improved security, you can create tokens with limited scopes that restrict access to only the operations needed. diff --git a/docs/ai-coder/ai-gateway/ai-gateway-proxy/setup.md b/docs/ai-coder/ai-gateway/ai-gateway-proxy/setup.md index cf033f0345..0fd17985c5 100644 --- a/docs/ai-coder/ai-gateway/ai-gateway-proxy/setup.md +++ b/docs/ai-coder/ai-gateway/ai-gateway-proxy/setup.md @@ -54,7 +54,8 @@ All other traffic is tunneled through without decryption. Intercepted requests are forwarded to the AI Gateway, configured via [`CODER_AI_GATEWAY_PROXY_TARGET`](../../../reference/cli/server.md#--ai-gateway-proxy-target). By default, this is the embedded AI Gateway at `/api/v2/ai-gateway`, and no configuration is needed. -To forward intercepted requests to an AI Gateway that is not embedded in this Coder deployment, set: +AI Gateway Proxy remains part of the `coder server` process when you [deploy AI Gateway as a standalone service](../standalone.md). +To forward intercepted requests to the standalone Gateway, set: ```sh CODER_AI_GATEWAY_PROXY_TARGET=https://ai-gateway.example.com/ diff --git a/docs/ai-coder/ai-gateway/auth.md b/docs/ai-coder/ai-gateway/auth.md index 95e589bf60..1cab40e8aa 100644 --- a/docs/ai-coder/ai-gateway/auth.md +++ b/docs/ai-coder/ai-gateway/auth.md @@ -4,25 +4,33 @@ > AI Gateway is part of [AI Governance](../ai-governance.md), which is > included with a Premium license. -AI Gateway authenticates clients with the same Coder API token -that a user already uses against the rest of the Coder API. -No separate AI Gateway login or credential is required. +AI Gateway uses different credentials for different kinds of connections: -Authenticating with a Coder token avoids distributing provider-specific API keys -(such as OpenAI or Anthropic keys) to individual users. +- AI clients use a Coder API token to authenticate with AI Gateway as a user. +- Standalone Gateway replicas use AI Gateway keys to connect to the Coder control plane. +- AI Gateway uses provider credentials configured by an administrator to authenticate to upstream AI providers. +- In Bring Your Own Key (BYOK) mode, a user also supplies a personal provider credential or subscription token. + +These credentials are not interchangeable. +A Gateway key does not authenticate an AI client, and a Coder API token does not authenticate a standalone replica. + +## Authenticate AI clients + +AI Gateway authenticates clients with the same Coder API token +that a user uses for the rest of the Coder API. +No separate AI Gateway login is required for client traffic. +For token creation, expiration, and revocation, refer to [Sessions and API tokens](../../admin/users/sessions-tokens.md). + +Authenticating with a Coder token avoids distributing centralized provider API keys, +such as OpenAI or Anthropic keys, to individual users. AI Gateway handles upstream credentials centrally and forwards each request to the configured provider on the user's behalf. -> [!NOTE] -> Only Coder-issued tokens can authenticate users to AI Gateway. -> AI Gateway will use provider-specific API keys to -> [authenticate against upstream AI services](./setup.md#configure-providers). +The exact environment variable or setting name differs between tools. +Refer to the list of [supported clients](./clients/index.md) and +your tool's documentation for details. -The exact environment variable or setting naming may differ from tool to tool. -Refer to the list of [supported clients](./clients/index.md), -and consult your tool's documentation for details. - -## Create a Coder API token +### Create a Coder API token You can generate a token from the Coder dashboard or the CLI. @@ -44,22 +52,17 @@ coder tokens create --lifetime 30d -n my-ai-token Use short lifetimes for automation and CI to limit the blast radius if a token leaks. -## Retrieve your session token +### Retrieve your session token -If you're logged in with the Coder CLI, you can retrieve your current session token -by using [`coder login token`](../../reference/cli/login_token.md): +If you're logged in with the Coder CLI, retrieve your current session token +with [`coder login token`](../../reference/cli/login_token.md): ```sh export ANTHROPIC_API_KEY=$(coder login token) export ANTHROPIC_BASE_URL="https://coder.example.com/api/v2/ai-gateway/anthropic" ``` -Alternatively, you can [generate a long-lived API token](../../admin/users/sessions-tokens.md#generate-a-long-lived-api-token-on-behalf-of-yourself) -from the Coder dashboard. - -For headless or service-account use, refer to [Headless authentication](../../admin/users/headless-auth.md). - -## AI Gateway Proxy authentication +### AI Gateway Proxy authentication For tools that don't support a configurable base URL, [AI Gateway Proxy](./ai-gateway-proxy/index.md) intercepts traffic and forwards it to AI Gateway. @@ -72,25 +75,98 @@ export HTTPS_PROXY="https://coder:$(coder login token)@:8888" The client machine also needs to trust the proxy's CA certificate. For full setup, refer to [AI Gateway Proxy setup](./ai-gateway-proxy/setup.md). +## Authenticate standalone Gateway replicas + +AI Gateway keys are scoped to the Coder deployment. +A [standalone AI Gateway](./standalone.md) uses one of these keys to connect to `coderd`. +Only the built-in Owner role can create, list, and delete these keys. +Coder custom roles are organization-scoped and cannot grant the site-level `ai_gateway_key` permissions. + +Create a key with a descriptive name: + +```sh +coder ai-gateway keys create standalone-production +``` + +The command displays the plaintext key once. +Save it immediately in your secret manager because Coder cannot retrieve it later. +Coder stores only a short prefix of the key for display and a SHA-256 hash for authentication, never the full secret. + +Names must be unique, 64 characters or fewer, and use only lowercase letters, numbers, and hyphens. +A name cannot start or end with a hyphen or contain consecutive hyphens. +Gateway keys do not expire and cannot be scoped or restricted. + +Configure the standalone process with either of the following options, but not both: + +- `CODER_AI_GATEWAY_KEY` or `--key` supplies the key directly. +- `CODER_AI_GATEWAY_KEY_FILE` or `--key-file` reads the key from a file. + +A user login and `CODER_SESSION_TOKEN` are not used by `coder ai-gateway start`. +The same Gateway key can authenticate multiple replicas. +Separate keys make it easier to rotate or revoke each deployment independently. + +List keys and the most recent heartbeat for each: + +```sh +coder ai-gateway keys list +``` + +A replica records a heartbeat when its control connection is established, then refreshes it every 60 seconds while that connection is active. +The heartbeat reports control-connection liveness rather than client request volume. +Coder stores one timestamp per key, so replicas that share a key cannot be distinguished. + +For usage and flags, refer to the generated CLI reference for [creating](../../reference/cli/ai-gateway_keys_create.md), [listing](../../reference/cli/ai-gateway_keys_list.md), and [deleting](../../reference/cli/ai-gateway_keys_delete.md) Gateway keys. + +### Rotate a Gateway key + +Rotate a key with a rolling restart. +Run more than 1 replica behind a load balancer so client traffic continues during the rollout: + +1. Create a new Gateway key. +1. Update the Kubernetes Secret, environment variable, or key file used by every replica. +1. Restart or roll out the standalone deployment so every replica uses the new key. +1. Verify readiness and confirm that the new key has a recent heartbeat. +1. Delete the old key. + +Delete a key by name or ID: + +```sh +coder ai-gateway keys delete standalone-production +``` + +Add `--yes` to skip the confirmation prompt in automation. + +Deleting a key rejects new connections immediately. +An established session closes when its next heartbeat detects the deletion, within 60 seconds. +The replica then tries to reconnect, receives HTTP 401, and treats that as fatal: the process exits non-zero rather than retrying. +Stop or update every replica before deleting its key for an orderly rotation. + +## Authenticate to upstream providers + +Administrators configure provider credentials in the Coder dashboard or with the [AI Providers API](../../reference/api/aiproviders.md). +Standalone replicas fetch this provider configuration from `coderd`, so do not copy centralized provider credentials into each standalone deployment. + +For provider setup and credential failover, refer to [Provider configuration](./providers.md). + ## Bring Your Own Key (BYOK) -In addition to centralized key management, AI Gateway supports **Bring Your Own Key** (BYOK) mode. -Users can provide their own LLM API keys or use provider subscriptions -(such as Claude Pro/Max or ChatGPT Plus/Pro), +In addition to centralized key management, AI Gateway supports Bring Your Own Key (BYOK) mode. +Users can provide their own LLM API keys or provider subscriptions, +such as Claude Pro or Max and ChatGPT Plus or Pro, while AI Gateway continues to provide observability and governance. ![BYOK authentication flow](../../images/aibridge/clients/byok_auth_flow.png) In BYOK mode, users need two credentials: -- A **Coder API token** to authenticate with AI Gateway. -- Their **own LLM credential** (personal API key or subscription token) +- A Coder API token to authenticate with AI Gateway. +- Their own LLM credential, such as a personal API key or subscription token, which AI Gateway forwards to the upstream provider. BYOK and centralized modes can be used together. When a user provides their own credential, AI Gateway forwards it directly. -When no user credential is present, AI Gateway uses the admin-configured provider key. -This approach offers centralized keys as a default, +When no user credential is present, AI Gateway uses the administrator-configured provider key. +This approach offers centralized keys as a default while allowing individual users to bring their own key. > [!NOTE] @@ -98,11 +174,11 @@ while allowing individual users to bring their own key. > is skipped. Coder Agents requests routed through AI Gateway are in-process control plane -requests, not external client requests that send their own AI Gateway bearer -token. Coder Agents use this same global BYOK setting. When BYOK is enabled, -users can save personal API keys for any enabled AI provider from the Agents -settings page. See -[Agents credential selection](../agents/models.md#credential-selection) +requests, not external client requests that send their own AI Gateway bearer token. +Coder Agents use the same global BYOK setting. +When BYOK is enabled, users can save personal API keys for any enabled AI provider +from the Agents settings page. +Refer to [Agents credential selection](../agents/models.md#credential-selection) for the Agents-specific behavior. Visit individual [client pages](./clients/index.md) for configuration details. @@ -110,21 +186,18 @@ Visit individual [client pages](./clients/index.md) for configuration details. ### Enable or disable BYOK BYOK is enabled by default. -Administrators can disable it using `--ai-gateway-allow-byok=false` or `CODER_AI_GATEWAY_ALLOW_BYOK=false`: +Administrators can disable it for the embedded Gateway with `--ai-gateway-allow-byok=false` or `CODER_AI_GATEWAY_ALLOW_BYOK=false`: ```sh coder server --ai-gateway-allow-byok=false ``` -When disabled, BYOK requests are rejected with a `403 Forbidden` response and only centralized key authentication is permitted. +For a standalone Gateway, set the option on each replica: -## Rotate or revoke a token +```sh +CODER_AI_GATEWAY_ALLOW_BYOK=false coder ai-gateway start +``` -To rotate a token without downtime: - -1. Create a new token with `coder tokens create`. -2. Update the client's configuration to use the new token. -3. Delete the old token from the dashboard or with `coder tokens rm `. - -Deleting a token immediately revokes access. -Deleting the user that owns a token revokes every token that user holds at the same time. +Each replica enforces its own local value, which governs only the traffic that replica handles. +The standalone process does not receive the `coderd` value, so set the option on every replica that should reject BYOK requests. +When disabled, BYOK requests are rejected with a `403 Forbidden` response, and only centralized provider credentials are permitted. diff --git a/docs/ai-coder/ai-gateway/clients/index.md b/docs/ai-coder/ai-gateway/clients/index.md index 5aa78d96d9..481aaf1884 100644 --- a/docs/ai-coder/ai-gateway/clients/index.md +++ b/docs/ai-coder/ai-gateway/clients/index.md @@ -26,6 +26,11 @@ The exact configuration method varies by client, some use environment variables, Replace `coder.example.com` with your actual Coder deployment URL. +If you run a [standalone AI Gateway](../standalone.md), point clients at the +Gateway endpoint and drop the `/api/v2/ai-gateway` prefix, for example +`https://ai-gateway.example.com/openai/v1` or +`https://ai-gateway.example.com/anthropic`. + ## Authentication For information about authenticating with AI Gateway, visit [AI Gateway Authentication](../auth.md). diff --git a/docs/ai-coder/ai-gateway/index.md b/docs/ai-coder/ai-gateway/index.md index 16081d1f86..e252b9becb 100644 --- a/docs/ai-coder/ai-gateway/index.md +++ b/docs/ai-coder/ai-gateway/index.md @@ -6,11 +6,12 @@ AI Gateway is a smart gateway for AI. It acts as an intermediary between your us and providers like OpenAI and Anthropic. By intercepting all the AI traffic between these clients and the upstream APIs, AI Gateway can record user prompts, token usage, and tool invocations. AI Gateway supports clients running inside or outside Coder workspaces. +You can run the Gateway inside the Coder control plane (`coderd`) or as a [standalone service](./standalone.md). AI Gateway solves 3 key problems: 1. **Centralized authn/z management**: no more issuing & managing API tokens for OpenAI/Anthropic usage. - Users use their Coder session or API tokens to authenticate with `coderd` (Coder control plane), and + Users use their Coder session or API tokens to authenticate with `coderd`, and `coderd` securely communicates with the upstream APIs on their behalf. 1. **Auditing and attribution**: all interactions with AI services, whether autonomous or human-initiated, will be audited and attributed back to a user. @@ -40,6 +41,7 @@ AI Gateway is best suited for organizations facing these centralized management ## Next steps - [Set up AI Gateway](./setup.md) on your Coder deployment +- Optionally, [deploy AI Gateway as a standalone service](./standalone.md) - [Configure AI clients](./clients/index.md) to use AI Gateway - [Configure MCP servers](./mcp.md) for tool access - [Audit AI sessions](./audit.md) diff --git a/docs/ai-coder/ai-gateway/mcp.md b/docs/ai-coder/ai-gateway/mcp.md index 50f1a99bc0..5ea0531c5b 100644 --- a/docs/ai-coder/ai-gateway/mcp.md +++ b/docs/ai-coder/ai-gateway/mcp.md @@ -33,7 +33,7 @@ CODER_EXTERNAL_AUTH_0_CLIENT_SECRET=... CODER_EXTERNAL_AUTH_0_MCP_URL=https://api.githubcopilot.com/mcp/ ``` -See the diagram in [Implementation Details](./reference.md#implementation-details) for more information. +Refer to the diagram in [Embedded Gateway](./reference.md#embedded-gateway) for more information. You can also control which tools are injected by using an allow and/or a deny regular expression on the tool names: @@ -63,7 +63,7 @@ AI Gateway marks automatically injected tools with a prefix `bmcp_` ("bridged MC ## Tool Injection -If a model decides to invoke a tool and it has a `bmcp_` prefix and AI Gateway has a connection with the related MCP server, it will invoke the tool. The tool result will be passed back to the upstream AI provider, and this will loop until the model has all of its required data. These inner loops are not relayed back to the client; all it sees is the result of this loop. See [Implementation Details](./reference.md#implementation-details). +If a model decides to invoke a tool and it has a `bmcp_` prefix and AI Gateway has a connection with the related MCP server, it will invoke the tool. The tool result will be passed back to the upstream AI provider, and this will loop until the model has all of its required data. These inner loops are not relayed back to the client. The client only sees the result of this loop. Refer to [Deployment topologies](./reference.md#deployment-topologies). In contrast, tools which are defined by the client (i.e. the [`Bash` tool](https://docs.claude.com/en/docs/claude-code/settings#tools-available-to-claude) defined by _Claude Code_) cannot be invoked by AI Gateway, and the tool call from the model will be relayed to the client, after which it will invoke the tool. diff --git a/docs/ai-coder/ai-gateway/monitoring.md b/docs/ai-coder/ai-gateway/monitoring.md index 6c50a88a9c..da45b643ee 100644 --- a/docs/ai-coder/ai-gateway/monitoring.md +++ b/docs/ai-coder/ai-gateway/monitoring.md @@ -10,44 +10,83 @@ AI Gateway records the last `user` prompt, token usage, model reasoning, and eve ![User Leaderboard](../../images/aibridge/grafana_user_leaderboard.png) -We provide an example Grafana dashboard that you can import as a starting point for your metrics. See [the Grafana dashboard README](../../../examples/monitoring/dashboards/grafana/aibridge/README.md). +Coder provides an example Grafana dashboard that you can import as a starting point for your metrics. +Refer to the [Grafana dashboard README](../../../examples/monitoring/dashboards/grafana/aibridge/README.md). These logs and metrics can be used to determine usage patterns, track costs, and evaluate tooling adoption. -## Provider metrics +## Prometheus metrics -AI Gateway (the in-process daemon) and AI Gateway Proxy (the external -proxy) each export Prometheus metrics describing the configured -provider pool and its reload loop. See -[Provider Configuration](./providers.md) for the lifecycle these -metrics describe. +The embedded Gateway and [standalone Gateway](./standalone.md) export the same AI Gateway request metrics. +Each process exports metrics for the traffic that it handles: -| Metric | Type | Labels | Purpose | -|--------------------------------------------------------------------------|---------|--------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------| -| `coder_ai_gateway_provider_info` | gauge | `provider_name`, `provider_type`, `status` | One series per configured provider. Value is always `1`; the `status` label (`enabled`, `disabled`, `error`) carries the alertable signal. | -| `coder_ai_gateway_providers_last_reload_timestamp_seconds` | gauge | | Unix timestamp of the last reload attempt, success or failure. | -| `coder_ai_gateway_providers_last_reload_success_timestamp_seconds` | gauge | | Unix timestamp of the last reload that successfully refreshed the pool. | -| `coder_ai_gateway_proxy_provider_info` | gauge | `provider_name`, `provider_type`, `status` | Same shape as `coder_ai_gateway_provider_info` but reported by the external proxy. | -| `coder_ai_gateway_proxy_providers_last_reload_timestamp_seconds` | gauge | | Last reload attempt timestamp in the external proxy. | -| `coder_ai_gateway_proxy_providers_last_reload_success_timestamp_seconds` | gauge | | Last successful reload timestamp in the external proxy. | -| `coder_ai_gateway_proxy_connect_sessions_total` | counter | `type` (`mitm`, `tunneled`) | CONNECT sessions established by the proxy. | -| `coder_ai_gateway_proxy_mitm_requests_total` | counter | `provider` | MITM requests handled. | -| `coder_ai_gateway_proxy_inflight_mitm_requests` | gauge | `provider` | In-flight MITM requests. | -| `coder_ai_gateway_proxy_mitm_responses_total` | counter | `code`, `provider` | MITM responses by HTTP status code. | +- The Coder control plane (`coderd`) Prometheus listener exports metrics for the embedded Gateway. +- Each standalone Gateway replica exports metrics from its own Prometheus listener. + +Refer to [provider configuration](./providers.md) for the provider reload lifecycle these metrics describe. + +| Metric | Type | Labels | Purpose | +|--------------------------------------------------------------------|-----------|----------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------| +| `coder_ai_gateway_interceptions_total` | counter | `client`, `initiator_id`, `method`, `model`, `provider`, `route`, `status` | Intercepted requests. | +| `coder_ai_gateway_interceptions_inflight` | gauge | `model`, `provider`, `route` | Intercepted requests currently being processed. | +| `coder_ai_gateway_interceptions_duration_seconds` | histogram | `model`, `provider` | Total intercepted request duration, including upstream processing. | +| `coder_ai_gateway_passthrough_total` | counter | `method`, `provider`, `route` | Requests passed through to an upstream provider without interception. | +| `coder_ai_gateway_prompts_total` | counter | `client`, `initiator_id`, `model`, `provider` | Prompts issued by users. | +| `coder_ai_gateway_tokens_total` | counter | `client`, `initiator_id`, `model`, `provider`, `type` | Tokens used by intercepted requests. | +| `coder_ai_gateway_injected_tool_invocations_total` | counter | `model`, `name`, `provider`, `server` | Invocations of MCP tools injected by AI Gateway. | +| `coder_ai_gateway_non_injected_tool_selections_total` | counter | `model`, `name`, `provider` | Tools selected by a model for the client to invoke. | +| `coder_ai_gateway_circuit_breaker_state` | gauge | `endpoint`, `model`, `provider` | Current circuit-breaker state: `0` for closed, `0.5` for half-open, and `1` for open. | +| `coder_ai_gateway_circuit_breaker_trips_total` | counter | `endpoint`, `model`, `provider` | Times a circuit breaker transitioned to the open state. | +| `coder_ai_gateway_circuit_breaker_rejects_total` | counter | `endpoint`, `model`, `provider` | Requests rejected because a circuit breaker was open. | +| `coder_ai_gateway_key_pool_state` | gauge | `provider`, `state` | Provider keys in each state: `valid`, `temporary`, or `permanent`. | +| `coder_ai_gateway_key_pool_state_transitions_total` | counter | `provider`, `reason` | Provider key state transitions during failover. | +| `coder_ai_gateway_key_pool_exhaustions_total` | counter | `outcome`, `provider` | Times a provider key pool had no usable key. | +| `coder_ai_gateway_key_pool_failover_attempts` | histogram | `provider` | Keys attempted before a request succeeded or exhausted the provider key pool. | +| `coder_ai_gateway_provider_info` | gauge | `provider_name`, `provider_type`, `status` | Build status of each configured provider, including disabled and errored ones. Value is always `1`; `status` is `enabled`, `disabled`, or `error`. | +| `coder_ai_gateway_providers_last_reload_timestamp_seconds` | gauge | | Unix timestamp of the last attempt to rebuild the Gateway provider pool. | +| `coder_ai_gateway_providers_last_reload_success_timestamp_seconds` | gauge | | Unix timestamp of the last successful rebuild of the Gateway provider pool. | + +Histograms also emit the standard `_bucket`, `_sum`, and `_count` series. + +### Cost control metrics + +Budget enforcement runs in `coderd`. +Cost control metrics are exported only from the `coderd` Prometheus listener. +Standalone replicas do not export them. + +| Metric | Type | Labels | Purpose | +|--------------------------------------------------------------------|-----------|---------------------|------------------------------------------------------------------------------------------| +| `coder_ai_gateway_cost_control_blocked_requests_total` | counter | `group_id` | AI requests blocked because the initiator's budget was exceeded. | +| `coder_ai_gateway_cost_control_blocked_users` | gauge | `group_id` | Users currently over their AI budget. | +| `coder_ai_gateway_cost_control_enforcement_duration_seconds` | histogram | `outcome` | Duration of AI budget enforcement checks. `outcome` is `allowed`, `blocked`, or `error`. | +| `coder_ai_gateway_cost_control_unpriced_token_usage_records_total` | counter | `model`, `provider` | Recorded token-usage records for which no model price was found. | + +### AI Gateway Proxy metrics + +AI Gateway Proxy exports metrics from the `coderd` Prometheus listener. + +| Metric | Type | Labels | Purpose | +|--------------------------------------------------------------------------|---------|--------------------------------------------|-----------------------------------------------------------------------------------------------------------------| +| `coder_ai_gateway_proxy_connect_sessions_total` | counter | `type` | CONNECT sessions established, classified as `mitm` or `tunneled`. | +| `coder_ai_gateway_proxy_mitm_requests_total` | counter | `provider` | MITM requests handled by AI Gateway Proxy. | +| `coder_ai_gateway_proxy_inflight_mitm_requests` | gauge | `provider` | MITM requests currently being processed. | +| `coder_ai_gateway_proxy_mitm_responses_total` | counter | `code`, `provider` | MITM responses by HTTP status code. | +| `coder_ai_gateway_proxy_provider_info` | gauge | `provider_name`, `provider_type`, `status` | Routing status of each configured provider. Value is always `1`; `status` is `enabled`, `disabled`, or `error`. | +| `coder_ai_gateway_proxy_providers_last_reload_timestamp_seconds` | gauge | | Unix timestamp of the last attempt to rebuild the proxy routing snapshot. | +| `coder_ai_gateway_proxy_providers_last_reload_success_timestamp_seconds` | gauge | | Unix timestamp of the last successful rebuild of the proxy routing snapshot. | + +Refer to the [Prometheus reference](../../admin/integrations/prometheus.md) for these metrics alongside the other metrics that Coder components export. + +### Metric name migration > [!IMPORTANT] -> The AI Gateway metric prefixes were renamed: `coder_aibridged_*` became -> `coder_ai_gateway_*` and `coder_aibridgeproxyd_*` became -> `coder_ai_gateway_proxy_*`. This rename covers every AI Gateway metric, -> including the interception, token, prompt, tool, and circuit-breaker counters -> listed in the [Prometheus reference](../../admin/integrations/prometheus.md). -> The legacy `coder_aibridged_*` and `coder_aibridgeproxyd_*` names are still -> emitted with identical values during the v2.35 and v2.36 deprecation window. -> They are planned for removal in v2.37. Migrate dashboards and alerts to the new -> names now. Do not relabel new names back to old names while legacy names are -> still emitted, because that creates duplicate legacy series in the same scrape. -> After legacy names are removed, use `metric_relabel_configs` only if you need a -> temporary compatibility bridge for dashboards that still use the old names: +> The embedded Gateway metric prefix changed from `coder_aibridged_*` to `coder_ai_gateway_*`, and the proxy prefix changed from `coder_aibridgeproxyd_*` to `coder_ai_gateway_proxy_*`. +> The embedded Gateway and AI Gateway Proxy emit the legacy names with identical values during the v2.35 and v2.36 deprecation window, and the legacy names are planned for removal in v2.37. +> The cost control metrics were added after the rename and have no legacy alias. +> The standalone Gateway emits only the current `coder_ai_gateway_*` names. +> Migrate dashboards and alerts to the new names. +> Do not relabel new names back to old names while both are emitted because this creates duplicate legacy series in the same scrape. +> After the legacy names are removed, use `metric_relabel_configs` only if you need a temporary compatibility bridge: > > ```yaml > metric_relabel_configs: @@ -67,28 +106,80 @@ metrics describe. Alert on any provider entering a non-`enabled` status: ```promql -sum by (provider_name, status) (coder_ai_gateway_provider_info{status!="enabled"}) > 0 +sum by (instance, provider_name, status) ( + coder_ai_gateway_provider_info{status!="enabled"} +) > 0 ``` -Alert when the reload loop is firing but failing to refresh the pool -for longer than a few minutes: +Alert when the provider reload loop is firing but failing to refresh the pool for longer than a few minutes: ```promql (coder_ai_gateway_providers_last_reload_timestamp_seconds - coder_ai_gateway_providers_last_reload_success_timestamp_seconds) > 300 ``` -Repeat the same query against `coder_ai_gateway_proxy_*` if you run the -external proxy. +Use the `coder_ai_gateway_proxy_*` metrics when you alert on AI Gateway Proxy. -## Structured Logging +## Standalone Gateway monitoring -AI Gateway can emit structured logs for every interception event to your -existing log pipeline. This is useful for exporting data to external SIEM or -observability platforms. See [Structured Logging](./setup.md#structured-logging) -in the setup guide for configuration and a full list of record types. +### Metrics listener -## Exporting Data +Enable the standalone metrics listener with `CODER_PROMETHEUS_ENABLE=true` or `--prometheus-enable`, and set its bind address with `CODER_PROMETHEUS_ADDRESS` or `--prometheus-address`. +The command default is `127.0.0.1:2112`. +The listener is unauthenticated, so expose it only to your monitoring network. + +In addition to the common `coder_ai_gateway_*` metrics, the standalone listener exports standard unprefixed `go_*`, `process_*`, and `promhttp_*` metrics for the standalone process and its metrics handler. + +### Kubernetes discovery + +The [AI Gateway Helm chart](../../../helm/ai-gateway/README.md#metrics) enables metrics and binds the listener to `0.0.0.0:2112` by default. +The chart exposes a named `metrics` container port, but it does not include this port in the data-plane Service or create monitoring discovery resources. +For Prometheus pod-based discovery, add scrape annotations to each Gateway pod: + +```yaml +coder: + podAnnotations: + prometheus.io/scrape: "true" + prometheus.io/port: "2112" +``` + +You can also configure a `PodMonitor` to select the chart's `app.kubernetes.io/name` and `app.kubernetes.io/instance` pod labels, or add dedicated selector labels with `coder.podLabels`. +Manage the `PodMonitor` separately or include it in the Helm release with `extraTemplates`. +To use a `ServiceMonitor`, create a separate Service that exposes the `metrics` container port because the chart's data-plane Service exposes only the HTTP traffic port. +Refer to the [Helm chart metrics configuration](../../../helm/ai-gateway/README.md#metrics) for the available chart settings. + +### Health and readiness + +A standalone AI Gateway exposes health endpoints on its data-plane listener: + +| Endpoint | Success condition | +|------------|-----------------------------------------------------------------------------------------------| +| `/healthz` | The HTTP listener is serving. | +| `/readyz` | The control connection to `coderd` is active and provider configuration has been initialized. | + +`/readyz` returns HTTP 503 until provider configuration is initialized and whenever the control connection to `coderd` is unavailable. +A `200 OK` response from `/healthz` only means the HTTP listener is accepting connections. It returns `200 OK` even when the control connection is down. +Both endpoints are unauthenticated, bypass the concurrency, rate limiting, and BYOK middleware, and do not create trace spans. + +The standalone Helm chart enables a `/healthz` liveness probe and a `/readyz` readiness probe by default. +The startup probe is disabled by default. + +## Logs + +### Standalone operational logs + +Standalone replicas use the standard Coder logging options. +Configure them on every replica or through `coder.env` in the AI Gateway Helm chart. +Refer to the [`coder ai-gateway start` logging options](../../reference/cli/ai-gateway_start.md#-l---log-filter) for configuration details. + +### Structured interception logs + +AI Gateway can emit a structured log for every interception record to an external SIEM or observability platform. +The `CODER_AI_GATEWAY_STRUCTURED_LOGGING` setting belongs to `coderd`, standalone Gateway does not consume it. +Standalone replicas send interception records to `coderd`, which writes the structured logs to the Coder server log output. +Refer to [structured logging](./setup.md#structured-logging) for configuration and record types. + +## Export data AI Gateway interception data can be exported for external analysis, compliance reporting, or integration with log aggregation systems. @@ -130,42 +221,48 @@ Available query filters: - `started_after` - Filter sessions after a timestamp - `started_before` - Filter sessions before a timestamp -See the [API documentation](../../reference/api/aigateway.md) for full details. +Refer to the [API documentation](../../reference/api/aigateway.md) for full details. -## Data Retention +## Data retention AI Gateway data is retained for **60 days by default**. Configure the retention period to balance storage costs with your organization's compliance and analysis needs. -For configuration options and details, see [Data Retention](./setup.md#data-retention) +For configuration options and details, refer to [Data Retention](./setup.md#data-retention) in the AI Gateway setup guide. ## Tracing -AI Gateway supports tracing via [OpenTelemetry](https://opentelemetry.io/), -providing visibility into request processing, upstream API calls, and MCP server -interactions. +AI Gateway supports tracing through [OpenTelemetry](https://opentelemetry.io/) for request processing, upstream API calls, and MCP server interactions. +Embedded Gateway spans are emitted by the `coder server` process. +Standalone spans are emitted independently by every replica with the service name `coder-ai-gateway`. -### Enabling Tracing +### Enable tracing -AI Gateway tracing is enabled when tracing is enabled for the Coder server. -To enable tracing set `CODER_TRACE_ENABLE` environment variable or -[--trace](https://coder.com/docs/reference/cli/server#--trace) CLI flag: +AI Gateway exports spans over OTLP/gRPC when you set `CODER_TRACE_ENABLE`, honoring `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT`. +The exporter always dials without TLS, so an `https://` endpoint is still contacted over plaintext gRPC. +`CODER_TRACE_HONEYCOMB_API_KEY` adds a Honeycomb exporter and works with or without `CODER_TRACE_ENABLE`. +Set only the Honeycomb key to export to Honeycomb alone, or set both to export to Honeycomb and an OTLP collector. + +The embedded and standalone Gateways share the same tracing options. +Refer to the [`coder server` tracing options](../../reference/cli/server.md#--trace) for the embedded Gateway and the [`coder ai-gateway start` tracing options](../../reference/cli/ai-gateway_start.md#--trace) for standalone replicas. +Configure tracing on every standalone process or through `coder.env` in the AI Gateway Helm chart. + +The following minimal configuration enables tracing and exports spans over OTLP/gRPC: ```sh export CODER_TRACE_ENABLE=true +export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://otel-collector:4317 ``` -```sh -coder server --trace -``` +In both deployment modes, each request to the Gateway's LLM API endpoint creates an HTTP request span, including requests that are passed through or rejected instead of intercepted. -### What is Traced +### Traced operations AI Gateway creates spans for the following operations: -| Span Name | Description | +| Span name | Description | |---------------------------------------------|------------------------------------------------------| | `CachedBridgePool.Acquire` | Acquiring a request bridge instance from the pool | | `Intercept` | Top-level span for processing an intercepted request | @@ -173,33 +270,33 @@ AI Gateway creates spans for the following operations: | `Intercept.ProcessRequest` | Processing the request through the bridge | | `Intercept.ProcessRequest.Upstream` | Forwarding the request to the upstream AI provider | | `Intercept.ProcessRequest.ToolCall` | Executing a tool call requested by the AI model | -| `Intercept.RecordInterception` | Recording creating interception record | -| `Intercept.RecordPromptUsage` | Recording prompt/message data | +| `Intercept.RecordInterception` | Creating the interception record | +| `Intercept.RecordPromptUsage` | Recording prompt and message data | | `Intercept.RecordTokenUsage` | Recording token consumption | -| `Intercept.RecordToolUsage` | Recording tool/function calls | +| `Intercept.RecordToolUsage` | Recording tool and function calls | +| `Intercept.RecordModelThought` | Recording model reasoning | | `Intercept.RecordInterceptionEnded` | Recording the interception as completed | +| `Passthrough` | Forwarding a non-intercepted provider request | | `ServerProxyManager.Init` | Initializing MCP server proxy connections | | `StreamableHTTPServerProxy.Init` | Setting up HTTP-based MCP server proxies | | `StreamableHTTPServerProxy.Init.fetchTools` | Fetching available tools from MCP servers | -Example trace of an interception using Jaeger backend: +Example trace of an interception using a Jaeger backend: ![Trace of interception](../../images/aibridge/jaeger_interception_trace.png) -### Capturing Logs in Traces +### Capture logs in traces > [!NOTE] > Enabling log capture may generate a large volume of trace events. -To include log messages as trace events, enable trace log capture -by setting `CODER_TRACE_LOGS` environment variable or using -[--trace-logs](https://coder.com/docs/reference/cli/server#--trace-logs) flag: +Set `CODER_TRACE_LOGS=true` to include log messages as trace events: ```sh export CODER_TRACE_ENABLE=true export CODER_TRACE_LOGS=true ``` -```sh -coder server --trace --trace-logs -``` +Log capture only applies to recording spans, so it requires tracing to be enabled through `CODER_TRACE_ENABLE` or a backend-specific exporter such as `CODER_TRACE_HONEYCOMB_API_KEY`. +Leave `CODER_TRACE_LOGS` unset to trace without capturing logs. +For standalone replicas, set both options on every process that should capture logs. diff --git a/docs/ai-coder/ai-gateway/providers.md b/docs/ai-coder/ai-gateway/providers.md index f1ffcf8f01..8e2535879d 100644 --- a/docs/ai-coder/ai-gateway/providers.md +++ b/docs/ai-coder/ai-gateway/providers.md @@ -38,6 +38,10 @@ After seeding, manage providers through the dashboard or API. A provider that has been edited or removed there is not recreated or overwritten from the environment on the next restart. +Seeding is a `coderd` operation. A [standalone Gateway](./standalone.md) +ignores the deprecated provider variables and fetches provider +configuration from `coderd`. + ## Provider types AI Gateway speaks two upstream API formats: the **OpenAI** format @@ -277,7 +281,7 @@ an API key. ## Provider lifecycle Every provider carries an explicit status, surfaced through the -[`provider_info`](./monitoring.md#provider-metrics) metric and the API: +[`provider_info`](./monitoring.md#prometheus-metrics) metric and the API: | Status | Meaning | Effect on requests | |------------|-------------------------------------------------------------------------------|--------------------------------------------------| @@ -299,11 +303,13 @@ attempt and each successful reload, exposed as Prometheus metrics: If you run the [external proxy](./ai-gateway-proxy/index.md), it exposes the same pair under the `coder_ai_gateway_proxy_` prefix. +Each [standalone Gateway](./standalone.md) replica reloads providers +independently. A growing gap between the attempt and success timestamps means reloads are firing but failing to apply. Alert on that gap rather than on a single failure, which may resolve on the next change. See -[Monitoring](./monitoring.md#provider-metrics) for the full metric list +[Monitoring](./monitoring.md#prometheus-metrics) for the full metric list and sample alert queries. ## Key failover diff --git a/docs/ai-coder/ai-gateway/reference.md b/docs/ai-coder/ai-gateway/reference.md index c1f3ae4377..68836e01ec 100644 --- a/docs/ai-coder/ai-gateway/reference.md +++ b/docs/ai-coder/ai-gateway/reference.md @@ -4,23 +4,75 @@ > AI Gateway is part of [AI Governance](../ai-governance.md), which is > included with a Premium license. -## Implementation Details +## Deployment topologies -`coderd` runs an in-memory instance of `aibridged`, whose logic is mostly contained in ../../../aibridge. In future releases we will support running external instances for higher throughput and complete memory isolation from `coderd`. +AI Gateway can run inside `coderd` or as a standalone data-plane service. +Both topologies run the same Gateway request handling and keep `coderd` as the source of truth for Coder API key validation, provider configuration, and AI session records. +They differ in how requests are routed to the Gateway. + +### Embedded Gateway + +By default, `coder server` runs an in-memory Gateway instance in the `coderd` process. +AI clients send requests to `/api/v2/ai-gateway//`. +The embedded Gateway uses the same control RPC as a standalone deployment, over an in-process transport rather than a network connection. +It does not use a Gateway key and does not negotiate an API version. + +The following diagram shows the embedded topology: ![AI Gateway implementation details](../../images/aibridge/aibridge-implementation-details.png) +### Standalone Gateway + +A [standalone deployment](./standalone.md) runs the AI traffic data plane outside the `coderd` process. +Each replica accepts client traffic, sends AI requests directly to upstream providers, and maintains a control connection to `coderd` using a [Gateway key](./standalone.md#create-a-gateway-key). + +The control connection carries: + +- Coder API key validation, which resolves each request to an active Coder user. +- AI budget checks, which reject requests from users over their spend limit. +- Provider configuration, plus a change signal when the provider set changes. +- AI session records. +- **Deprecated**: the configuration and access tokens used by [injected MCP](./mcp.md). + +Standalone replicas do not own authoritative database state. +They keep ephemeral provider snapshots, request caches, provider key pools, and metrics in memory, and emit their own logs and traces. +Each replica writes its own [API dumps](./setup.md#api-dumps) to its own local disk when dumps are enabled. + +`coderd` remains required for standalone operation. +A replica becomes unready when its control connection is unavailable, even if its HTTP listener remains healthy. +AI Gateway Proxy remains part of `coder server` and can forward its intercepted traffic to either the embedded Gateway or a standalone endpoint. + +## Version compatibility + +The control connection between a standalone replica and `coderd` is versioned. +The current version is defined in [`coderd/aibridged/proto/version.go`](https://github.com/coder/coder/blob/main/coderd/aibridged/proto/version.go). + +`coderd` validates the version that a standalone replica advertises before it accepts the control connection. +Compatibility follows these rules: + +- The Gateway and `coderd` major versions must match. +- The Gateway minor version must be less than or equal to the `coderd` minor version. +- `coderd` rejects a standalone Gateway that advertises a newer minor version. + +A rejected replica receives an HTTP 400 response that reports the `client_api_version` and `server_api_version` values. +Coder build versions are not the compatibility criterion. + +For upgrade and rollback ordering, refer to [Version compatibility](./standalone.md#version-compatibility) in the standalone deployment guide. + ## Supported APIs -API support is broken down into two categories: +API support is divided into two categories: -- **Intercepted**: requests are intercepted, audited, and augmented - full AI Gateway functionality -- **Passthrough**: requests are proxied directly to the upstream, no auditing or augmentation takes place +- **Intercepted**: Requests are intercepted, audited, and augmented. +- **Passthrough**: Requests are proxied directly to the upstream provider without auditing or augmentation. Where relevant, both streaming and non-streaming requests are supported. +Paths are relative to the provider's base URL, such as `https://ai-gateway.example.com/openai/v1` or `https://ai-gateway.example.com/anthropic`. ### OpenAI +The OpenAI provider also serves the Azure OpenAI, Google, OpenRouter, Vercel, and OpenAI-compatible provider types. + #### Intercepted - [`/v1/chat/completions`](https://platform.openai.com/docs/api-reference/chat/create) @@ -28,18 +80,44 @@ Where relevant, both streaming and non-streaming requests are supported. #### Passthrough +- [`/v1/conversations(/*)`](https://platform.openai.com/docs/api-reference/conversations) - [`/v1/models(/*)`](https://platform.openai.com/docs/api-reference/models/list) +- [`/v1/responses/*`](https://platform.openai.com/docs/api-reference/responses/get) + +The legacy [`/v1/completions`](https://platform.openai.com/docs/api-reference/completions) API is deprecated and is not passed through. ### Anthropic +The Anthropic provider also serves the AWS Bedrock provider type. + #### Intercepted - [`/v1/messages`](https://docs.claude.com/en/api/messages) #### Passthrough +- [`/v1/messages/count_tokens`](https://docs.claude.com/en/api/messages-count-tokens) - [`/v1/models(/*)`](https://docs.claude.com/en/api/models-list) +- `/api/event_logging/*` + +### GitHub Copilot + +#### Intercepted + +- `/chat/completions` +- `/responses` +- `/v1/messages` + +#### Passthrough + +- `/models(/*)` +- `/agents/*` +- `/mcp/*` +- `/.well-known/*` + +Any route that is not listed above returns `404`. ## Troubleshooting -To report a bug, file a feature request, or view a list of known issues, please visit our [GitHub repository](https://github.com/coder/coder/issues). If you encounter issues with AI Gateway, please reach out to us via [Discord](https://discord.gg/coder). +To report a bug, file a feature request, or review known issues, visit the [Coder GitHub repository](https://github.com/coder/coder/issues). +For help with AI Gateway, visit the [Coder Discord](https://discord.gg/coder). diff --git a/docs/ai-coder/ai-gateway/setup.md b/docs/ai-coder/ai-gateway/setup.md index ebcb6f6cd7..91f48c120e 100644 --- a/docs/ai-coder/ai-gateway/setup.md +++ b/docs/ai-coder/ai-gateway/setup.md @@ -1,6 +1,9 @@ # Setup -AI Gateway runs inside the Coder control plane (`coderd`), requiring no separate compute to deploy or scale. Once enabled, `coderd` runs the `aibridged` in-memory and brokers traffic to your configured AI providers on behalf of authenticated users. +By default, AI Gateway runs inside the Coder control plane (`coderd`) and requires no separate compute. +In embedded mode, `coderd` runs the Gateway in memory and brokers traffic to your configured AI providers on behalf of authenticated users. + +If AI traffic needs dedicated compute, independent scaling, or a separate network endpoint, you can [deploy AI Gateway as a standalone service](./standalone.md). > [!NOTE] > Since v2.34, provider environment variables and flags are deprecated. @@ -11,8 +14,10 @@ AI Gateway runs inside the Coder control plane (`coderd`), requiring no separate ## Activation -AI Gateway must be enabled in deployment config before users can authenticate -to it. +The AI Gateway feature must be enabled in the Coder deployment configuration before +embedded or standalone Gateway instances can serve authenticated traffic. + +_AI Gateway is enabled by default as of v2.34._ ```sh export CODER_AI_GATEWAY_ENABLED=true @@ -21,7 +26,9 @@ coder server coder server --ai-gateway-enabled=true ``` -_AI Gateway is enabled by default as of v2.34._ +A standalone process does not read `CODER_AI_GATEWAY_ENABLED` from its own environment. +However, this setting must remain enabled on `coderd`. +It is required for Gateway key management endpoints to work and for standalone replicas to connect to the control plane. ## Configure Providers @@ -84,6 +91,9 @@ with `/var/lib/coder/ai-gateway-dumps` configured writes to Sensitive headers are redacted before dumps are written. Leave the value empty to disable dumping. +Each [standalone Gateway](./standalone.md) replica accepts the same API dump +settings and writes dumps to its own local disk. + > [!WARNING] > API dumps are intended for short diagnostic sessions only. Dump files contain > raw request and response data, which may include proprietary or sensitive @@ -138,6 +148,9 @@ stderr) or [`--log-json`](../../reference/cli/server.md#--log-json). For machine ingestion, set `--log-json` to a file path or `/dev/stderr` so that records are emitted as JSON. +This setting belongs to `coderd`. +A [standalone Gateway](./standalone.md) does not consume it. + Filter for AI Gateway records in your logging pipeline by matching on the `"interception log"` message. Each log line includes a `record_type` field that indicates the kind of event captured: diff --git a/docs/ai-coder/ai-gateway/standalone.md b/docs/ai-coder/ai-gateway/standalone.md new file mode 100644 index 0000000000..08fe70c4e6 --- /dev/null +++ b/docs/ai-coder/ai-gateway/standalone.md @@ -0,0 +1,319 @@ +# Deploy AI Gateway as a standalone service + +> [!NOTE] +> AI Gateway requires the [AI Governance Add-On](../ai-governance.md). +> As of Coder v2.32, deployments without the add-on will not be able to +> access AI Gateway. + +When AI traffic needs dedicated compute, independent scaling, or a separate network endpoint, you can deploy AI Gateway separately from the Coder control plane (`coderd`). + +A standalone AI Gateway serves client traffic on its own listener and maintains a control connection to `coderd`. +`coderd` continues to manage authentication, authorization, provider configuration, and AI session records. + +## Before you begin + +Standalone AI Gateway requires: + +- Coder v2.36.0 or later. +- A Coder license with the [AI Governance Add-On](../ai-governance.md). +- AI Gateway enabled on the Coder control plane with `CODER_AI_GATEWAY_ENABLED=true` or `--ai-gateway-enabled=true`. +- The full Coder image or a Coder binary that includes the `coder ai-gateway start` command. + +`coder ai-gateway start` does not read `CODER_AI_GATEWAY_ENABLED` from the standalone process. +However, this setting must remain enabled on `coderd` for Gateway key management and standalone control connections. + +## Create a Gateway key + +Each standalone replica uses a Gateway key to authenticate and establish its control connection to `coderd`. +Gateway key management requires site-level `ai_gateway_key` permissions, which only the built-in Owner role includes. +Coder custom roles are organization-scoped and cannot grant site-level permissions. + +Log in to the Coder CLI as an Owner or another user with these permissions, then create a dedicated key for the standalone deployment: + +```sh +coder login https://coder.example.com +coder ai-gateway keys create standalone-production +``` + +Save the key from the command output immediately because Coder does not display the plaintext value again. + +A single key can authenticate multiple replicas in the same deployment. +For independent rotation and revocation, use a separate key for each standalone deployment or environment. + +## Start a standalone process + +Set the Coder URL, Gateway key, and listener address, then start the Gateway: + +```sh +export CODER_URL=https://coder.example.com +export CODER_AI_GATEWAY_KEY='' +export CODER_AI_GATEWAY_HTTP_ADDRESS=0.0.0.0:4001 +coder ai-gateway start +``` + +Use `CODER_AI_GATEWAY_KEY_FILE` instead of `CODER_AI_GATEWAY_KEY` to read the key from a file. +The standalone process does not require a user login or `CODER_SESSION_TOKEN` after you provide the Gateway key. + +The listener defaults to `127.0.0.1:4001`, which accepts connections only from the local host. +Set `CODER_AI_GATEWAY_HTTP_ADDRESS` to a routable address, as shown above, before other hosts or pods can reach the Gateway. + +The standalone Gateway fetches provider configuration from `coderd`. +Configure at least one [AI provider](./providers.md) in Coder before sending provider traffic through the Gateway. +The standalone Gateway does not use the deprecated [provider seed variables](./providers.md#database-management-of-providers). + +The listener uses HTTP by default. +Set both `CODER_AI_GATEWAY_TLS_CERT_FILE` and `CODER_AI_GATEWAY_TLS_KEY_FILE` to terminate TLS in the process. +For all command options, refer to [`coder ai-gateway start`](../../reference/cli/ai-gateway_start.md). + +## Run AI Gateway in Kubernetes + +To run the standalone Gateway as a Kubernetes workload, provide the same environment variables and network access to the Coder URL. +You can manage the workload with your own Kubernetes manifests or use the provided Helm chart. +The chart configures the Deployment, probes, and Service, plus an optional Ingress or `HTTPRoute`. + +To use the Kubernetes examples, install `kubectl` and configure access to the cluster where AI Gateway will run. +Install Helm if you use the chart. + +Create a namespace and store the Gateway key in a Kubernetes Secret: + +```sh +kubectl create namespace coder-ai-gateway +kubectl create secret generic coder-ai-gateway-key \ + --namespace coder-ai-gateway \ + --from-literal=key='' +``` + +### Configure the Helm chart + +Create `ai-gateway-values.yaml` with at least the Coder URL and key Secret: + +```yaml +coder: + env: + - name: CODER_URL + value: https://coder.example.com + +aigateway: + keySecret: + name: coder-ai-gateway-key +``` + +The Gateway must be able to reach the Coder URL from every replica. +For an HTTPS URL signed by a private CA, set `aigateway.coderTLS.caSecret.name` to the Secret holding the CA bundle. +If Coder requires client mTLS, also set `aigateway.coderTLS.clientSecret.name`. + +The chart defaults to one replica, a `ClusterIP` Service, and a data-plane listener on port 4001. +It also enables a separate Prometheus listener on port 2112. +Ingress and `HTTPRoute` are disabled by default. + +For the complete chart configuration, including private CAs, mTLS, Ingress, Kubernetes Gateway API, and additional manifests, refer to the [AI Gateway Helm chart README](../../../helm/ai-gateway/README.md). + +### Install the Helm chart + +Install the chart version that matches the Coder control plane from GitHub Container Registry: + +```sh +helm install ai-gateway \ + oci://ghcr.io/coder/chart/coder-ai-gateway \ + --namespace coder-ai-gateway \ + --values ai-gateway-values.yaml \ + --version '' +``` + +Chart versions omit the leading `v`, so use `2.36.0` rather than `v2.36.0`. + +When you install the chart directly from a Git checkout, set `coder.image.tag` explicitly and install `./helm/ai-gateway`. +Released chart packages select the matching Coder image by default. +If you need to run different `coderd` and Gateway versions, refer to [Version compatibility](#version-compatibility). + +### Validate the Helm deployment + +Wait for the Deployment to become available: + +```sh +kubectl rollout status deployment/coder-ai-gateway \ + --namespace coder-ai-gateway \ + --timeout=5m +``` + +The chart configures the following probes by default: + +- The liveness probe requests `/healthz` and verifies that the HTTP listener is serving. +- The readiness probe requests `/readyz` and verifies that the Gateway is connected to `coderd` and has completed its initial provider load. +- The startup probe is disabled by default. + +`/healthz` can return HTTP 200 while `/readyz` returns HTTP 503. +This occurs while the initial connection to `coderd` is being established, before the initial provider load, and whenever the control connection drops. +A replica can only become ready after its initial provider load succeeds, and it returns to ready whenever the control connection recovers. +Kubernetes removes an unready replica from the Service, so clients normally stop reaching it. +A request that still arrives waits for the control connection instead of failing. + +Port-forward the Service to inspect both endpoints: + +```sh +kubectl port-forward \ + --namespace coder-ai-gateway \ + service/coder-ai-gateway \ + 4001:80 +``` + +In another terminal, run: + +```sh +curl --fail http://127.0.0.1:4001/healthz +curl --fail http://127.0.0.1:4001/readyz +``` + +Both endpoints return HTTP 200 with an empty body when the replica is serving and ready. + +List Gateway keys and verify that the key has a recent heartbeat: + +```sh +coder ai-gateway keys list +``` + +The `LAST HEARTBEAT AT` column holds the timestamp. +The first heartbeat is recorded when the replica connects. +An active control connection updates the timestamp every 60 seconds. +Coder stores one timestamp per key, so a recent heartbeat on a shared key does not confirm that every replica is connected. +Check `/readyz` on each replica to verify individual health. + +## Route traffic to the standalone Gateway + +The standalone deployment does not move traffic automatically. +Configure AI Gateway Proxy and direct AI clients to send requests to the standalone endpoint. + +### AI Gateway Proxy + +[AI Gateway Proxy](./ai-gateway-proxy/index.md) remains part of the `coder server` process. +Configure its target to use the standalone Service, Ingress, `HTTPRoute`, or load balancer. + +If you installed the Helm chart, get the exact in-cluster Service URL from the chart notes: + +```sh +helm get notes ai-gateway --namespace coder-ai-gateway +``` + +If you use your own Kubernetes manifests, use the URL of the Service you created. +Set the Service URL on the Coder control plane, for example: + +```sh +CODER_AI_GATEWAY_PROXY_TARGET=http://coder-ai-gateway.coder-ai-gateway.svc.cluster.local:80 +``` + +Restart or upgrade the Coder deployment after changing this setting. +The proxy appends the provider name and request path to the target, so do not include query parameters. + +### Direct AI clients + +Clients that use `/api/v2/ai-gateway//` continue to reach the embedded Gateway. +To route these clients to the standalone Gateway, replace the Coder access URL and embedded route prefix with the standalone endpoint while keeping the provider path. + +For example, if the Gateway is reachable at `https://ai-gateway.example.com` and the providers are configured, configure Claude Code with: + +```sh +ANTHROPIC_BASE_URL=https://ai-gateway.example.com/anthropic +``` + +Configure an OpenAI-compatible client with: + +```sh +OPENAI_BASE_URL=https://ai-gateway.example.com/openai/v1 +``` + +The standalone listener also accepts the equivalent `/api/v2/ai-gateway//` paths for compatibility. +For full per-client configuration examples, refer to [Client Configuration](./clients/index.md). + +### Expose the standalone AI Gateway + +If clients cannot access the in-cluster Service, expose the Gateway through an Ingress, an `HTTPRoute`, or a load balancer Service. +In the provided Helm chart, set: + +```yaml +service: + type: LoadBalancer +``` + +Use TLS to protect credentials and AI traffic whenever they cross an untrusted network. +Prefer TLS termination at the Ingress or Kubernetes Gateway when your platform already manages certificates there. +To terminate TLS in the AI Gateway process, set `aigateway.listenerTLS.name` to an existing TLS Secret. + +If `CODER_AI_GATEWAY_PROXY_TARGET` uses HTTPS with a private CA, add that CA to the trust store of the `coderd` pods before changing the proxy target. +For the Coder Helm chart, you can mount a CA bundle with `coder.certs.secrets`. +This trust configuration is separate from `aigateway.coderTLS.caSecret`, which configures the connection from the standalone Gateway to `coderd`. + +## Scale replicas + +Run multiple replicas behind a Service or load balancer. +In the provided Helm chart, set: + +```yaml +coder: + replicaCount: 3 +``` + +The Service balances requests across ready replicas. +Each replica maintains its own control connection, local caches, concurrency limits, request metrics, logs, and traces. +All replicas fetch provider configuration from the same Coder deployment and write AI session records through `coderd`. +Cost control metrics are registered only on `coderd`, so `coder_ai_gateway_cost_control_*` series never appear on a replica's Prometheus listener. +Budget enforcement also runs in `coderd` over the control connection rather than per replica. +Sticky load balancing can improve cache efficiency, but it is not required for correctness. + +When you enable [API dumps](./setup.md#api-dumps), each replica writes dumps to its own local disk. +Use persistent storage if you need those files to survive pod replacement. + +The chart requests 1 CPU and 1 GiB of memory per replica and does not set resource limits. +Measure production traffic before changing requests, limits, or `CODER_AI_GATEWAY_MAX_CONCURRENCY`. +The chart does not create a Horizontal Pod Autoscaler or PodDisruptionBudget. + +## Migrate from the embedded Gateway + +Use a gradual cutover so that you can validate the standalone data plane before moving production traffic: + +1. Upgrade `coderd` to the target Coder version. +1. Keep `CODER_AI_GATEWAY_ENABLED=true` on `coderd`. +1. Create a dedicated Gateway key. +1. Deploy one standalone replica without changing client or proxy routing. +1. Verify `/readyz`, the key heartbeat, metrics, logs, and traces. +1. Send a test request directly to the standalone Gateway and confirm that the AI session appears in Coder. +1. Set `CODER_AI_GATEWAY_PROXY_TARGET` to the standalone endpoint. +1. Update direct client base URLs that should use the standalone endpoint. +1. Scale the standalone deployment after the canary path is stable. + +Keep `CODER_AI_GATEWAY_ENABLED=true` on `coderd` after the cutover. +The setting is required for standalone control connections and Coder features that use the in-process Gateway. + +## Roll back to the embedded Gateway + +To return traffic to the embedded Gateway: + +1. Remove `CODER_AI_GATEWAY_PROXY_TARGET`, or set it to `/api/v2/ai-gateway`. +1. Restart or upgrade `coderd` so the proxy target change takes effect. +1. Restore direct client base URLs to `/api/v2/ai-gateway//`. +1. Verify requests and AI session records through the embedded route. +1. Scale down or uninstall the standalone deployment. +1. Delete the standalone Gateway key after all replicas have stopped. + +Delete the key with: + +```sh +coder ai-gateway keys delete standalone-production +``` + +Deleting a key prevents new connections immediately. +An established connection closes when its next heartbeat detects the deletion, within 60 seconds. +The replica then tries to reconnect, receives HTTP 401, and treats that as fatal: the process exits non-zero rather than retrying. +Stop the deployment before deleting its key. + +## Version compatibility + +The standalone Gateway image does not need to match the `coderd` image version exactly. +The components can connect when their AI Gateway API versions are compatible, and `coderd` checks compatibility whenever a Gateway replica connects. + +A replica may run at the same API version as `coderd` or an earlier minor version of the same major version, but never a newer one. +Sequence changes that move both components: + +- To upgrade, upgrade `coderd` first, then roll out the standalone Gateway release. +- To roll back, roll back the standalone Gateway first, then roll back `coderd`. + +Running the same Coder release on `coderd` and every standalone replica is the simplest way to stay compatible, but it is not required. diff --git a/docs/install/kubernetes.md b/docs/install/kubernetes.md index 8954a2b131..80a6fa6cd9 100644 --- a/docs/install/kubernetes.md +++ b/docs/install/kubernetes.md @@ -197,6 +197,14 @@ helm upgrade coder coder-v2/coder \ -f values.yaml ``` +## Standalone AI Gateway Chart + +Coder also publishes a chart, `oci://ghcr.io/coder/chart/coder-ai-gateway`, +that runs [AI Gateway](../ai-coder/ai-gateway/index.md) as its own Deployment +alongside the control plane. Use it when you want to scale AI traffic +independently of `coderd`. For installation and configuration, visit +[Standalone AI Gateway](../ai-coder/ai-gateway/standalone.md). + ## Coder Observability Chart Use the [Observability Helm chart](https://github.com/coder/observability) for a diff --git a/docs/manifest.json b/docs/manifest.json index 1de411ee04..56e75d4cd2 100644 --- a/docs/manifest.json +++ b/docs/manifest.json @@ -1195,7 +1195,7 @@ }, { "title": "AI Gateway", - "description": "Govern and observe AI coding traffic with AI Gateway, an LLM proxy built into the Coder control plane.", + "description": "Govern and observe AI coding traffic with AI Gateway.", "path": "./ai-coder/ai-gateway/index.md", "icon_path": "./images/icons/api.svg", "state": ["premium"], @@ -1206,6 +1206,12 @@ "path": "./ai-coder/ai-gateway/setup.md", "state": ["premium"] }, + { + "title": "Standalone Deployment", + "description": "Deploy and operate AI Gateway as a standalone service", + "path": "./ai-coder/ai-gateway/standalone.md", + "state": ["ai governance add-on"] + }, { "title": "Authentication", "description": "Authenticate AI coding tools against AI Gateway using your Coder API token.", @@ -1331,7 +1337,7 @@ }, { "title": "Reference", - "description": "Technical reference for AI Gateway, including the aibridged component that runs inside coderd.", + "description": "Technical reference for AI Gateway, including deployment topologies, version compatibility, and supported APIs.", "path": "./ai-coder/ai-gateway/reference.md", "state": ["premium"] }, diff --git a/helm/ai-gateway/README.md b/helm/ai-gateway/README.md index 40a5643f9c..a1bce8edcb 100644 --- a/helm/ai-gateway/README.md +++ b/helm/ai-gateway/README.md @@ -3,7 +3,7 @@ This chart deploys the Coder AI Gateway as a standalone Kubernetes Deployment. The Gateway connects to Coder using `CODER_URL` and an AI Gateway key. To forward proxied AI traffic to the standalone Gateway, configure the Coder AI Gateway -Proxy (`aibridgeproxyd`) after installing the chart. +Proxy after installing the chart. The chart does not create credentials or TLS Secrets. diff --git a/scripts/metricsdocgen/metrics b/scripts/metricsdocgen/metrics index 70ae90436c..6464f6ff94 100644 --- a/scripts/metricsdocgen/metrics +++ b/scripts/metricsdocgen/metrics @@ -195,7 +195,7 @@ coder_ai_gateway_interceptions_total{client="Codex",initiator_id="95f6752b-08cc- coder_ai_gateway_key_pool_state{provider="openai",state="valid"} 2 coder_ai_gateway_key_pool_state{provider="openai",state="temporary"} 0 coder_ai_gateway_key_pool_state{provider="openai",state="permanent"} 0 -# HELP coder_ai_gateway_key_pool_state_transitions_total The number of API key state transitions during failover (reason: rate_limited, unauthorized, forbidden). +# HELP coder_ai_gateway_key_pool_state_transitions_total The number of API key state transitions during failover (reason: rate_limited, unauthorized). # TYPE coder_ai_gateway_key_pool_state_transitions_total counter coder_ai_gateway_key_pool_state_transitions_total{provider="openai",reason="rate_limited"} 1 # HELP coder_ai_gateway_key_pool_exhaustions_total The number of times the key pool was exhausted with no usable key (outcome: rate_limited, auth_failed).