docs: add standalone AI Gateway docs (#27592)

Documents standalone AI Gateway deployment, Gateway key authentication,
monitoring, and the updated embedded vs standalone topology in the AI
Gateway docs.

---------

Co-authored-by: Cian Johnston <cian@coder.com>
This commit is contained in:
Paweł Banaszewski
2026-07-29 19:38:22 +02:00
committed by GitHub
co-authored by Cian Johnston
parent 6c42309ccb
commit 18128b7b52
18 changed files with 773 additions and 136 deletions
+10 -1
View File
@@ -72,6 +72,13 @@ scrape_configs:
apps: "coder"
```
If you run a [standalone AI Gateway](../../ai-coder/ai-gateway/standalone.md),
each replica exports its own metrics on its own listener. Its Helm chart uses the
same `0.0.0.0:2112` default as the `coder` chart, but sets up no scrape
discovery. Refer to
[AI Gateway monitoring](../../ai-coder/ai-gateway/monitoring.md#kubernetes-discovery)
for more details.
To use the Kubernetes Prometheus operator to scrape metrics, you will need to
create a `ServiceMonitor` in your Coder deployment namespace. The following is
an example `ServiceMonitor`.
@@ -102,6 +109,8 @@ You must first enable `coderd_agentstats_*` with the flag
`CODER_PROMETHEUS_COLLECT_AGENT_STATS` before they can be retrieved from the
deployment. They will always be available from the agent.
The `coder_ai_gateway_cost_control_*` metrics are exported only by `coderd`.
<!-- Code generated by 'make docs/admin/integrations/prometheus.md'. DO NOT EDIT -->
| Name | Type | Description | Labels |
@@ -124,7 +133,7 @@ deployment. They will always be available from the agent.
| `coder_ai_gateway_key_pool_exhaustions_total` | counter | The number of times the key pool was exhausted with no usable key (outcome: rate_limited, auth_failed). | `outcome` `provider` |
| `coder_ai_gateway_key_pool_failover_attempts` | histogram | The number of keys attempted before success or exhaustion, per interception for bridged requests and per request for passthrough requests. | `provider` |
| `coder_ai_gateway_key_pool_state` | gauge | The number of keys currently in each state (state: valid, temporary, permanent). | `provider` `state` |
| `coder_ai_gateway_key_pool_state_transitions_total` | counter | The number of API key state transitions during failover (reason: rate_limited, unauthorized, forbidden). | `provider` `reason` |
| `coder_ai_gateway_key_pool_state_transitions_total` | counter | The number of API key state transitions during failover (reason: rate_limited, unauthorized). | `provider` `reason` |
| `coder_ai_gateway_non_injected_tool_selections_total` | counter | The number of times an AI model selected a tool to be invoked by the client. | `model` `name` `provider` |
| `coder_ai_gateway_passthrough_total` | counter | The count of requests which were not intercepted but passed through to the upstream. | `method` `provider` `route` |
| `coder_ai_gateway_prompts_total` | counter | The number of prompts issued by users (initiators). | `client` `initiator_id` `model` `provider` |