feat: remove native chat usage limits in favor of AI Gateway budgets (#27329)

## Stack Context

This stack makes AI Gateway data and budgets the source of truth for AI
spend controls.

1. Re-back the per-chat cost endpoint with AI Gateway data (#27328,
merged).
2. **This PR:** remove native chat usage limits.
3. Remove native chat cost tracking and its dedicated admin UI (#27330).

## Summary

Removes the native usage-limit API, SDK types, SQL, and chat enforcement
for deployment, user, and group chat limits. Compact AI Gateway budget
indicators remain in the Agents sidebar, user menu, and group settings.
Gateway budget rejections and provider quota failures continue to
classify as usage-limit errors, including a 409 response for synchronous
title generation.

Budget-period labels now use the API's UTC boundaries, so users see the
same dates in every browser timezone. The documentation explains the AI
Gateway replacement, its licensing requirements, and the differences
from native limits.

No schema is dropped in this release. The usage-limit table, index, user
and group columns, constraints, audit mappings, and generated scan
fields remain for mixed-version rolling upgrades. #27600 tracks their
removal after the compatibility window.

## Breaking change

Native day, week, and month chat spend limits are removed and are not
migrated. AI Gateway budgets are month-based, group-scoped with per-user
overrides, and require the AI Gateway entitlement. Deployments without
that entitlement no longer have chat spend enforcement.

> Mux prepared this PR on Mike's behalf.
This commit is contained in:
Michael Suchacz
2026-08-04 11:36:49 +02:00
committed by GitHub
parent a2287d6739
commit f0e6ac64b3
78 changed files with 611 additions and 6850 deletions
@@ -39,9 +39,8 @@ Once the experiment is enabled, configure the advisor under **AI Settings** >
| Advisor model | Use chat model | Optional dedicated chat model config for the advisor. When unset, the advisor reuses the root agent's model. |
| Reasoning effort | Model default | Overrides the selected advisor model's reasoning effort. Available only when the model supports selectable effort. |
The advisor is not available in plan mode or to subagents. Failed advisor
invocations refund the per-turn budget, and advisor calls are not metered
against the root chat's usage limit.
The advisor is not available in plan mode or to subagents.
Failed advisor invocations refund the per-turn budget.
The same configuration is available at:
@@ -116,12 +116,15 @@ none, if the template does not define any).
### Spend management
Administrators can set spend limits to cap LLM usage per user within a rolling
time period, with per-user and per-group overrides. The cost tracking dashboard
provides visibility into per-user spending, token consumption, and per-model
breakdowns.
AI Gateway budgets cap each user's AI spend, including Coder Agents chats, over a monthly period.
Coder sets budgets per group, and the deployment policy selects the group with the largest spend limit when a user belongs to several budgeted groups.
A per-user override takes priority over all group budgets.
See [Spend Management](./usage-insights.md) for details.
Budgets are the only spend cap for Coder Agents chats.
Chats no longer enforce a separate limit of their own, and existing native limit values are not migrated to budgets.
Budget controls in the Coder UI, the group budget endpoints, and the AI spend status endpoints all require a license that includes AI Gateway.
Refer to [Spend Management](./usage-insights.md) for details.
### Git providers
@@ -158,10 +161,8 @@ For chat debug logging (not experiment-gated), see [Chat debug logging](./chat-d
## Where we are headed
The controls above cover providers, models, system prompts, templates, MCP
servers, usage limits, and data retention. We are continuing to invest in platform controls
based on what we hear from customers deploying agents in regulated and
enterprise environments.
The controls above cover providers, models, system prompts, templates, MCP servers, AI Gateway budgets, and data retention.
We are continuing to invest in platform controls based on what feedback we get from customers deploying agents in regulated and enterprise environments.
### Infrastructure-level enforcement
@@ -1,90 +1,55 @@
# Spend Management
Coder provides admin-only controls for monitoring and controlling agent
spend: usage limits and cost tracking.
Coder controls agent spend with AI Gateway budgets, and surfaces the resulting spend to both admins and users.
## Usage limits
## Budgets
Navigate to **Agents** > **Settings** > **Manage Agents** > **Spend**.
Coder Agents spend is controlled by AI Gateway budgets, which cap all AI Gateway usage (including Coder Agents chats) per user over the budget period.
Usage limits cap how much each user can spend on LLM usage within a rolling
time period. When enabled, the system checks the user's current spend before
processing each chat message.
- **Group budgets**: set a budget for a group from the group's settings page.
The deployment budget policy resolves which group budget applies when a user belongs to multiple budgeted groups, which defaults to the group with the highest budget.
- **Per-user overrides**: set a custom budget for an individual user, attributed to one of their groups.
Per-user overrides take priority over group budgets.
### Configuration
The deployment flags `--ai-budget-policy` and `--ai-budget-period` currently
support only `highest` and `month`. `highest` selects the group with the
largest spend limit, and `month` resets spend at the start of each UTC
calendar month. A user who belongs to no budgeted group falls back to the
Everyone group, which has no budget unless one is set for it. There is no
deployment-wide budget amount, and a configured spend limit cannot exceed
$1,000,000 per member per period.
- **Enable/disable toggle** — master on/off for the entire limit system.
- **Period** — `day`, `week`, or `month`. Periods are UTC-aligned: midnight
UTC for daily, Monday start for weekly, first of the month for monthly.
- **Default limit** — deployment-wide default in dollars. Applies to all
users who do not have a more specific override. Leave unset for no limit.
- **Per-user overrides** — set a custom dollar limit for an individual user.
Takes highest priority.
- **Per-group overrides** — set a limit for a group. When a user belongs to
multiple groups, the lowest group limit applies.
> [!IMPORTANT]
> Budget controls in the Coder UI, the group budget endpoints (`/api/v2/groups/{group}/ai/budget`), and the AI spend status and reporting endpoints all require the AI Gateway entitlement.
> No experiment is needed.
>
> Native chat usage limits are removed from the application.
> Existing native limit values are not migrated to AI Gateway budgets and are no longer enforced.
> Configure AI Gateway budgets separately.
### Priority hierarchy
The system resolves a user's effective limit in this order:
1. Individual user override (highest priority)
1. Minimum group limit across all of the user's groups
1. Global default limit
1. No limit (if limits are disabled or no value is configured)
The API reference documents how to [get](../../../reference/api/enterprise.md#get-group-ai-budget), [upsert](../../../reference/api/enterprise.md#upsert-group-ai-budget), and [delete](../../../reference/api/enterprise.md#delete-group-ai-budget) a group budget.
### Enforcement
- Checked before each chat message is processed.
- When current spend meets or exceeds the limit, the chat returns a
**409 Conflict** response and the message is blocked.
- Fail-open: if the limit query itself fails, the message is allowed
through.
- Brief overage is possible when concurrent messages are in flight, because
cost is determined only after the LLM returns.
- The AI Gateway checks the user's current spend before forwarding each request.
When spend meets or exceeds the budget, the request is rejected and the chat shows a terminal error explaining that the budget was exceeded.
- Brief overage is possible when concurrent requests are in flight, because cost is recorded after the LLM responds.
### User-facing status
Users can view their own spend status, including whether a limit is active,
their effective limit, current spend, and when the current period resets.
The usage indicator on the Agents page and the summary in the user menu both show the signed-in user's current AI spend, their budget, and the period reset date.
Both appear only when the deployment has the AI Gateway entitlement.
## Spend visibility
Spend is shown where it is actionable:
- **Agents page and user menu**: the signed-in user's spend against their budget, as described previously.
- **Group settings**: each member's spend against the group's budget, for admins who can manage the group.
- **Chat summary panel**: the cost of one chat tree, on a chat's Summary tab.
A subagent reports the total for its whole tree, including the chat that started it.
- **Agents** > **Settings** > **Manage Agents** > **Spend**: deployment-wide chat cost per user, with per-user drill-down.
> [!NOTE]
> The admin configuration page shows the count of models without pricing
> data. Models missing pricing cannot be tracked accurately against limits.
## Cost tracking
Navigate to **Agents** > **Settings** > **Manage Agents** > **Spend**.
This view shows deployment-wide LLM chat costs with per-user drill-down.
### Top-level view
A per-user rollup table with the following columns:
| Column | Description |
|--------------------|-------------------------------------|
| Total cost | Aggregate dollar spend for the user |
| Messages | Number of chat messages sent |
| Chats | Number of distinct chat sessions |
| Input tokens | Total input tokens consumed |
| Output tokens | Total output tokens consumed |
| Cache read tokens | Tokens served from cache |
| Cache write tokens | Tokens written to cache |
The table supports date range filtering (default: last 30 days), search by
name or username, and pagination.
### Per-user detail view
Select a user to see:
- **Summary cards** — total cost, token breakdowns, and message counts.
- **Usage limit progress** — if a limit is active, a color-coded progress
bar shows current spend relative to the limit.
- **Per-model breakdown** — table of costs and token usage by model.
- **Per-chat breakdown** — table of costs and token usage by chat session.
> [!NOTE]
> Automatic title generation uses lightweight models, such as Claude Haiku or GPT-4o
> Mini. Its token usage is not counted towards usage limits or shown in usage
> summaries.
> Per-chat cost comes from AI Gateway records, which are pruned according to `--ai-gateway-retention` (60 days by default).
> A chat for which gateway records have been pruned reports no cost.