mirror of
https://github.com/coder/coder.git
synced 2026-09-24 15:04:27 +08:00
feat: remove native chat usage limits in favor of AI Gateway budgets (#27329)
## Stack Context This stack makes AI Gateway data and budgets the source of truth for AI spend controls. 1. Re-back the per-chat cost endpoint with AI Gateway data (#27328, merged). 2. **This PR:** remove native chat usage limits. 3. Remove native chat cost tracking and its dedicated admin UI (#27330). ## Summary Removes the native usage-limit API, SDK types, SQL, and chat enforcement for deployment, user, and group chat limits. Compact AI Gateway budget indicators remain in the Agents sidebar, user menu, and group settings. Gateway budget rejections and provider quota failures continue to classify as usage-limit errors, including a 409 response for synchronous title generation. Budget-period labels now use the API's UTC boundaries, so users see the same dates in every browser timezone. The documentation explains the AI Gateway replacement, its licensing requirements, and the differences from native limits. No schema is dropped in this release. The usage-limit table, index, user and group columns, constraints, audit mappings, and generated scan fields remain for mixed-version rolling upgrades. #27600 tracks their removal after the compatibility window. ## Breaking change Native day, week, and month chat spend limits are removed and are not migrated. AI Gateway budgets are month-based, group-scoped with per-user overrides, and require the AI Gateway entitlement. Deployments without that entitlement no longer have chat spend enforcement. > Mux prepared this PR on Mike's behalf.
This commit is contained in:
@@ -39,9 +39,8 @@ Once the experiment is enabled, configure the advisor under **AI Settings** >
|
||||
| Advisor model | Use chat model | Optional dedicated chat model config for the advisor. When unset, the advisor reuses the root agent's model. |
|
||||
| Reasoning effort | Model default | Overrides the selected advisor model's reasoning effort. Available only when the model supports selectable effort. |
|
||||
|
||||
The advisor is not available in plan mode or to subagents. Failed advisor
|
||||
invocations refund the per-turn budget, and advisor calls are not metered
|
||||
against the root chat's usage limit.
|
||||
The advisor is not available in plan mode or to subagents.
|
||||
Failed advisor invocations refund the per-turn budget.
|
||||
|
||||
The same configuration is available at:
|
||||
|
||||
|
||||
@@ -116,12 +116,15 @@ none, if the template does not define any).
|
||||
|
||||
### Spend management
|
||||
|
||||
Administrators can set spend limits to cap LLM usage per user within a rolling
|
||||
time period, with per-user and per-group overrides. The cost tracking dashboard
|
||||
provides visibility into per-user spending, token consumption, and per-model
|
||||
breakdowns.
|
||||
AI Gateway budgets cap each user's AI spend, including Coder Agents chats, over a monthly period.
|
||||
Coder sets budgets per group, and the deployment policy selects the group with the largest spend limit when a user belongs to several budgeted groups.
|
||||
A per-user override takes priority over all group budgets.
|
||||
|
||||
See [Spend Management](./usage-insights.md) for details.
|
||||
Budgets are the only spend cap for Coder Agents chats.
|
||||
Chats no longer enforce a separate limit of their own, and existing native limit values are not migrated to budgets.
|
||||
Budget controls in the Coder UI, the group budget endpoints, and the AI spend status endpoints all require a license that includes AI Gateway.
|
||||
|
||||
Refer to [Spend Management](./usage-insights.md) for details.
|
||||
|
||||
### Git providers
|
||||
|
||||
@@ -158,10 +161,8 @@ For chat debug logging (not experiment-gated), see [Chat debug logging](./chat-d
|
||||
|
||||
## Where we are headed
|
||||
|
||||
The controls above cover providers, models, system prompts, templates, MCP
|
||||
servers, usage limits, and data retention. We are continuing to invest in platform controls
|
||||
based on what we hear from customers deploying agents in regulated and
|
||||
enterprise environments.
|
||||
The controls above cover providers, models, system prompts, templates, MCP servers, AI Gateway budgets, and data retention.
|
||||
We are continuing to invest in platform controls based on what feedback we get from customers deploying agents in regulated and enterprise environments.
|
||||
|
||||
### Infrastructure-level enforcement
|
||||
|
||||
|
||||
@@ -1,90 +1,55 @@
|
||||
# Spend Management
|
||||
|
||||
Coder provides admin-only controls for monitoring and controlling agent
|
||||
spend: usage limits and cost tracking.
|
||||
Coder controls agent spend with AI Gateway budgets, and surfaces the resulting spend to both admins and users.
|
||||
|
||||
## Usage limits
|
||||
## Budgets
|
||||
|
||||
Navigate to **Agents** > **Settings** > **Manage Agents** > **Spend**.
|
||||
Coder Agents spend is controlled by AI Gateway budgets, which cap all AI Gateway usage (including Coder Agents chats) per user over the budget period.
|
||||
|
||||
Usage limits cap how much each user can spend on LLM usage within a rolling
|
||||
time period. When enabled, the system checks the user's current spend before
|
||||
processing each chat message.
|
||||
- **Group budgets**: set a budget for a group from the group's settings page.
|
||||
The deployment budget policy resolves which group budget applies when a user belongs to multiple budgeted groups, which defaults to the group with the highest budget.
|
||||
- **Per-user overrides**: set a custom budget for an individual user, attributed to one of their groups.
|
||||
Per-user overrides take priority over group budgets.
|
||||
|
||||
### Configuration
|
||||
The deployment flags `--ai-budget-policy` and `--ai-budget-period` currently
|
||||
support only `highest` and `month`. `highest` selects the group with the
|
||||
largest spend limit, and `month` resets spend at the start of each UTC
|
||||
calendar month. A user who belongs to no budgeted group falls back to the
|
||||
Everyone group, which has no budget unless one is set for it. There is no
|
||||
deployment-wide budget amount, and a configured spend limit cannot exceed
|
||||
$1,000,000 per member per period.
|
||||
|
||||
- **Enable/disable toggle** — master on/off for the entire limit system.
|
||||
- **Period** — `day`, `week`, or `month`. Periods are UTC-aligned: midnight
|
||||
UTC for daily, Monday start for weekly, first of the month for monthly.
|
||||
- **Default limit** — deployment-wide default in dollars. Applies to all
|
||||
users who do not have a more specific override. Leave unset for no limit.
|
||||
- **Per-user overrides** — set a custom dollar limit for an individual user.
|
||||
Takes highest priority.
|
||||
- **Per-group overrides** — set a limit for a group. When a user belongs to
|
||||
multiple groups, the lowest group limit applies.
|
||||
> [!IMPORTANT]
|
||||
> Budget controls in the Coder UI, the group budget endpoints (`/api/v2/groups/{group}/ai/budget`), and the AI spend status and reporting endpoints all require the AI Gateway entitlement.
|
||||
> No experiment is needed.
|
||||
>
|
||||
> Native chat usage limits are removed from the application.
|
||||
> Existing native limit values are not migrated to AI Gateway budgets and are no longer enforced.
|
||||
> Configure AI Gateway budgets separately.
|
||||
|
||||
### Priority hierarchy
|
||||
|
||||
The system resolves a user's effective limit in this order:
|
||||
|
||||
1. Individual user override (highest priority)
|
||||
1. Minimum group limit across all of the user's groups
|
||||
1. Global default limit
|
||||
1. No limit (if limits are disabled or no value is configured)
|
||||
The API reference documents how to [get](../../../reference/api/enterprise.md#get-group-ai-budget), [upsert](../../../reference/api/enterprise.md#upsert-group-ai-budget), and [delete](../../../reference/api/enterprise.md#delete-group-ai-budget) a group budget.
|
||||
|
||||
### Enforcement
|
||||
|
||||
- Checked before each chat message is processed.
|
||||
- When current spend meets or exceeds the limit, the chat returns a
|
||||
**409 Conflict** response and the message is blocked.
|
||||
- Fail-open: if the limit query itself fails, the message is allowed
|
||||
through.
|
||||
- Brief overage is possible when concurrent messages are in flight, because
|
||||
cost is determined only after the LLM returns.
|
||||
- The AI Gateway checks the user's current spend before forwarding each request.
|
||||
When spend meets or exceeds the budget, the request is rejected and the chat shows a terminal error explaining that the budget was exceeded.
|
||||
- Brief overage is possible when concurrent requests are in flight, because cost is recorded after the LLM responds.
|
||||
|
||||
### User-facing status
|
||||
|
||||
Users can view their own spend status, including whether a limit is active,
|
||||
their effective limit, current spend, and when the current period resets.
|
||||
The usage indicator on the Agents page and the summary in the user menu both show the signed-in user's current AI spend, their budget, and the period reset date.
|
||||
Both appear only when the deployment has the AI Gateway entitlement.
|
||||
|
||||
## Spend visibility
|
||||
|
||||
Spend is shown where it is actionable:
|
||||
|
||||
- **Agents page and user menu**: the signed-in user's spend against their budget, as described previously.
|
||||
- **Group settings**: each member's spend against the group's budget, for admins who can manage the group.
|
||||
- **Chat summary panel**: the cost of one chat tree, on a chat's Summary tab.
|
||||
A subagent reports the total for its whole tree, including the chat that started it.
|
||||
- **Agents** > **Settings** > **Manage Agents** > **Spend**: deployment-wide chat cost per user, with per-user drill-down.
|
||||
|
||||
> [!NOTE]
|
||||
> The admin configuration page shows the count of models without pricing
|
||||
> data. Models missing pricing cannot be tracked accurately against limits.
|
||||
|
||||
## Cost tracking
|
||||
|
||||
Navigate to **Agents** > **Settings** > **Manage Agents** > **Spend**.
|
||||
|
||||
This view shows deployment-wide LLM chat costs with per-user drill-down.
|
||||
|
||||
### Top-level view
|
||||
|
||||
A per-user rollup table with the following columns:
|
||||
|
||||
| Column | Description |
|
||||
|--------------------|-------------------------------------|
|
||||
| Total cost | Aggregate dollar spend for the user |
|
||||
| Messages | Number of chat messages sent |
|
||||
| Chats | Number of distinct chat sessions |
|
||||
| Input tokens | Total input tokens consumed |
|
||||
| Output tokens | Total output tokens consumed |
|
||||
| Cache read tokens | Tokens served from cache |
|
||||
| Cache write tokens | Tokens written to cache |
|
||||
|
||||
The table supports date range filtering (default: last 30 days), search by
|
||||
name or username, and pagination.
|
||||
|
||||
### Per-user detail view
|
||||
|
||||
Select a user to see:
|
||||
|
||||
- **Summary cards** — total cost, token breakdowns, and message counts.
|
||||
- **Usage limit progress** — if a limit is active, a color-coded progress
|
||||
bar shows current spend relative to the limit.
|
||||
- **Per-model breakdown** — table of costs and token usage by model.
|
||||
- **Per-chat breakdown** — table of costs and token usage by chat session.
|
||||
|
||||
> [!NOTE]
|
||||
> Automatic title generation uses lightweight models, such as Claude Haiku or GPT-4o
|
||||
> Mini. Its token usage is not counted towards usage limits or shown in usage
|
||||
> summaries.
|
||||
> Per-chat cost comes from AI Gateway records, which are pruned according to `--ai-gateway-retention` (60 days by default).
|
||||
> A chat for which gateway records have been pruned reports no cost.
|
||||
|
||||
Reference in New Issue
Block a user