feat: remove native chat cost tracking in favor of AI Gateway cost data (#27330)

## Stack Context

This stack makes AI Gateway data and budgets the source of truth for AI
spend controls.

1. Re-back the per-chat cost endpoint with AI Gateway data (#27328,
merged).
2. Remove native chat usage limits (#27329, merged).
3. **This PR, now based on `main`:** remove native chat cost tracking
and its dedicated admin UI.

## Summary

Removes native per-message price calculation, model pricing fields, cost
persistence, aggregate cost queries, and admin cost API types. It also
deletes the Analytics and Spend pages plus their legacy redirects. The
AI Gateway-backed per-chat cost row and compact budget indicators
remain.

The spend documentation is renamed to `spend-management.md` and updated
for the remaining surfaces, group budget APIs, CSV export, upgrade
handling for native pricing and cost history, and the absence of a
deployment-wide spend dashboard. The per-chat cost API documents that
data follows AI Gateway retention and reports zero after all matching
requests are purged.

No schema is dropped in this release. `chat_messages.total_cost_micros`
remains nullable and unwritten so replicas from the previous release can
continue inserting messages during rolling upgrades. #27600 tracks
removal after the compatibility window.

> Mux prepared this PR on Mike's behalf.
This commit is contained in:
Michael Suchacz
2026-08-04 12:27:38 +02:00
committed by GitHub
parent 0b8b48913f
commit 6b8f820493
72 changed files with 159 additions and 5303 deletions
+1 -2
View File
@@ -224,8 +224,7 @@ sub-agent delegation, and complex multi-step work can consume significant
token volume. Consider:
- Starting with a single model to establish a cost baseline.
- Setting per-model token pricing under **Admin settings** > **AI** >
**Models** (Input Price, Output Price) to track spend.
- Capping spend with [AI Gateway budgets](./platform-controls/spend-management.md).
- Monitoring provider dashboards for usage trends during the evaluation.
### Pilot with a small group
-4
View File
@@ -192,10 +192,6 @@ These options apply to all providers:
| Top K | Limits token selection to the top K candidates. |
| Presence Penalty | Penalizes tokens that have already appeared in the conversation. |
| Frequency Penalty | Penalizes tokens proportional to how often they have appeared. |
| Input Price | Optional USD price metadata for input tokens, recorded per 1M tokens. |
| Output Price | Optional USD price metadata for output tokens, recorded per 1M tokens. |
| Cache Read Price | Optional USD price metadata for cache read tokens, recorded per 1M tokens. |
| Cache Write Price | Optional USD price metadata for cache creation/write tokens, recorded per 1M tokens. |
### Provider-specific options
@@ -124,7 +124,7 @@ Budgets are the only spend cap for Coder Agents chats.
Chats no longer enforce a separate limit of their own, and existing native limit values are not migrated to budgets.
Budget controls in the Coder UI, the group budget endpoints, and the AI spend status endpoints all require a license that includes AI Gateway.
Refer to [Spend Management](./usage-insights.md) for details.
Refer to [Spend management](./spend-management.md) for details.
### Git providers
@@ -1,4 +1,4 @@
# Spend Management
# Spend management (Premium)
Coder controls agent spend with AI Gateway budgets, and surfaces the resulting spend to both admins and users.
@@ -23,8 +23,17 @@ $1,000,000 per member per period.
> Budget controls in the Coder UI, the group budget endpoints (`/api/v2/groups/{group}/ai/budget`), and the AI spend status and reporting endpoints all require the AI Gateway entitlement.
> No experiment is needed.
>
> Native chat usage limits are removed from the application.
> Native chat usage limits and native cost tracking are removed from the application.
> Existing native limit values are not migrated to AI Gateway budgets and are no longer enforced.
> Configured per-model prices and historical native cost totals are also not migrated to AI Gateway.
> Before upgrading, record any per-model prices you need from **Admin settings** > **AI** > **Models**.
> The old cost endpoints default `start_date` to 30 days before the request and `end_date` to the request time, so choose explicit RFC 3339 UTC values that cover all history you need.
> Fetch `/api/experimental/chats/cost/users?start_date=<start>&end_date=<end>&limit=100&offset=0` and save the response.
> After each page, stop when `offset + users.length >= count`; otherwise, increase `offset` by 100 and fetch the next page.
> For every `users[].user_id` across those pages, save `/api/experimental/chats/cost/{user_id}/summary?start_date=<start>&end_date=<end>` with the same dates.
> Each summary contains the user's totals plus `by_model` and `by_chat` breakdowns.
> After upgrading, the native **Spend** page, per-model pricing fields, and aggregate cost endpoints are unavailable.
> Historical `chat_messages.total_cost_micros` values remain in the database temporarily for rolling upgrade compatibility, but AI Gateway reports do not include or reconstruct them.
> Configure AI Gateway budgets separately.
The API reference documents how to [get](../../../reference/api/enterprise.md#get-group-ai-budget), [upsert](../../../reference/api/enterprise.md#upsert-group-ai-budget), and [delete](../../../reference/api/enterprise.md#delete-group-ai-budget) a group budget.
@@ -40,16 +49,30 @@ The API reference documents how to [get](../../../reference/api/enterprise.md#ge
The usage indicator on the Agents page and the summary in the user menu both show the signed-in user's current AI spend, their budget, and the period reset date.
Both appear only when the deployment has the AI Gateway entitlement.
## Spend visibility
## Spend details
Coder has no dedicated deployment-wide spend dashboard.
Spend is shown where it is actionable:
- **Agents page and user menu**: the signed-in user's spend against their budget, as described previously.
- **Group settings**: each member's spend against the group's budget, for admins who can manage the group.
- **Chat summary panel**: the cost of one chat tree, on a chat's Summary tab.
A subagent reports the total for its whole tree, including the chat that started it.
- **Agents** > **Settings** > **Manage Agents** > **Spend**: deployment-wide chat cost per user, with per-user drill-down.
> [!NOTE]
> Per-chat cost comes from AI Gateway records, which are pruned according to `--ai-gateway-retention` (60 days by default).
> A chat for which gateway records have been pruned reports no cost.
Organization administrators can export per-user, per-group, per-model, and per-provider spend to CSV:
```sh
curl -X GET "https://coder.example.com/api/v2/organizations/$ORGANIZATION/ai/spend/export" \
-H "Coder-Session-Token: $CODER_SESSION_TOKEN"
```
A successful response has the `Content-Type` header `text/csv; charset=utf-8` and starts with this CSV header:
```csv
user_id,username,group_id,group_name,organization_id,organization_name,model,provider,provider_name,input_tokens,output_tokens,cache_read_tokens,cache_write_tokens,cost_micros,period_start,period_end
```
The AI Gateway [sessions views](../../ai-gateway/audit.md#navigating-the-ui) show per-request token usage, which is the input to those costs rather than the costs themselves.
AI Gateway data is subject to its own [retention period](../../ai-gateway/monitoring.md#data-retention), 60 days by default, which is configured independently of chat retention.
Spend for requests older than that period is no longer reported, so a chat for which gateway records have been pruned reports no cost.
+3 -3
View File
@@ -1084,10 +1084,10 @@
"state": ["beta"]
},
{
"title": "Spend Management",
"title": "Spend management",
"description": "Cap Coder Agents spend with AI Gateway budgets and track the resulting spend.",
"path": "./ai-coder/agents/platform-controls/usage-insights.md",
"state": ["beta"]
"path": "./ai-coder/agents/platform-controls/spend-management.md",
"state": ["premium"]
},
{
"title": "Git Providers",