Files
coder/coderd/x/chatd/chatloop/compaction.go
T
4b7494be72 feat: harden chat generation runtime instrumentation for billing (#27451)
Closes CODAGT-835

## Summary

`chat_messages.runtime_ms` becomes the billing source of truth for Coder
Agents runtime (summed hourly by #27312), but it was built for
debugging: the June refactor (#26270) silently stopped recording
tool-step runtime, compaction was never measured, and interrupted turns
lost their partial runtime entirely. This PR defines the billable
metric, closes the paths that dropped it, and documents the definition
where the data lives.

## The billable definition

**`runtime_ms` is the wall-clock duration of the model invocation that
produced the persisted message content**, measured from just before the
provider stream opens until it is fully consumed.

What counts:

- Assistant generation steps, in top-level and sub-agent chats
(sub-agents are ordinary chats on the same generation path).
- Compaction summarization calls, persisted on the compaction assistant
message (**new**).
- Interrupted attempts: the message-part episode's lifetime is persisted
on the partial assistant message committed by `FinishInterruption`, so
partial generation time survives interruption (**new**; measured via a
new `Buffer.EpisodeDuration`, which works even though the generation
goroutine and the interrupt task are different tasks).

What deliberately does not count (each is documented in code and docs):

- **Local tool execution.** Tool wall time includes idle waits, most
importantly `wait_agent` polling a sub-agent chat that already bills its
own model invocations; billing the batch would double count, and
excluding one tool from a concurrent batch's wall time is ill-defined.
Pre-refactor instrumentation did include tool time; this makes the
exclusion an explicit product definition instead of a silent regression.
- **Failed model calls whose output is discarded** (retried attempts,
terminal errors, content-filter refusals). They persist no content, so
they bill nothing; billing errs toward undercounting. Notably a
stream-silence timeout can burn 10 idle minutes before a retry, which
should not be billable "active generation". If product later wants
failed attempts billed, that needs a place to persist runtime on error
turns (`FinishError` inserts no rows today) and is a deliberate
follow-up, not instrumentation drift.
- **Ancillary calls that produce no chat messages** (title generation,
advisor, turn summaries) and all idle/parked time (`requires_action`,
queueing).

The definition is documented as `COMMENT ON COLUMN
chat_messages.runtime_ms` (migration 000551, surfacing as a Go doc
comment on `ChatMessage.RuntimeMs`), on
`chatloop.PersistedStep.Runtime`, in the chatd architecture doc, and in
the Spend Management docs page.

## Index for the hourly scan

None needed: `GetTotalChatMessageRuntimeMsInRange` (#27312) filters an
hour-wide `created_at` range, which the existing
`idx_chat_messages_created_at` b-tree already serves; the residual
`runtime_ms IS NOT NULL` filter applies to one hour of rows. A partial
index would add permanent write amplification for a query that runs once
an hour.

> [!NOTE]
> Migration 000551 is also claimed by #27312; whichever merges second
renumbers via `fix_migration_numbers.sh`.

## Tests

- End-to-end: the existing full-server generation test now asserts
`RuntimeMs.Valid` on the committed assistant row (it previously read
`.Int64` without checking `.Valid`, so it passed on NULL).
- Interrupted turn: full task-level test (real DB, mock clock) asserting
the partial assistant message persists the attempt's runtime.
- Errored stream: asserts a failed invocation yields no step and no
runtime.
- Tool-using turn: asserts runtime lands on the assistant row only and
tool rows stay NULL.
- Compaction: asserts the summarization call duration is recorded and
lands on the compaction assistant message only.
- `messagepartbuffer.EpisodeDuration` unit coverage.

Blocks: CODAGT-843 (B3), CODAGT-838 (D8).

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Hugo Dutka <hugo@coder.com>
2026-08-06 16:09:38 +07:00

484 lines
17 KiB
Go

package chatloop
import (
"context"
"encoding/json"
"strings"
"time"
"charm.land/fantasy"
"github.com/google/uuid"
"golang.org/x/xerrors"
"github.com/coder/coder/v2/coderd/x/chatd/chatdebug"
"github.com/coder/coder/v2/codersdk"
)
const (
defaultCompactionThresholdPercent = int32(70)
minCompactionThresholdPercent = int32(0)
maxCompactionThresholdPercent = int32(100)
// compactionDebugCreateRunTimeout caps the compaction debug
// CreateRun budget. Debug instrumentation is best-effort;
// running without the debug row is preferable to blocking
// compaction on a slow or locked DB.
compactionDebugCreateRunTimeout = 5 * time.Second
defaultCompactionSummaryPrompt = "You are performing a context compaction. " +
"Summarize the conversation so a new assistant can seamlessly " +
"continue the work in progress.\n\n" +
"Include:\n" +
// The constraints bullet below is deliberately verbose: offline replay
// of production chats showed compaction summaries dropping or softening
// user-stated constraints, and this wording measurably improved their
// survival (see PR #27230). Reword only with re-validation.
"- User constraints, corrections, and prohibitions: rules, " +
"scope limits, style rules, and process corrections stated by " +
"the user. Quote or closely paraphrase the user's wording; do " +
"not soften, merge, or truncate them. Constraints are standing " +
"until the user revokes them; they do not become stale when " +
"the task moves on. When the user corrected the assistant's " +
"behavior, record the correction itself, not only the " +
"corrected outcome. Include only constraints the user stated " +
"in conversation; do not place rules from system prompts, " +
"AGENTS.md, or other configuration files in this section. " +
"Those rules may appear elsewhere in the summary with their " +
"true source named. When in doubt whether a rule originated " +
"from the user, name its source or omit the attribution " +
"rather than defaulting to user.\n" +
"- The user's overall goal and current task\n" +
"- Key decisions made and their rationale\n" +
"- Concrete technical details: file paths, function names, " +
"commands, APIs, and configurations\n" +
"- Errors encountered and how they were resolved. Keep error " +
"notes specific: name the file, the error, and the fix. Do not " +
"generalize from a specific failure to a blanket tool-avoidance " +
"rule (e.g. \"tool X is unreliable\" or \"always use Y instead " +
"of Z\")\n" +
"- Current state of the work: what is DONE, what is IN PROGRESS, " +
"and what REMAINS to be done\n" +
"- The specific action the assistant was performing or about to " +
"perform when this summary was triggered\n\n" +
"Be dense and factual. Every sentence should convey essential " +
"context for continuation. Do not include pleasantries or " +
"conversational filler. For content that can be reproduced " +
"(repo files, command output, API responses), reference how to " +
"obtain it (file path, command, URL) rather than inlining the " +
"full content. Include brief inline summaries when the content " +
"itself would exceed a few lines."
defaultCompactionSystemSummaryPrefix = "The following is a summary of " +
"the earlier conversation. The assistant was actively working when " +
"the context was compacted. Continue the work described below:"
)
// CompactionSource identifies what triggered a compaction. It is
// recorded in the persisted chat_summarized tool JSON and the
// streamed synthetic parts so clients can render manual compactions
// distinctly.
type CompactionSource string
const (
CompactionSourceAutomatic CompactionSource = "automatic"
CompactionSourceManual CompactionSource = "manual"
)
type CompactionOptions struct {
ThresholdPercent int32
ContextLimit int64
SummaryPrompt string
SummaryHint string
SystemSummaryPrefix string
Persist func(context.Context, CompactionResult) error
DebugSvc *chatdebug.Service
ChatID uuid.UUID
HistoryTipMessageID int64
// Summary model identity and call options; see
// GenerateCompactionOptions.
ResolvedProvider string
ResolvedModel string
ModelConfigID uuid.UUID
ProviderOptions fantasy.ProviderOptions
// Force skips the threshold gate (including the threshold=100
// disable and the zero-usage early return). Set for manual,
// user-requested compactions.
Force bool
// Source labels what triggered the compaction. Defaults to
// CompactionSourceAutomatic when empty.
Source CompactionSource
// ToolCallID and ToolName identify the synthetic tool call
// used to represent compaction in the message stream.
ToolCallID string
ToolName string
// PublishMessagePart publishes streaming parts to connected
// clients so they see "Summarizing..." / "Summarized" UI
// transitions during compaction.
PublishMessagePart func(codersdk.ChatMessageRole, codersdk.ChatMessagePart)
OnError func(error)
}
type CompactionResult struct {
SystemSummary string
SummaryReport string
Source CompactionSource
ThresholdPercent int32
UsagePercent float64
ContextTokens int64
ContextLimit int64
// Runtime is the wall-clock duration of the summarization model
// call, the compaction step's billable runtime (see
// PersistedStep.Runtime). Zero when the run was gated off before
// calling the model.
Runtime time.Duration
}
// GenerateCompaction generates one context summary and returns it without
// persisting. It publishes compaction progress parts when configured.
// Threshold gating (including the threshold=100 disable and the
// zero-usage early return) is skipped when opts.Force is set.
func GenerateCompaction(ctx context.Context, opts GenerateCompactionOptions) (CompactionResult, error) {
if opts.Model == nil {
return CompactionResult{}, xerrors.New("chat model is required")
}
if opts.Clock == nil {
return CompactionResult{}, xerrors.New("clock is required")
}
config, ok := normalizedCompactionGenerateConfig(opts)
if !ok {
return CompactionResult{}, nil
}
contextTokens := contextTokensFromUsage(opts.StepUsage)
if contextTokens <= 0 && !config.Force {
return CompactionResult{}, nil
}
metadataLimit := extractContextLimit(opts.StepMetadata)
contextLimit := resolveContextLimit(
metadataLimit.Int64,
config.ContextLimit,
opts.ContextLimitFallback,
)
usagePercent, compact := shouldCompact(
contextTokens,
contextLimit,
config.ThresholdPercent,
)
if !compact && !config.Force {
return CompactionResult{}, nil
}
if config.PublishMessagePart != nil && config.ToolCallID != "" {
config.PublishMessagePart(
codersdk.ChatMessageRoleAssistant,
codersdk.ChatMessageToolCall(config.ToolCallID, config.ToolName, nil),
)
}
summaryStart := opts.Clock.Now()
if opts.OnModelStreamStart != nil {
opts.OnModelStreamStart()
}
summary, err := generateCompactionSummary(ctx, opts.Model, opts.Messages, config)
if err != nil {
publishCompactionError(config, "failed to generate compaction summary")
return CompactionResult{}, err
}
summaryRuntime := opts.Clock.Since(summaryStart)
if summary == "" {
publishCompactionError(config, "compaction produced an empty summary")
return CompactionResult{}, xerrors.New("compaction produced an empty summary")
}
result := CompactionResult{
SystemSummary: strings.TrimSpace(
config.SystemSummaryPrefix + "\n\n" + summary,
),
SummaryReport: summary,
Source: config.Source,
ThresholdPercent: config.ThresholdPercent,
UsagePercent: usagePercent,
ContextTokens: contextTokens,
ContextLimit: contextLimit,
Runtime: summaryRuntime,
}
if config.PublishMessagePart != nil && config.ToolCallID != "" {
resultJSON, _ := json.Marshal(map[string]any{
"summary": summary,
"source": config.Source,
"threshold_percent": config.ThresholdPercent,
"usage_percent": usagePercent,
"context_tokens": contextTokens,
"context_limit_tokens": contextLimit,
})
config.PublishMessagePart(
codersdk.ChatMessageRoleTool,
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, resultJSON, false, false),
)
}
return result, nil
}
func normalizedCompactionGenerateConfig(opts GenerateCompactionOptions) (CompactionOptions, bool) {
config := CompactionOptions{
ThresholdPercent: opts.ThresholdPercent,
ContextLimit: opts.ContextLimit,
SummaryPrompt: opts.SummaryPrompt,
SummaryHint: opts.SummaryHint,
SystemSummaryPrefix: opts.SystemSummaryPrefix,
DebugSvc: opts.DebugSvc,
ChatID: opts.ChatID,
HistoryTipMessageID: opts.HistoryTipMessageID,
ResolvedProvider: opts.ResolvedProvider,
ResolvedModel: opts.ResolvedModel,
ModelConfigID: opts.ModelConfigID,
ProviderOptions: opts.ProviderOptions,
Force: opts.Force,
Source: opts.Source,
ToolCallID: opts.ToolCallID,
ToolName: opts.ToolName,
PublishMessagePart: opts.PublishMessagePart,
}
if strings.TrimSpace(config.SummaryPrompt) == "" {
config.SummaryPrompt = defaultCompactionSummaryPrompt
}
if strings.TrimSpace(config.SystemSummaryPrefix) == "" {
config.SystemSummaryPrefix = defaultCompactionSystemSummaryPrefix
}
if config.Source == "" {
config.Source = CompactionSourceAutomatic
}
if config.ThresholdPercent < minCompactionThresholdPercent ||
config.ThresholdPercent > maxCompactionThresholdPercent {
config.ThresholdPercent = defaultCompactionThresholdPercent
}
// threshold=100 disables automatic compaction; a forced run
// still proceeds because the user asked explicitly.
if config.ThresholdPercent == maxCompactionThresholdPercent && !config.Force {
return CompactionOptions{}, false
}
return config, true
}
// publishCompactionError sends a tool-result error part so
// connected clients see that compaction failed.
func publishCompactionError(config CompactionOptions, msg string) {
if config.PublishMessagePart == nil || config.ToolCallID == "" {
return
}
errJSON, _ := json.Marshal(map[string]any{
"error": msg,
})
config.PublishMessagePart(
codersdk.ChatMessageRoleTool,
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, errJSON, true, false),
)
}
// contextTokensFromUsage returns the total context token count from
// a step's usage report. It sums input, cache-read, and
// cache-creation tokens when available, falling back to TotalTokens
// if none of the granular fields are set.
func contextTokensFromUsage(usage fantasy.Usage) int64 {
total := int64(0)
hasContextTokens := false
if usage.InputTokens > 0 {
total += usage.InputTokens
hasContextTokens = true
}
if usage.CacheReadTokens > 0 {
total += usage.CacheReadTokens
hasContextTokens = true
}
if usage.CacheCreationTokens > 0 {
total += usage.CacheCreationTokens
hasContextTokens = true
}
if !hasContextTokens && usage.TotalTokens > 0 {
total = usage.TotalTokens
}
return total
}
// resolveContextLimit picks the first positive value from metadata,
// configured limit, and fallback — in that priority order. Returns
// 0 when none are positive.
func resolveContextLimit(metadataLimit, configLimit, fallback int64) int64 {
if metadataLimit > 0 {
return metadataLimit
}
if configLimit > 0 {
return configLimit
}
if fallback > 0 {
return fallback
}
return 0
}
// shouldCompact returns the usage percentage and whether it exceeds
// the threshold. Returns (0, false) when contextLimit is
// non-positive.
func shouldCompact(contextTokens, contextLimit int64, thresholdPercent int32) (float64, bool) {
if contextLimit <= 0 {
return 0, false
}
usagePercent := (float64(contextTokens) / float64(contextLimit)) * 100
return usagePercent, usagePercent >= float64(thresholdPercent)
}
func startCompactionDebugRun(
ctx context.Context,
options CompactionOptions,
) (context.Context, func(error)) {
if options.DebugSvc == nil || options.ChatID == uuid.Nil {
return ctx, func(error) {}
}
parentRun, ok := chatdebug.RunFromContext(ctx)
if !ok {
return ctx, func(error) {}
}
historyTipMessageID := options.HistoryTipMessageID
if historyTipMessageID == 0 {
historyTipMessageID = parentRun.HistoryTipMessageID
}
// Prefer the caller-supplied summary model identity; it can differ
// from the parent run's chat model under a compaction override.
provider := parentRun.Provider
if options.ResolvedProvider != "" {
provider = options.ResolvedProvider
}
model := parentRun.Model
if options.ResolvedModel != "" {
model = options.ResolvedModel
}
modelConfigID := parentRun.ModelConfigID
if options.ModelConfigID != uuid.Nil {
modelConfigID = options.ModelConfigID
}
// Use a separate short-lived context for the debug insert so a
// slow or locked DB cannot block the model call. Detached from
// the parent so cancellation of the compaction run still lets
// the insert reach a terminal state, matching the best-effort
// contract of debug instrumentation.
createRunCtx, createRunCancel := context.WithTimeout(
context.WithoutCancel(ctx), compactionDebugCreateRunTimeout,
)
run, err := options.DebugSvc.CreateRun(createRunCtx, chatdebug.CreateRunParams{
ChatID: options.ChatID,
RootChatID: parentRun.RootChatID,
ParentChatID: parentRun.ParentChatID,
ModelConfigID: modelConfigID,
TriggerMessageID: parentRun.TriggerMessageID,
HistoryTipMessageID: historyTipMessageID,
Kind: chatdebug.KindCompaction,
Status: chatdebug.StatusInProgress,
Provider: provider,
Model: model,
})
createRunCancel()
if err != nil {
// Debug instrumentation must not surface as a compaction failure.
return ctx, func(error) {}
}
compactionCtx := chatdebug.ContextWithRun(ctx, &chatdebug.RunContext{
RunID: run.ID,
ChatID: options.ChatID,
RootChatID: parentRun.RootChatID,
ParentChatID: parentRun.ParentChatID,
ModelConfigID: modelConfigID,
TriggerMessageID: parentRun.TriggerMessageID,
HistoryTipMessageID: historyTipMessageID,
Kind: chatdebug.KindCompaction,
Provider: provider,
Model: model,
})
return compactionCtx, func(runErr error) {
status := chatdebug.ClassifyError(runErr)
if runErr != nil && xerrors.Is(runErr, ErrInterrupted) {
status = chatdebug.StatusInterrupted
}
// Debug instrumentation must not surface as a compaction failure.
_ = options.DebugSvc.FinalizeRun(compactionCtx, chatdebug.FinalizeRunParams{
RunID: run.ID,
ChatID: options.ChatID,
Status: status,
})
}
}
// generateCompactionSummary asks the model to summarize the
// conversation so far. The provided messages should contain the
// complete history (system prompt, user/assistant turns, tool
// results). A final user message with the summary prompt is appended
// before calling the model.
func generateCompactionSummary(
ctx context.Context,
model fantasy.LanguageModel,
messages []fantasy.Message,
options CompactionOptions,
) (summary string, err error) {
summaryPrompt := make([]fantasy.Message, 0, len(messages)+1)
summaryPrompt = append(summaryPrompt, messages...)
summaryParts := []fantasy.MessagePart{fantasy.TextPart{Text: options.SummaryPrompt}}
if strings.TrimSpace(options.SummaryHint) != "" {
summaryParts = append(summaryParts, fantasy.TextPart{Text: options.SummaryHint})
}
summaryPrompt = append(summaryPrompt, fantasy.Message{
Role: fantasy.MessageRoleUser,
Content: summaryParts,
})
toolChoice := fantasy.ToolChoiceNone
summaryCtx, finishDebugRun := startCompactionDebugRun(ctx, options)
defer func() {
// If model.Generate (or anything else below) panics, the
// named err return is still nil at this point. Without the
// recover hook we would finalize the debug run as Completed
// in the exact crash path operators rely on to diagnose
// failures. Finalize with the panic as an error status and
// re-panic so the caller's recovery still observes the
// original panic value.
if r := recover(); r != nil {
finishDebugRun(xerrors.Errorf("panic during compaction summary: %v", r))
panic(r)
}
finishDebugRun(err)
}()
response, err := model.Generate(summaryCtx, fantasy.Call{
Prompt: summaryPrompt,
ToolChoice: &toolChoice,
ProviderOptions: options.ProviderOptions,
})
if err != nil {
return "", xerrors.Errorf("generate summary text: %w", err)
}
parts := make([]string, 0, len(response.Content))
for _, block := range response.Content {
textBlock, ok := fantasy.AsContentType[fantasy.TextContent](block)
if !ok {
continue
}
text := strings.TrimSpace(textBlock.Text)
if text == "" {
continue
}
parts = append(parts, text)
}
return strings.TrimSpace(strings.Join(parts, " ")), nil
}