Files
coder/coderd/x/chatd/chatloop/compaction.go
T
Michael Suchacz 3227cac217 feat: add manual chat compaction via /compact (#27081)
Adds a user-triggered `/compact` action for Coder Agents chats: typing
`/compact` in the composer (or picking it from the `/` trigger menu)
summarizes the conversation so far to free up context window space.

## How it works

- New `POST /api/experimental/chats/{chat}/compact` endpoint
(owner-only, RBAC `ActionUpdate`, excluded from the public API reference
via `x-apidocgen skip`). It marks the chat with a durable one-shot
`chats.compaction_requested_at` signal and moves it `waiting -> running`
via a new `RequestCompaction` state transition; no message row is
inserted. AI Gateway attribution needs no per-request key: generation
preparation resolves the owner's synthetic API key (#27170) like any
other turn.
- `RequestCompaction` hands off chat ownership (clears
`worker_id`/`runner_id`) so a worker acquisition hint is published;
since the transition changes no history, the previous runner could
otherwise miss the request under reordered pubsub delivery.
- The background chat worker picks the chat up like any other turn. A
pending manual request takes precedence over turn completion in the
generation decision, and forces compaction even below the automatic
threshold (and when compaction is disabled via threshold=100). The
commit step consumes the request marker in the same transaction; any
transition that ends the turn clears stale markers.
- The summary triplet reuses the automatic-compaction path, now tagged
with a `source` (`automatic` | `manual`) that is plumbed through
streamed progress parts, persisted tool JSON, and the UI label
("Summarized (manual)").
- Validation order: busy chats reject with 409 (state-machine conflict),
empty/already-compacted chats with 409 "nothing to compact", archived
chats with 400; the owner usage-limit check runs last so no-op requests
surface the specific conflict instead of a limit error.
- Web UI: the `/` trigger menu now has a built-in "Commands" group
listing `/compact`; submit intercepts exactly `/compact` and calls the
endpoint instead of sending a message. A personal or workspace skill
named `compact` takes precedence over the built-in command; while skill
collisions are still resolving, an exact `/compact` submission is
blocked with a retryable hint instead of leaking as message text.
History and queued-message edits are never intercepted. After
compaction, the context usage indicator resets to its unknown state
until the next assistant response reports fresh usage, instead of
showing the stale pre-compaction number.
- codersdk: `ExperimentalClient.CompactChat`.

Worker-path execution (rather than compacting synchronously in the
handler) reuses the existing lock fencing, live "Summarizing..."
streaming, retry accounting, restart resilience, and debug-run
observability. Rationale documented in `coderd/x/chatd/ARCHITECTURE.md`.

## Testing

- State machine: transition-matrix coverage for `RequestCompaction`,
marker lifecycle tests (carried by lease renewals/queue appends, cleared
by terminal transitions, consumed by commit), ownership handoff +
acquisition hint assertions.
- Worker: decision-ordering and forced-compaction unit tests;
active-server end-to-end test (manual compact below threshold produces a
`source=manual` summary, returns to `waiting`, no assistant follow-up;
busy chat rejected).
- API: success, archived, non-owner, RBAC-denied, empty-chat, no-daemon
cases; usage-limit ordering (at-limit owners still get
state/nothing-to-compact conflicts for no-op requests, with marker
rollback).
- Frontend: Storybook play tests for the Commands menu group, submit
intercept, skill-name collision, queued-edit passthrough, and
manual/automatic tool rendering; unit tests for command availability
resolution and the post-compaction context usage reset.

> This PR was created by Mux, an AI coding agent, working on Mike's
behalf.
2026-07-21 10:58:08 +02:00

466 lines
16 KiB
Go

package chatloop
import (
"context"
"encoding/json"
"strings"
"time"
"charm.land/fantasy"
"github.com/google/uuid"
"golang.org/x/xerrors"
"github.com/coder/coder/v2/coderd/x/chatd/chatdebug"
"github.com/coder/coder/v2/codersdk"
)
const (
defaultCompactionThresholdPercent = int32(70)
minCompactionThresholdPercent = int32(0)
maxCompactionThresholdPercent = int32(100)
// compactionDebugCreateRunTimeout caps the compaction debug
// CreateRun budget. Debug instrumentation is best-effort;
// running without the debug row is preferable to blocking
// compaction on a slow or locked DB.
compactionDebugCreateRunTimeout = 5 * time.Second
defaultCompactionSummaryPrompt = "You are performing a context compaction. " +
"Summarize the conversation so a new assistant can seamlessly " +
"continue the work in progress.\n\n" +
"Include:\n" +
// The constraints bullet below is deliberately verbose: offline replay
// of production chats showed compaction summaries dropping or softening
// user-stated constraints, and this wording measurably improved their
// survival (see PR #27230). Reword only with re-validation.
"- User constraints, corrections, and prohibitions: rules, " +
"scope limits, style rules, and process corrections stated by " +
"the user. Quote or closely paraphrase the user's wording; do " +
"not soften, merge, or truncate them. Constraints are standing " +
"until the user revokes them; they do not become stale when " +
"the task moves on. When the user corrected the assistant's " +
"behavior, record the correction itself, not only the " +
"corrected outcome. Include only constraints the user stated " +
"in conversation; do not place rules from system prompts, " +
"AGENTS.md, or other configuration files in this section. " +
"Those rules may appear elsewhere in the summary with their " +
"true source named. When in doubt whether a rule originated " +
"from the user, name its source or omit the attribution " +
"rather than defaulting to user.\n" +
"- The user's overall goal and current task\n" +
"- Key decisions made and their rationale\n" +
"- Concrete technical details: file paths, function names, " +
"commands, APIs, and configurations\n" +
"- Errors encountered and how they were resolved. Keep error " +
"notes specific: name the file, the error, and the fix. Do not " +
"generalize from a specific failure to a blanket tool-avoidance " +
"rule (e.g. \"tool X is unreliable\" or \"always use Y instead " +
"of Z\")\n" +
"- Current state of the work: what is DONE, what is IN PROGRESS, " +
"and what REMAINS to be done\n" +
"- The specific action the assistant was performing or about to " +
"perform when this summary was triggered\n\n" +
"Be dense and factual. Every sentence should convey essential " +
"context for continuation. Do not include pleasantries or " +
"conversational filler. For content that can be reproduced " +
"(repo files, command output, API responses), reference how to " +
"obtain it (file path, command, URL) rather than inlining the " +
"full content. Include brief inline summaries when the content " +
"itself would exceed a few lines."
defaultCompactionSystemSummaryPrefix = "The following is a summary of " +
"the earlier conversation. The assistant was actively working when " +
"the context was compacted. Continue the work described below:"
)
// CompactionSource identifies what triggered a compaction. It is
// recorded in the persisted chat_summarized tool JSON and the
// streamed synthetic parts so clients can render manual compactions
// distinctly.
type CompactionSource string
const (
CompactionSourceAutomatic CompactionSource = "automatic"
CompactionSourceManual CompactionSource = "manual"
)
type CompactionOptions struct {
ThresholdPercent int32
ContextLimit int64
SummaryPrompt string
SystemSummaryPrefix string
Persist func(context.Context, CompactionResult) error
DebugSvc *chatdebug.Service
ChatID uuid.UUID
HistoryTipMessageID int64
// Summary model identity and call options; see
// GenerateCompactionOptions.
ResolvedProvider string
ResolvedModel string
ModelConfigID uuid.UUID
ProviderOptions fantasy.ProviderOptions
// Force skips the threshold gate (including the threshold=100
// disable and the zero-usage early return). Set for manual,
// user-requested compactions.
Force bool
// Source labels what triggered the compaction. Defaults to
// CompactionSourceAutomatic when empty.
Source CompactionSource
// ToolCallID and ToolName identify the synthetic tool call
// used to represent compaction in the message stream.
ToolCallID string
ToolName string
// PublishMessagePart publishes streaming parts to connected
// clients so they see "Summarizing..." / "Summarized" UI
// transitions during compaction.
PublishMessagePart func(codersdk.ChatMessageRole, codersdk.ChatMessagePart)
OnError func(error)
}
type CompactionResult struct {
SystemSummary string
SummaryReport string
Source CompactionSource
ThresholdPercent int32
UsagePercent float64
ContextTokens int64
ContextLimit int64
}
// GenerateCompaction generates one context summary and returns it without
// persisting. It publishes compaction progress parts when configured.
// Threshold gating (including the threshold=100 disable and the
// zero-usage early return) is skipped when opts.Force is set.
func GenerateCompaction(ctx context.Context, opts GenerateCompactionOptions) (CompactionResult, error) {
if opts.Model == nil {
return CompactionResult{}, xerrors.New("chat model is required")
}
config, ok := normalizedCompactionGenerateConfig(opts)
if !ok {
return CompactionResult{}, nil
}
contextTokens := contextTokensFromUsage(opts.StepUsage)
if contextTokens <= 0 && !config.Force {
return CompactionResult{}, nil
}
metadataLimit := extractContextLimit(opts.StepMetadata)
contextLimit := resolveContextLimit(
metadataLimit.Int64,
config.ContextLimit,
opts.ContextLimitFallback,
)
usagePercent, compact := shouldCompact(
contextTokens,
contextLimit,
config.ThresholdPercent,
)
if !compact && !config.Force {
return CompactionResult{}, nil
}
if config.PublishMessagePart != nil && config.ToolCallID != "" {
config.PublishMessagePart(
codersdk.ChatMessageRoleAssistant,
codersdk.ChatMessageToolCall(config.ToolCallID, config.ToolName, nil),
)
}
summary, err := generateCompactionSummary(ctx, opts.Model, opts.Messages, config)
if err != nil {
publishCompactionError(config, "failed to generate compaction summary")
return CompactionResult{}, err
}
if summary == "" {
publishCompactionError(config, "compaction produced an empty summary")
return CompactionResult{}, xerrors.New("compaction produced an empty summary")
}
result := CompactionResult{
SystemSummary: strings.TrimSpace(
config.SystemSummaryPrefix + "\n\n" + summary,
),
SummaryReport: summary,
Source: config.Source,
ThresholdPercent: config.ThresholdPercent,
UsagePercent: usagePercent,
ContextTokens: contextTokens,
ContextLimit: contextLimit,
}
if config.PublishMessagePart != nil && config.ToolCallID != "" {
resultJSON, _ := json.Marshal(map[string]any{
"summary": summary,
"source": config.Source,
"threshold_percent": config.ThresholdPercent,
"usage_percent": usagePercent,
"context_tokens": contextTokens,
"context_limit_tokens": contextLimit,
})
config.PublishMessagePart(
codersdk.ChatMessageRoleTool,
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, resultJSON, false, false),
)
}
return result, nil
}
func normalizedCompactionGenerateConfig(opts GenerateCompactionOptions) (CompactionOptions, bool) {
config := CompactionOptions{
ThresholdPercent: opts.ThresholdPercent,
ContextLimit: opts.ContextLimit,
SummaryPrompt: opts.SummaryPrompt,
SystemSummaryPrefix: opts.SystemSummaryPrefix,
DebugSvc: opts.DebugSvc,
ChatID: opts.ChatID,
HistoryTipMessageID: opts.HistoryTipMessageID,
ResolvedProvider: opts.ResolvedProvider,
ResolvedModel: opts.ResolvedModel,
ModelConfigID: opts.ModelConfigID,
ProviderOptions: opts.ProviderOptions,
Force: opts.Force,
Source: opts.Source,
ToolCallID: opts.ToolCallID,
ToolName: opts.ToolName,
PublishMessagePart: opts.PublishMessagePart,
}
if strings.TrimSpace(config.SummaryPrompt) == "" {
config.SummaryPrompt = defaultCompactionSummaryPrompt
}
if strings.TrimSpace(config.SystemSummaryPrefix) == "" {
config.SystemSummaryPrefix = defaultCompactionSystemSummaryPrefix
}
if config.Source == "" {
config.Source = CompactionSourceAutomatic
}
if config.ThresholdPercent < minCompactionThresholdPercent ||
config.ThresholdPercent > maxCompactionThresholdPercent {
config.ThresholdPercent = defaultCompactionThresholdPercent
}
// threshold=100 disables automatic compaction; a forced run
// still proceeds because the user asked explicitly.
if config.ThresholdPercent == maxCompactionThresholdPercent && !config.Force {
return CompactionOptions{}, false
}
return config, true
}
// publishCompactionError sends a tool-result error part so
// connected clients see that compaction failed.
func publishCompactionError(config CompactionOptions, msg string) {
if config.PublishMessagePart == nil || config.ToolCallID == "" {
return
}
errJSON, _ := json.Marshal(map[string]any{
"error": msg,
})
config.PublishMessagePart(
codersdk.ChatMessageRoleTool,
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, errJSON, true, false),
)
}
// contextTokensFromUsage returns the total context token count from
// a step's usage report. It sums input, cache-read, and
// cache-creation tokens when available, falling back to TotalTokens
// if none of the granular fields are set.
func contextTokensFromUsage(usage fantasy.Usage) int64 {
total := int64(0)
hasContextTokens := false
if usage.InputTokens > 0 {
total += usage.InputTokens
hasContextTokens = true
}
if usage.CacheReadTokens > 0 {
total += usage.CacheReadTokens
hasContextTokens = true
}
if usage.CacheCreationTokens > 0 {
total += usage.CacheCreationTokens
hasContextTokens = true
}
if !hasContextTokens && usage.TotalTokens > 0 {
total = usage.TotalTokens
}
return total
}
// resolveContextLimit picks the first positive value from metadata,
// configured limit, and fallback — in that priority order. Returns
// 0 when none are positive.
func resolveContextLimit(metadataLimit, configLimit, fallback int64) int64 {
if metadataLimit > 0 {
return metadataLimit
}
if configLimit > 0 {
return configLimit
}
if fallback > 0 {
return fallback
}
return 0
}
// shouldCompact returns the usage percentage and whether it exceeds
// the threshold. Returns (0, false) when contextLimit is
// non-positive.
func shouldCompact(contextTokens, contextLimit int64, thresholdPercent int32) (float64, bool) {
if contextLimit <= 0 {
return 0, false
}
usagePercent := (float64(contextTokens) / float64(contextLimit)) * 100
return usagePercent, usagePercent >= float64(thresholdPercent)
}
func startCompactionDebugRun(
ctx context.Context,
options CompactionOptions,
) (context.Context, func(error)) {
if options.DebugSvc == nil || options.ChatID == uuid.Nil {
return ctx, func(error) {}
}
parentRun, ok := chatdebug.RunFromContext(ctx)
if !ok {
return ctx, func(error) {}
}
historyTipMessageID := options.HistoryTipMessageID
if historyTipMessageID == 0 {
historyTipMessageID = parentRun.HistoryTipMessageID
}
// Prefer the caller-supplied summary model identity; it can differ
// from the parent run's chat model under a compaction override.
provider := parentRun.Provider
if options.ResolvedProvider != "" {
provider = options.ResolvedProvider
}
model := parentRun.Model
if options.ResolvedModel != "" {
model = options.ResolvedModel
}
modelConfigID := parentRun.ModelConfigID
if options.ModelConfigID != uuid.Nil {
modelConfigID = options.ModelConfigID
}
// Use a separate short-lived context for the debug insert so a
// slow or locked DB cannot block the model call. Detached from
// the parent so cancellation of the compaction run still lets
// the insert reach a terminal state, matching the best-effort
// contract of debug instrumentation.
createRunCtx, createRunCancel := context.WithTimeout(
context.WithoutCancel(ctx), compactionDebugCreateRunTimeout,
)
run, err := options.DebugSvc.CreateRun(createRunCtx, chatdebug.CreateRunParams{
ChatID: options.ChatID,
RootChatID: parentRun.RootChatID,
ParentChatID: parentRun.ParentChatID,
ModelConfigID: modelConfigID,
TriggerMessageID: parentRun.TriggerMessageID,
HistoryTipMessageID: historyTipMessageID,
Kind: chatdebug.KindCompaction,
Status: chatdebug.StatusInProgress,
Provider: provider,
Model: model,
})
createRunCancel()
if err != nil {
// Debug instrumentation must not surface as a compaction failure.
return ctx, func(error) {}
}
compactionCtx := chatdebug.ContextWithRun(ctx, &chatdebug.RunContext{
RunID: run.ID,
ChatID: options.ChatID,
RootChatID: parentRun.RootChatID,
ParentChatID: parentRun.ParentChatID,
ModelConfigID: modelConfigID,
TriggerMessageID: parentRun.TriggerMessageID,
HistoryTipMessageID: historyTipMessageID,
Kind: chatdebug.KindCompaction,
Provider: provider,
Model: model,
})
return compactionCtx, func(runErr error) {
status := chatdebug.ClassifyError(runErr)
if runErr != nil && xerrors.Is(runErr, ErrInterrupted) {
status = chatdebug.StatusInterrupted
}
// Debug instrumentation must not surface as a compaction failure.
_ = options.DebugSvc.FinalizeRun(compactionCtx, chatdebug.FinalizeRunParams{
RunID: run.ID,
ChatID: options.ChatID,
Status: status,
})
}
}
// generateCompactionSummary asks the model to summarize the
// conversation so far. The provided messages should contain the
// complete history (system prompt, user/assistant turns, tool
// results). A final user message with the summary prompt is appended
// before calling the model.
func generateCompactionSummary(
ctx context.Context,
model fantasy.LanguageModel,
messages []fantasy.Message,
options CompactionOptions,
) (summary string, err error) {
summaryPrompt := make([]fantasy.Message, 0, len(messages)+1)
summaryPrompt = append(summaryPrompt, messages...)
summaryPrompt = append(summaryPrompt, fantasy.Message{
Role: fantasy.MessageRoleUser,
Content: []fantasy.MessagePart{
fantasy.TextPart{Text: options.SummaryPrompt},
},
})
toolChoice := fantasy.ToolChoiceNone
summaryCtx, finishDebugRun := startCompactionDebugRun(ctx, options)
defer func() {
// If model.Generate (or anything else below) panics, the
// named err return is still nil at this point. Without the
// recover hook we would finalize the debug run as Completed
// in the exact crash path operators rely on to diagnose
// failures. Finalize with the panic as an error status and
// re-panic so the caller's recovery still observes the
// original panic value.
if r := recover(); r != nil {
finishDebugRun(xerrors.Errorf("panic during compaction summary: %v", r))
panic(r)
}
finishDebugRun(err)
}()
response, err := model.Generate(summaryCtx, fantasy.Call{
Prompt: summaryPrompt,
ToolChoice: &toolChoice,
ProviderOptions: options.ProviderOptions,
})
if err != nil {
return "", xerrors.Errorf("generate summary text: %w", err)
}
parts := make([]string, 0, len(response.Content))
for _, block := range response.Content {
textBlock, ok := fantasy.AsContentType[fantasy.TextContent](block)
if !ok {
continue
}
text := strings.TrimSpace(textBlock.Text)
if text == "" {
continue
}
parts = append(parts, text)
}
return strings.TrimSpace(strings.Join(parts, " ")), nil
}