mirror of
https://github.com/coder/coder.git
synced 2026-09-23 14:03:57 +08:00
Adds a user-triggered `/compact` action for Coder Agents chats: typing
`/compact` in the composer (or picking it from the `/` trigger menu)
summarizes the conversation so far to free up context window space.
## How it works
- New `POST /api/experimental/chats/{chat}/compact` endpoint
(owner-only, RBAC `ActionUpdate`, excluded from the public API reference
via `x-apidocgen skip`). It marks the chat with a durable one-shot
`chats.compaction_requested_at` signal and moves it `waiting -> running`
via a new `RequestCompaction` state transition; no message row is
inserted. AI Gateway attribution needs no per-request key: generation
preparation resolves the owner's synthetic API key (#27170) like any
other turn.
- `RequestCompaction` hands off chat ownership (clears
`worker_id`/`runner_id`) so a worker acquisition hint is published;
since the transition changes no history, the previous runner could
otherwise miss the request under reordered pubsub delivery.
- The background chat worker picks the chat up like any other turn. A
pending manual request takes precedence over turn completion in the
generation decision, and forces compaction even below the automatic
threshold (and when compaction is disabled via threshold=100). The
commit step consumes the request marker in the same transaction; any
transition that ends the turn clears stale markers.
- The summary triplet reuses the automatic-compaction path, now tagged
with a `source` (`automatic` | `manual`) that is plumbed through
streamed progress parts, persisted tool JSON, and the UI label
("Summarized (manual)").
- Validation order: busy chats reject with 409 (state-machine conflict),
empty/already-compacted chats with 409 "nothing to compact", archived
chats with 400; the owner usage-limit check runs last so no-op requests
surface the specific conflict instead of a limit error.
- Web UI: the `/` trigger menu now has a built-in "Commands" group
listing `/compact`; submit intercepts exactly `/compact` and calls the
endpoint instead of sending a message. A personal or workspace skill
named `compact` takes precedence over the built-in command; while skill
collisions are still resolving, an exact `/compact` submission is
blocked with a retryable hint instead of leaking as message text.
History and queued-message edits are never intercepted. After
compaction, the context usage indicator resets to its unknown state
until the next assistant response reports fresh usage, instead of
showing the stale pre-compaction number.
- codersdk: `ExperimentalClient.CompactChat`.
Worker-path execution (rather than compacting synchronously in the
handler) reuses the existing lock fencing, live "Summarizing..."
streaming, retry accounting, restart resilience, and debug-run
observability. Rationale documented in `coderd/x/chatd/ARCHITECTURE.md`.
## Testing
- State machine: transition-matrix coverage for `RequestCompaction`,
marker lifecycle tests (carried by lease renewals/queue appends, cleared
by terminal transitions, consumed by commit), ownership handoff +
acquisition hint assertions.
- Worker: decision-ordering and forced-compaction unit tests;
active-server end-to-end test (manual compact below threshold produces a
`source=manual` summary, returns to `waiting`, no assistant follow-up;
busy chat rejected).
- API: success, archived, non-owner, RBAC-denied, empty-chat, no-daemon
cases; usage-limit ordering (at-limit owners still get
state/nothing-to-compact conflicts for no-op requests, with marker
rollback).
- Frontend: Storybook play tests for the Commands menu group, submit
intercept, skill-name collision, queued-edit passthrough, and
manual/automatic tool rendering; unit tests for command availability
resolution and the post-compaction context usage reset.
> This PR was created by Mux, an AI coding agent, working on Mike's
behalf.
466 lines
16 KiB
Go
466 lines
16 KiB
Go
package chatloop
|
|
|
|
import (
|
|
"context"
|
|
"encoding/json"
|
|
"strings"
|
|
"time"
|
|
|
|
"charm.land/fantasy"
|
|
"github.com/google/uuid"
|
|
"golang.org/x/xerrors"
|
|
|
|
"github.com/coder/coder/v2/coderd/x/chatd/chatdebug"
|
|
"github.com/coder/coder/v2/codersdk"
|
|
)
|
|
|
|
const (
|
|
defaultCompactionThresholdPercent = int32(70)
|
|
minCompactionThresholdPercent = int32(0)
|
|
maxCompactionThresholdPercent = int32(100)
|
|
|
|
// compactionDebugCreateRunTimeout caps the compaction debug
|
|
// CreateRun budget. Debug instrumentation is best-effort;
|
|
// running without the debug row is preferable to blocking
|
|
// compaction on a slow or locked DB.
|
|
compactionDebugCreateRunTimeout = 5 * time.Second
|
|
|
|
defaultCompactionSummaryPrompt = "You are performing a context compaction. " +
|
|
"Summarize the conversation so a new assistant can seamlessly " +
|
|
"continue the work in progress.\n\n" +
|
|
"Include:\n" +
|
|
// The constraints bullet below is deliberately verbose: offline replay
|
|
// of production chats showed compaction summaries dropping or softening
|
|
// user-stated constraints, and this wording measurably improved their
|
|
// survival (see PR #27230). Reword only with re-validation.
|
|
"- User constraints, corrections, and prohibitions: rules, " +
|
|
"scope limits, style rules, and process corrections stated by " +
|
|
"the user. Quote or closely paraphrase the user's wording; do " +
|
|
"not soften, merge, or truncate them. Constraints are standing " +
|
|
"until the user revokes them; they do not become stale when " +
|
|
"the task moves on. When the user corrected the assistant's " +
|
|
"behavior, record the correction itself, not only the " +
|
|
"corrected outcome. Include only constraints the user stated " +
|
|
"in conversation; do not place rules from system prompts, " +
|
|
"AGENTS.md, or other configuration files in this section. " +
|
|
"Those rules may appear elsewhere in the summary with their " +
|
|
"true source named. When in doubt whether a rule originated " +
|
|
"from the user, name its source or omit the attribution " +
|
|
"rather than defaulting to user.\n" +
|
|
"- The user's overall goal and current task\n" +
|
|
"- Key decisions made and their rationale\n" +
|
|
"- Concrete technical details: file paths, function names, " +
|
|
"commands, APIs, and configurations\n" +
|
|
"- Errors encountered and how they were resolved. Keep error " +
|
|
"notes specific: name the file, the error, and the fix. Do not " +
|
|
"generalize from a specific failure to a blanket tool-avoidance " +
|
|
"rule (e.g. \"tool X is unreliable\" or \"always use Y instead " +
|
|
"of Z\")\n" +
|
|
"- Current state of the work: what is DONE, what is IN PROGRESS, " +
|
|
"and what REMAINS to be done\n" +
|
|
"- The specific action the assistant was performing or about to " +
|
|
"perform when this summary was triggered\n\n" +
|
|
"Be dense and factual. Every sentence should convey essential " +
|
|
"context for continuation. Do not include pleasantries or " +
|
|
"conversational filler. For content that can be reproduced " +
|
|
"(repo files, command output, API responses), reference how to " +
|
|
"obtain it (file path, command, URL) rather than inlining the " +
|
|
"full content. Include brief inline summaries when the content " +
|
|
"itself would exceed a few lines."
|
|
defaultCompactionSystemSummaryPrefix = "The following is a summary of " +
|
|
"the earlier conversation. The assistant was actively working when " +
|
|
"the context was compacted. Continue the work described below:"
|
|
)
|
|
|
|
// CompactionSource identifies what triggered a compaction. It is
|
|
// recorded in the persisted chat_summarized tool JSON and the
|
|
// streamed synthetic parts so clients can render manual compactions
|
|
// distinctly.
|
|
type CompactionSource string
|
|
|
|
const (
|
|
CompactionSourceAutomatic CompactionSource = "automatic"
|
|
CompactionSourceManual CompactionSource = "manual"
|
|
)
|
|
|
|
type CompactionOptions struct {
|
|
ThresholdPercent int32
|
|
ContextLimit int64
|
|
SummaryPrompt string
|
|
SystemSummaryPrefix string
|
|
Persist func(context.Context, CompactionResult) error
|
|
DebugSvc *chatdebug.Service
|
|
ChatID uuid.UUID
|
|
HistoryTipMessageID int64
|
|
|
|
// Summary model identity and call options; see
|
|
// GenerateCompactionOptions.
|
|
ResolvedProvider string
|
|
ResolvedModel string
|
|
ModelConfigID uuid.UUID
|
|
ProviderOptions fantasy.ProviderOptions
|
|
|
|
// Force skips the threshold gate (including the threshold=100
|
|
// disable and the zero-usage early return). Set for manual,
|
|
// user-requested compactions.
|
|
Force bool
|
|
// Source labels what triggered the compaction. Defaults to
|
|
// CompactionSourceAutomatic when empty.
|
|
Source CompactionSource
|
|
|
|
// ToolCallID and ToolName identify the synthetic tool call
|
|
// used to represent compaction in the message stream.
|
|
ToolCallID string
|
|
ToolName string
|
|
|
|
// PublishMessagePart publishes streaming parts to connected
|
|
// clients so they see "Summarizing..." / "Summarized" UI
|
|
// transitions during compaction.
|
|
PublishMessagePart func(codersdk.ChatMessageRole, codersdk.ChatMessagePart)
|
|
|
|
OnError func(error)
|
|
}
|
|
|
|
type CompactionResult struct {
|
|
SystemSummary string
|
|
SummaryReport string
|
|
Source CompactionSource
|
|
ThresholdPercent int32
|
|
UsagePercent float64
|
|
ContextTokens int64
|
|
ContextLimit int64
|
|
}
|
|
|
|
// GenerateCompaction generates one context summary and returns it without
|
|
// persisting. It publishes compaction progress parts when configured.
|
|
// Threshold gating (including the threshold=100 disable and the
|
|
// zero-usage early return) is skipped when opts.Force is set.
|
|
func GenerateCompaction(ctx context.Context, opts GenerateCompactionOptions) (CompactionResult, error) {
|
|
if opts.Model == nil {
|
|
return CompactionResult{}, xerrors.New("chat model is required")
|
|
}
|
|
config, ok := normalizedCompactionGenerateConfig(opts)
|
|
if !ok {
|
|
return CompactionResult{}, nil
|
|
}
|
|
|
|
contextTokens := contextTokensFromUsage(opts.StepUsage)
|
|
if contextTokens <= 0 && !config.Force {
|
|
return CompactionResult{}, nil
|
|
}
|
|
metadataLimit := extractContextLimit(opts.StepMetadata)
|
|
contextLimit := resolveContextLimit(
|
|
metadataLimit.Int64,
|
|
config.ContextLimit,
|
|
opts.ContextLimitFallback,
|
|
)
|
|
usagePercent, compact := shouldCompact(
|
|
contextTokens,
|
|
contextLimit,
|
|
config.ThresholdPercent,
|
|
)
|
|
if !compact && !config.Force {
|
|
return CompactionResult{}, nil
|
|
}
|
|
|
|
if config.PublishMessagePart != nil && config.ToolCallID != "" {
|
|
config.PublishMessagePart(
|
|
codersdk.ChatMessageRoleAssistant,
|
|
codersdk.ChatMessageToolCall(config.ToolCallID, config.ToolName, nil),
|
|
)
|
|
}
|
|
|
|
summary, err := generateCompactionSummary(ctx, opts.Model, opts.Messages, config)
|
|
if err != nil {
|
|
publishCompactionError(config, "failed to generate compaction summary")
|
|
return CompactionResult{}, err
|
|
}
|
|
if summary == "" {
|
|
publishCompactionError(config, "compaction produced an empty summary")
|
|
return CompactionResult{}, xerrors.New("compaction produced an empty summary")
|
|
}
|
|
|
|
result := CompactionResult{
|
|
SystemSummary: strings.TrimSpace(
|
|
config.SystemSummaryPrefix + "\n\n" + summary,
|
|
),
|
|
SummaryReport: summary,
|
|
Source: config.Source,
|
|
ThresholdPercent: config.ThresholdPercent,
|
|
UsagePercent: usagePercent,
|
|
ContextTokens: contextTokens,
|
|
ContextLimit: contextLimit,
|
|
}
|
|
if config.PublishMessagePart != nil && config.ToolCallID != "" {
|
|
resultJSON, _ := json.Marshal(map[string]any{
|
|
"summary": summary,
|
|
"source": config.Source,
|
|
"threshold_percent": config.ThresholdPercent,
|
|
"usage_percent": usagePercent,
|
|
"context_tokens": contextTokens,
|
|
"context_limit_tokens": contextLimit,
|
|
})
|
|
config.PublishMessagePart(
|
|
codersdk.ChatMessageRoleTool,
|
|
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, resultJSON, false, false),
|
|
)
|
|
}
|
|
return result, nil
|
|
}
|
|
|
|
func normalizedCompactionGenerateConfig(opts GenerateCompactionOptions) (CompactionOptions, bool) {
|
|
config := CompactionOptions{
|
|
ThresholdPercent: opts.ThresholdPercent,
|
|
ContextLimit: opts.ContextLimit,
|
|
SummaryPrompt: opts.SummaryPrompt,
|
|
SystemSummaryPrefix: opts.SystemSummaryPrefix,
|
|
DebugSvc: opts.DebugSvc,
|
|
ChatID: opts.ChatID,
|
|
HistoryTipMessageID: opts.HistoryTipMessageID,
|
|
ResolvedProvider: opts.ResolvedProvider,
|
|
ResolvedModel: opts.ResolvedModel,
|
|
ModelConfigID: opts.ModelConfigID,
|
|
ProviderOptions: opts.ProviderOptions,
|
|
Force: opts.Force,
|
|
Source: opts.Source,
|
|
ToolCallID: opts.ToolCallID,
|
|
ToolName: opts.ToolName,
|
|
PublishMessagePart: opts.PublishMessagePart,
|
|
}
|
|
if strings.TrimSpace(config.SummaryPrompt) == "" {
|
|
config.SummaryPrompt = defaultCompactionSummaryPrompt
|
|
}
|
|
if strings.TrimSpace(config.SystemSummaryPrefix) == "" {
|
|
config.SystemSummaryPrefix = defaultCompactionSystemSummaryPrefix
|
|
}
|
|
if config.Source == "" {
|
|
config.Source = CompactionSourceAutomatic
|
|
}
|
|
if config.ThresholdPercent < minCompactionThresholdPercent ||
|
|
config.ThresholdPercent > maxCompactionThresholdPercent {
|
|
config.ThresholdPercent = defaultCompactionThresholdPercent
|
|
}
|
|
// threshold=100 disables automatic compaction; a forced run
|
|
// still proceeds because the user asked explicitly.
|
|
if config.ThresholdPercent == maxCompactionThresholdPercent && !config.Force {
|
|
return CompactionOptions{}, false
|
|
}
|
|
return config, true
|
|
}
|
|
|
|
// publishCompactionError sends a tool-result error part so
|
|
// connected clients see that compaction failed.
|
|
func publishCompactionError(config CompactionOptions, msg string) {
|
|
if config.PublishMessagePart == nil || config.ToolCallID == "" {
|
|
return
|
|
}
|
|
errJSON, _ := json.Marshal(map[string]any{
|
|
"error": msg,
|
|
})
|
|
config.PublishMessagePart(
|
|
codersdk.ChatMessageRoleTool,
|
|
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, errJSON, true, false),
|
|
)
|
|
}
|
|
|
|
// contextTokensFromUsage returns the total context token count from
|
|
// a step's usage report. It sums input, cache-read, and
|
|
// cache-creation tokens when available, falling back to TotalTokens
|
|
// if none of the granular fields are set.
|
|
func contextTokensFromUsage(usage fantasy.Usage) int64 {
|
|
total := int64(0)
|
|
hasContextTokens := false
|
|
|
|
if usage.InputTokens > 0 {
|
|
total += usage.InputTokens
|
|
hasContextTokens = true
|
|
}
|
|
if usage.CacheReadTokens > 0 {
|
|
total += usage.CacheReadTokens
|
|
hasContextTokens = true
|
|
}
|
|
if usage.CacheCreationTokens > 0 {
|
|
total += usage.CacheCreationTokens
|
|
hasContextTokens = true
|
|
}
|
|
if !hasContextTokens && usage.TotalTokens > 0 {
|
|
total = usage.TotalTokens
|
|
}
|
|
|
|
return total
|
|
}
|
|
|
|
// resolveContextLimit picks the first positive value from metadata,
|
|
// configured limit, and fallback — in that priority order. Returns
|
|
// 0 when none are positive.
|
|
func resolveContextLimit(metadataLimit, configLimit, fallback int64) int64 {
|
|
if metadataLimit > 0 {
|
|
return metadataLimit
|
|
}
|
|
if configLimit > 0 {
|
|
return configLimit
|
|
}
|
|
if fallback > 0 {
|
|
return fallback
|
|
}
|
|
return 0
|
|
}
|
|
|
|
// shouldCompact returns the usage percentage and whether it exceeds
|
|
// the threshold. Returns (0, false) when contextLimit is
|
|
// non-positive.
|
|
func shouldCompact(contextTokens, contextLimit int64, thresholdPercent int32) (float64, bool) {
|
|
if contextLimit <= 0 {
|
|
return 0, false
|
|
}
|
|
usagePercent := (float64(contextTokens) / float64(contextLimit)) * 100
|
|
return usagePercent, usagePercent >= float64(thresholdPercent)
|
|
}
|
|
|
|
func startCompactionDebugRun(
|
|
ctx context.Context,
|
|
options CompactionOptions,
|
|
) (context.Context, func(error)) {
|
|
if options.DebugSvc == nil || options.ChatID == uuid.Nil {
|
|
return ctx, func(error) {}
|
|
}
|
|
|
|
parentRun, ok := chatdebug.RunFromContext(ctx)
|
|
if !ok {
|
|
return ctx, func(error) {}
|
|
}
|
|
|
|
historyTipMessageID := options.HistoryTipMessageID
|
|
if historyTipMessageID == 0 {
|
|
historyTipMessageID = parentRun.HistoryTipMessageID
|
|
}
|
|
|
|
// Prefer the caller-supplied summary model identity; it can differ
|
|
// from the parent run's chat model under a compaction override.
|
|
provider := parentRun.Provider
|
|
if options.ResolvedProvider != "" {
|
|
provider = options.ResolvedProvider
|
|
}
|
|
model := parentRun.Model
|
|
if options.ResolvedModel != "" {
|
|
model = options.ResolvedModel
|
|
}
|
|
modelConfigID := parentRun.ModelConfigID
|
|
if options.ModelConfigID != uuid.Nil {
|
|
modelConfigID = options.ModelConfigID
|
|
}
|
|
|
|
// Use a separate short-lived context for the debug insert so a
|
|
// slow or locked DB cannot block the model call. Detached from
|
|
// the parent so cancellation of the compaction run still lets
|
|
// the insert reach a terminal state, matching the best-effort
|
|
// contract of debug instrumentation.
|
|
createRunCtx, createRunCancel := context.WithTimeout(
|
|
context.WithoutCancel(ctx), compactionDebugCreateRunTimeout,
|
|
)
|
|
run, err := options.DebugSvc.CreateRun(createRunCtx, chatdebug.CreateRunParams{
|
|
ChatID: options.ChatID,
|
|
RootChatID: parentRun.RootChatID,
|
|
ParentChatID: parentRun.ParentChatID,
|
|
ModelConfigID: modelConfigID,
|
|
TriggerMessageID: parentRun.TriggerMessageID,
|
|
HistoryTipMessageID: historyTipMessageID,
|
|
Kind: chatdebug.KindCompaction,
|
|
Status: chatdebug.StatusInProgress,
|
|
Provider: provider,
|
|
Model: model,
|
|
})
|
|
createRunCancel()
|
|
if err != nil {
|
|
// Debug instrumentation must not surface as a compaction failure.
|
|
return ctx, func(error) {}
|
|
}
|
|
|
|
compactionCtx := chatdebug.ContextWithRun(ctx, &chatdebug.RunContext{
|
|
RunID: run.ID,
|
|
ChatID: options.ChatID,
|
|
RootChatID: parentRun.RootChatID,
|
|
ParentChatID: parentRun.ParentChatID,
|
|
ModelConfigID: modelConfigID,
|
|
TriggerMessageID: parentRun.TriggerMessageID,
|
|
HistoryTipMessageID: historyTipMessageID,
|
|
Kind: chatdebug.KindCompaction,
|
|
Provider: provider,
|
|
Model: model,
|
|
})
|
|
|
|
return compactionCtx, func(runErr error) {
|
|
status := chatdebug.ClassifyError(runErr)
|
|
if runErr != nil && xerrors.Is(runErr, ErrInterrupted) {
|
|
status = chatdebug.StatusInterrupted
|
|
}
|
|
// Debug instrumentation must not surface as a compaction failure.
|
|
_ = options.DebugSvc.FinalizeRun(compactionCtx, chatdebug.FinalizeRunParams{
|
|
RunID: run.ID,
|
|
ChatID: options.ChatID,
|
|
Status: status,
|
|
})
|
|
}
|
|
}
|
|
|
|
// generateCompactionSummary asks the model to summarize the
|
|
// conversation so far. The provided messages should contain the
|
|
// complete history (system prompt, user/assistant turns, tool
|
|
// results). A final user message with the summary prompt is appended
|
|
// before calling the model.
|
|
func generateCompactionSummary(
|
|
ctx context.Context,
|
|
model fantasy.LanguageModel,
|
|
messages []fantasy.Message,
|
|
options CompactionOptions,
|
|
) (summary string, err error) {
|
|
summaryPrompt := make([]fantasy.Message, 0, len(messages)+1)
|
|
summaryPrompt = append(summaryPrompt, messages...)
|
|
summaryPrompt = append(summaryPrompt, fantasy.Message{
|
|
Role: fantasy.MessageRoleUser,
|
|
Content: []fantasy.MessagePart{
|
|
fantasy.TextPart{Text: options.SummaryPrompt},
|
|
},
|
|
})
|
|
toolChoice := fantasy.ToolChoiceNone
|
|
|
|
summaryCtx, finishDebugRun := startCompactionDebugRun(ctx, options)
|
|
defer func() {
|
|
// If model.Generate (or anything else below) panics, the
|
|
// named err return is still nil at this point. Without the
|
|
// recover hook we would finalize the debug run as Completed
|
|
// in the exact crash path operators rely on to diagnose
|
|
// failures. Finalize with the panic as an error status and
|
|
// re-panic so the caller's recovery still observes the
|
|
// original panic value.
|
|
if r := recover(); r != nil {
|
|
finishDebugRun(xerrors.Errorf("panic during compaction summary: %v", r))
|
|
panic(r)
|
|
}
|
|
finishDebugRun(err)
|
|
}()
|
|
|
|
response, err := model.Generate(summaryCtx, fantasy.Call{
|
|
Prompt: summaryPrompt,
|
|
ToolChoice: &toolChoice,
|
|
ProviderOptions: options.ProviderOptions,
|
|
})
|
|
if err != nil {
|
|
return "", xerrors.Errorf("generate summary text: %w", err)
|
|
}
|
|
|
|
parts := make([]string, 0, len(response.Content))
|
|
for _, block := range response.Content {
|
|
textBlock, ok := fantasy.AsContentType[fantasy.TextContent](block)
|
|
if !ok {
|
|
continue
|
|
}
|
|
text := strings.TrimSpace(textBlock.Text)
|
|
if text == "" {
|
|
continue
|
|
}
|
|
parts = append(parts, text)
|
|
}
|
|
return strings.TrimSpace(strings.Join(parts, " ")), nil
|
|
}
|