mirror of
https://github.com/coder/coder.git
synced 2026-09-23 22:20:22 +08:00
Closes CODAGT-835 ## Summary `chat_messages.runtime_ms` becomes the billing source of truth for Coder Agents runtime (summed hourly by #27312), but it was built for debugging: the June refactor (#26270) silently stopped recording tool-step runtime, compaction was never measured, and interrupted turns lost their partial runtime entirely. This PR defines the billable metric, closes the paths that dropped it, and documents the definition where the data lives. ## The billable definition **`runtime_ms` is the wall-clock duration of the model invocation that produced the persisted message content**, measured from just before the provider stream opens until it is fully consumed. What counts: - Assistant generation steps, in top-level and sub-agent chats (sub-agents are ordinary chats on the same generation path). - Compaction summarization calls, persisted on the compaction assistant message (**new**). - Interrupted attempts: the message-part episode's lifetime is persisted on the partial assistant message committed by `FinishInterruption`, so partial generation time survives interruption (**new**; measured via a new `Buffer.EpisodeDuration`, which works even though the generation goroutine and the interrupt task are different tasks). What deliberately does not count (each is documented in code and docs): - **Local tool execution.** Tool wall time includes idle waits, most importantly `wait_agent` polling a sub-agent chat that already bills its own model invocations; billing the batch would double count, and excluding one tool from a concurrent batch's wall time is ill-defined. Pre-refactor instrumentation did include tool time; this makes the exclusion an explicit product definition instead of a silent regression. - **Failed model calls whose output is discarded** (retried attempts, terminal errors, content-filter refusals). They persist no content, so they bill nothing; billing errs toward undercounting. Notably a stream-silence timeout can burn 10 idle minutes before a retry, which should not be billable "active generation". If product later wants failed attempts billed, that needs a place to persist runtime on error turns (`FinishError` inserts no rows today) and is a deliberate follow-up, not instrumentation drift. - **Ancillary calls that produce no chat messages** (title generation, advisor, turn summaries) and all idle/parked time (`requires_action`, queueing). The definition is documented as `COMMENT ON COLUMN chat_messages.runtime_ms` (migration 000551, surfacing as a Go doc comment on `ChatMessage.RuntimeMs`), on `chatloop.PersistedStep.Runtime`, in the chatd architecture doc, and in the Spend Management docs page. ## Index for the hourly scan None needed: `GetTotalChatMessageRuntimeMsInRange` (#27312) filters an hour-wide `created_at` range, which the existing `idx_chat_messages_created_at` b-tree already serves; the residual `runtime_ms IS NOT NULL` filter applies to one hour of rows. A partial index would add permanent write amplification for a query that runs once an hour. > [!NOTE] > Migration 000551 is also claimed by #27312; whichever merges second renumbers via `fix_migration_numbers.sh`. ## Tests - End-to-end: the existing full-server generation test now asserts `RuntimeMs.Valid` on the committed assistant row (it previously read `.Int64` without checking `.Valid`, so it passed on NULL). - Interrupted turn: full task-level test (real DB, mock clock) asserting the partial assistant message persists the attempt's runtime. - Errored stream: asserts a failed invocation yields no step and no runtime. - Tool-using turn: asserts runtime lands on the assistant row only and tool rows stay NULL. - Compaction: asserts the summarization call duration is recorded and lands on the compaction assistant message only. - `messagepartbuffer.EpisodeDuration` unit coverage. Blocks: CODAGT-843 (B3), CODAGT-838 (D8). --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Hugo Dutka <hugo@coder.com>
484 lines
17 KiB
Go
484 lines
17 KiB
Go
package chatloop
|
|
|
|
import (
|
|
"context"
|
|
"encoding/json"
|
|
"strings"
|
|
"time"
|
|
|
|
"charm.land/fantasy"
|
|
"github.com/google/uuid"
|
|
"golang.org/x/xerrors"
|
|
|
|
"github.com/coder/coder/v2/coderd/x/chatd/chatdebug"
|
|
"github.com/coder/coder/v2/codersdk"
|
|
)
|
|
|
|
const (
|
|
defaultCompactionThresholdPercent = int32(70)
|
|
minCompactionThresholdPercent = int32(0)
|
|
maxCompactionThresholdPercent = int32(100)
|
|
|
|
// compactionDebugCreateRunTimeout caps the compaction debug
|
|
// CreateRun budget. Debug instrumentation is best-effort;
|
|
// running without the debug row is preferable to blocking
|
|
// compaction on a slow or locked DB.
|
|
compactionDebugCreateRunTimeout = 5 * time.Second
|
|
|
|
defaultCompactionSummaryPrompt = "You are performing a context compaction. " +
|
|
"Summarize the conversation so a new assistant can seamlessly " +
|
|
"continue the work in progress.\n\n" +
|
|
"Include:\n" +
|
|
// The constraints bullet below is deliberately verbose: offline replay
|
|
// of production chats showed compaction summaries dropping or softening
|
|
// user-stated constraints, and this wording measurably improved their
|
|
// survival (see PR #27230). Reword only with re-validation.
|
|
"- User constraints, corrections, and prohibitions: rules, " +
|
|
"scope limits, style rules, and process corrections stated by " +
|
|
"the user. Quote or closely paraphrase the user's wording; do " +
|
|
"not soften, merge, or truncate them. Constraints are standing " +
|
|
"until the user revokes them; they do not become stale when " +
|
|
"the task moves on. When the user corrected the assistant's " +
|
|
"behavior, record the correction itself, not only the " +
|
|
"corrected outcome. Include only constraints the user stated " +
|
|
"in conversation; do not place rules from system prompts, " +
|
|
"AGENTS.md, or other configuration files in this section. " +
|
|
"Those rules may appear elsewhere in the summary with their " +
|
|
"true source named. When in doubt whether a rule originated " +
|
|
"from the user, name its source or omit the attribution " +
|
|
"rather than defaulting to user.\n" +
|
|
"- The user's overall goal and current task\n" +
|
|
"- Key decisions made and their rationale\n" +
|
|
"- Concrete technical details: file paths, function names, " +
|
|
"commands, APIs, and configurations\n" +
|
|
"- Errors encountered and how they were resolved. Keep error " +
|
|
"notes specific: name the file, the error, and the fix. Do not " +
|
|
"generalize from a specific failure to a blanket tool-avoidance " +
|
|
"rule (e.g. \"tool X is unreliable\" or \"always use Y instead " +
|
|
"of Z\")\n" +
|
|
"- Current state of the work: what is DONE, what is IN PROGRESS, " +
|
|
"and what REMAINS to be done\n" +
|
|
"- The specific action the assistant was performing or about to " +
|
|
"perform when this summary was triggered\n\n" +
|
|
"Be dense and factual. Every sentence should convey essential " +
|
|
"context for continuation. Do not include pleasantries or " +
|
|
"conversational filler. For content that can be reproduced " +
|
|
"(repo files, command output, API responses), reference how to " +
|
|
"obtain it (file path, command, URL) rather than inlining the " +
|
|
"full content. Include brief inline summaries when the content " +
|
|
"itself would exceed a few lines."
|
|
defaultCompactionSystemSummaryPrefix = "The following is a summary of " +
|
|
"the earlier conversation. The assistant was actively working when " +
|
|
"the context was compacted. Continue the work described below:"
|
|
)
|
|
|
|
// CompactionSource identifies what triggered a compaction. It is
|
|
// recorded in the persisted chat_summarized tool JSON and the
|
|
// streamed synthetic parts so clients can render manual compactions
|
|
// distinctly.
|
|
type CompactionSource string
|
|
|
|
const (
|
|
CompactionSourceAutomatic CompactionSource = "automatic"
|
|
CompactionSourceManual CompactionSource = "manual"
|
|
)
|
|
|
|
type CompactionOptions struct {
|
|
ThresholdPercent int32
|
|
ContextLimit int64
|
|
SummaryPrompt string
|
|
SummaryHint string
|
|
SystemSummaryPrefix string
|
|
Persist func(context.Context, CompactionResult) error
|
|
DebugSvc *chatdebug.Service
|
|
ChatID uuid.UUID
|
|
HistoryTipMessageID int64
|
|
|
|
// Summary model identity and call options; see
|
|
// GenerateCompactionOptions.
|
|
ResolvedProvider string
|
|
ResolvedModel string
|
|
ModelConfigID uuid.UUID
|
|
ProviderOptions fantasy.ProviderOptions
|
|
|
|
// Force skips the threshold gate (including the threshold=100
|
|
// disable and the zero-usage early return). Set for manual,
|
|
// user-requested compactions.
|
|
Force bool
|
|
// Source labels what triggered the compaction. Defaults to
|
|
// CompactionSourceAutomatic when empty.
|
|
Source CompactionSource
|
|
|
|
// ToolCallID and ToolName identify the synthetic tool call
|
|
// used to represent compaction in the message stream.
|
|
ToolCallID string
|
|
ToolName string
|
|
|
|
// PublishMessagePart publishes streaming parts to connected
|
|
// clients so they see "Summarizing..." / "Summarized" UI
|
|
// transitions during compaction.
|
|
PublishMessagePart func(codersdk.ChatMessageRole, codersdk.ChatMessagePart)
|
|
|
|
OnError func(error)
|
|
}
|
|
|
|
type CompactionResult struct {
|
|
SystemSummary string
|
|
SummaryReport string
|
|
Source CompactionSource
|
|
ThresholdPercent int32
|
|
UsagePercent float64
|
|
ContextTokens int64
|
|
ContextLimit int64
|
|
// Runtime is the wall-clock duration of the summarization model
|
|
// call, the compaction step's billable runtime (see
|
|
// PersistedStep.Runtime). Zero when the run was gated off before
|
|
// calling the model.
|
|
Runtime time.Duration
|
|
}
|
|
|
|
// GenerateCompaction generates one context summary and returns it without
|
|
// persisting. It publishes compaction progress parts when configured.
|
|
// Threshold gating (including the threshold=100 disable and the
|
|
// zero-usage early return) is skipped when opts.Force is set.
|
|
func GenerateCompaction(ctx context.Context, opts GenerateCompactionOptions) (CompactionResult, error) {
|
|
if opts.Model == nil {
|
|
return CompactionResult{}, xerrors.New("chat model is required")
|
|
}
|
|
if opts.Clock == nil {
|
|
return CompactionResult{}, xerrors.New("clock is required")
|
|
}
|
|
config, ok := normalizedCompactionGenerateConfig(opts)
|
|
if !ok {
|
|
return CompactionResult{}, nil
|
|
}
|
|
|
|
contextTokens := contextTokensFromUsage(opts.StepUsage)
|
|
if contextTokens <= 0 && !config.Force {
|
|
return CompactionResult{}, nil
|
|
}
|
|
metadataLimit := extractContextLimit(opts.StepMetadata)
|
|
contextLimit := resolveContextLimit(
|
|
metadataLimit.Int64,
|
|
config.ContextLimit,
|
|
opts.ContextLimitFallback,
|
|
)
|
|
usagePercent, compact := shouldCompact(
|
|
contextTokens,
|
|
contextLimit,
|
|
config.ThresholdPercent,
|
|
)
|
|
if !compact && !config.Force {
|
|
return CompactionResult{}, nil
|
|
}
|
|
|
|
if config.PublishMessagePart != nil && config.ToolCallID != "" {
|
|
config.PublishMessagePart(
|
|
codersdk.ChatMessageRoleAssistant,
|
|
codersdk.ChatMessageToolCall(config.ToolCallID, config.ToolName, nil),
|
|
)
|
|
}
|
|
|
|
summaryStart := opts.Clock.Now()
|
|
if opts.OnModelStreamStart != nil {
|
|
opts.OnModelStreamStart()
|
|
}
|
|
summary, err := generateCompactionSummary(ctx, opts.Model, opts.Messages, config)
|
|
if err != nil {
|
|
publishCompactionError(config, "failed to generate compaction summary")
|
|
return CompactionResult{}, err
|
|
}
|
|
summaryRuntime := opts.Clock.Since(summaryStart)
|
|
if summary == "" {
|
|
publishCompactionError(config, "compaction produced an empty summary")
|
|
return CompactionResult{}, xerrors.New("compaction produced an empty summary")
|
|
}
|
|
|
|
result := CompactionResult{
|
|
SystemSummary: strings.TrimSpace(
|
|
config.SystemSummaryPrefix + "\n\n" + summary,
|
|
),
|
|
SummaryReport: summary,
|
|
Source: config.Source,
|
|
ThresholdPercent: config.ThresholdPercent,
|
|
UsagePercent: usagePercent,
|
|
ContextTokens: contextTokens,
|
|
ContextLimit: contextLimit,
|
|
Runtime: summaryRuntime,
|
|
}
|
|
if config.PublishMessagePart != nil && config.ToolCallID != "" {
|
|
resultJSON, _ := json.Marshal(map[string]any{
|
|
"summary": summary,
|
|
"source": config.Source,
|
|
"threshold_percent": config.ThresholdPercent,
|
|
"usage_percent": usagePercent,
|
|
"context_tokens": contextTokens,
|
|
"context_limit_tokens": contextLimit,
|
|
})
|
|
config.PublishMessagePart(
|
|
codersdk.ChatMessageRoleTool,
|
|
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, resultJSON, false, false),
|
|
)
|
|
}
|
|
return result, nil
|
|
}
|
|
|
|
func normalizedCompactionGenerateConfig(opts GenerateCompactionOptions) (CompactionOptions, bool) {
|
|
config := CompactionOptions{
|
|
ThresholdPercent: opts.ThresholdPercent,
|
|
ContextLimit: opts.ContextLimit,
|
|
SummaryPrompt: opts.SummaryPrompt,
|
|
SummaryHint: opts.SummaryHint,
|
|
SystemSummaryPrefix: opts.SystemSummaryPrefix,
|
|
DebugSvc: opts.DebugSvc,
|
|
ChatID: opts.ChatID,
|
|
HistoryTipMessageID: opts.HistoryTipMessageID,
|
|
ResolvedProvider: opts.ResolvedProvider,
|
|
ResolvedModel: opts.ResolvedModel,
|
|
ModelConfigID: opts.ModelConfigID,
|
|
ProviderOptions: opts.ProviderOptions,
|
|
Force: opts.Force,
|
|
Source: opts.Source,
|
|
ToolCallID: opts.ToolCallID,
|
|
ToolName: opts.ToolName,
|
|
PublishMessagePart: opts.PublishMessagePart,
|
|
}
|
|
if strings.TrimSpace(config.SummaryPrompt) == "" {
|
|
config.SummaryPrompt = defaultCompactionSummaryPrompt
|
|
}
|
|
if strings.TrimSpace(config.SystemSummaryPrefix) == "" {
|
|
config.SystemSummaryPrefix = defaultCompactionSystemSummaryPrefix
|
|
}
|
|
if config.Source == "" {
|
|
config.Source = CompactionSourceAutomatic
|
|
}
|
|
if config.ThresholdPercent < minCompactionThresholdPercent ||
|
|
config.ThresholdPercent > maxCompactionThresholdPercent {
|
|
config.ThresholdPercent = defaultCompactionThresholdPercent
|
|
}
|
|
// threshold=100 disables automatic compaction; a forced run
|
|
// still proceeds because the user asked explicitly.
|
|
if config.ThresholdPercent == maxCompactionThresholdPercent && !config.Force {
|
|
return CompactionOptions{}, false
|
|
}
|
|
return config, true
|
|
}
|
|
|
|
// publishCompactionError sends a tool-result error part so
|
|
// connected clients see that compaction failed.
|
|
func publishCompactionError(config CompactionOptions, msg string) {
|
|
if config.PublishMessagePart == nil || config.ToolCallID == "" {
|
|
return
|
|
}
|
|
errJSON, _ := json.Marshal(map[string]any{
|
|
"error": msg,
|
|
})
|
|
config.PublishMessagePart(
|
|
codersdk.ChatMessageRoleTool,
|
|
codersdk.ChatMessageToolResult(config.ToolCallID, config.ToolName, errJSON, true, false),
|
|
)
|
|
}
|
|
|
|
// contextTokensFromUsage returns the total context token count from
|
|
// a step's usage report. It sums input, cache-read, and
|
|
// cache-creation tokens when available, falling back to TotalTokens
|
|
// if none of the granular fields are set.
|
|
func contextTokensFromUsage(usage fantasy.Usage) int64 {
|
|
total := int64(0)
|
|
hasContextTokens := false
|
|
|
|
if usage.InputTokens > 0 {
|
|
total += usage.InputTokens
|
|
hasContextTokens = true
|
|
}
|
|
if usage.CacheReadTokens > 0 {
|
|
total += usage.CacheReadTokens
|
|
hasContextTokens = true
|
|
}
|
|
if usage.CacheCreationTokens > 0 {
|
|
total += usage.CacheCreationTokens
|
|
hasContextTokens = true
|
|
}
|
|
if !hasContextTokens && usage.TotalTokens > 0 {
|
|
total = usage.TotalTokens
|
|
}
|
|
|
|
return total
|
|
}
|
|
|
|
// resolveContextLimit picks the first positive value from metadata,
|
|
// configured limit, and fallback — in that priority order. Returns
|
|
// 0 when none are positive.
|
|
func resolveContextLimit(metadataLimit, configLimit, fallback int64) int64 {
|
|
if metadataLimit > 0 {
|
|
return metadataLimit
|
|
}
|
|
if configLimit > 0 {
|
|
return configLimit
|
|
}
|
|
if fallback > 0 {
|
|
return fallback
|
|
}
|
|
return 0
|
|
}
|
|
|
|
// shouldCompact returns the usage percentage and whether it exceeds
|
|
// the threshold. Returns (0, false) when contextLimit is
|
|
// non-positive.
|
|
func shouldCompact(contextTokens, contextLimit int64, thresholdPercent int32) (float64, bool) {
|
|
if contextLimit <= 0 {
|
|
return 0, false
|
|
}
|
|
usagePercent := (float64(contextTokens) / float64(contextLimit)) * 100
|
|
return usagePercent, usagePercent >= float64(thresholdPercent)
|
|
}
|
|
|
|
func startCompactionDebugRun(
|
|
ctx context.Context,
|
|
options CompactionOptions,
|
|
) (context.Context, func(error)) {
|
|
if options.DebugSvc == nil || options.ChatID == uuid.Nil {
|
|
return ctx, func(error) {}
|
|
}
|
|
|
|
parentRun, ok := chatdebug.RunFromContext(ctx)
|
|
if !ok {
|
|
return ctx, func(error) {}
|
|
}
|
|
|
|
historyTipMessageID := options.HistoryTipMessageID
|
|
if historyTipMessageID == 0 {
|
|
historyTipMessageID = parentRun.HistoryTipMessageID
|
|
}
|
|
|
|
// Prefer the caller-supplied summary model identity; it can differ
|
|
// from the parent run's chat model under a compaction override.
|
|
provider := parentRun.Provider
|
|
if options.ResolvedProvider != "" {
|
|
provider = options.ResolvedProvider
|
|
}
|
|
model := parentRun.Model
|
|
if options.ResolvedModel != "" {
|
|
model = options.ResolvedModel
|
|
}
|
|
modelConfigID := parentRun.ModelConfigID
|
|
if options.ModelConfigID != uuid.Nil {
|
|
modelConfigID = options.ModelConfigID
|
|
}
|
|
|
|
// Use a separate short-lived context for the debug insert so a
|
|
// slow or locked DB cannot block the model call. Detached from
|
|
// the parent so cancellation of the compaction run still lets
|
|
// the insert reach a terminal state, matching the best-effort
|
|
// contract of debug instrumentation.
|
|
createRunCtx, createRunCancel := context.WithTimeout(
|
|
context.WithoutCancel(ctx), compactionDebugCreateRunTimeout,
|
|
)
|
|
run, err := options.DebugSvc.CreateRun(createRunCtx, chatdebug.CreateRunParams{
|
|
ChatID: options.ChatID,
|
|
RootChatID: parentRun.RootChatID,
|
|
ParentChatID: parentRun.ParentChatID,
|
|
ModelConfigID: modelConfigID,
|
|
TriggerMessageID: parentRun.TriggerMessageID,
|
|
HistoryTipMessageID: historyTipMessageID,
|
|
Kind: chatdebug.KindCompaction,
|
|
Status: chatdebug.StatusInProgress,
|
|
Provider: provider,
|
|
Model: model,
|
|
})
|
|
createRunCancel()
|
|
if err != nil {
|
|
// Debug instrumentation must not surface as a compaction failure.
|
|
return ctx, func(error) {}
|
|
}
|
|
|
|
compactionCtx := chatdebug.ContextWithRun(ctx, &chatdebug.RunContext{
|
|
RunID: run.ID,
|
|
ChatID: options.ChatID,
|
|
RootChatID: parentRun.RootChatID,
|
|
ParentChatID: parentRun.ParentChatID,
|
|
ModelConfigID: modelConfigID,
|
|
TriggerMessageID: parentRun.TriggerMessageID,
|
|
HistoryTipMessageID: historyTipMessageID,
|
|
Kind: chatdebug.KindCompaction,
|
|
Provider: provider,
|
|
Model: model,
|
|
})
|
|
|
|
return compactionCtx, func(runErr error) {
|
|
status := chatdebug.ClassifyError(runErr)
|
|
if runErr != nil && xerrors.Is(runErr, ErrInterrupted) {
|
|
status = chatdebug.StatusInterrupted
|
|
}
|
|
// Debug instrumentation must not surface as a compaction failure.
|
|
_ = options.DebugSvc.FinalizeRun(compactionCtx, chatdebug.FinalizeRunParams{
|
|
RunID: run.ID,
|
|
ChatID: options.ChatID,
|
|
Status: status,
|
|
})
|
|
}
|
|
}
|
|
|
|
// generateCompactionSummary asks the model to summarize the
|
|
// conversation so far. The provided messages should contain the
|
|
// complete history (system prompt, user/assistant turns, tool
|
|
// results). A final user message with the summary prompt is appended
|
|
// before calling the model.
|
|
func generateCompactionSummary(
|
|
ctx context.Context,
|
|
model fantasy.LanguageModel,
|
|
messages []fantasy.Message,
|
|
options CompactionOptions,
|
|
) (summary string, err error) {
|
|
summaryPrompt := make([]fantasy.Message, 0, len(messages)+1)
|
|
summaryPrompt = append(summaryPrompt, messages...)
|
|
summaryParts := []fantasy.MessagePart{fantasy.TextPart{Text: options.SummaryPrompt}}
|
|
if strings.TrimSpace(options.SummaryHint) != "" {
|
|
summaryParts = append(summaryParts, fantasy.TextPart{Text: options.SummaryHint})
|
|
}
|
|
summaryPrompt = append(summaryPrompt, fantasy.Message{
|
|
Role: fantasy.MessageRoleUser,
|
|
Content: summaryParts,
|
|
})
|
|
toolChoice := fantasy.ToolChoiceNone
|
|
|
|
summaryCtx, finishDebugRun := startCompactionDebugRun(ctx, options)
|
|
defer func() {
|
|
// If model.Generate (or anything else below) panics, the
|
|
// named err return is still nil at this point. Without the
|
|
// recover hook we would finalize the debug run as Completed
|
|
// in the exact crash path operators rely on to diagnose
|
|
// failures. Finalize with the panic as an error status and
|
|
// re-panic so the caller's recovery still observes the
|
|
// original panic value.
|
|
if r := recover(); r != nil {
|
|
finishDebugRun(xerrors.Errorf("panic during compaction summary: %v", r))
|
|
panic(r)
|
|
}
|
|
finishDebugRun(err)
|
|
}()
|
|
|
|
response, err := model.Generate(summaryCtx, fantasy.Call{
|
|
Prompt: summaryPrompt,
|
|
ToolChoice: &toolChoice,
|
|
ProviderOptions: options.ProviderOptions,
|
|
})
|
|
if err != nil {
|
|
return "", xerrors.Errorf("generate summary text: %w", err)
|
|
}
|
|
|
|
parts := make([]string, 0, len(response.Content))
|
|
for _, block := range response.Content {
|
|
textBlock, ok := fantasy.AsContentType[fantasy.TextContent](block)
|
|
if !ok {
|
|
continue
|
|
}
|
|
text := strings.TrimSpace(textBlock.Text)
|
|
if text == "" {
|
|
continue
|
|
}
|
|
parts = append(parts, text)
|
|
}
|
|
return strings.TrimSpace(strings.Join(parts, " ")), nil
|
|
}
|