mirror of
https://github.com/coder/coder.git
synced 2026-09-24 15:04:27 +08:00
fix(chatd): preserve context.Canceled in persistStep during shutdown (#22890)
## Problem
When a chat worker shuts down gracefully (e.g. Kubernetes pod SIGTERM)
while a tool is executing (like `wait_agent` polling for a subagent),
the chat gets stuck in `waiting` status forever — no other worker will
pick it up.
### Root Cause
`persistStep` in `chatd.go` unconditionally returned
`chatloop.ErrInterrupted` for **any** canceled context:
```go
if persistCtx.Err() != nil {
return chatloop.ErrInterrupted // BUG: doesn't check WHY the context was canceled
}
```
During shutdown, the context cause is `context.Canceled` (not
`ErrInterrupted`). But because `persistStep` returned `ErrInterrupted`,
the error handling in `processChat` hit the `ErrInterrupted` check first
(line 2011) and set status to `waiting` — the `isShutdownCancellation`
check (line 2017) was never reached:
```go
// Checked FIRST — matches because persistStep returned ErrInterrupted
if errors.Is(err, chatloop.ErrInterrupted) {
status = database.ChatStatusWaiting // Stuck forever
return
}
// NEVER REACHED during shutdown
if isShutdownCancellation(ctx, chatCtx, err) {
status = database.ChatStatusPending // Would have been correct
return
}
```
### Trigger scenario (from production logs)
1. Chat spawns a subagent via `spawn_agent`, then calls `wait_agent`
2. `wait_agent` blocks in `awaitSubagentCompletion` polling loop
3. Worker pod receives SIGTERM → `Close()` cancels server context
4. Context cancellation propagates to `awaitSubagentCompletion` →
returns `context.Canceled`
5. Tool execution completes, `persistStep` is called with canceled
context
6. `persistStep` returns `ErrInterrupted` (wrong!) → status set to
`waiting` (stuck!)
## Fix
Check `context.Cause()` before deciding which error to return:
```go
if persistCtx.Err() != nil {
if errors.Is(context.Cause(persistCtx), chatloop.ErrInterrupted) {
return chatloop.ErrInterrupted // Intentional interruption
}
return persistCtx.Err() // Shutdown → context.Canceled
}
```
This preserves `context.Canceled` for shutdown, allowing
`isShutdownCancellation` to match and set status to `pending` so another
worker retries the chat.
## Test
Added `TestRun_ShutdownDuringToolExecutionReturnsContextCanceled` which:
1. Streams a tool call to a blocking tool (simulating `wait_agent`)
2. Cancels the server context (simulating shutdown) while the tool
blocks
3. Verifies `Run` returns `context.Canceled`, NOT `ErrInterrupted`
This commit is contained in:
+14
-7
@@ -2216,14 +2216,21 @@ func (p *Server) runChat(
|
||||
modelConfigContextLimit := modelConfig.ContextLimit
|
||||
|
||||
persistStep := func(persistCtx context.Context, step chatloop.PersistedStep) error {
|
||||
// If the chat context has been canceled (e.g. by an
|
||||
// EditMessage call), bail out before inserting any
|
||||
// messages. This closes the race window between
|
||||
// EditMessage committing its transaction (which deletes
|
||||
// messages after the edit point) and the cancellation
|
||||
// propagating to the processing loop.
|
||||
// If the chat context has been canceled, bail out before
|
||||
// inserting any messages. We distinguish the cause so that
|
||||
// the caller can tell an intentional interruption (e.g.
|
||||
// EditMessage, user stop) from a server shutdown:
|
||||
// - ErrInterrupted cause → return ErrInterrupted
|
||||
// (processChat sets status = waiting).
|
||||
// - Any other cause (e.g. context.Canceled during
|
||||
// Close()) → return the original context error so
|
||||
// isShutdownCancellation can match and set status =
|
||||
// pending, allowing another replica to retry.
|
||||
if persistCtx.Err() != nil {
|
||||
return chatloop.ErrInterrupted
|
||||
if errors.Is(context.Cause(persistCtx), chatloop.ErrInterrupted) {
|
||||
return chatloop.ErrInterrupted
|
||||
}
|
||||
return persistCtx.Err()
|
||||
}
|
||||
|
||||
// Split the step content into assistant blocks and tool
|
||||
|
||||
Reference in New Issue
Block a user