Files
coder/coderd/x
Ethan bfe64d6355 test(coderd): unskip chatd notification flow tests (#26366)
Fixes coder/internal#1519
Fixes CODAGT-353

These nine tests were skipped pending a chatd notification flow refactor
that would let workers distinguish stale control `NOTIFY` messages from
real interrupts. They now pass consistently, so this drops the `t.Skip`
calls and the now-stale `TODO(CODAGT-353)` blocks.

While rerunning the package after unskipping them,
`TestNewReplicaRecoversStaleChatFromDeadReplica` also surfaced as flaky
on `main` because it asserted transient ownership state. This PR keeps
that server-level test as a stable end-to-end recovery check and adds a
deterministic worker-level stale reacquisition test so we still directly
cover lease takeover behavior.

## Tests unskipped

- `coderd`: `TestPatchChatMessage/ChangesModel`
- `coderd/x/chatd`:
  - `TestExploreChatSendMessageCannotMutateMCPSnapshot`
  - `TestAutoPromoteQueuedMessagesPreservesPerTurnModelOrder`
  - `TestSignalWakeSendMessage`
  - `TestAdvisorChainMode_SnapshotKeepsFullHistory`
  - `TestOpenAIResponsesNoStaleWebSearchReplay`
  - `TestOpenAIResponsesFullReplayPairsReasoningAndWebSearch`
  - `TestOpenAIResponsesChainModeSkipsWhenLocalCallPending`
  - `TestOpenAIResponsesChainModeStillFiresForProviderExecutedOnly`

## Stale recovery follow-up

- `coderd/x/chatd`: `TestNewReplicaRecoversStaleChatFromDeadReplica` now
waits for the stable `waiting` and unowned end state after recovery.
- `coderd/x/chatd`: `TestWorker_ReacquiresStaleOwnedChat` blocks the
runner after reacquisition and directly asserts the new worker
ownership, new runner ID, and fresh heartbeat.

I stress-ran the stale recovery tests locally with repeated plain and
race runs, and re-ran the nine unskipped tests plus
`TestPatchChatMessage/ChangesModel` after these follow-up changes.
2026-06-15 20:24:13 +10:00
..