mirror of
https://github.com/coder/coder.git
synced 2026-09-22 05:05:20 +08:00
Fixes coder/internal#1519 Fixes CODAGT-353 These nine tests were skipped pending a chatd notification flow refactor that would let workers distinguish stale control `NOTIFY` messages from real interrupts. They now pass consistently, so this drops the `t.Skip` calls and the now-stale `TODO(CODAGT-353)` blocks. While rerunning the package after unskipping them, `TestNewReplicaRecoversStaleChatFromDeadReplica` also surfaced as flaky on `main` because it asserted transient ownership state. This PR keeps that server-level test as a stable end-to-end recovery check and adds a deterministic worker-level stale reacquisition test so we still directly cover lease takeover behavior. ## Tests unskipped - `coderd`: `TestPatchChatMessage/ChangesModel` - `coderd/x/chatd`: - `TestExploreChatSendMessageCannotMutateMCPSnapshot` - `TestAutoPromoteQueuedMessagesPreservesPerTurnModelOrder` - `TestSignalWakeSendMessage` - `TestAdvisorChainMode_SnapshotKeepsFullHistory` - `TestOpenAIResponsesNoStaleWebSearchReplay` - `TestOpenAIResponsesFullReplayPairsReasoningAndWebSearch` - `TestOpenAIResponsesChainModeSkipsWhenLocalCallPending` - `TestOpenAIResponsesChainModeStillFiresForProviderExecutedOnly` ## Stale recovery follow-up - `coderd/x/chatd`: `TestNewReplicaRecoversStaleChatFromDeadReplica` now waits for the stable `waiting` and unowned end state after recovery. - `coderd/x/chatd`: `TestWorker_ReacquiresStaleOwnedChat` blocks the runner after reacquisition and directly asserts the new worker ownership, new runner ID, and fresh heartbeat. I stress-ran the stale recovery tests locally with repeated plain and race runs, and re-ran the nine unskipped tests plus `TestPatchChatMessage/ChangesModel` after these follow-up changes.