fix: fix graceful disconnect in DialWorkspaceAgent (#11993)

I noticed in testing that the CLI wasn't correctly sending the disconnect message when it shuts down, and thus agents are seeing this as a "lost" peer, rather than a "disconnected" one. 

What was happening is that we just used a single context for everything from the netconn to the RPCs, and when the context was canceled we failed to send the disconnect message due to canceled context.

So, this PR splits things into two contexts, with a graceful one set to last up to 1 second longer than the main one.
This commit is contained in:
Spike Curtis
2024-02-05 14:01:37 +04:00
committed by GitHub
parent bb99cb7d2b
commit e5ba586e30
3 changed files with 190 additions and 39 deletions
+2 -1
View File
@@ -134,6 +134,7 @@ func (c *remoteCoordination) Close() (retErr error) {
if err != nil {
return xerrors.Errorf("send disconnect: %w", err)
}
c.logger.Debug(context.Background(), "sent disconnect")
return nil
}
@@ -167,7 +168,7 @@ func (c *remoteCoordination) respLoop() {
}
}
// NewRemoteCoordination uses the provided protocol to coordinate the provided coordinee (usually a
// NewRemoteCoordination uses the provided protocol to coordinate the provided coordinatee (usually a
// Conn). If the tunnelTarget is not uuid.Nil, then we add a tunnel to the peer (i.e. we are acting as
// a client---agents should NOT set this!).
func NewRemoteCoordination(logger slog.Logger,