fix: reuse shared tailnet for coderd-hosted MCP workspace tools (#24460)

## Problem

Coderd can expose an MCP server at `/api/experimental/mcp/http` (we have
this enabled on dogfood). Its workspace tools dialed agents through a
per-call client-side tailnet stack. Every tool call re-created a
WireGuard device, netstack, magicsock + UDP sockets, DERP connection,
coordinator websocket, and their goroutines — in a process that already
runs a long-lived shared tailnet. The duplicate stacks drove up resource
usage under load.

## Fix

Route this server's tool calls through the existing shared tailnet, so
none of those transports are reconstructed per call. Closing an
`AgentConn` now releases a tunnel reference instead of tearing down a
transport.

## Potential follow-up

`coder exp mcp server` still builds a fresh tailnet per call. It pays
per-call latency and causes coordinator/DERP churn. A shared CLI tailnet
is more involved — unlike coderd, the CLI has no existing shared tailnet
to reuse, so it would need a new long-lived client-side tailnet with
reconnect, sleep/wake, and idle-destination handling. There's less
motivation to optimize this, given the client-side MCP does not compete
for resources with coderd.

Closes CODAGT-199

> Generated by mux, but reviewed by a human
This commit is contained in:
Ethan
2026-04-21 11:37:10 +10:00
committed by GitHub
parent 1203f625b7
commit 181e103201
8 changed files with 283 additions and 99 deletions
+4
View File
@@ -175,6 +175,10 @@ func (c *Client) AgentConnectionInfo(ctx context.Context, agentID uuid.UUID) (Ag
return connInfo, json.NewDecoder(res.Body).Decode(&connInfo)
}
// AgentConnFunc returns a new connection to the specified agent. If release is
// non-nil, callers must invoke it after they are done with the AgentConn.
type AgentConnFunc func(ctx context.Context, agentID uuid.UUID) (conn AgentConn, release func(), err error)
// @typescript-ignore DialAgentOptions
type DialAgentOptions struct {
Logger slog.Logger