Files
teleport/lib/utils/workpool
Trent Clarke 97c18fa1a9 Restart entire node on tunnel collapse (#8102)
Fixes #7606, where a node doesn't notice when the tunnel port changes. 

Imagine you have a cluster with a node connected in via a tunnel through a proxy `proxy.example.com` on port `3024`

Now change the proxy config so that `tunnel_public_address` is `proxy.example.com:4024`. You either restart the proxy, or reload the proxy config with a `SIGHUP`.

...and then the node
  a) loses its connection to auth (because the tunnel is gone), and 
  b)  _doesn't reconnect_, because even though the proxy address hasn't changed,
      the node has cached the old tunnel_public_address and keeps trying to connect
      to that.

You can always manually restart the node to have it reconnect, but that would be a pain if you have thousands of nodes.

In order to not have to manually restart all nodes, this change implements a check for a connection failures to the auth server, and re-starts the node if there are multiple connection failures in a given period of time. The check as-implemented piggybacks on the node's "common.rotate" service, which can already restart the node in certain circumstances, and uses the success of the periodic rotation sync as a proxy for the health of the node's connection to the auth server.

See-Also: #7606
2021-11-17 12:01:48 +11:00
..