Files
coder/docs/install/upgrade-best-practices.md
T
Nick Vigilante e458692cb8 refactor(docs): convert absolute coder/coder blob/tree/main links to relative (DOCS-351) (#26341)
Closes [DOCS-351](https://linear.app/codercom/issue/DOCS-351).

> [!WARNING]
> **DO NOT MERGE** until
[DOCS-349](https://linear.app/codercom/issue/DOCS-349)
([coder.com#877](https://github.com/coder/coder.com/pull/877)) has
shipped to production and baked for at least one Vercel cycle.
>
> Without DOCS-349, the relative links in this PR resolve to broken
docs-route URLs (`/docs/helm/coder/values.yaml` -> 404) instead of
GitHub URLs tagged with the displayed docs version. DOCS-349 fixes the
rewriter to classify these as GitHub blob/tree URLs with the page's
resolved ref.

## TL;DR

Converts 121 absolute
`https://github.com/coder/coder/(blob|tree)/main/<path>` links across 39
docs markdown files to relative paths. After this lands AND DOCS-349
deploys, every one of these links will follow the displayed docs version
(mainline tag on bare URLs, explicit tag on `/@vX.Y.Z/`, `main` on
`/@main/`) instead of always pointing to `main`.

## Why

Today a reader on `/docs/@v2.30.0/install/docker` follows a
`compose.yaml` link and arrives at `main`'s `compose.yaml`, which
doesn't necessarily match what the docs page describes. Helm values,
Terraform templates, and source-code references in particular drift
across versions. The fix is to let the coder.com rewriter substitute the
page's resolved ref into the URL; that only works on relative links.

## Example payoff (post-DOCS-349)

| URL | Today (absolute, always `main`) | After (relative + rewriter) |
|---|---|---|
| `/docs/install/docker` |
`https://github.com/coder/coder/blob/main/compose.yaml` |
`https://github.com/coder/coder/blob/v2.34.1/compose.yaml` (today's
mainline) |
| `/docs/@v2.30.0/install/docker` | same as above |
`https://github.com/coder/coder/blob/v2.30.0/compose.yaml` |
| `/docs/@main/install/docker` | same as above |
`https://github.com/coder/coder/blob/main/compose.yaml` |

## Scope

- **121 conversions** across **39 files**.
- Verb breakdown: `tree/main` (directories) and `blob/main` (files),
both flipped to relative paths.
- Line anchors (`#L23-L24`) and query strings preserved verbatim.
- Conversion is mechanical: relative path computed from the doc file's
directory to the target via `os.path.relpath`. Any path starting at the
same directory or below gets a `./` prefix; otherwise `../` chains.

## Rebased on main

The branch was rebased onto `main` after the DOCS-350 hotfix
([#26339](https://github.com/coder/coder/pull/26339)) merged. The hotfix
repointed 3 `docs-backend-contrib-guide` refs in `backend.md` to `main`,
which then needed the same `main` -> relative conversion this PR is
doing for the other 121 links. The conflict was resolved by reapplying
the mechanical conversion to `backend.md` after taking the hotfix's
content. Net result: those 3 links land here as relative, same as
everything else. New HEAD `3f501cb622`.

## Inline fix folded in: dead `nix` link

- `docs/about/contributing/CONTRIBUTING.md:7` -> `../../../nix`

The original absolute URL `https://github.com/coder/coder/tree/main/nix`
already returned 404 today. Repointed to `flake.nix` (modern Nix
entrypoint, what the prose "Nix environment" semantically refers to).
Closes [DOCS-357](https://linear.app/codercom/issue/DOCS-357) here since
the `check-docs` Linkspector job surfaced it during rebase; cheaper to
fix inline than in a separate single-line PR.

## Out of scope (filed separately)

- [DOCS-350](https://linear.app/codercom/issue/DOCS-350): 3 dead
`docs-backend-contrib-guide` branch refs in `backend.md`
([#26339](https://github.com/coder/coder/pull/26339), merged).
- [DOCS-352](https://linear.app/codercom/issue/DOCS-352): 10 SHA-pinned
`(blob|tree)/<sha>` links pending intent review.
- [DOCS-355](https://linear.app/codercom/issue/DOCS-355): code-server
analog (4 absolute `(blob|tree)/main` links in `coder/code-server`).
- [DOCS-356](https://linear.app/codercom/issue/DOCS-356): 2 upstream
content bugs in `coder/code-server/docs/CONTRIBUTING.md` (independent of
this PR).


## Not triggering `/coder-agents-review`

Docs-only edit; per `AGENTS.md` the bot review is reserved for
product/CI changes.

## Pre-mortem

| Concern | Mitigation |
|---|---|
| Merging before DOCS-349 deploys regresses ~120 currently-working links
into 404s on coder.com | Clear DO-NOT-MERGE banner; tracked as blocker
in Linear. |
| Relative path computed incorrectly (off-by-one `..`) | Verified all
114 newly-relative non-md/non-image paths resolve to existing files in
the repo (only exception is the pre-existing dead `nix` link above). |
| Line anchors stripped during conversion | Preserved by the
substitution regex; verified `#L<n>-L<m>` cases in `airgap.md` and
`speed-up-templates.md`. |
| Future code reorgs change file locations | Relative links will start
pointing to nothing. Same failure mode as absolute links pointing to
renamed files; can be caught with a future link-checker job. |

## Validation

```
$ grep -rE 'github\.com/coder/coder/(blob|tree)/main' docs --include="*.md" | wc -l
0
$ git diff --stat origin/main | tail -1
39 files changed, 118 insertions(+), 118 deletions(-)
```

114 newly-relative paths verified to resolve to existing repo files
(Python `os.path.exists` check on each computed target).

<details>
<summary>Decision log + planning context</summary>

**Why relative over `(blob|tree)/{{currentDocsVersion}}/...`
templating**: relative paths require zero markdown-system support and
zero upstream churn beyond this one PR. Templating would require a
preprocessor on `coder.com` side AND a convention upstream authors have
to remember; relative paths just work in a plain editor and
`github.com`'s own renderer too.

**Why `./` prefix on same-directory targets**: makes the conversion
grep-able later (`grep -E '\((\.\./|\./)'`).

**Why preserve `#L<n>-L<m>` anchors verbatim**: the anchor is meaningful
to the linked file's content, not to the URL form; keeping it as-is
preserves authorial intent. If the file later changes such that the line
range drifts, that's a different problem the SHA-pin audit
([DOCS-352](https://linear.app/codercom/issue/DOCS-352)) will surface.

</details>

---

*Generated by Coder Agents on @nickvigilante's behalf.*





## Drive-by external link fix folded in

`docs/about/contributing/CONTRIBUTING.md:296` cited
`https://reflectoring.io/meaningful-commit-messages/` which is returning
HTTP 503 (the host appears to be down site-wide right now). `check-docs`
Linkspector flagged it after the rebase. Replaced with
`https://cbea.ms/git-commit/` (Chris Beams' canonical "If applied, this
commit will..." article, confirmed 200), which is the original source of
the rule the prose recites anyway.
2026-06-22 11:39:12 -04:00

7.8 KiB

Upgrading Best Practices

This guide provides best practices for upgrading Coder, along with troubleshooting steps for common issues encountered during upgrades, particularly with database migrations in high availability (HA) deployments.

Before you upgrade

Tip

To check your current Coder version, use coder version from the CLI, check the bottom-right of the Coder dashboard, or query the /api/v2/buildinfo endpoint. See the version command for details.

  • Schedule upgrades during off-peak hours. Upgrades can cause a noticeable disruption to the developer experience. Plan your maintenance window when the fewest developers are actively using their workspaces.
  • The larger the version jump, the more migrations will run. If you are upgrading across multiple minor versions, expect longer migration times.
  • Large upgrades should complete in minutes (typically 4-7 minutes). If your upgrade is taking significantly longer, there may be an issue requiring investigation.
  • Check for known issues affecting your upgrade path. Some version upgrades have known issues that may require a larger maintenance window or additional steps. For example, upgrades from v2.26.0 to v2.27.8 may encounter issues with the api_keys table—upgrading to v2.26.6 first can help mitigate this. Contact Coder support for guidance on your specific upgrade path.

Pre-upgrade strategy for Kubernetes HA deployments

Standard Kubernetes rolling updates may fail when exclusive database locks are required because old replicas keep connections open. For production deployments running multiple replicas (HA), active connections from existing pods can prevent the new pod from acquiring necessary locks.

  1. Scale down before upgrading: Before running helm upgrade, scale your Coder deployment down to eliminate database connection contention from existing pods.

    • Scale to zero for a clean cutover with no active database connections when the upgrade starts. This momentarily ensures no application access to the database, allowing migrations to acquire locks immediately:

      kubectl scale deployment coder --replicas=0
      
    • Scale to one if you prefer to minimize downtime. This keeps one pod running but eliminates contention from multiple replicas:

      kubectl scale deployment coder --replicas=1
      
  2. Perform upgrade: Run your standard Helm upgrade command. When scaling to zero, this will bring up a fresh pod that can run migrations without competing for database locks.

  3. Scale back: Once the upgrade is healthy, scale back to your desired replica count.

Kubernetes liveness probes and long-running migrations

Liveness probes can cause pods to be killed during long-running database migrations. Starting with Coder v2.30.0, liveness probes are disabled by default in the Helm chart.

This change was made because:

  • Liveness probes can kill pods during legitimate long-running migrations
  • If a Coder pod becomes unresponsive (due to a deadlock, etc.), it's better to investigate the issue rather than have Kubernetes silently restart the pod

If you have enabled liveness probes in your deployment and observe pods restarting with CrashLoopBackOff during an upgrade, the liveness probe may be killing the pod prematurely.

Diagnosing liveness probe issues

To confirm whether Kubernetes is killing pods due to liveness probe failures, check the Kubernetes events and pod logs:

# Check events for the Coder deployment
kubectl get events --field-selector involvedObject.name=coder -n <namespace>

# Check pod logs for migration progress
kubectl logs -l app.kubernetes.io/name=coder -n <namespace> --previous

Look for events indicating Liveness probe failed or Container coder failed liveness probe, will be restarted.

If you have liveness probes enabled and experience issues during upgrades, disable them before upgrading:

kubectl edit deployment coder

Remove the livenessProbe section entirely, then proceed with the upgrade.

Note

For versions prior to v2.30.0, liveness probes were enabled by default. You can disable them by editing the Deployment directly with kubectl edit deployment coder or by using a ConfigMap override. See the Helm chart values for configuration options available in v2.30.0+.

Workaround steps

  1. Remove or adjust liveness probes: Temporarily remove the livenessProbe from your Deployment configuration to prevent Kubernetes from restarting the pod during migrations.

  2. Isolate the migration: Ensure all extra replica sets are shut down. If you have clear evidence of database locks from old pods, scale the deployment to 1 replica to prevent old pods from holding locks on the tables being upgraded.

  3. Clear database locks: Monitor database activity. If the migration remains blocked by locks despite scaling down, you may need to manually terminate existing connections. See Recovering from failed database migrations below for instructions.

Recovering from failed database migrations

If an upgrade gets stuck in a restart loop due to database locks:

  1. Scale to zero: Scale the Coder deployment to 0 to stop all application activity.

    kubectl scale deployment coder --replicas=0
    
  2. Clear connections: Terminate existing connections to the Coder database to release any lingering locks. This PostgreSQL command drops all active connections to the database:

    Caution

    This command is intrusive and should be used as a last resort. Contact Coder support before running destructive database commands in production. SQL commands may vary depending on your PostgreSQL version and configuration.

    SELECT pg_terminate_backend(pid)
    FROM pg_stat_activity
    WHERE datname = 'coder'
    AND pid <> pg_backend_pid();
    
  3. Check schema migrations: Verify the level of upgrade and check if dirty is true. If this has progressed, this now indicates your current Coder installation state.

    Note

    The SQL commands below are for informational purposes. If you are unsure about querying your database directly, contact Coder support for assistance.

    SELECT * FROM schema_migrations;
    
  4. Ensure image version: Confirm the Deployment image is set to the appropriate version (old or new, depending on the database migration state found in step 3). Match your tag in the migrations directory to the value in the schema_migrations output.

  5. Resume the upgrade: Follow the pre-upgrade strategy to scale back up and continue the upgrade process.

When to contact support

If you encounter any of the following issues, contact Coder support:

  • Locking issues that cannot be mitigated by the steps in this guide
  • Migrations taking significantly longer than expected (more than 15 minutes) without evidence of lock contention—this may indicate database resource constraints requiring investigation
  • Resource consumption issues (excessive memory, CPU, or OOM kills) during upgrades
  • Any other upgrade problems not covered by this documentation

When contacting support, please collect and provide:

  • coderd logs with details on the stages where the upgrade stalled
  • PostgreSQL logs if available
  • The Coder versions involved (source and target)
  • Your deployment configuration (number of replicas, resource limits)