From ccf0b348726370bcf42adab28957d153b1c9eb4a Mon Sep 17 00:00:00 2001 From: Spike Curtis Date: Tue, 21 Oct 2025 08:48:21 +0400 Subject: [PATCH] docs: create WIP 10k scale doc (#20213) Adds a new document for our ongoing efforts achieving 10k user scale. The content is caveated as work in progress, but represents what we have tested so far. closes: https://github.com/coder/internal/issues/1025 --- .../validated-architectures/10k-users.md | 108 ++++++++++++++++++ .../validated-architectures/index.md | 2 + docs/manifest.json | 5 + 3 files changed, 115 insertions(+) create mode 100644 docs/admin/infrastructure/validated-architectures/10k-users.md diff --git a/docs/admin/infrastructure/validated-architectures/10k-users.md b/docs/admin/infrastructure/validated-architectures/10k-users.md new file mode 100644 index 0000000000..e641371118 --- /dev/null +++ b/docs/admin/infrastructure/validated-architectures/10k-users.md @@ -0,0 +1,108 @@ +# Reference Architecture: up to 10,000 users + +> [!CAUTION] +> This page is a work in progress. +> +> We are actively testing different load profiles for this user target and will be updating +> recommendations. Use these recommendations as a starting point, but monitor your cluster resource +> utilization and adjust. + +The 10,000 users architecture targets large-scale enterprises with development +teams in multiple geographic regions. + +**Geographic Distribution**: For these tests we deploy on 3 cloud-managed Kubernetes clusters in +the following regions: + +1. USA - Primary - Coderd collocated with the PostgreSQL database deployment. +2. Europe - Workspace Proxies +3. Asia - Workspace Proxies + +**High Availability**: Typically, such scale requires a fully-managed HA +PostgreSQL service, and all Coder observability features enabled for operational +purposes. + +**Observability**: Deploy monitoring solutions to gather Prometheus metrics and +visualize them with Grafana to gain detailed insights into infrastructure and +application behavior. This allows operators to respond quickly to incidents and +continuously improve the reliability and performance of the platform. + +## Testing Methodology + +### Workspace Network Traffic + +6000 concurrent workspaces (2000 per region), each sending 10 kB/s application traffic. + +Test procedure: + +1. Create workspaces. This happens simultaneously in each region with 200 provisioners (and thus 600 concurrent builds). +2. Wait 5 minutes to establish baselines for metrics. +3. Generate 10 kB/s traffic to each workspace (originating within the same region & cluster). + +After, we examine the Coderd, Workspace Proxy, and Database metrics to look for issues. + +### API Request Traffic + +To be determined. + +## Hardware recommendations + +### Coderd + +These are deployed in the Primary region only. + +| vCPU Limit | Memory Limit | Replicas | GCP Node Pool Machine Type | +|----------------|--------------|----------|----------------------------| +| 4 vCPU (4000m) | 12 GiB | 10 | `c2d-standard-16` | + +### Provisioners + +These are deployed in each of the 3 regions. + +| vCPU Limit | Memory Limit | Replicas | GCP Node Pool Machine Type | +|-----------------|--------------|----------|----------------------------| +| 0.1 vCPU (100m) | 1 GiB | 200 | `c2d-standard-16` | + +**Footnotes**: + +- Each provisioner handles a single concurrent build, so this configuration implies 200 concurrent + workspace builds per region. +- Provisioners are run as a separate Kubernetes Deployment from Coderd, although they may + share the same node pool. +- Separate provisioners into different namespaces in favor of zero-trust or + multi-cloud deployments. + +### Workspace Proxies + +These are deployed in the non-Primary regions only. + +| vCPU Limit | Memory Limit | Replicas | GCP Node Pool Machine Type | +|----------------|--------------|----------|----------------------------| +| 4 vCPU (4000m) | 12 GiB | 10 | `c2d-standard-16` | + +**Footnotes**: + +- Our testing implies this is somewhat overspecced for the loads we have tried. We are in process of revising these numbers. + +### Workspaces + +These numbers are for each of the 3 regions. We recommend that you use a separate node pool for user Workspaces. + +| Users | Node capacity | Replicas | GCP | AWS | Azure | +|-------------|----------------------|-------------------------------|------------------|--------------|-------------------| +| Up to 3,000 | 8 vCPU, 32 GB memory | 256 nodes, 12 workspaces each | `t2d-standard-8` | `m5.2xlarge` | `Standard_D8s_v3` | + +**Footnotes**: + +- Assumed that a workspace user needs 2 GB memory to perform +- Maximum number of Kubernetes workspace pods per node: 256 +- As workspace nodes can be distributed between regions, on-premises networks + and cloud areas, consider different namespaces in favor of zero-trust or + multi-cloud deployments. + +### Database nodes + +We conducted our test using the `db-custom-16-61440` tier on Google Cloud SQL. + +**Footnotes**: + +- This database tier was only just able to keep up with 600 concurrent builds in our tests. diff --git a/docs/admin/infrastructure/validated-architectures/index.md b/docs/admin/infrastructure/validated-architectures/index.md index 6bd18f7f3c..59602f22bc 100644 --- a/docs/admin/infrastructure/validated-architectures/index.md +++ b/docs/admin/infrastructure/validated-architectures/index.md @@ -220,6 +220,8 @@ For sizing recommendations, see the below reference architectures: - [Up to 3,000 users](3k-users.md) +- DRAFT: [Up to 10,000 users](10k-users.md) + ### AWS Instance Types For production AWS deployments, we recommend using non-burstable instance types, diff --git a/docs/manifest.json b/docs/manifest.json index 023479cad2..5b01b6e9da 100644 --- a/docs/manifest.json +++ b/docs/manifest.json @@ -391,6 +391,11 @@ "title": "Up to 3,000 Users", "description": "Enterprise-scale architecture recommendations for Coder deployments that support up to 3,000 users", "path": "./admin/infrastructure/validated-architectures/3k-users.md" + }, + { + "title": "Up to 10,000 Users", + "description": "Enterprise-scale architecture recommendations for Coder deployments that support up to 10,000 users", + "path": "./admin/infrastructure/validated-architectures/10k-users.md" } ] },