diff --git a/docs/admin/infrastructure/scale-testing.md b/docs/admin/infrastructure/scale-testing.md index de36131531..bbfbb35309 100644 --- a/docs/admin/infrastructure/scale-testing.md +++ b/docs/admin/infrastructure/scale-testing.md @@ -100,6 +100,8 @@ Database: - [Up to 2,000 users](./validated-architectures/2k-users.md) - [Up to 3,000 users](./validated-architectures/3k-users.md) +- +- [Up to 10,000 users](./validated-architectures/10k-users.md) ## Hardware recommendation diff --git a/docs/admin/infrastructure/validated-architectures/10k-users.md b/docs/admin/infrastructure/validated-architectures/10k-users.md new file mode 100644 index 0000000000..0def894d72 --- /dev/null +++ b/docs/admin/infrastructure/validated-architectures/10k-users.md @@ -0,0 +1,97 @@ +# Reference Architecture: up to 10,000 users + +The 10,000 users architecture targets enterprises with an extremely large global workforce of technical professionals or +applications requiring lots of simultaneous workspaces (for example, Agentic AI). + +The recommendations on this page apply to deployments with up to the following limits. If your needs +exceed any of these limits, consider increasing deployment resources. + +| Users | Concurrent Running Workspaces | Concurrent Builds | +|-------|-------------------------------|-------------------| +| 10000 | 6000 | 600 | + +**Observability**: Deploy monitoring solutions to gather Prometheus metrics and +visualize them with Grafana to gain detailed insights into infrastructure and +application behavior. This allows operators to respond quickly to incidents and +continuously improve the reliability and performance of the platform. + +## Hardware recommendations + +### Coderd + +| vCPU | Memory | Replicas | +|------|--------|----------| +| 4 | 12 GB | 10 | + +**Notes**: + +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `4000m` + - Set Memory request and limit to `12Gi` +- Coderd does not typically benefit from high performance disks like SSDs (unless you are co-locating provisioners). +- Coderd instances should be deployed in the same region as the database. + +### Workspace Proxies + +If you choose to deploy workspaces in multiple geographic regions, provision +[Workspace Proxies](../../networking/workspace-proxies.md) in each region. + +| vCPU | Memory | Replicas | +|------|--------|----------| +| 4 | 12 GB | 10 | + +**Notes**: + +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `4000m` + - Set Memory request and limit to `12Gi` +- Workspace Proxies do not typically benefit from high performance disks like SSDs. + +### Provisioners + +| vCPU | Memory | Replicas | +|------|--------|----------| +| 1 | 1 GB | 180 | + +**Notes**: + +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `1000m` + - Set Memory request and limit to `1Gi` +- If deploying on virtual machines, stack up to 30 provisioners per machine with a commensurate amount of memory and CPU. +- Provisioners benefit from high performance disks like SSDs. +- [Do not run provisioners on Coderd nodes](../../provisioners/index.md#disable-built-in-provisioners) at this scale. +- If deploying workspaces to multiple clouds or multiple Kubernetes clusters, divide the provisioner replicas among the + clouds or clusters according to expected usage. + +### Database + +| vCPU | Memory | Replicas | +|------|--------|----------| +| 64 | 240 GB | 1 | + +**Notes**: + +- "General purpose" virtual machines, such as the M8-series in AWS work well. +- Deploy in the same region as `coderd` + +### Workspaces + +The following resource requirements are for the Coder Workspace Agent, which runs alongside your end users work, and as +such should be interpreted as the _bare minimum_ requirements for a Coder workspace. Size your workspaces to fit the use +case your users will be undertaking. If in doubt, chose sizes based on the development environments your users are +migrating from onto Coder. + +| vCPU | Memory | +|------|--------| +| 0.1 | 128 MB | + +## Footnotes for AWS instance types + +- For production deployments, we recommend using non-burstable instance types, + such as `m5` or `c5`, instead of burstable instances, such as `t3`. + Burstable instances can experience significant performance degradation once + CPU credits are exhausted, leading to poor user experience under sustained load. diff --git a/docs/admin/infrastructure/validated-architectures/1k-users.md b/docs/admin/infrastructure/validated-architectures/1k-users.md index 540f9bb24b..af848f329b 100644 --- a/docs/admin/infrastructure/validated-architectures/1k-users.md +++ b/docs/admin/infrastructure/validated-architectures/1k-users.md @@ -43,7 +43,7 @@ architectural tier](./2k-users.md). - If deploying on Kubernetes: - Set CPU request and limit to `1000m` - Set Memory request and limit to `1Gi` -- If deploying on virtual machines, stack up to 30 provisioners per machine with a cummensurate amount of memory and CPU. +- If deploying on virtual machines, stack up to 30 provisioners per machine with a commensurate amount of memory and CPU. - Provisioners benefit from high performance disks like SSDs. - For small deployments (ca. 100 users, 10 concurrent workspace builds), it is acceptable to deploy provisioners on `coderd` nodes. diff --git a/docs/admin/infrastructure/validated-architectures/2k-users.md b/docs/admin/infrastructure/validated-architectures/2k-users.md index b989effdba..fb0dd5a21e 100644 --- a/docs/admin/infrastructure/validated-architectures/2k-users.md +++ b/docs/admin/infrastructure/validated-architectures/2k-users.md @@ -5,55 +5,94 @@ suggesting a growing user base or expanding operations. This setup is well-suited for mid-sized companies experiencing growth or for universities seeking to accommodate their expanding user populations. -Users can be evenly distributed between 2 regions or be attached to different -clusters. +The recommendations on this page apply to deployments with up to the following limits. If your needs +exceed any of these limits, consider increasing deployment resources or moving to the [next-higher +architectural tier](./3k-users.md). -**Target load**: API: up to 300 RPS +| Users | Concurrent Running Workspaces | Concurrent Builds | +|-------|-------------------------------|-------------------| +| 2000 | 1200 | 120 | -**High Availability**: The mode is _enabled_; multiple replicas provide higher -deployment reliability under load. +**Observability**: Deploy monitoring solutions to gather Prometheus metrics and +visualize them with Grafana to gain detailed insights into infrastructure and +application behavior. This allows operators to respond quickly to incidents and +continuously improve the reliability and performance of the platform. ## Hardware recommendations -### Coderd nodes +### Coderd -| Users | Node capacity | Replicas | GCP | AWS | Azure | -|-------------|----------------------|------------------------|-----------------|-------------|-------------------| -| Up to 2,000 | 4 vCPU, 16 GB memory | 2 nodes, 1 coderd each | `n1-standard-4` | `m5.xlarge` | `Standard_D4s_v3` | +| vCPU | Memory | Replicas | +|------|--------|----------| +| 4 | 12 GB | 3 | -### Provisioner nodes +**Notes**: -| Users | Node capacity | Replicas | GCP | AWS | Azure | -|-------------|----------------------|-------------------------------|------------------|--------------|-------------------| -| Up to 2,000 | 8 vCPU, 32 GB memory | 4 nodes, 30 provisioners each | `t2d-standard-8` | `c5.2xlarge` | `Standard_D8s_v3` | +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `4000m` + - Set Memory request and limit to `12Gi` +- Coderd does not typically benefit from high performance disks like SSDs (unless you are co-locating provisioners). +- Coderd instances should be deployed in the same region as the database. -**Footnotes**: +### Workspace Proxies -- An external provisioner is deployed as Kubernetes pod. -- It is not recommended to run provisioner daemons on `coderd` nodes. -- Consider separating provisioners into different namespaces in favor of - zero-trust or multi-cloud deployments. +If you choose to deploy workspaces in multiple geographic regions, provision +[Workspace Proxies](../../networking/workspace-proxies.md) in each region. -### Workspace nodes +| vCPU | Memory | Replicas | +|------|--------|----------| +| 4 | 12 GB | 3 | -| Users | Node capacity | Replicas | GCP | AWS | Azure | -|-------------|----------------------|-------------------------------|------------------|--------------|-------------------| -| Up to 2,000 | 8 vCPU, 32 GB memory | 128 nodes, 16 workspaces each | `t2d-standard-8` | `m5.2xlarge` | `Standard_D8s_v3` | +**Notes**: -**Footnotes**: +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `4000m` + - Set Memory request and limit to `12Gi` +- Workspace Proxies do not typically benefit from high performance disks like SSDs. -- Assumed that a workspace user needs 2 GB memory to perform -- Maximum number of Kubernetes workspace pods per node: 256 -- Nodes can be distributed in 2 regions, not necessarily evenly split, depending - on developer team sizes +### Provisioners -### Database node +| vCPU | Memory | Replicas | +|------|--------|----------| +| 1 | 1 GB | 120 | -| Users | Node capacity | Storage | GCP | AWS | Azure | -|-------------|----------------------|---------|---------------------|----------------|-------------------| -| Up to 2,000 | 4 vCPU, 16 GB memory | 1 TB | `db-custom-4-15360` | `db.m5.xlarge` | `Standard_D4s_v3` | +**Notes**: -**Footnotes for AWS instance types**: +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `1000m` + - Set Memory request and limit to `1Gi` +- If deploying on virtual machines, stack up to 30 provisioners per machine with a commensurate amount of memory and CPU. +- Provisioners benefit from high performance disks like SSDs. +- [Do not run provisioners on Coderd nodes](../../provisioners/index.md#disable-built-in-provisioners) at this scale. +- If deploying workspaces to multiple clouds or multiple Kubernetes clusters, divide the provisioner replicas among the + clouds or clusters according to expected usage. + +### Database + +| vCPU | Memory | Replicas | +|------|--------|----------| +| 16 | 60 GB | 1 | + +**Notes**: + +- "General purpose" virtual machines, such as the M8-series in AWS work well. +- Deploy in the same region as `coderd` + +### Workspaces + +The following resource requirements are for the Coder Workspace Agent, which runs alongside your end users work, and as +such should be interpreted as the _bare minimum_ requirements for a Coder workspace. Size your workspaces to fit the use +case your users will be undertaking. If in doubt, chose sizes based on the development environments your users are +migrating from onto Coder. + +| vCPU | Memory | +|------|--------| +| 0.1 | 128 MB | + +## Footnotes for AWS instance types - For production deployments, we recommend using non-burstable instance types, such as `m5` or `c5`, instead of burstable instances, such as `t3`. diff --git a/docs/admin/infrastructure/validated-architectures/3k-users.md b/docs/admin/infrastructure/validated-architectures/3k-users.md index 12165496b2..21aa7916ec 100644 --- a/docs/admin/infrastructure/validated-architectures/3k-users.md +++ b/docs/admin/infrastructure/validated-architectures/3k-users.md @@ -3,11 +3,13 @@ The 3,000 users architecture targets large-scale enterprises, possibly with on-premises network and cloud deployments. -**Target load**: API: up to 550 RPS +The recommendations on this page apply to deployments with up to the following limits. If your needs +exceed any of these limits, consider increasing deployment resources or moving to the [next-higher +architectural tier](./10k-users.md). -**High Availability**: Typically, such scale requires a fully-managed HA -PostgreSQL service, and all Coder observability features enabled for operational -purposes. +| Users | Concurrent Running Workspaces | Concurrent Builds | +|-------|-------------------------------|-------------------| +| 3000 | 1800 | 180 | **Observability**: Deploy monitoring solutions to gather Prometheus metrics and visualize them with Grafana to gain detailed insights into infrastructure and @@ -16,47 +18,79 @@ continuously improve the reliability and performance of the platform. ## Hardware recommendations -### Coderd nodes +### Coderd -| Users | Node capacity | Replicas | GCP | AWS | Azure | -|-------------|----------------------|-----------------------|-----------------|-------------|-------------------| -| Up to 3,000 | 8 vCPU, 32 GB memory | 4 node, 1 coderd each | `n1-standard-4` | `m5.xlarge` | `Standard_D4s_v3` | +| vCPU | Memory | Replicas | +|------|--------|----------| +| 4 | 12 GB | 4 | -### Provisioner nodes +**Notes**: -| Users | Node capacity | Replicas | GCP | AWS | Azure | -|-------------|----------------------|-------------------------------|------------------|--------------|-------------------| -| Up to 3,000 | 8 vCPU, 32 GB memory | 8 nodes, 30 provisioners each | `t2d-standard-8` | `c5.2xlarge` | `Standard_D8s_v3` | +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `4000m` + - Set Memory request and limit to `12Gi` +- Coderd does not typically benefit from high performance disks like SSDs (unless you are co-locating provisioners). +- Coderd instances should be deployed in the same region as the database. -**Footnotes**: +### Workspace Proxies -- An external provisioner is deployed as Kubernetes pod. -- It is strongly discouraged to run provisioner daemons on `coderd` nodes at - this level of scale. -- Separate provisioners into different namespaces in favor of zero-trust or - multi-cloud deployments. +If you choose to deploy workspaces in multiple geographic regions, provision +[Workspace Proxies](../../networking/workspace-proxies.md) in each region. -### Workspace nodes +| vCPU | Memory | Replicas | +|------|--------|----------| +| 4 | 12 GB | 4 | -| Users | Node capacity | Replicas | GCP | AWS | Azure | -|-------------|----------------------|-------------------------------|------------------|--------------|-------------------| -| Up to 3,000 | 8 vCPU, 32 GB memory | 256 nodes, 12 workspaces each | `t2d-standard-8` | `m5.2xlarge` | `Standard_D8s_v3` | +**Notes**: -**Footnotes**: +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `4000m` + - Set Memory request and limit to `12Gi` +- Workspace Proxies do not typically benefit from high performance disks like SSDs. -- Assumed that a workspace user needs 2 GB memory to perform -- Maximum number of Kubernetes workspace pods per node: 256 -- As workspace nodes can be distributed between regions, on-premises networks - and cloud areas, consider different namespaces in favor of zero-trust or - multi-cloud deployments. +### Provisioners -### Database node +| vCPU | Memory | Replicas | +|------|--------|----------| +| 1 | 1 GB | 180 | -| Users | Node capacity | Storage | GCP | AWS | Azure | -|-------------|----------------------|---------|---------------------|-----------------|-------------------| -| Up to 3,000 | 8 vCPU, 32 GB memory | 1.5 TB | `db-custom-8-30720` | `db.m5.2xlarge` | `Standard_D8s_v3` | +**Notes**: -**Footnotes for AWS instance types**: +- "General purpose" virtual machines, such as N4-series in GCP or M8-series in AWS work well. +- If deploying on Kubernetes: + - Set CPU request and limit to `1000m` + - Set Memory request and limit to `1Gi` +- If deploying on virtual machines, stack up to 30 provisioners per machine with a commensurate amount of memory and CPU. +- Provisioners benefit from high performance disks like SSDs. +- [Do not run provisioners on Coderd nodes](../../provisioners/index.md#disable-built-in-provisioners) at this scale. +- If deploying workspaces to multiple clouds or multiple Kubernetes clusters, divide the provisioner replicas among the + clouds or clusters according to expected usage. + +### Database + +| vCPU | Memory | Replicas | +|------|--------|----------| +| 32 | 120 GB | 1 | + +**Notes**: + +- "General purpose" virtual machines, such as the M8-series in AWS work well. +- Deploy in the same region as `coderd` + +### Workspaces + +The following resource requirements are for the Coder Workspace Agent, which runs alongside your end users work, and as +such should be interpreted as the _bare minimum_ requirements for a Coder workspace. Size your workspaces to fit the use +case your users will be undertaking. If in doubt, chose sizes based on the development environments your users are +migrating from onto Coder. + +| vCPU | Memory | +|------|--------| +| 0.1 | 128 MB | + +## Footnotes for AWS instance types - For production deployments, we recommend using non-burstable instance types, such as `m5` or `c5`, instead of burstable instances, such as `t3`. diff --git a/docs/admin/infrastructure/validated-architectures/index.md b/docs/admin/infrastructure/validated-architectures/index.md index 6bd18f7f3c..b2426e2d3f 100644 --- a/docs/admin/infrastructure/validated-architectures/index.md +++ b/docs/admin/infrastructure/validated-architectures/index.md @@ -220,6 +220,8 @@ For sizing recommendations, see the below reference architectures: - [Up to 3,000 users](3k-users.md) +- [Up to 10,000 users](10k-users.md) + ### AWS Instance Types For production AWS deployments, we recommend using non-burstable instance types, diff --git a/docs/manifest.json b/docs/manifest.json index 87931ab6a4..8f6531c81e 100644 --- a/docs/manifest.json +++ b/docs/manifest.json @@ -411,6 +411,11 @@ "title": "Up to 3,000 Users", "description": "Enterprise-scale architecture recommendations for Coder deployments that support up to 3,000 users", "path": "./admin/infrastructure/validated-architectures/3k-users.md" + }, + { + "title": "Up to 10,000 Users", + "description": "Enterprise-scale architecture recommendations for Coder deployments that support up to 10,000 users", + "path": "./admin/infrastructure/validated-architectures/10k-users.md" } ] },