docs: add new scaling doc to best practices section (#15904)

[preview](https://coder.com/docs/@bp-scaling-coder/tutorials/best-practices/scale-coder)

---------

Co-authored-by: Spike Curtis <spike@coder.com>
This commit is contained in:
Edward Angert
2025-01-21 15:02:02 -05:00
committed by GitHub
co-authored by Spike Curtis
parent 0fa6b3df13
commit 02d0650ae8
4 changed files with 393 additions and 39 deletions
+21 -16
View File
@@ -5,7 +5,7 @@ without compromising service. This process encompasses infrastructure setup,
traffic projections, and aggressive testing to identify and mitigate potential
bottlenecks.
A dedicated Kubernetes cluster for Coder is recommended to configure, host and
A dedicated Kubernetes cluster for Coder is recommended to configure, host, and
manage Coder workloads. Kubernetes provides container orchestration
capabilities, allowing Coder to efficiently deploy, scale, and manage workspaces
across a distributed infrastructure. This ensures high availability, fault
@@ -13,27 +13,29 @@ tolerance, and scalability for Coder deployments. Coder is deployed on this
cluster using the
[Helm chart](../../install/kubernetes.md#4-install-coder-with-helm).
For more information about scaling, see our [Coder scaling best practices](../../tutorials/best-practices/scale-coder.md).
## Methodology
Our scale tests include the following stages:
1. Prepare environment: create expected users and provision workspaces.
2. SSH connections: establish user connections with agents, verifying their
1. SSH connections: establish user connections with agents, verifying their
ability to echo back received content.
3. Web Terminal: verify the PTY connection used for communication with Web
1. Web Terminal: verify the PTY connection used for communication with Web
Terminal.
4. Workspace application traffic: assess the handling of user connections with
1. Workspace application traffic: assess the handling of user connections with
specific workspace apps, confirming their capability to echo back received
content effectively.
5. Dashboard evaluation: verify the responsiveness and stability of Coder
1. Dashboard evaluation: verify the responsiveness and stability of Coder
dashboards under varying load conditions. This is achieved by simulating user
interactions using instances of headless Chromium browsers.
6. Cleanup: delete workspaces and users created in step 1.
1. Cleanup: delete workspaces and users created in step 1.
## Infrastructure and setup requirements
@@ -54,13 +56,16 @@ channel for IDEs with VS Code and JetBrains plugins.
The basic setup of scale tests environment involves:
1. Scale tests runner (32 vCPU, 128 GB RAM)
2. Coder: 2 replicas (4 vCPU, 16 GB RAM)
3. Database: 1 instance (2 vCPU, 32 GB RAM)
4. Provisioner: 50 instances (0.5 vCPU, 512 MB RAM)
1. Coder: 2 replicas (4 vCPU, 16 GB RAM)
1. Database: 1 instance (2 vCPU, 32 GB RAM)
1. Provisioner: 50 instances (0.5 vCPU, 512 MB RAM)
The test is deemed successful if users did not experience interruptions in their
workflows, `coderd` did not crash or require restarts, and no other internal
errors were observed.
The test is deemed successful if:
- Users did not experience interruptions in their
workflows,
- `coderd` did not crash or require restarts, and
- No other internal errors were observed.
## Traffic Projections
@@ -90,11 +95,11 @@ Database:
## Available reference architectures
[Up to 1,000 users](./validated-architectures/1k-users.md)
- [Up to 1,000 users](./validated-architectures/1k-users.md)
[Up to 2,000 users](./validated-architectures/2k-users.md)
- [Up to 2,000 users](./validated-architectures/2k-users.md)
[Up to 3,000 users](./validated-architectures/3k-users.md)
- [Up to 3,000 users](./validated-architectures/3k-users.md)
## Hardware recommendation
@@ -107,7 +112,7 @@ guidance on optimal configurations. A reasonable approach involves using scaling
formulas based on factors like CPU, memory, and the number of users.
While the minimum requirements specify 1 CPU core and 2 GB of memory per
`coderd` replica, it is recommended to allocate additional resources depending
`coderd` replica, we recommend that you allocate additional resources depending
on the workload size to ensure deployment stability.
#### CPU and memory usage