mirror of
https://github.com/coder/coder.git
synced 2026-09-24 15:04:27 +08:00
docs: move architecture to top level (#13722)
This commit is contained in:
@@ -1,51 +0,0 @@
|
||||
# Reference Architecture: up to 1,000 users
|
||||
|
||||
The 1,000 users architecture is designed to cover a wide range of workflows.
|
||||
Examples of subjects that might utilize this architecture include medium-sized
|
||||
tech startups, educational units, or small to mid-sized enterprises.
|
||||
|
||||
**Target load**: API: up to 180 RPS
|
||||
|
||||
**High Availability**: non-essential for small deployments
|
||||
|
||||
## Hardware recommendations
|
||||
|
||||
### Coderd nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | ------------------- | ------------------- | --------------- | ---------- | ----------------- |
|
||||
| Up to 1,000 | 2 vCPU, 8 GB memory | 1-2 / 1 coderd each | `n1-standard-2` | `t3.large` | `Standard_D2s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- For small deployments (ca. 100 users, 10 concurrent workspace builds), it is
|
||||
acceptable to deploy provisioners on `coderd` nodes.
|
||||
|
||||
### Provisioner nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ------------------------------ | ---------------- | ------------ | ----------------- |
|
||||
| Up to 1,000 | 8 vCPU, 32 GB memory | 2 nodes / 30 provisioners each | `t2d-standard-8` | `t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- An external provisioner is deployed as Kubernetes pod.
|
||||
|
||||
### Workspace nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ----------------------- | ---------------- | ------------ | ----------------- |
|
||||
| Up to 1,000 | 8 vCPU, 32 GB memory | 64 / 16 workspaces each | `t2d-standard-8` | `t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- Assumed that a workspace user needs at minimum 2 GB memory to perform. We
|
||||
recommend against over-provisioning memory for developer workloads, as this my
|
||||
lead to OOMKiller invocations.
|
||||
- Maximum number of Kubernetes workspace pods per node: 256
|
||||
|
||||
### Database nodes
|
||||
|
||||
| Users | Node capacity | Replicas | Storage | GCP | AWS | Azure |
|
||||
| ----------- | ------------------- | -------- | ------- | ------------------ | ------------- | ----------------- |
|
||||
| Up to 1,000 | 2 vCPU, 8 GB memory | 1 | 512 GB | `db-custom-2-7680` | `db.t3.large` | `Standard_D2s_v3` |
|
||||
@@ -1,59 +0,0 @@
|
||||
# Reference Architecture: up to 2,000 users
|
||||
|
||||
In the 2,000 users architecture, there is a moderate increase in traffic,
|
||||
suggesting a growing user base or expanding operations. This setup is
|
||||
well-suited for mid-sized companies experiencing growth or for universities
|
||||
seeking to accommodate their expanding user populations.
|
||||
|
||||
Users can be evenly distributed between 2 regions or be attached to different
|
||||
clusters.
|
||||
|
||||
**Target load**: API: up to 300 RPS
|
||||
|
||||
**High Availability**: The mode is _enabled_; multiple replicas provide higher
|
||||
deployment reliability under load.
|
||||
|
||||
## Hardware recommendations
|
||||
|
||||
### Coderd nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ----------------------- | --------------- | ----------- | ----------------- |
|
||||
| Up to 2,000 | 4 vCPU, 16 GB memory | 2 nodes / 1 coderd each | `n1-standard-4` | `t3.xlarge` | `Standard_D4s_v3` |
|
||||
|
||||
### Provisioner nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ------------------------------ | ---------------- | ------------ | ----------------- |
|
||||
| Up to 2,000 | 8 vCPU, 32 GB memory | 4 nodes / 30 provisioners each | `t2d-standard-8` | `t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- An external provisioner is deployed as Kubernetes pod.
|
||||
- It is not recommended to run provisioner daemons on `coderd` nodes.
|
||||
- Consider separating provisioners into different namespaces in favor of
|
||||
zero-trust or multi-cloud deployments.
|
||||
|
||||
### Workspace nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ------------------------ | ---------------- | ------------ | ----------------- |
|
||||
| Up to 2,000 | 8 vCPU, 32 GB memory | 128 / 16 workspaces each | `t2d-standard-8` | `t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- Assumed that a workspace user needs 2 GB memory to perform
|
||||
- Maximum number of Kubernetes workspace pods per node: 256
|
||||
- Nodes can be distributed in 2 regions, not necessarily evenly split, depending
|
||||
on developer team sizes
|
||||
|
||||
### Database nodes
|
||||
|
||||
| Users | Node capacity | Replicas | Storage | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | -------- | ------- | ------------------- | -------------- | ----------------- |
|
||||
| Up to 2,000 | 4 vCPU, 16 GB memory | 1 | 1 TB | `db-custom-4-15360` | `db.t3.xlarge` | `Standard_D4s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- Consider adding more replicas if the workspace activity is higher than 500
|
||||
workspace builds per day or to achieve higher RPS.
|
||||
@@ -1,62 +0,0 @@
|
||||
# Reference Architecture: up to 3,000 users
|
||||
|
||||
The 3,000 users architecture targets large-scale enterprises, possibly with
|
||||
on-premises network and cloud deployments.
|
||||
|
||||
**Target load**: API: up to 550 RPS
|
||||
|
||||
**High Availability**: Typically, such scale requires a fully-managed HA
|
||||
PostgreSQL service, and all Coder observability features enabled for operational
|
||||
purposes.
|
||||
|
||||
**Observability**: Deploy monitoring solutions to gather Prometheus metrics and
|
||||
visualize them with Grafana to gain detailed insights into infrastructure and
|
||||
application behavior. This allows operators to respond quickly to incidents and
|
||||
continuously improve the reliability and performance of the platform.
|
||||
|
||||
## Hardware recommendations
|
||||
|
||||
### Coderd nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ----------------- | --------------- | ----------- | ----------------- |
|
||||
| Up to 3,000 | 8 vCPU, 32 GB memory | 4 / 1 coderd each | `n1-standard-4` | `t3.xlarge` | `Standard_D4s_v3` |
|
||||
|
||||
### Provisioner nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ------------------------ | ---------------- | ------------ | ----------------- |
|
||||
| Up to 3,000 | 8 vCPU, 32 GB memory | 8 / 30 provisioners each | `t2d-standard-8` | `t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- An external provisioner is deployed as Kubernetes pod.
|
||||
- It is strongly discouraged to run provisioner daemons on `coderd` nodes at
|
||||
this level of scale.
|
||||
- Separate provisioners into different namespaces in favor of zero-trust or
|
||||
multi-cloud deployments.
|
||||
|
||||
### Workspace nodes
|
||||
|
||||
| Users | Node capacity | Replicas | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | ------------------------------ | ---------------- | ------------ | ----------------- |
|
||||
| Up to 3,000 | 8 vCPU, 32 GB memory | 256 nodes / 12 workspaces each | `t2d-standard-8` | `t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- Assumed that a workspace user needs 2 GB memory to perform
|
||||
- Maximum number of Kubernetes workspace pods per node: 256
|
||||
- As workspace nodes can be distributed between regions, on-premises networks
|
||||
and cloud areas, consider different namespaces in favor of zero-trust or
|
||||
multi-cloud deployments.
|
||||
|
||||
### Database nodes
|
||||
|
||||
| Users | Node capacity | Replicas | Storage | GCP | AWS | Azure |
|
||||
| ----------- | -------------------- | -------- | ------- | ------------------- | --------------- | ----------------- |
|
||||
| Up to 3,000 | 8 vCPU, 32 GB memory | 2 | 1.5 TB | `db-custom-8-30720` | `db.t3.2xlarge` | `Standard_D8s_v3` |
|
||||
|
||||
**Footnotes**:
|
||||
|
||||
- Consider adding more replicas if the workspace activity is higher than 1500
|
||||
workspace builds per day or to achieve higher RPS.
|
||||
@@ -1,396 +0,0 @@
|
||||
# Architecture
|
||||
|
||||
The Coder deployment model is flexible and offers various components that
|
||||
platform administrators can deploy and scale depending on their use case. This
|
||||
page describes possible deployments, challenges, and risks associated with them.
|
||||
|
||||
## Primary components
|
||||
|
||||
### coderd
|
||||
|
||||
_coderd_ is the service created by running `coder server`. It is a thin API that
|
||||
connects workspaces, provisioners and users. _coderd_ stores its state in
|
||||
Postgres and is the only service that communicates with Postgres.
|
||||
|
||||
It offers:
|
||||
|
||||
- Dashboard (UI)
|
||||
- HTTP API
|
||||
- Dev URLs (HTTP reverse proxy to workspaces)
|
||||
- Workspace Web Applications (e.g for easy access to `code-server`)
|
||||
- Agent registration
|
||||
|
||||
### provisionerd
|
||||
|
||||
_provisionerd_ is the execution context for infrastructure modifying providers.
|
||||
At the moment, the only provider is Terraform (running `terraform`).
|
||||
|
||||
By default, the Coder server runs multiple provisioner daemons.
|
||||
[External provisioners](../provisioners.md) can be added for security or
|
||||
scalability purposes.
|
||||
|
||||
### Agents
|
||||
|
||||
An agent is the Coder service that runs within a user's remote workspace. It
|
||||
provides a consistent interface for coderd and clients to communicate with
|
||||
workspaces regardless of operating system, architecture, or cloud.
|
||||
|
||||
It offers the following services along with much more:
|
||||
|
||||
- SSH
|
||||
- Port forwarding
|
||||
- Liveness checks
|
||||
- `startup_script` automation
|
||||
|
||||
Templates are responsible for
|
||||
[creating and running agents](../../templates/index.md#coder-agent) within
|
||||
workspaces.
|
||||
|
||||
### Service Bundling
|
||||
|
||||
While _coderd_ and Postgres can be orchestrated independently, our default
|
||||
installation paths bundle them all together into one system service. It's
|
||||
perfectly fine to run a production deployment this way, but there are certain
|
||||
situations that necessitate decomposition:
|
||||
|
||||
- Reducing global client latency (distribute coderd and centralize database)
|
||||
- Achieving greater availability and efficiency (horizontally scale individual
|
||||
services)
|
||||
|
||||
### Workspaces
|
||||
|
||||
At the highest level, a workspace is a set of cloud resources. These resources
|
||||
can be VMs, Kubernetes clusters, storage buckets, or whatever else Terraform
|
||||
lets you dream up.
|
||||
|
||||
The resources that run the agent are described as _computational resources_,
|
||||
while those that don't are called _peripheral resources_.
|
||||
|
||||
Each resource may also be _persistent_ or _ephemeral_ depending on whether
|
||||
they're destroyed on workspace stop.
|
||||
|
||||
## Deployment models
|
||||
|
||||
### Single region architecture
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
This architecture consists of a single load balancer, several _coderd_ replicas,
|
||||
and _Coder workspaces_ deployed in the same region.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
- Deploy at least one _coderd_ replica per availability zone with _coderd_
|
||||
instances and provisioners. High availability is recommended but not essential
|
||||
for small deployments.
|
||||
- Single replica deployment is a special case that can address a
|
||||
tiny/small/proof-of-concept installation on a single virtual machine. If you
|
||||
are serving more than 100 users/workspaces, you should add more replicas.
|
||||
|
||||
**Coder workspace**
|
||||
|
||||
- For small deployments consider a lightweight workspace runtime like the
|
||||
[Sysbox](https://github.com/nestybox/sysbox) container runtime. Learn more how
|
||||
to enable
|
||||
[docker-in-docker using Sysbox](https://asciinema.org/a/kkTmOxl8DhEZiM2fLZNFlYzbo?speed=2).
|
||||
|
||||
**HA Database**
|
||||
|
||||
- Monitor node status and resource utilization metrics.
|
||||
- Implement robust backup and disaster recovery strategies to protect against
|
||||
data loss.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Load balancer**
|
||||
|
||||
- Distributes and load balances traffic from agents and clients to _Coder
|
||||
Server_ replicas across availability zones.
|
||||
- Layer 7 load balancing. The load balancer can decrypt SSL traffic, and
|
||||
re-encrypt using an internal certificate.
|
||||
- Session persistence (sticky sessions) can be disabled as _coderd_ instances
|
||||
are stateless.
|
||||
- WebSocket and long-lived connections must be supported.
|
||||
|
||||
**Single sign-on**
|
||||
|
||||
- Integrate with existing Single Sign-On (SSO) solutions used within the
|
||||
organization via the supported OAuth 2.0 or OpenID Connect standards.
|
||||
- Learn more about [Authentication in Coder](../auth.md).
|
||||
|
||||
### Multi-region architecture
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
This architecture is for globally distributed developer teams using Coder
|
||||
workspaces on daily basis. It features a single load balancer with regionally
|
||||
deployed _Workspace Proxies_, several _coderd_ replicas, and _Coder workspaces_
|
||||
provisioned in different regions.
|
||||
|
||||
Note: The _multi-region architecture_ assumes the same deployment principles as
|
||||
the _single region architecture_, but it extends them to multi region deployment
|
||||
with workspace proxies. Proxies are deployed in regions closest to developers to
|
||||
offer the fastest developer experience.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Workspace proxy**
|
||||
|
||||
- Workspace proxy offers developers the option to establish a fast relay
|
||||
connection when accessing their workspace via SSH, a workspace application, or
|
||||
port forwarding.
|
||||
- Dashboard connections, API calls (e.g. _list workspaces_) are not served over
|
||||
proxies.
|
||||
- Proxies do not establish connections to the database.
|
||||
- Proxy instances do not share authentication tokens between one another.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Proxy load balancer**
|
||||
|
||||
- Distributes and load balances workspace relay traffic in a single region
|
||||
across availability zones.
|
||||
- Layer 7 load balancing. The load balancer can decrypt SSL traffic, and
|
||||
re-encrypt using internal certificate.
|
||||
- Session persistence (sticky sessions) can be disabled as _coderd_ instances
|
||||
are stateless.
|
||||
- WebSocket and long-lived connections must be supported.
|
||||
|
||||
### Multi-cloud architecture
|
||||
|
||||
By distributing Coder workspaces across different cloud providers, organizations
|
||||
can mitigate the risk of downtime caused by provider-specific outages or
|
||||
disruptions. Additionally, multi-cloud deployment enables organizations to
|
||||
leverage the unique features and capabilities offered by each cloud provider,
|
||||
such as region availability and pricing models.
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
The deployment model comprises:
|
||||
|
||||
- `coderd` instances deployed within a single region of the same cloud provider,
|
||||
with replicas strategically distributed across availability zones.
|
||||
- Workspace provisioners deployed in each cloud, communicating with `coderd`
|
||||
instances.
|
||||
- Workspace proxies running in the same locations as provisioners to optimize
|
||||
user connections to workspaces for maximum speed.
|
||||
|
||||
Due to the relatively large overhead of cross-regional communication, it is not
|
||||
advised to set up multi-cloud control planes. It is recommended to keep coderd
|
||||
replicas and the database within the same cloud-provider and region.
|
||||
|
||||
Note: The _multi-cloud architecture_ follows the deployment principles outlined
|
||||
in the _multi-region architecture_. However, it adapts component selection based
|
||||
on the specific cloud provider. Developers can initiate workspaces based on the
|
||||
nearest region and technical specifications provided by the cloud providers.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Workspace provisioner**
|
||||
|
||||
- _Security recommendation_: Create a long, random pre-shared key (PSK) and add
|
||||
it to the regional secret store, so that local _provisionerd_ can access it.
|
||||
Remember to distribute it using safe, encrypted communication channel. The PSK
|
||||
must also be added to the _coderd_ configuration.
|
||||
|
||||
**Workspace proxy**
|
||||
|
||||
- _Security recommendation_: Use `coder` CLI to create
|
||||
[authentication tokens for every workspace proxy](../workspace-proxies.md#requirements),
|
||||
and keep them in regional secret stores. Remember to distribute them using
|
||||
safe, encrypted communication channel.
|
||||
|
||||
**Managed database**
|
||||
|
||||
- For AWS: _Amazon RDS for PostgreSQL_
|
||||
- For Azure: _Azure Database for PostgreSQL - Flexible Server_
|
||||
- For GCP: _Cloud SQL for PostgreSQL_
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Kubernetes platform (optional)**
|
||||
|
||||
- For AWS: _Amazon Elastic Kubernetes Service_
|
||||
- For Azure: _Azure Kubernetes Service_
|
||||
- For GCP: _Google Kubernetes Engine_
|
||||
|
||||
See how to deploy
|
||||
[Coder on Azure Kubernetes Service](https://github.com/ericpaulsen/coder-aks).
|
||||
|
||||
Learn more about [security requirements](../../install/kubernetes.md) for
|
||||
deploying Coder on Kubernetes.
|
||||
|
||||
**Load balancer**
|
||||
|
||||
- For AWS:
|
||||
- _AWS Network Load Balancer_
|
||||
- Level 4 load balancing
|
||||
- For Kubernetes deployment: annotate service with
|
||||
`service.beta.kubernetes.io/aws-load-balancer-type: "nlb"`, preserve the
|
||||
client source IP with `externalTrafficPolicy: Local`
|
||||
- _AWS Classic Load Balancer_
|
||||
- Level 7 load balancing
|
||||
- For Kubernetes deployment: set `sessionAffinity` to `None`
|
||||
- For Azure:
|
||||
- _Azure Load Balancer_
|
||||
- Level 7 load balancing
|
||||
- Azure Application Gateway
|
||||
- Deploy Azure Application Gateway when more advanced traffic routing
|
||||
policies are needed for Kubernetes applications.
|
||||
- Take advantage of features such as WebSocket support and TLS termination
|
||||
provided by Azure Application Gateway, enhancing the capabilities of
|
||||
Kubernetes deployments on Azure.
|
||||
- For GCP:
|
||||
- _Cloud Load Balancing_ with SSL load balancer:
|
||||
- Layer 4 load balancing, SSL enabled
|
||||
- _Cloud Load Balancing_ with HTTPS load balancer:
|
||||
- Layer 7 load balancing
|
||||
- For Kubernetes deployment: annotate service (with ingress enabled) with
|
||||
`kubernetes.io/ingress.class: "gce"`, leverage the `NodePort` service
|
||||
type.
|
||||
- Note: HTTP load balancer rejects DERP upgrade, Coder will fallback to
|
||||
WebSockets
|
||||
|
||||
**Single sign-on**
|
||||
|
||||
- For AWS:
|
||||
[AWS IAM Identity Center](https://docs.aws.amazon.com/singlesignon/latest/userguide/what-is.html)
|
||||
- For Azure:
|
||||
[Microsoft Entra ID Sign-On](https://learn.microsoft.com/en-us/entra/identity/app-proxy/)
|
||||
- For GCP:
|
||||
[Google Cloud Identity Platform](https://cloud.google.com/architecture/identity/single-sign-on)
|
||||
|
||||
### Air-gapped architecture
|
||||
|
||||
The air-gapped deployment model refers to the setup of Coder's development
|
||||
environment within a restricted network environment that lacks internet
|
||||
connectivity. This deployment model is often required for organizations with
|
||||
strict security policies or those operating in isolated environments, such as
|
||||
government agencies or certain enterprise setups.
|
||||
|
||||
The key features of the air-gapped architecture include:
|
||||
|
||||
- _Offline installation_: Deploy workspaces without relying on an external
|
||||
internet connection.
|
||||
- _Isolated package/plugin repositories_: Depend on local repositories for
|
||||
software installation, updates, and security patches.
|
||||
- _Secure data transfer_: Enable encrypted communication channels and robust
|
||||
access controls to safeguard sensitive information.
|
||||
|
||||
Learn more about [offline deployments](../../install/offline.md) of Coder.
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
The deployment model includes:
|
||||
|
||||
- _Workspace provisioners_ with direct access to self-hosted package and plugin
|
||||
repositories and restricted internet access.
|
||||
- _Mirror of Terraform Registry_ with multiple versions of Terraform plugins.
|
||||
- _Certificate Authority_ with all TLS certificates to build secure
|
||||
communication channels.
|
||||
|
||||
The model is compatible with various infrastructure models, enabling deployment
|
||||
across multiple regions and diverse cloud platforms.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Workspace provisioner**
|
||||
|
||||
- Includes Terraform binary in the container or system image.
|
||||
- Checks out Terraform plugins from self-hosted _Registry_ mirror.
|
||||
- Deploys workspace images stored in the self-hosted _Container Registry_.
|
||||
|
||||
**Coder server**
|
||||
|
||||
- Update checks are disabled (`CODER_UPDATE_CHECK=false`).
|
||||
- Telemetry data is not collected (`CODER_TELEMETRY_ENABLE=false`).
|
||||
- Direct connections are not possible, workspace traffic is relayed through
|
||||
control plane's DERP proxy.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Self-hosted Database**
|
||||
|
||||
- In the air-gapped deployment model, _Coderd_ instance is unable to download
|
||||
Postgres binaries from the internet, so external database must be provided.
|
||||
|
||||
**Container Registry**
|
||||
|
||||
- Since the _Registry_ is isolated from the internet, platform engineers are
|
||||
responsible for maintaining Workspace container images and conducting periodic
|
||||
updates of base Docker images.
|
||||
- It is recommended to keep [Dev Containers](../../templates/dev-containers.md)
|
||||
up to date with the latest released
|
||||
[Envbuilder](https://github.com/coder/envbuilder) runtime.
|
||||
|
||||
**Mirror of Terraform Registry**
|
||||
|
||||
- Stores all necessary Terraform plugin dependencies, ensuring successful
|
||||
workspace provisioning and maintenance without internet access.
|
||||
- Platform engineers are responsible for periodically updating the mirrored
|
||||
Terraform plugins, including
|
||||
[terraform-provider-coder](https://github.com/coder/terraform-provider-coder).
|
||||
|
||||
**Certificate Authority**
|
||||
|
||||
- Manages and issues TLS certificates to facilitate secure communication
|
||||
channels within the infrastructure.
|
||||
|
||||
### Dev Containers
|
||||
|
||||
Note: _Dev containers_ are at early stage and considered experimental at the
|
||||
moment.
|
||||
|
||||
This architecture enhances a Coder workspace with a
|
||||
[development container](https://containers.dev/) setup built using the
|
||||
[envbuilder](https://github.com/coder/envbuilder) project. Workspace users have
|
||||
the flexibility to extend generic, base developer environments with custom,
|
||||
project-oriented [features](https://containers.dev/features) without requiring
|
||||
platform administrators to push altered Docker images.
|
||||
|
||||
Learn more about
|
||||
[Dev containers support](https://coder.com/docs/v2/latest/templates/dev-containers)
|
||||
in Coder.
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
The deployment model includes:
|
||||
|
||||
- _Workspace_ built using Coder template with _envbuilder_ enabled to set up the
|
||||
developer environment accordingly to the dev container spec.
|
||||
- _Container Registry_ for Docker images used by _envbuilder_, maintained by
|
||||
Coder platform engineers or developer productivity engineers.
|
||||
|
||||
Since this model is strictly focused on workspace nodes, it does not affect the
|
||||
setup of regional infrastructure. It can be deployed alongside other deployment
|
||||
models, in multiple regions, or across various cloud platforms.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Coder workspace**
|
||||
|
||||
- Docker and Kubernetes based templates are supported.
|
||||
- The `docker_container` resource uses `ghcr.io/coder/envbuilder` as the base
|
||||
image.
|
||||
|
||||
_Envbuilder_ checks out the base Docker image from the container registry and
|
||||
installs selected features as specified in the `devcontainer.json` on top.
|
||||
Eventually, it starts the container with the developer environment.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Container Registry (optional)**
|
||||
|
||||
- Workspace nodes need access to the Container Registry to check out images. To
|
||||
shorten the provisioning time, it is recommended to deploy registry mirrors in
|
||||
the same region as the workspace nodes.
|
||||
@@ -1,363 +0,0 @@
|
||||
# Coder Validated Architecture
|
||||
|
||||
Many customers operate Coder in complex organizational environments, consisting
|
||||
of multiple business units, agencies, and/or subsidiaries. This can lead to
|
||||
numerous Coder deployments, due to discrepancies in regulatory compliance, data
|
||||
sovereignty, and level of funding across groups. The Coder Validated
|
||||
Architecture (CVA) prescribes a Kubernetes-based deployment approach, enabling
|
||||
your organization to deploy a stable Coder instance that is easier to maintain
|
||||
and troubleshoot.
|
||||
|
||||
The following sections will detail the components of the Coder Validated
|
||||
Architecture, provide guidance on how to configure and deploy these components,
|
||||
and offer insights into how to maintain and troubleshoot your Coder environment.
|
||||
|
||||
- [General concepts](#general-concepts)
|
||||
- [Kubernetes Infrastructure](#kubernetes-infrastructure)
|
||||
- [PostgreSQL Database](#postgresql-database)
|
||||
- [Operational readiness](#operational-readiness)
|
||||
|
||||
## Who is this document for?
|
||||
|
||||
This guide targets the following personas. It assumes a basic understanding of
|
||||
cloud/on-premise computing, containerization, and the Coder platform.
|
||||
|
||||
| Role | Description |
|
||||
| ------------------------- | ------------------------------------------------------------------------------ |
|
||||
| Platform Engineers | Responsible for deploying, operating the Coder deployment and infrastructure |
|
||||
| Enterprise Architects | Responsible for architecting Coder deployments to meet enterprise requirements |
|
||||
| Managed Service Providers | Entities that deploy and run Coder software as a service for customers |
|
||||
|
||||
## CVA Guidance
|
||||
|
||||
| CVA provides: | CVA does not provide: |
|
||||
| ---------------------------------------------- | ---------------------------------------------------------------------------------------- |
|
||||
| Single and multi-region K8s deployment options | Prescribing OS, or cloud vs. on-premise |
|
||||
| Reference architectures for up to 3,000 users | An approval of your architecture; the CVA solely provides recommendations and guidelines |
|
||||
| Best practices for building a Coder deployment | Recommendations for every possible deployment scenario |
|
||||
|
||||
> For higher level design principles and architectural best practices, see
|
||||
> Coder's
|
||||
> [Well-Architected Framework](https://coder.com/blog/coder-well-architected-framework).
|
||||
|
||||
## General concepts
|
||||
|
||||
This section outlines core concepts and terminology essential for understanding
|
||||
Coder's architecture and deployment strategies.
|
||||
|
||||
### Administrator
|
||||
|
||||
An administrator is a user role within the Coder platform with elevated
|
||||
privileges. Admins have access to administrative functions such as user
|
||||
management, template definitions, insights, and deployment configuration.
|
||||
|
||||
### Coder control plane
|
||||
|
||||
Coder's control plane, also known as _coderd_, is the main service recommended
|
||||
for deployment with multiple replicas to ensure high availability. It provides
|
||||
an API for managing workspaces and templates, and serves the dashboard UI. In
|
||||
addition, each _coderd_ replica hosts 3 Terraform [provisioners](#provisioner)
|
||||
by default.
|
||||
|
||||
### User
|
||||
|
||||
A [user](../users.md) is an individual who utilizes the Coder platform to
|
||||
develop, test, and deploy applications using workspaces. Users can select
|
||||
available templates to provision workspaces. They interact with Coder using the
|
||||
web interface, the CLI tool, or directly calling API methods.
|
||||
|
||||
### Workspace
|
||||
|
||||
A [workspace](../../workspaces.md) refers to an isolated development environment
|
||||
where users can write, build, and run code. Workspaces are fully configurable
|
||||
and can be tailored to specific project requirements, providing developers with
|
||||
a consistent and efficient development environment. Workspaces can be
|
||||
autostarted and autostopped, enabling efficient resource management.
|
||||
|
||||
Users can connect to workspaces using SSH or via workspace applications like
|
||||
`code-server`, facilitating collaboration and remote access. Additionally,
|
||||
workspaces can be parameterized, allowing users to customize settings and
|
||||
configurations based on their unique needs. Workspaces are instantiated using
|
||||
Coder templates and deployed on resources created by provisioners.
|
||||
|
||||
### Template
|
||||
|
||||
A [template](../../templates/index.md) in Coder is a predefined configuration
|
||||
for creating workspaces. Templates streamline the process of workspace creation
|
||||
by providing pre-configured settings, tooling, and dependencies. They are built
|
||||
by template administrators on top of Terraform, allowing for efficient
|
||||
management of infrastructure resources. Additionally, templates can utilize
|
||||
Coder modules to leverage existing features shared with other templates,
|
||||
enhancing flexibility and consistency across deployments. Templates describe
|
||||
provisioning rules for infrastructure resources offered by Terraform providers.
|
||||
|
||||
### Workspace Proxy
|
||||
|
||||
A [workspace proxy](../workspace-proxies.md) serves as a relay connection option
|
||||
for developers connecting to their workspace over SSH, a workspace app, or
|
||||
through port forwarding. It helps reduce network latency for geo-distributed
|
||||
teams by minimizing the distance network traffic needs to travel. Notably,
|
||||
workspace proxies do not handle dashboard connections or API calls.
|
||||
|
||||
### Provisioner
|
||||
|
||||
Provisioners in Coder execute Terraform during workspace and template builds.
|
||||
While the platform includes built-in provisioner daemons by default, there are
|
||||
advantages to employing external provisioners. These external daemons provide
|
||||
secure build environments and reduce server load, improving performance and
|
||||
scalability. Each provisioner can handle a single concurrent workspace build,
|
||||
allowing for efficient resource allocation and workload management.
|
||||
|
||||
### Registry
|
||||
|
||||
The [Coder Registry](https://registry.coder.com) is a platform where you can
|
||||
find starter templates and _Modules_ for various cloud services and platforms.
|
||||
|
||||
Templates help create self-service development environments using
|
||||
Terraform-defined infrastructure, while _Modules_ simplify template creation by
|
||||
providing common features like workspace applications, third-party integrations,
|
||||
or helper scripts.
|
||||
|
||||
Please note that the Registry is a hosted service and isn't available for
|
||||
offline use.
|
||||
|
||||
## Kubernetes Infrastructure
|
||||
|
||||
Kubernetes is the recommended, and supported platform for deploying Coder in the
|
||||
enterprise. It is the hosting platform of choice for a large majority of Coder's
|
||||
Fortune 500 customers, and it is the platform in which we build and test against
|
||||
here at Coder.
|
||||
|
||||
### General recommendations
|
||||
|
||||
In general, it is recommended to deploy Coder into its own respective cluster,
|
||||
separate from production applications. Keep in mind that Coder runs development
|
||||
workloads, so the cluster should be deployed as such, without production-level
|
||||
configurations.
|
||||
|
||||
### Compute
|
||||
|
||||
Deploy your Kubernetes cluster with two node groups, one for Coder's control
|
||||
plane, and another for user workspaces (if you intend on leveraging K8s for
|
||||
end-user compute).
|
||||
|
||||
#### Control plane nodes
|
||||
|
||||
The Coder control plane node group must be static, to prevent scale down events
|
||||
from dropping pods, and thus dropping user connections to the dashboard UI and
|
||||
their workspaces.
|
||||
|
||||
Coder's Helm Chart supports
|
||||
[defining nodeSelectors, affinities, and tolerations](https://github.com/coder/coder/blob/e96652ebbcdd7554977594286b32015115c3f5b6/helm/coder/values.yaml#L221-L249)
|
||||
to schedule the control plane pods on the appropriate node group.
|
||||
|
||||
#### Workspace nodes
|
||||
|
||||
Coder workspaces can be deployed either as Pods or Deployments in Kubernetes.
|
||||
See our
|
||||
[example Kubernetes workspace template](https://github.com/coder/coder/tree/main/examples/templates/kubernetes).
|
||||
Configure the workspace node group to be auto-scaling, to dynamically allocate
|
||||
compute as users start/stop workspaces at the beginning and end of their day.
|
||||
Set nodeSelectors, affinities, and tolerations in Coder templates to assign
|
||||
workspaces to the given node group:
|
||||
|
||||
```hcl
|
||||
resource "kubernetes_deployment" "coder" {
|
||||
spec {
|
||||
template {
|
||||
metadata {
|
||||
labels = {
|
||||
app = "coder-workspace"
|
||||
}
|
||||
}
|
||||
|
||||
spec {
|
||||
affinity {
|
||||
pod_anti_affinity {
|
||||
preferred_during_scheduling_ignored_during_execution {
|
||||
weight = 1
|
||||
pod_affinity_term {
|
||||
label_selector {
|
||||
match_expressions {
|
||||
key = "app.kubernetes.io/instance"
|
||||
operator = "In"
|
||||
values = ["coder-workspace"]
|
||||
}
|
||||
}
|
||||
topology_key = # add your node group label here
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
tolerations {
|
||||
# Add your tolerations here
|
||||
}
|
||||
|
||||
node_selector {
|
||||
# Add your node selectors here
|
||||
}
|
||||
|
||||
container {
|
||||
image = "coder-workspace:latest"
|
||||
name = "dev"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Node sizing
|
||||
|
||||
For sizing recommendations, see the below reference architectures:
|
||||
|
||||
- [Up to 1,000 users](1k-users.md)
|
||||
|
||||
- [Up to 2,000 users](2k-users.md)
|
||||
|
||||
- [Up to 3,000 users](3k-users.md)
|
||||
|
||||
### Networking
|
||||
|
||||
It is likely your enterprise deploys Kubernetes clusters with various networking
|
||||
restrictions. With this in mind, Coder requires the following connectivity:
|
||||
|
||||
- Egress from workspace compute to the Coder control plane pods
|
||||
- Egress from control plane pods to Coder's PostgreSQL database
|
||||
- Egress from control plane pods to git and package repositories
|
||||
- Ingress from user devices to the control plane Load Balancer or Ingress
|
||||
controller
|
||||
|
||||
We recommend configuring your network policies in accordance with the above.
|
||||
Note that Coder workspaces do not require any ports to be open.
|
||||
|
||||
### Storage
|
||||
|
||||
If running Coder workspaces as Kubernetes Pods or Deployments, you will need to
|
||||
assign persistent storage. We recommend leveraging a
|
||||
[supported Container Storage Interface (CSI) driver](https://kubernetes-csi.github.io/docs/drivers.html)
|
||||
in your cluster, with Dynamic Provisioning and read/write, to provide on-demand
|
||||
storage to end-user workspaces.
|
||||
|
||||
The following Kubernetes volume types have been validated by Coder internally,
|
||||
and/or by our customers:
|
||||
|
||||
- [PersistentVolumeClaim](https://kubernetes.io/docs/concepts/storage/volumes/#persistentvolumeclaim)
|
||||
- [NFS](https://kubernetes.io/docs/concepts/storage/volumes/#nfs)
|
||||
- [subPath](https://kubernetes.io/docs/concepts/storage/volumes/#using-subpath)
|
||||
- [cephfs](https://kubernetes.io/docs/concepts/storage/volumes/#cephfs)
|
||||
|
||||
Our
|
||||
[example Kubernetes workspace template](https://github.com/coder/coder/blob/5b9a65e5c137232351381fc337d9784bc9aeecfc/examples/templates/kubernetes/main.tf#L191-L219)
|
||||
provisions a PersistentVolumeClaim block storage device, attached to the
|
||||
Deployment.
|
||||
|
||||
It is not recommended to mount volumes from the host node(s) into workspaces,
|
||||
for security and reliability purposes. The below volume types are _not_
|
||||
recommended for use with Coder:
|
||||
|
||||
- [Local](https://kubernetes.io/docs/concepts/storage/volumes/#local)
|
||||
- [hostPath](https://kubernetes.io/docs/concepts/storage/volumes/#hostpath)
|
||||
|
||||
Not that Coder's control plane filesystem is ephemeral, so no persistent storage
|
||||
is required.
|
||||
|
||||
## PostgreSQL database
|
||||
|
||||
Coder requires access to an external PostgreSQL database to store user data,
|
||||
workspace state, template files, and more. Depending on the scale of the
|
||||
user-base, workspace activity, and High Availability requirements, the amount of
|
||||
CPU and memory resources required by Coder's database may differ.
|
||||
|
||||
### Disaster recovery
|
||||
|
||||
Prepare internal scripts for dumping and restoring your database. We recommend
|
||||
scheduling regular database backups, especially before upgrading Coder to a new
|
||||
release. Coder does not support downgrades without initially restoring the
|
||||
database to the prior version.
|
||||
|
||||
### Performance efficiency
|
||||
|
||||
We highly recommend deploying the PostgreSQL instance in the same region (and if
|
||||
possible, same availability zone) as the Coder server to optimize for low
|
||||
latency connections. We recommend keeping latency under 10ms between the Coder
|
||||
server and database.
|
||||
|
||||
When determining scaling requirements, take into account the following
|
||||
considerations:
|
||||
|
||||
- `2 vCPU x 8 GB RAM x 512 GB storage`: A baseline for database requirements for
|
||||
Coder deployment with less than 1000 users, and low activity level (30% active
|
||||
users). This capacity should be sufficient to support 100 external
|
||||
provisioners.
|
||||
- Storage size depends on user activity, workspace builds, log verbosity,
|
||||
overhead on database encryption, etc.
|
||||
- Allocate two additional CPU core to the database instance for every 1000
|
||||
active users.
|
||||
- Enable High Availability mode for database engine for large scale deployments.
|
||||
|
||||
If you enable [database encryption](../encryption.md) in Coder, consider
|
||||
allocating an additional CPU core to every `coderd` replica.
|
||||
|
||||
#### Resource utilization guidelines
|
||||
|
||||
Below are general recommendations for sizing your PostgreSQL instance:
|
||||
|
||||
- Increase number of vCPU if CPU utilization or database latency is high.
|
||||
- Allocate extra memory if database performance is poor, CPU utilization is low,
|
||||
and memory utilization is high.
|
||||
- Utilize faster disk options (higher IOPS) such as SSDs or NVMe drives for
|
||||
optimal performance enhancement and possibly reduce database load.
|
||||
|
||||
## Operational readiness
|
||||
|
||||
Operational readiness in Coder is about ensuring that everything is set up
|
||||
correctly before launching a platform into production. It involves making sure
|
||||
that the service is reliable, secure, and easily scales accordingly to user-base
|
||||
needs. Operational readiness is crucial because it helps prevent issues that
|
||||
could affect workspace users experience once the platform is live.
|
||||
|
||||
### Helm Chart Configuration
|
||||
|
||||
1. Reference our [Helm chart values file](../../../helm/coder/values.yaml) and
|
||||
identify the required values for deployment.
|
||||
1. Create a `values.yaml` and add it to your version control system.
|
||||
1. Determine the necessary environment variables. Here is the
|
||||
[full list of supported server environment variables](../../cli/server.md).
|
||||
1. Follow our documented
|
||||
[steps for installing Coder via Helm](../../install/kubernetes.md).
|
||||
|
||||
### Template configuration
|
||||
|
||||
1. Establish dedicated accounts for users with the _Template Administrator_
|
||||
role.
|
||||
1. Maintain Coder templates using
|
||||
[version control](../../templates/change-management.md).
|
||||
1. Consider implementing a GitOps workflow to automatically push new template
|
||||
versions into Coder from git. For example, on Github, you can use the
|
||||
[Update Coder Template](https://github.com/marketplace/actions/update-coder-template)
|
||||
action.
|
||||
1. Evaluate enabling
|
||||
[automatic template updates](../../templates/general-settings.md#require-automatic-updates-enterprise)
|
||||
upon workspace startup.
|
||||
|
||||
### Observability
|
||||
|
||||
1. Enable the Prometheus endpoint (environment variable:
|
||||
`CODER_PROMETHEUS_ENABLE`).
|
||||
1. Deploy the
|
||||
[Coder Observability bundle](https://github.com/coder/observability) to
|
||||
leverage pre-configured dashboards, alerts, and runbooks for monitoring
|
||||
Coder. This includes integrations between Prometheus, Grafana, Loki, and
|
||||
Alertmanager.
|
||||
1. Review the [Prometheus response](../prometheus.md) and set up alarms on
|
||||
selected metrics.
|
||||
|
||||
### User support
|
||||
|
||||
1. Incorporate [support links](../appearance.md#support-links) into internal
|
||||
documentation accessible from the user context menu. Ensure that hyperlinks
|
||||
are valid and lead to up-to-date materials.
|
||||
1. Encourage the use of `coder support bundle` to allow workspace users to
|
||||
generate and provide network-related diagnostic data.
|
||||
@@ -90,11 +90,11 @@ Database:
|
||||
|
||||
## Available reference architectures
|
||||
|
||||
[Up to 1,000 users](../architectures/1k-users.md)
|
||||
[Up to 1,000 users](../../architecture/1k-users.md)
|
||||
|
||||
[Up to 2,000 users](../architectures/2k-users.md)
|
||||
[Up to 2,000 users](../../architecture/2k-users.md)
|
||||
|
||||
[Up to 3,000 users](../architectures/3k-users.md)
|
||||
[Up to 3,000 users](../../architecture/3k-users.md)
|
||||
|
||||
## Hardware recommendation
|
||||
|
||||
|
||||
@@ -6,15 +6,15 @@ infrastructure. For scale-testing Kubernetes clusters we recommend to install
|
||||
and use the dedicated Coder template,
|
||||
[scaletest-runner](https://github.com/coder/coder/tree/main/scaletest/templates/scaletest-runner).
|
||||
|
||||
Learn more about [Coder’s architecture](../architectures/architecture.md) and
|
||||
Learn more about [Coder’s architecture](../../architecture/architecture.md) and
|
||||
our [scale-testing methodology](scale-testing.md).
|
||||
|
||||
## Recent scale tests
|
||||
|
||||
> Note: the below information is for reference purposes only, and are not
|
||||
> intended to be used as guidelines for infrastructure sizing. Review the
|
||||
> [Reference Architectures](../architectures/validated-arch.md#node-sizing) for
|
||||
> hardware sizing recommendations.
|
||||
> [Reference Architectures](../../architecture/validated-arch.md#node-sizing)
|
||||
> for hardware sizing recommendations.
|
||||
|
||||
| Environment | Coder CPU | Coder RAM | Coder Replicas | Database | Users | Concurrent builds | Concurrent connections (Terminal/SSH) | Coder Version | Last tested |
|
||||
| ---------------- | --------- | --------- | -------------- | ----------------- | ----- | ----------------- | ------------------------------------- | ------------- | ------------ |
|
||||
|
||||
Reference in New Issue
Block a user