mirror of
https://github.com/coder/coder.git
synced 2026-09-22 05:05:20 +08:00
docs: add validated architecture (#13561)
* docs: add validated architecture * make: fmt * formatting * fix 404s * fix 404s pt 2 * fix 404s pt 3
This commit is contained in:
@@ -0,0 +1,396 @@
|
||||
# Architecture
|
||||
|
||||
The Coder deployment model is flexible and offers various components that
|
||||
platform administrators can deploy and scale depending on their use case. This
|
||||
page describes possible deployments, challenges, and risks associated with them.
|
||||
|
||||
## Primary components
|
||||
|
||||
### coderd
|
||||
|
||||
_coderd_ is the service created by running `coder server`. It is a thin API that
|
||||
connects workspaces, provisioners and users. _coderd_ stores its state in
|
||||
Postgres and is the only service that communicates with Postgres.
|
||||
|
||||
It offers:
|
||||
|
||||
- Dashboard (UI)
|
||||
- HTTP API
|
||||
- Dev URLs (HTTP reverse proxy to workspaces)
|
||||
- Workspace Web Applications (e.g for easy access to `code-server`)
|
||||
- Agent registration
|
||||
|
||||
### provisionerd
|
||||
|
||||
_provisionerd_ is the execution context for infrastructure modifying providers.
|
||||
At the moment, the only provider is Terraform (running `terraform`).
|
||||
|
||||
By default, the Coder server runs multiple provisioner daemons.
|
||||
[External provisioners](../provisioners.md) can be added for security or
|
||||
scalability purposes.
|
||||
|
||||
### Agents
|
||||
|
||||
An agent is the Coder service that runs within a user's remote workspace. It
|
||||
provides a consistent interface for coderd and clients to communicate with
|
||||
workspaces regardless of operating system, architecture, or cloud.
|
||||
|
||||
It offers the following services along with much more:
|
||||
|
||||
- SSH
|
||||
- Port forwarding
|
||||
- Liveness checks
|
||||
- `startup_script` automation
|
||||
|
||||
Templates are responsible for
|
||||
[creating and running agents](../../templates/index.md#coder-agent) within
|
||||
workspaces.
|
||||
|
||||
### Service Bundling
|
||||
|
||||
While _coderd_ and Postgres can be orchestrated independently, our default
|
||||
installation paths bundle them all together into one system service. It's
|
||||
perfectly fine to run a production deployment this way, but there are certain
|
||||
situations that necessitate decomposition:
|
||||
|
||||
- Reducing global client latency (distribute coderd and centralize database)
|
||||
- Achieving greater availability and efficiency (horizontally scale individual
|
||||
services)
|
||||
|
||||
### Workspaces
|
||||
|
||||
At the highest level, a workspace is a set of cloud resources. These resources
|
||||
can be VMs, Kubernetes clusters, storage buckets, or whatever else Terraform
|
||||
lets you dream up.
|
||||
|
||||
The resources that run the agent are described as _computational resources_,
|
||||
while those that don't are called _peripheral resources_.
|
||||
|
||||
Each resource may also be _persistent_ or _ephemeral_ depending on whether
|
||||
they're destroyed on workspace stop.
|
||||
|
||||
## Deployment models
|
||||
|
||||
### Single region architecture
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
This architecture consists of a single load balancer, several _coderd_ replicas,
|
||||
and _Coder workspaces_ deployed in the same region.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
- Deploy at least one _coderd_ replica per availability zone with _coderd_
|
||||
instances and provisioners. High availability is recommended but not essential
|
||||
for small deployments.
|
||||
- Single replica deployment is a special case that can address a
|
||||
tiny/small/proof-of-concept installation on a single virtual machine. If you
|
||||
are serving more than 100 users/workspaces, you should add more replicas.
|
||||
|
||||
**Coder workspace**
|
||||
|
||||
- For small deployments consider a lightweight workspace runtime like the
|
||||
[Sysbox](https://github.com/nestybox/sysbox) container runtime. Learn more how
|
||||
to enable
|
||||
[docker-in-docker using Sysbox](https://asciinema.org/a/kkTmOxl8DhEZiM2fLZNFlYzbo?speed=2).
|
||||
|
||||
**HA Database**
|
||||
|
||||
- Monitor node status and resource utilization metrics.
|
||||
- Implement robust backup and disaster recovery strategies to protect against
|
||||
data loss.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Load balancer**
|
||||
|
||||
- Distributes and load balances traffic from agents and clients to _Coder
|
||||
Server_ replicas across availability zones.
|
||||
- Layer 7 load balancing. The load balancer can decrypt SSL traffic, and
|
||||
re-encrypt using an internal certificate.
|
||||
- Session persistence (sticky sessions) can be disabled as _coderd_ instances
|
||||
are stateless.
|
||||
- WebSocket and long-lived connections must be supported.
|
||||
|
||||
**Single sign-on**
|
||||
|
||||
- Integrate with existing Single Sign-On (SSO) solutions used within the
|
||||
organization via the supported OAuth 2.0 or OpenID Connect standards.
|
||||
- Learn more about [Authentication in Coder](../auth.md).
|
||||
|
||||
### Multi-region architecture
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
This architecture is for globally distributed developer teams using Coder
|
||||
workspaces on daily basis. It features a single load balancer with regionally
|
||||
deployed _Workspace Proxies_, several _coderd_ replicas, and _Coder workspaces_
|
||||
provisioned in different regions.
|
||||
|
||||
Note: The _multi-region architecture_ assumes the same deployment principles as
|
||||
the _single region architecture_, but it extends them to multi region deployment
|
||||
with workspace proxies. Proxies are deployed in regions closest to developers to
|
||||
offer the fastest developer experience.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Workspace proxy**
|
||||
|
||||
- Workspace proxy offers developers the option to establish a fast relay
|
||||
connection when accessing their workspace via SSH, a workspace application, or
|
||||
port forwarding.
|
||||
- Dashboard connections, API calls (e.g. _list workspaces_) are not served over
|
||||
proxies.
|
||||
- Proxies do not establish connections to the database.
|
||||
- Proxy instances do not share authentication tokens between one another.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Proxy load balancer**
|
||||
|
||||
- Distributes and load balances workspace relay traffic in a single region
|
||||
across availability zones.
|
||||
- Layer 7 load balancing. The load balancer can decrypt SSL traffic, and
|
||||
re-encrypt using internal certificate.
|
||||
- Session persistence (sticky sessions) can be disabled as _coderd_ instances
|
||||
are stateless.
|
||||
- WebSocket and long-lived connections must be supported.
|
||||
|
||||
### Multi-cloud architecture
|
||||
|
||||
By distributing Coder workspaces across different cloud providers, organizations
|
||||
can mitigate the risk of downtime caused by provider-specific outages or
|
||||
disruptions. Additionally, multi-cloud deployment enables organizations to
|
||||
leverage the unique features and capabilities offered by each cloud provider,
|
||||
such as region availability and pricing models.
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
The deployment model comprises:
|
||||
|
||||
- `coderd` instances deployed within a single region of the same cloud provider,
|
||||
with replicas strategically distributed across availability zones.
|
||||
- Workspace provisioners deployed in each cloud, communicating with `coderd`
|
||||
instances.
|
||||
- Workspace proxies running in the same locations as provisioners to optimize
|
||||
user connections to workspaces for maximum speed.
|
||||
|
||||
Due to the relatively large overhead of cross-regional communication, it is not
|
||||
advised to set up multi-cloud control planes. It is recommended to keep coderd
|
||||
replicas and the database within the same cloud-provider and region.
|
||||
|
||||
Note: The _multi-cloud architecture_ follows the deployment principles outlined
|
||||
in the _multi-region architecture_. However, it adapts component selection based
|
||||
on the specific cloud provider. Developers can initiate workspaces based on the
|
||||
nearest region and technical specifications provided by the cloud providers.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Workspace provisioner**
|
||||
|
||||
- _Security recommendation_: Create a long, random pre-shared key (PSK) and add
|
||||
it to the regional secret store, so that local _provisionerd_ can access it.
|
||||
Remember to distribute it using safe, encrypted communication channel. The PSK
|
||||
must also be added to the _coderd_ configuration.
|
||||
|
||||
**Workspace proxy**
|
||||
|
||||
- _Security recommendation_: Use `coder` CLI to create
|
||||
[authentication tokens for every workspace proxy](../workspace-proxies.md#requirements),
|
||||
and keep them in regional secret stores. Remember to distribute them using
|
||||
safe, encrypted communication channel.
|
||||
|
||||
**Managed database**
|
||||
|
||||
- For AWS: _Amazon RDS for PostgreSQL_
|
||||
- For Azure: _Azure Database for PostgreSQL - Flexible Server_
|
||||
- For GCP: _Cloud SQL for PostgreSQL_
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Kubernetes platform (optional)**
|
||||
|
||||
- For AWS: _Amazon Elastic Kubernetes Service_
|
||||
- For Azure: _Azure Kubernetes Service_
|
||||
- For GCP: _Google Kubernetes Engine_
|
||||
|
||||
See how to deploy
|
||||
[Coder on Azure Kubernetes Service](https://github.com/ericpaulsen/coder-aks).
|
||||
|
||||
Learn more about [security requirements](../../install/kubernetes.md) for
|
||||
deploying Coder on Kubernetes.
|
||||
|
||||
**Load balancer**
|
||||
|
||||
- For AWS:
|
||||
- _AWS Network Load Balancer_
|
||||
- Level 4 load balancing
|
||||
- For Kubernetes deployment: annotate service with
|
||||
`service.beta.kubernetes.io/aws-load-balancer-type: "nlb"`, preserve the
|
||||
client source IP with `externalTrafficPolicy: Local`
|
||||
- _AWS Classic Load Balancer_
|
||||
- Level 7 load balancing
|
||||
- For Kubernetes deployment: set `sessionAffinity` to `None`
|
||||
- For Azure:
|
||||
- _Azure Load Balancer_
|
||||
- Level 7 load balancing
|
||||
- Azure Application Gateway
|
||||
- Deploy Azure Application Gateway when more advanced traffic routing
|
||||
policies are needed for Kubernetes applications.
|
||||
- Take advantage of features such as WebSocket support and TLS termination
|
||||
provided by Azure Application Gateway, enhancing the capabilities of
|
||||
Kubernetes deployments on Azure.
|
||||
- For GCP:
|
||||
- _Cloud Load Balancing_ with SSL load balancer:
|
||||
- Layer 4 load balancing, SSL enabled
|
||||
- _Cloud Load Balancing_ with HTTPS load balancer:
|
||||
- Layer 7 load balancing
|
||||
- For Kubernetes deployment: annotate service (with ingress enabled) with
|
||||
`kubernetes.io/ingress.class: "gce"`, leverage the `NodePort` service
|
||||
type.
|
||||
- Note: HTTP load balancer rejects DERP upgrade, Coder will fallback to
|
||||
WebSockets
|
||||
|
||||
**Single sign-on**
|
||||
|
||||
- For AWS:
|
||||
[AWS IAM Identity Center](https://docs.aws.amazon.com/singlesignon/latest/userguide/what-is.html)
|
||||
- For Azure:
|
||||
[Microsoft Entra ID Sign-On](https://learn.microsoft.com/en-us/entra/identity/app-proxy/)
|
||||
- For GCP:
|
||||
[Google Cloud Identity Platform](https://cloud.google.com/architecture/identity/single-sign-on)
|
||||
|
||||
### Air-gapped architecture
|
||||
|
||||
The air-gapped deployment model refers to the setup of Coder's development
|
||||
environment within a restricted network environment that lacks internet
|
||||
connectivity. This deployment model is often required for organizations with
|
||||
strict security policies or those operating in isolated environments, such as
|
||||
government agencies or certain enterprise setups.
|
||||
|
||||
The key features of the air-gapped architecture include:
|
||||
|
||||
- _Offline installation_: Deploy workspaces without relying on an external
|
||||
internet connection.
|
||||
- _Isolated package/plugin repositories_: Depend on local repositories for
|
||||
software installation, updates, and security patches.
|
||||
- _Secure data transfer_: Enable encrypted communication channels and robust
|
||||
access controls to safeguard sensitive information.
|
||||
|
||||
Learn more about [offline deployments](../../install/offline.md) of Coder.
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
The deployment model includes:
|
||||
|
||||
- _Workspace provisioners_ with direct access to self-hosted package and plugin
|
||||
repositories and restricted internet access.
|
||||
- _Mirror of Terraform Registry_ with multiple versions of Terraform plugins.
|
||||
- _Certificate Authority_ with all TLS certificates to build secure
|
||||
communication channels.
|
||||
|
||||
The model is compatible with various infrastructure models, enabling deployment
|
||||
across multiple regions and diverse cloud platforms.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Workspace provisioner**
|
||||
|
||||
- Includes Terraform binary in the container or system image.
|
||||
- Checks out Terraform plugins from self-hosted _Registry_ mirror.
|
||||
- Deploys workspace images stored in the self-hosted _Container Registry_.
|
||||
|
||||
**Coder server**
|
||||
|
||||
- Update checks are disabled (`CODER_UPDATE_CHECK=false`).
|
||||
- Telemetry data is not collected (`CODER_TELEMETRY_ENABLE=false`).
|
||||
- Direct connections are not possible, workspace traffic is relayed through
|
||||
control plane's DERP proxy.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Self-hosted Database**
|
||||
|
||||
- In the air-gapped deployment model, _Coderd_ instance is unable to download
|
||||
Postgres binaries from the internet, so external database must be provided.
|
||||
|
||||
**Container Registry**
|
||||
|
||||
- Since the _Registry_ is isolated from the internet, platform engineers are
|
||||
responsible for maintaining Workspace container images and conducting periodic
|
||||
updates of base Docker images.
|
||||
- It is recommended to keep [Dev Containers](../../templates/dev-containers.md)
|
||||
up to date with the latest released
|
||||
[Envbuilder](https://github.com/coder/envbuilder) runtime.
|
||||
|
||||
**Mirror of Terraform Registry**
|
||||
|
||||
- Stores all necessary Terraform plugin dependencies, ensuring successful
|
||||
workspace provisioning and maintenance without internet access.
|
||||
- Platform engineers are responsible for periodically updating the mirrored
|
||||
Terraform plugins, including
|
||||
[terraform-provider-coder](https://github.com/coder/terraform-provider-coder).
|
||||
|
||||
**Certificate Authority**
|
||||
|
||||
- Manages and issues TLS certificates to facilitate secure communication
|
||||
channels within the infrastructure.
|
||||
|
||||
### Dev Containers
|
||||
|
||||
Note: _Dev containers_ are at early stage and considered experimental at the
|
||||
moment.
|
||||
|
||||
This architecture enhances a Coder workspace with a
|
||||
[development container](https://containers.dev/) setup built using the
|
||||
[envbuilder](https://github.com/coder/envbuilder) project. Workspace users have
|
||||
the flexibility to extend generic, base developer environments with custom,
|
||||
project-oriented [features](https://containers.dev/features) without requiring
|
||||
platform administrators to push altered Docker images.
|
||||
|
||||
Learn more about
|
||||
[Dev containers support](https://coder.com/docs/v2/latest/templates/dev-containers)
|
||||
in Coder.
|
||||
|
||||

|
||||
|
||||
#### Components
|
||||
|
||||
The deployment model includes:
|
||||
|
||||
- _Workspace_ built using Coder template with _envbuilder_ enabled to set up the
|
||||
developer environment accordingly to the dev container spec.
|
||||
- _Container Registry_ for Docker images used by _envbuilder_, maintained by
|
||||
Coder platform engineers or developer productivity engineers.
|
||||
|
||||
Since this model is strictly focused on workspace nodes, it does not affect the
|
||||
setup of regional infrastructure. It can be deployed alongside other deployment
|
||||
models, in multiple regions, or across various cloud platforms.
|
||||
|
||||
##### Workload resources
|
||||
|
||||
**Coder workspace**
|
||||
|
||||
- Docker and Kubernetes based templates are supported.
|
||||
- The `docker_container` resource uses `ghcr.io/coder/envbuilder` as the base
|
||||
image.
|
||||
|
||||
_Envbuilder_ checks out the base Docker image from the container registry and
|
||||
installs selected features as specified in the `devcontainer.json` on top.
|
||||
Eventually, it starts the container with the developer environment.
|
||||
|
||||
##### Workload supporting resources
|
||||
|
||||
**Container Registry (optional)**
|
||||
|
||||
- Workspace nodes need access to the Container Registry to check out images. To
|
||||
shorten the provisioning time, it is recommended to deploy registry mirrors in
|
||||
the same region as the workspace nodes.
|
||||
@@ -1,90 +1,4 @@
|
||||
# Reference Architectures
|
||||
|
||||
This document provides prescriptive solutions and reference architectures to
|
||||
support successful deployments of up to 3000 users and outlines at a high-level
|
||||
the methodology currently used to scale-test Coder.
|
||||
|
||||
## General concepts
|
||||
|
||||
This section outlines core concepts and terminology essential for understanding
|
||||
Coder's architecture and deployment strategies.
|
||||
|
||||
### Administrator
|
||||
|
||||
An administrator is a user role within the Coder platform with elevated
|
||||
privileges. Admins have access to administrative functions such as user
|
||||
management, template definitions, insights, and deployment configuration.
|
||||
|
||||
### Coder
|
||||
|
||||
Coder, also known as _coderd_, is the main service recommended for deployment
|
||||
with multiple replicas to ensure high availability. It provides an API for
|
||||
managing workspaces and templates. Each _coderd_ replica has the capability to
|
||||
host multiple [provisioners](#provisioner).
|
||||
|
||||
### User
|
||||
|
||||
A user is an individual who utilizes the Coder platform to develop, test, and
|
||||
deploy applications using workspaces. Users can select available templates to
|
||||
provision workspaces. They interact with Coder using the web interface, the CLI
|
||||
tool, or directly calling API methods.
|
||||
|
||||
### Workspace
|
||||
|
||||
A workspace refers to an isolated development environment where users can write,
|
||||
build, and run code. Workspaces are fully configurable and can be tailored to
|
||||
specific project requirements, providing developers with a consistent and
|
||||
efficient development environment. Workspaces can be autostarted and
|
||||
autostopped, enabling efficient resource management.
|
||||
|
||||
Users can connect to workspaces using SSH or via workspace applications like
|
||||
`code-server`, facilitating collaboration and remote access. Additionally,
|
||||
workspaces can be parameterized, allowing users to customize settings and
|
||||
configurations based on their unique needs. Workspaces are instantiated using
|
||||
Coder templates and deployed on resources created by provisioners.
|
||||
|
||||
### Template
|
||||
|
||||
A template in Coder is a predefined configuration for creating workspaces.
|
||||
Templates streamline the process of workspace creation by providing
|
||||
pre-configured settings, tooling, and dependencies. They are built by template
|
||||
administrators on top of Terraform, allowing for efficient management of
|
||||
infrastructure resources. Additionally, templates can utilize Coder modules to
|
||||
leverage existing features shared with other templates, enhancing flexibility
|
||||
and consistency across deployments. Templates describe provisioning rules for
|
||||
infrastructure resources offered by Terraform providers.
|
||||
|
||||
### Workspace Proxy
|
||||
|
||||
A workspace proxy serves as a relay connection option for developers connecting
|
||||
to their workspace over SSH, a workspace app, or through port forwarding. It
|
||||
helps reduce network latency for geo-distributed teams by minimizing the
|
||||
distance network traffic needs to travel. Notably, workspace proxies do not
|
||||
handle dashboard connections or API calls.
|
||||
|
||||
### Provisioner
|
||||
|
||||
Provisioners in Coder execute Terraform during workspace and template builds.
|
||||
While the platform includes built-in provisioner daemons by default, there are
|
||||
advantages to employing external provisioners. These external daemons provide
|
||||
secure build environments and reduce server load, improving performance and
|
||||
scalability. Each provisioner can handle a single concurrent workspace build,
|
||||
allowing for efficient resource allocation and workload management.
|
||||
|
||||
### Registry
|
||||
|
||||
The Coder Registry is a platform where you can find starter templates and
|
||||
_Modules_ for various cloud services and platforms.
|
||||
|
||||
Templates help create self-service development environments using
|
||||
Terraform-defined infrastructure, while _Modules_ simplify template creation by
|
||||
providing common features like workspace applications, third-party integrations,
|
||||
or helper scripts.
|
||||
|
||||
Please note that the Registry is a hosted service and isn't available for
|
||||
offline use.
|
||||
|
||||
## Scale-testing methodology
|
||||
## Scale Testing
|
||||
|
||||
Scaling Coder involves planning and testing to ensure it can handle more load
|
||||
without compromising service. This process encompasses infrastructure setup,
|
||||
@@ -95,7 +9,7 @@ A dedicated Kubernetes cluster for Coder is Kubernetes cluster specifically
|
||||
configured to host and manage Coder workloads. Kubernetes provides container
|
||||
orchestration capabilities, allowing Coder to efficiently deploy, scale, and
|
||||
manage workspaces across a distributed infrastructure. This ensures high
|
||||
availability, fault tolerance, and scalability for Coder deployments. Code is
|
||||
availability, fault tolerance, and scalability for Coder deployments. Coder is
|
||||
deployed on this cluster using the
|
||||
[Helm chart](../../install/kubernetes.md#install-coder-with-helm).
|
||||
|
||||
@@ -315,96 +229,3 @@ Scaling down workspace nodes to zero is not recommended, as it will result in
|
||||
longer wait times for workspace provisioning by users. However, this may be
|
||||
necessary for workspaces with special resource requirements (e.g. GPUs) that
|
||||
incur significant cost overheads.
|
||||
|
||||
### Data plane: External database
|
||||
|
||||
While running in production, Coder requires a access to an external PostgreSQL
|
||||
database. Depending on the scale of the user-base, workspace activity, and High
|
||||
Availability requirements, the amount of CPU and memory resources required by
|
||||
Coder's database may differ.
|
||||
|
||||
#### Scaling formula
|
||||
|
||||
When determining scaling requirements, take into account the following
|
||||
considerations:
|
||||
|
||||
- `2 vCPU x 8 GB RAM x 512 GB storage`: A baseline for database requirements for
|
||||
Coder deployment with less than 1000 users, and low activity level (30% active
|
||||
users). This capacity should be sufficient to support 100 external
|
||||
provisioners.
|
||||
- Storage size depends on user activity, workspace builds, log verbosity,
|
||||
overhead on database encryption, etc.
|
||||
- Allocate two additional CPU core to the database instance for every 1000
|
||||
active users.
|
||||
- Enable _High Availability_ mode for database engine for large scale
|
||||
deployments.
|
||||
|
||||
If you enable [database encryption](../encryption.md) in Coder, consider
|
||||
allocating an additional CPU core to every `coderd` replica.
|
||||
|
||||
#### Performance optimization guidelines
|
||||
|
||||
We provide the following general recommendations for PostgreSQL settings:
|
||||
|
||||
- Increase number of vCPU if CPU utilization or database latency is high.
|
||||
- Allocate extra memory if database performance is poor, CPU utilization is low,
|
||||
and memory utilization is high.
|
||||
- Utilize faster disk options (higher IOPS) such as SSDs or NVMe drives for
|
||||
optimal performance enhancement and possibly reduce database load.
|
||||
|
||||
## Operational readiness
|
||||
|
||||
Operational readiness in Coder is about ensuring that everything is set up
|
||||
correctly before launching a platform into production. It involves making sure
|
||||
that the service is reliable, secure, and easily scales accordingly to user-base
|
||||
needs. Operational readiness is crucial because it helps prevent issues that
|
||||
could affect workspace users experience once the platform is live.
|
||||
|
||||
Learn about Coder design principles and architectural best practices described
|
||||
in the
|
||||
[Well-Architected Framework](https://coder.com/blog/coder-well-architected-framework).
|
||||
|
||||
### Configuration
|
||||
|
||||
1. Identify the required Helm values for configuration.
|
||||
1. Create `values.yaml` and add it to a version control system. _Note:_ it is
|
||||
highly recommended that you create a custom `values.yaml` as opposed to
|
||||
copying the entire default values.
|
||||
1. Determine the necessary environment variables.
|
||||
|
||||
### Template configuration
|
||||
|
||||
1. Establish a dedicated user account for the _Template Administrator_.
|
||||
1. Maintain Coder templates using version control.
|
||||
1. Consider implementing a GitOps workflow to automatically push new template.
|
||||
For example, on Github, you can use the
|
||||
[Update Coder Template](https://github.com/marketplace/actions/update-coder-template)
|
||||
action.
|
||||
1. Evaluate enabling automatic template updates upon workspace startup.
|
||||
|
||||
### Deployment
|
||||
|
||||
1. Leverage automation tooling to automate deployment and upgrades of Coder.
|
||||
|
||||
### Observability
|
||||
|
||||
1. Enable the Prometheus endpoint (environment variable:
|
||||
`CODER_PROMETHEUS_ENABLE`).
|
||||
1. Deploy a visual monitoring system such as Grafana for metrics visualization.
|
||||
1. Deploy a centralized logs aggregation solution to collect and monitor
|
||||
application logs.
|
||||
1. Review the [Prometheus response](../prometheus.md) and set up alarms on
|
||||
selected metrics.
|
||||
|
||||
### Database backups
|
||||
|
||||
1. Prepare internal scripts for dumping and restoring databases.
|
||||
1. Schedule regular database backups, especially before release upgrades.
|
||||
|
||||
### User support
|
||||
|
||||
1. Incorporate [support links](../appearance.md#support-links) into internal
|
||||
documentation accessible from the user context menu. Ensure that hyperlinks
|
||||
are valid and lead to up-to-date materials.
|
||||
1. Encourage the use of `coder support bundle` to allow workspace users to
|
||||
generate and provide network-related diagnostic data.
|
||||
@@ -0,0 +1,363 @@
|
||||
# Coder Validated Architecture
|
||||
|
||||
Many customers operate Coder in complex organizational environments, consisting
|
||||
of multiple business units, agencies, and/or subsidiaries. This can lead to
|
||||
numerous Coder deployments, caused by discrepancies in regulatory compliance,
|
||||
data sovereignty, and level of funding across groups. The Coder Validated
|
||||
Architecture (CVA) prescribes a Kubernetes-based deployment approach, enabling
|
||||
your organization to deploy a stable Coder instance that is easier to maintain
|
||||
and troubleshoot.
|
||||
|
||||
The following sections will detail the components of the Coder Validated
|
||||
Architecture, provide guidance on how to configure and deploy these components,
|
||||
and offer insights into how to maintain and troubleshoot your Coder environment.
|
||||
|
||||
- [General concepts](#general-concepts)
|
||||
- [Kubernetes Infrastructure](#kubernetes-infrastructure)
|
||||
- [PostgreSQL Database](#postgresql-database)
|
||||
- [Operational readiness](#operational-readiness)
|
||||
|
||||
## Who is this document for?
|
||||
|
||||
This guide targets the following personas. It assumes a basic understanding of
|
||||
cloud/on-premise computing, containerization, and the Coder platform.
|
||||
|
||||
| Role | Description |
|
||||
| ------------------------- | ------------------------------------------------------------------------------ |
|
||||
| Platform Engineers | Responsible for deploying, operating the Coder deployment and infrastructure |
|
||||
| Enterprise Architects | Responsible for architecting Coder deployments to meet enterprise requirements |
|
||||
| Managed Service Providers | Entities that deploy and run Coder software as a service for customers |
|
||||
|
||||
## CVA Guidance
|
||||
|
||||
| CVA provides: | CVA does not provide: |
|
||||
| ---------------------------------------------- | ---------------------------------------------------------------------------------------- |
|
||||
| Single and multi-region K8s deployment options | Prescribing OS, or cloud vs. on-premise |
|
||||
| Reference architectures for up to 3,000 users | An approval of your architecture; the CVA solely provides recommendations and guidelines |
|
||||
| Best practices for building a Coder deployment | Recommendations for every possible deployment scenario |
|
||||
|
||||
> For higher level design principles and architectural best practices, see
|
||||
> Coder's
|
||||
> [Well-Architected Framework](https://coder.com/blog/coder-well-architected-framework).
|
||||
|
||||
## General concepts
|
||||
|
||||
This section outlines core concepts and terminology essential for understanding
|
||||
Coder's architecture and deployment strategies.
|
||||
|
||||
### Administrator
|
||||
|
||||
An administrator is a user role within the Coder platform with elevated
|
||||
privileges. Admins have access to administrative functions such as user
|
||||
management, template definitions, insights, and deployment configuration.
|
||||
|
||||
### Coder control plane
|
||||
|
||||
Coder's control plane, also known as _coderd_, is the main service recommended
|
||||
for deployment with multiple replicas to ensure high availability. It provides
|
||||
an API for managing workspaces and templates, and serves the dashboard UI. In
|
||||
addition, each _coderd_ replica hosts 3 Terraform [provisioners](#provisioner)
|
||||
by default.
|
||||
|
||||
### User
|
||||
|
||||
A [user](../users.md) is an individual who utilizes the Coder platform to
|
||||
develop, test, and deploy applications using workspaces. Users can select
|
||||
available templates to provision workspaces. They interact with Coder using the
|
||||
web interface, the CLI tool, or directly calling API methods.
|
||||
|
||||
### Workspace
|
||||
|
||||
A [workspace](../../workspaces.md) refers to an isolated development environment
|
||||
where users can write, build, and run code. Workspaces are fully configurable
|
||||
and can be tailored to specific project requirements, providing developers with
|
||||
a consistent and efficient development environment. Workspaces can be
|
||||
autostarted and autostopped, enabling efficient resource management.
|
||||
|
||||
Users can connect to workspaces using SSH or via workspace applications like
|
||||
`code-server`, facilitating collaboration and remote access. Additionally,
|
||||
workspaces can be parameterized, allowing users to customize settings and
|
||||
configurations based on their unique needs. Workspaces are instantiated using
|
||||
Coder templates and deployed on resources created by provisioners.
|
||||
|
||||
### Template
|
||||
|
||||
A [template](../../templates/index.md) in Coder is a predefined configuration
|
||||
for creating workspaces. Templates streamline the process of workspace creation
|
||||
by providing pre-configured settings, tooling, and dependencies. They are built
|
||||
by template administrators on top of Terraform, allowing for efficient
|
||||
management of infrastructure resources. Additionally, templates can utilize
|
||||
Coder modules to leverage existing features shared with other templates,
|
||||
enhancing flexibility and consistency across deployments. Templates describe
|
||||
provisioning rules for infrastructure resources offered by Terraform providers.
|
||||
|
||||
### Workspace Proxy
|
||||
|
||||
A [workspace proxy](../workspace-proxies.md) serves as a relay connection option
|
||||
for developers connecting to their workspace over SSH, a workspace app, or
|
||||
through port forwarding. It helps reduce network latency for geo-distributed
|
||||
teams by minimizing the distance network traffic needs to travel. Notably,
|
||||
workspace proxies do not handle dashboard connections or API calls.
|
||||
|
||||
### Provisioner
|
||||
|
||||
Provisioners in Coder execute Terraform during workspace and template builds.
|
||||
While the platform includes built-in provisioner daemons by default, there are
|
||||
advantages to employing external provisioners. These external daemons provide
|
||||
secure build environments and reduce server load, improving performance and
|
||||
scalability. Each provisioner can handle a single concurrent workspace build,
|
||||
allowing for efficient resource allocation and workload management.
|
||||
|
||||
### Registry
|
||||
|
||||
The [Coder Registry](https://registry.coder.com) is a platform where you can
|
||||
find starter templates and _Modules_ for various cloud services and platforms.
|
||||
|
||||
Templates help create self-service development environments using
|
||||
Terraform-defined infrastructure, while _Modules_ simplify template creation by
|
||||
providing common features like workspace applications, third-party integrations,
|
||||
or helper scripts.
|
||||
|
||||
Please note that the Registry is a hosted service and isn't available for
|
||||
offline use.
|
||||
|
||||
## Kubernetes Infrastructure
|
||||
|
||||
Kubernetes is the recommended, and supported platform for deploying Coder in the
|
||||
enterprise. It is the hosting platform of choice for a large majority of Coder's
|
||||
Fortune 500 customers, and it is the platform in which we build and test against
|
||||
here at Coder.
|
||||
|
||||
### General recommendations
|
||||
|
||||
In general, it is recommended to deploy Coder into its own respective cluster,
|
||||
separate from production applications. Keep in mind that Coder runs development
|
||||
workloads, so the cluster should be deployed as such, without production-level
|
||||
configurations.
|
||||
|
||||
### Compute
|
||||
|
||||
Deploy your Kubernetes cluster with two node groups, one for Coder's control
|
||||
plane, and another for user workspaces (if you intend on leveraging K8s for
|
||||
end-user compute).
|
||||
|
||||
#### Control plane nodes
|
||||
|
||||
The Coder control plane node group must be static, to prevent scale down events
|
||||
from dropping pods, and thus dropping user connections to the dashboard UI and
|
||||
their workspaces.
|
||||
|
||||
Coder's Helm Chart supports
|
||||
[defining nodeSelectors, affinities, and tolerations](https://github.com/coder/coder/blob/e96652ebbcdd7554977594286b32015115c3f5b6/helm/coder/values.yaml#L221-L249)
|
||||
to schedule the control plane pods on the appropriate node group.
|
||||
|
||||
#### Workspace nodes
|
||||
|
||||
Coder workspaces can be deployed either as Pods or Deployments in Kubernetes.
|
||||
See our
|
||||
[example Kubernetes workspace template](https://github.com/coder/coder/tree/main/examples/templates/kubernetes).
|
||||
Configure the workspace node group to be auto-scaling, to dynamically allocate
|
||||
compute as users start/stop workspaces at the beginning and end of their day.
|
||||
Set nodeSelectors, affinities, and tolerations in Coder templates to assign
|
||||
workspaces to the given node group:
|
||||
|
||||
```hcl
|
||||
resource "kubernetes_deployment" "coder" {
|
||||
spec {
|
||||
template {
|
||||
metadata {
|
||||
labels = {
|
||||
app = "coder-workspace"
|
||||
}
|
||||
}
|
||||
|
||||
spec {
|
||||
affinity {
|
||||
pod_anti_affinity {
|
||||
preferred_during_scheduling_ignored_during_execution {
|
||||
weight = 1
|
||||
pod_affinity_term {
|
||||
label_selector {
|
||||
match_expressions {
|
||||
key = "app.kubernetes.io/instance"
|
||||
operator = "In"
|
||||
values = ["coder-workspace"]
|
||||
}
|
||||
}
|
||||
topology_key = # add your node group label here
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
tolerations {
|
||||
# Add your tolerations here
|
||||
}
|
||||
|
||||
node_selector {
|
||||
# Add your node selectors here
|
||||
}
|
||||
|
||||
container {
|
||||
image = "coder-workspace:latest"
|
||||
name = "dev"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Node sizing
|
||||
|
||||
For sizing recommendations, see the below reference architectures:
|
||||
|
||||
- [Up to 1,000 users](1k-users.md)
|
||||
|
||||
- [Up to 2,000 users](2k-users.md)
|
||||
|
||||
- [Up to 3,000 users](3k-users.md)
|
||||
|
||||
### Networking
|
||||
|
||||
It is likely your enterprise deploys Kubernetes clusters with various networking
|
||||
restrictions. With this in mind, Coder requires the following connectivity:
|
||||
|
||||
- Egress from workspace compute to the Coder control plane pods
|
||||
- Egress from control plane pods to Coder's PostgreSQL database
|
||||
- Egress from control plane pods to git and package repositories
|
||||
- Ingress from user devices to the control plane Load Balancer or Ingress
|
||||
controller
|
||||
|
||||
We recommend configuring your network policies in accordance with the above.
|
||||
Note that Coder workspaces do not require any ports to be open.
|
||||
|
||||
### Storage
|
||||
|
||||
If running Coder workspaces as Kubernetes Pods or Deployments, you will need to
|
||||
assign persistent storage. We recommend leveraging a
|
||||
[supported Container Storage Interface (CSI) driver](https://kubernetes-csi.github.io/docs/drivers.html)
|
||||
in your cluster, with Dynamic Provisioning and read/write, to provide on-demand
|
||||
storage to end-user workspaces.
|
||||
|
||||
The following Kubernetes volume types have been validated by Coder internally,
|
||||
and/or by our customers:
|
||||
|
||||
- [PersistentVolumeClaim](https://kubernetes.io/docs/concepts/storage/volumes/#persistentvolumeclaim)
|
||||
- [NFS](https://kubernetes.io/docs/concepts/storage/volumes/#nfs)
|
||||
- [subPath](https://kubernetes.io/docs/concepts/storage/volumes/#using-subpath)
|
||||
- [cephfs](https://kubernetes.io/docs/concepts/storage/volumes/#cephfs)
|
||||
|
||||
Our
|
||||
[example Kubernetes workspace template](https://github.com/coder/coder/blob/5b9a65e5c137232351381fc337d9784bc9aeecfc/examples/templates/kubernetes/main.tf#L191-L219)
|
||||
provisions a PersistentVolumeClaim block storage device, attached to the
|
||||
Deployment.
|
||||
|
||||
It is not recommended to mount volumes from the host node(s) into workspaces,
|
||||
for security and reliability purposes. The below volume types are _not_
|
||||
recommended for use with Coder:
|
||||
|
||||
- [Local](https://kubernetes.io/docs/concepts/storage/volumes/#local)
|
||||
- [hostPath](https://kubernetes.io/docs/concepts/storage/volumes/#hostpath)
|
||||
|
||||
Not that Coder's control plane filesystem is ephemeral, so no persistent storage
|
||||
is required.
|
||||
|
||||
## PostgreSQL database
|
||||
|
||||
Coder requires access to an external PostgreSQL database to store user data,
|
||||
workspace state, template files, and more. Depending on the scale of the
|
||||
user-base, workspace activity, and High Availability requirements, the amount of
|
||||
CPU and memory resources required by Coder's database may differ.
|
||||
|
||||
### Disaster recovery
|
||||
|
||||
Prepare internal scripts for dumping and restoring your database. We recommend
|
||||
scheduling regular database backups, especially before upgrading Coder to a new
|
||||
release. Coder does not support downgrades without initially restoring the
|
||||
database to the prior version.
|
||||
|
||||
### Performance efficiency
|
||||
|
||||
We highly recommend deploying the PostgreSQL instance in the same region (and if
|
||||
possible, same availability zone) as the Coder server to optimize for low
|
||||
latency connections. We recommend keeping latency under 10ms between the Coder
|
||||
server and database.
|
||||
|
||||
When determining scaling requirements, take into account the following
|
||||
considerations:
|
||||
|
||||
- `2 vCPU x 8 GB RAM x 512 GB storage`: A baseline for database requirements for
|
||||
Coder deployment with less than 1000 users, and low activity level (30% active
|
||||
users). This capacity should be sufficient to support 100 external
|
||||
provisioners.
|
||||
- Storage size depends on user activity, workspace builds, log verbosity,
|
||||
overhead on database encryption, etc.
|
||||
- Allocate two additional CPU core to the database instance for every 1000
|
||||
active users.
|
||||
- Enable High Availability mode for database engine for large scale deployments.
|
||||
|
||||
If you enable [database encryption](../encryption.md) in Coder, consider
|
||||
allocating an additional CPU core to every `coderd` replica.
|
||||
|
||||
#### Resource utilization guidelines
|
||||
|
||||
Below are general recommendations for sizing your PostgreSQL instance:
|
||||
|
||||
- Increase number of vCPU if CPU utilization or database latency is high.
|
||||
- Allocate extra memory if database performance is poor, CPU utilization is low,
|
||||
and memory utilization is high.
|
||||
- Utilize faster disk options (higher IOPS) such as SSDs or NVMe drives for
|
||||
optimal performance enhancement and possibly reduce database load.
|
||||
|
||||
## Operational readiness
|
||||
|
||||
Operational readiness in Coder is about ensuring that everything is set up
|
||||
correctly before launching a platform into production. It involves making sure
|
||||
that the service is reliable, secure, and easily scales accordingly to user-base
|
||||
needs. Operational readiness is crucial because it helps prevent issues that
|
||||
could affect workspace users experience once the platform is live.
|
||||
|
||||
### Helm Chart Configuration
|
||||
|
||||
1. Reference our [Helm chart values file](../../../helm/coder/values.yaml) and
|
||||
identify the required values for deployment.
|
||||
1. Create a `values.yaml` and add it to your version control system.
|
||||
1. Determine the necessary environment variables. Here is the
|
||||
[full list of supported server environment variables](../../cli/server.md).
|
||||
1. Follow our documented
|
||||
[steps for installing Coder via Helm](../../install/kubernetes.md).
|
||||
|
||||
### Template configuration
|
||||
|
||||
1. Establish dedicated accounts for users with the _Template Administrator_
|
||||
role.
|
||||
1. Maintain Coder templates using
|
||||
[version control](../../templates/change-management.md).
|
||||
1. Consider implementing a GitOps workflow to automatically push new template
|
||||
versions into Coder from git. For example, on Github, you can use the
|
||||
[Update Coder Template](https://github.com/marketplace/actions/update-coder-template)
|
||||
action.
|
||||
1. Evaluate enabling
|
||||
[automatic template updates](../../templates/general-settings.md#require-automatic-updates-enterprise)
|
||||
upon workspace startup.
|
||||
|
||||
### Observability
|
||||
|
||||
1. Enable the Prometheus endpoint (environment variable:
|
||||
`CODER_PROMETHEUS_ENABLE`).
|
||||
1. Deploy the
|
||||
[Coder Observability bundle](https://github.com/coder/observability) to
|
||||
leverage pre-configured dashboards, alerts, and runbooks for monitoring
|
||||
Coder. This includes integrations between Prometheus, Grafana, Loki, and
|
||||
Alertmanager.
|
||||
1. Review the [Prometheus response](../prometheus.md) and set up alarms on
|
||||
selected metrics.
|
||||
|
||||
### User support
|
||||
|
||||
1. Incorporate [support links](../appearance.md#support-links) into internal
|
||||
documentation accessible from the user context menu. Ensure that hyperlinks
|
||||
are valid and lead to up-to-date materials.
|
||||
1. Encourage the use of `coder support bundle` to allow workspace users to
|
||||
generate and provide network-related diagnostic data.
|
||||
+4
-4
@@ -4,15 +4,15 @@ infrastructure. For scale-testing Kubernetes clusters we recommend to install
|
||||
and use the dedicated Coder template,
|
||||
[scaletest-runner](https://github.com/coder/coder/tree/main/scaletest/templates/scaletest-runner).
|
||||
|
||||
Learn more about [Coder’s architecture](../about/architecture.md) and our
|
||||
[scale-testing methodology](architectures/index.md#scale-testing-methodology).
|
||||
Learn more about [Coder’s architecture](architectures/architecture.md) and our
|
||||
[scale-testing methodology](architectures/scale-testing.md).
|
||||
|
||||
## Recent scale tests
|
||||
|
||||
> Note: the below information is for reference purposes only, and are not
|
||||
> intended to be used as guidelines for infrastructure sizing. Review the
|
||||
> [Reference Architectures](architectures/index.md) for hardware sizing
|
||||
> recommendations.
|
||||
> [Reference Architectures](architectures/validated-arch.md#node-sizing) for
|
||||
> hardware sizing recommendations.
|
||||
|
||||
| Environment | Coder CPU | Coder RAM | Coder Replicas | Database | Users | Concurrent builds | Concurrent connections (Terminal/SSH) | Coder Version | Last tested |
|
||||
| ---------------- | --------- | --------- | -------------- | ----------------- | ----- | ----------------- | ------------------------------------- | ------------- | ------------ |
|
||||
|
||||
Reference in New Issue
Block a user