From ec3d5bdd451422dc7fea4b1c5c88007ae5078949 Mon Sep 17 00:00:00 2001 From: Gaurav Trivedi Date: Tue, 28 Jul 2026 00:44:04 +0530 Subject: [PATCH 1/2] chore: Plan as top-level Antora module + Optimize etcd procedures (JTBD) Plan module: - Create modules/plan/ with nav, pages, examples, and images - Move 4 pages from admin-guide and install modules - Break running-at-scale.adoc into 3 focused concept pages: cluster-capacity-limits, etcd-storage-limits, multi-cluster-deployments - Add page-aliases for backward-compatible URLs Optimize module: - Add 3 procedure pages to reduce etcd storage pressure: configuring-automatic-cleanup-of-inactive-workspaces (DevWorkspace Pruner) disabling-workspace-ca-bundle-mount (ca-certs-merged ConfigMaps) disabling-copied-csvs (OLM namespace-scoped CSV copies) - Update optimize/nav.adoc with "Reduce etcd pressure at scale" section Plan nav links to Optimize procedures via cross-module xrefs. Zero content loss. Matches downstream devspaces-documentation structure. Co-authored-by: Cursor --- antora.yml | 1 + .../partials/proc_configuring-routes.adoc | 2 +- modules/discover/nav.adoc | 1 - ...-running-at-scale-calculate-resources.adoc | 6 - modules/install/nav.adoc | 8 +- ...-on-amazon-elastic-kubernetes-service.adoc | 2 +- ...talling-che-on-minikube-keycloak-oidc.adoc | 2 +- .../proc_installing-che-on-minikube.adoc | 2 +- ...installing-che-on-openshift-using-cli.adoc | 2 +- ...alling-che-on-red-hat-openshift-local.adoc | 2 +- ...che-on-the-virtual-kubernetes-cluster.adoc | 2 +- .../install/pages/proc_uninstalling-che.adoc | 4 +- modules/install/pages/running-at-scale.adoc | 208 ------------------ ...r-custom-resource-during-installation.adoc | 2 +- ...g-images-for-a-restricted-environment.adoc | 2 +- modules/optimize/nav.adoc | 4 + ...omatic-cleanup-of-inactive-workspaces.adoc | 61 +++++ .../optimize/pages/disabling-copied-csvs.adoc | 49 +++++ .../disabling-workspace-ca-bundle-mount.adoc | 49 +++++ ...ctl-management-tool-on-linux-or-macos.adoc | 0 ...the-chectl-management-tool-on-windows.adoc | 0 .../snip_che-example-devfile-disclaimer.adoc | 0 ...-running-at-scale-calculate-resources.adoc | 6 + .../snip_che-supported-platforms.adoc | 6 +- .../snip_che-multi-cluster.png | Bin modules/plan/nav.adoc | 11 + ...calculating-che-resource-requirements.adoc | 10 +- .../plan/pages/cluster-capacity-limits.adoc | 58 +++++ modules/plan/pages/etcd-storage-limits.adoc | 26 +++ ...installing-the-chectl-management-tool.adoc | 4 +- .../plan/pages/multi-cluster-deployments.adoc | 66 ++++++ .../pages/supported-platforms.adoc | 8 +- ...-upgrading-the-chectl-management-tool.adoc | 2 +- ...ing-che-using-the-cli-management-tool.adoc | 2 +- 34 files changed, 367 insertions(+), 241 deletions(-) delete mode 100644 modules/install/examples/snip_che-running-at-scale-calculate-resources.adoc delete mode 100644 modules/install/pages/running-at-scale.adoc create mode 100644 modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc create mode 100644 modules/optimize/pages/disabling-copied-csvs.adoc create mode 100644 modules/optimize/pages/disabling-workspace-ca-bundle-mount.adoc rename modules/{install => plan}/examples/proc_che-installing-the-chectl-management-tool-on-linux-or-macos.adoc (100%) rename modules/{install => plan}/examples/proc_che-installing-the-chectl-management-tool-on-windows.adoc (100%) create mode 100644 modules/plan/examples/snip_che-example-devfile-disclaimer.adoc create mode 100644 modules/plan/examples/snip_che-running-at-scale-calculate-resources.adoc rename modules/{discover => plan}/examples/snip_che-supported-platforms.adoc (69%) rename modules/{install => plan}/images/running-at-scale/snip_che-multi-cluster.png (100%) create mode 100644 modules/plan/nav.adoc rename modules/{install => plan}/pages/calculating-che-resource-requirements.adoc (93%) create mode 100644 modules/plan/pages/cluster-capacity-limits.adoc create mode 100644 modules/plan/pages/etcd-storage-limits.adoc rename modules/{install => plan}/pages/installing-the-chectl-management-tool.adoc (76%) create mode 100644 modules/plan/pages/multi-cluster-deployments.adoc rename modules/{discover => plan}/pages/supported-platforms.adoc (69%) diff --git a/antora.yml b/antora.yml index f019768006..b27f3960fb 100644 --- a/antora.yml +++ b/antora.yml @@ -6,6 +6,7 @@ prerelease: true start_page: discover:what-is-che.adoc nav: - modules/discover/nav.adoc + - modules/plan/nav.adoc - modules/install/nav.adoc - modules/get-started-admin/nav.adoc - modules/get-started-user/nav.adoc diff --git a/modules/administration-guide/partials/proc_configuring-routes.adoc b/modules/administration-guide/partials/proc_configuring-routes.adoc index 8e74a48c60..9321b02fdc 100644 --- a/modules/administration-guide/partials/proc_configuring-routes.adoc +++ b/modules/administration-guide/partials/proc_configuring-routes.adoc @@ -8,7 +8,7 @@ You can configure labels, annotations, and domains for OpenShift Route to work w * An active `oc` session with administrative permissions to the OpenShift cluster. See link:https://docs.openshift.com/container-platform/{ocp4-ver}/cli_reference/openshift_cli/getting-started-cli.html[Getting started with the OpenShift CLI]. -* `{prod-cli}`. See: xref:install:installing-the-chectl-management-tool.adoc[]. +* `{prod-cli}`. See: xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure diff --git a/modules/discover/nav.adoc b/modules/discover/nav.adoc index f9b1d99683..a447a51698 100644 --- a/modules/discover/nav.adoc +++ b/modules/discover/nav.adoc @@ -14,5 +14,4 @@ *** xref:che-devfile-registry.adoc[] *** xref:plugin-registry.adoc[] ** xref:user-workspaces.adoc[] -* xref:supported-platforms.adoc[] * xref:roles-and-tasks.adoc[] diff --git a/modules/install/examples/snip_che-running-at-scale-calculate-resources.adoc b/modules/install/examples/snip_che-running-at-scale-calculate-resources.adoc deleted file mode 100644 index 8fd9b6f4a6..0000000000 --- a/modules/install/examples/snip_che-running-at-scale-calculate-resources.adoc +++ /dev/null @@ -1,6 +0,0 @@ -:_content-type: SNIPPET - -[NOTE] -==== -You can find more details about calculating resource requirements in the link:https://eclipse.dev/che/docs/stable/administration-guide/calculating-che-resource-requirements/[official documentation]. -==== \ No newline at end of file diff --git a/modules/install/nav.adoc b/modules/install/nav.adoc index 9e9e9ec674..4b0155b6a4 100644 --- a/modules/install/nav.adoc +++ b/modules/install/nav.adoc @@ -2,10 +2,10 @@ * xref:con_installation-overview.adoc[] // Requirements * Requirements -** xref:administration-guide:supported-platforms.adoc[] -** xref:calculating-che-resource-requirements.adoc[] -** xref:running-at-scale.adoc[] -** xref:installing-the-chectl-management-tool.adoc[] +** xref:plan:supported-platforms.adoc[] +** xref:plan:calculating-che-resource-requirements.adoc[] +** xref:plan:cluster-capacity-limits.adoc[] +** xref:plan:installing-the-chectl-management-tool.adoc[] // Installing * Deploy using GitOps ** xref:proc_deploying-che-using-gitops.adoc[] diff --git a/modules/install/pages/installing-che-on-amazon-elastic-kubernetes-service.adoc b/modules/install/pages/installing-che-on-amazon-elastic-kubernetes-service.adoc index 10f0df07ed..280cdb88a2 100644 --- a/modules/install/pages/installing-che-on-amazon-elastic-kubernetes-service.adoc +++ b/modules/install/pages/installing-che-on-amazon-elastic-kubernetes-service.adoc @@ -15,7 +15,7 @@ include::partial$snip_persona-admin.adoc[] * You have `helm` installed. See link:https://helm.sh/docs/intro/install/[Installing Helm]. -* You have `{prod-cli}` installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have `{prod-cli}` installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. * You have the `aws` CLI installed. See link:https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html[AWS CLI install and update instructions]. diff --git a/modules/install/pages/proc_installing-che-on-minikube-keycloak-oidc.adoc b/modules/install/pages/proc_installing-che-on-minikube-keycloak-oidc.adoc index cf252b28bb..fabd47bb74 100644 --- a/modules/install/pages/proc_installing-che-on-minikube-keycloak-oidc.adoc +++ b/modules/install/pages/proc_installing-che-on-minikube-keycloak-oidc.adoc @@ -25,7 +25,7 @@ Single-node {kubernetes} clusters are suited only for testing or development. Do * You have `{orch-cli}` installed. See link:https://kubernetes.io/docs/tasks/tools/#kubectl[Installing `{orch-cli}`]. -* You have `{prod-cli}` installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have `{prod-cli}` installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure diff --git a/modules/install/pages/proc_installing-che-on-minikube.adoc b/modules/install/pages/proc_installing-che-on-minikube.adoc index fb1e774de7..83ea91b832 100644 --- a/modules/install/pages/proc_installing-che-on-minikube.adoc +++ b/modules/install/pages/proc_installing-che-on-minikube.adoc @@ -23,7 +23,7 @@ Single-node {kubernetes} clusters are suited only for testing or development. Do * You have `{orch-cli}` installed. See link:https://kubernetes.io/docs/tasks/tools/#kubectl[Installing `{orch-cli}`]. -* You have `{prod-cli}` installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have `{prod-cli}` installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure diff --git a/modules/install/pages/proc_installing-che-on-openshift-using-cli.adoc b/modules/install/pages/proc_installing-che-on-openshift-using-cli.adoc index ddf716aa18..57b5caed08 100644 --- a/modules/install/pages/proc_installing-che-on-openshift-using-cli.adoc +++ b/modules/install/pages/proc_installing-che-on-openshift-using-cli.adoc @@ -18,7 +18,7 @@ include::partial$snip_persona-admin.adoc[] * You have an active `{orch-cli}` session with administrative permissions to the {orch-name} cluster. See link:https://docs.openshift.com/container-platform/{ocp4-ver}/cli_reference/openshift_cli/getting-started-cli.html[Getting started with the OpenShift CLI]. -* You have the `{prod-cli}` management tool installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have the `{prod-cli}` management tool installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure diff --git a/modules/install/pages/proc_installing-che-on-red-hat-openshift-local.adoc b/modules/install/pages/proc_installing-che-on-red-hat-openshift-local.adoc index c5b5e45af8..d3b6e2cbe1 100644 --- a/modules/install/pages/proc_installing-che-on-red-hat-openshift-local.adoc +++ b/modules/install/pages/proc_installing-che-on-red-hat-openshift-local.adoc @@ -16,7 +16,7 @@ include::partial$snip_persona-admin.adoc[] * You have an active `{orch-cli}` session with administrative permissions to the {orch-name} cluster. See link:https://docs.openshift.com/container-platform/{ocp4-ver}/cli_reference/openshift_cli/getting-started-cli.html[Getting started with the OpenShift CLI]. -* You have `{prod-cli}` installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have `{prod-cli}` installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. * You have a running instance of {rh-os-local}. See link:https://developers.redhat.com/products/openshift-local/overview[{rh-os-local} overview]. diff --git a/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc b/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc index 6995b3ea67..d139cb8868 100644 --- a/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc +++ b/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc @@ -22,7 +22,7 @@ include::partial$snip_persona-admin.adoc[] * You have `kubelogin` installed. See link:https://github.com/int128/kubelogin[Installing kubelogin]. -* You have `{prod-cli}` installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have `{prod-cli}` installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. * You have an active `{orch-cli}` session with administrative permissions to the destination {orch-name} cluster. diff --git a/modules/install/pages/proc_uninstalling-che.adoc b/modules/install/pages/proc_uninstalling-che.adoc index cfa65ffec6..0ecd9bf989 100644 --- a/modules/install/pages/proc_uninstalling-che.adoc +++ b/modules/install/pages/proc_uninstalling-che.adoc @@ -21,7 +21,7 @@ Uninstalling {prod-short} removes all {prod-short}-related user data. * You have an active `{orch-cli}` session with administrative permissions to the {orch-name} cluster. See link:https://docs.openshift.com/container-platform/{ocp4-ver}/cli_reference/openshift_cli/getting-started-cli.html[Getting started with the OpenShift CLI]. -* You have the `{prod-cli}` management tool installed. See xref:installing-the-chectl-management-tool.adoc[]. +* You have the `{prod-cli}` management tool installed. See xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure @@ -58,4 +58,4 @@ The expected output is `NotFound`. [role="_additional-resources"] .Additional resources -* xref:installing-the-chectl-management-tool.adoc[] +* xref:plan:installing-the-chectl-management-tool.adoc[] diff --git a/modules/install/pages/running-at-scale.adoc b/modules/install/pages/running-at-scale.adoc deleted file mode 100644 index 37a6bdcaeb..0000000000 --- a/modules/install/pages/running-at-scale.adoc +++ /dev/null @@ -1,208 +0,0 @@ -:_content-type: CONCEPT -:description: Scaling cloud development environments (CDEs) to thousands of concurrent workspaces imposes high infrastructure demands on etcd storage, Operator memory, and worker node capacity. This topic covers the bottlenecks, tested maximums, and architectural patterns, including multi-cluster deployments, that address these challenges. -:keywords: install, scale, infrastructure, workload, scalability, CDE, cloud -:navtitle: Scalability limits and multi-cluster deployments -:page-aliases: administration-guide:running-at-scale.adoc, plan:running-at-scale.adoc - -[id="running-at-scale"] -= Scalability limits and multi-cluster deployments - -[role="_abstract"] -Scaling cloud development environments (CDEs) to thousands of concurrent workspaces imposes high infrastructure demands on etcd storage, Operator memory, and worker node capacity. This topic covers the bottlenecks, tested maximums, and architectural patterns, including multi-cluster deployments, that address these challenges. - -include::partial$snip_persona-admin.adoc[] - -Such a scale imposes high infrastructure demands and introduces potential bottlenecks that can impact performance and stability. Addressing these challenges requires meticulous planning, strategic architectural choices, monitoring, and continuous optimization. - -CDE workloads are particularly complex to scale. The underlying IDE solutions, such as Visual Studio Code - Open Source ("Code - OSS") or JetBrains Gateway, are designed as single-user applications, not as multitenant services. - -== Tested cluster maximums that constrain scaling - -While there is no strict limit on the number of resources in a {kubernetes} cluster, there are certain considerations for large clusters to remember. - -{ocp}, a certified distribution of {kubernetes}, provides a set of tested maximums for various resources. These maximums can serve as an initial guideline for planning your environment: - -.{ocp} tested cluster maximums -[%header,format=csv] -|=== -Resource type, Tested maximum -Number of nodes,2000 -Number of pods,150000 -Number of pods per node,2500 -Number of namespace,10000 -Number of services,10000 -Number of secrets,80000 -Number of config maps,90000 -|=== - -For more details on {ocp} tested object maximums, see Additional resources. - -For example, it is generally not recommended to have more than 10,000 namespaces due to potential performance and management overhead. In {prod}, each user is allocated a namespace. If you expect the user base to be large, consider spreading workloads across multiple "fit-for-purpose" clusters and potentially using solutions for multi-cluster orchestration. - -== How workspace size determines cluster capacity - -When deploying {prod} on {kubernetes}, accurately calculate the resource requirements for each CDE, including memory and CPU or GPU needs. This determines the right sizing of the cluster. In general, the CDE size is limited by and cannot be bigger than the worker node size. - -The resource requirements for CDEs can vary significantly based on the specific workloads and configurations. A simple CDE might require only a few hundred megabytes of memory. A more complex one might need several gigabytes of memory and multiple CPU cores. - -For details about calculating resource requirements, see Additional resources. - -== Why etcd is the primary scaling bottleneck - -The primary datastore of {kubernetes} cluster configuration and state is etcd. It holds information about nodes, pods, services, and custom resources. - -As a distributed key-value store, etcd does not scale well past a certain threshold. As the size of etcd grows, so does the load on the cluster, risking its stability. - -[IMPORTANT] -==== -The default etcd size is 2 GB, and the recommended maximum is 8 GB. Exceeding the maximum limit can make the {kubernetes} cluster unstable and unresponsive. Even though the data stored in a `ConfigMap` cannot exceed 1 MiB by design, a few thousand relatively large `ConfigMap` objects can overload etcd storage. -==== - -== How large Kubernetes objects strain etcd - -The size of the objects stored in etcd is also a critical factor. Each object consumes space, and as the number of objects increases, the overall size of etcd grows. The larger the object, the more space it takes. For example, etcd can be overloaded with only a few thousand large {kubernetes} objects. - -In the context of {prod}, by default the Operator creates and manages the 'ca-certs-merged' ConfigMap, which contains the Certificate Authorities (CAs) bundle, in every user namespace. With a large number of Transport Layer Security (TLS) certificates in the cluster, this results in additional etcd usage. - -To disable mounting the CA bundle by using the `ConfigMap` under the `/etc/pki/ca-trust/extracted/pem` path, configure the `CheCluster` Custom Resource by setting the `disableWorkspaceCaBundleMount` property to `true`. With this configuration, only custom certificates are mounted under the path `/public-certs`: - -[source,yaml] ----- -spec: - devEnvironments: - trustedCerts: - disableWorkspaceCaBundleMount: true ----- - -== How DevWorkspace objects affect etcd storage - -For large {kubernetes} deployments, particularly those involving a high number of custom resources such as `DevWorkspace` objects, which represent CDEs, etcd can become a significant performance bottleneck. - -[IMPORTANT] -==== -Based on the load testing for 6,000 `DevWorkspace` objects, storage consumption for etcd was approximately 2.5GB. -==== - -Starting from {devworkspace} Operator version 0.34.0, you can configure a pruner that automatically cleans up `DevWorkspace` objects that were not in use for a certain period of time. To set the pruner up, configure the `DevWorkspaceOperatorConfig` object as follows: - -[source,yaml] ----- -apiVersion: controller.devfile.io/v1alpha1 -kind: DevWorkspaceOperatorConfig -metadata: - name: devworkspace-operator-config - namespace: crw -config: - workspace: - cleanupCronJob: - enabled: true - dryRun: false - retainTime: 2592000 - schedule: “0 0 1 * *” ----- - -retainTime:: By default, if a workspace was not started for more than 30 days, it is marked for deletion. - -schedule:: By default, the pruner runs once per month. - -== Reduce etcd usage by disabling Copied CSVs - -When an Operator is installed by the Operator Lifecycle Manager (OLM), a stripped-down copy of its CSV is created in every {namespace} the Operator watches. These “Copied CSVs” communicate which controllers are reconciling resource events in a given namespace. - -On large clusters with hundreds or thousands of namespaces, Copied CSVs consume an unsustainable amount of resources, including OLM memory, etcd storage, and network bandwidth. To remove the CSVs copied to every namespace, configure the `OLMConfig` object: - -[source,yaml] ----- -apiVersion: operators.coreos.com/v1 -kind: OLMConfig -metadata: - name: cluster -spec: - features: - disableCopiedCSVs: true ----- - -For more information about the `disableCopiedCSVs` feature, see Additional resources. - -In clusters with many namespaces and cluster-wide Operators, Copied CSVs increase etcd storage usage and memory consumption. Disabling Copied CSVs significantly reduces the data stored in etcd and improves cluster performance and stability. - -Disabling Copied CSVs also reduces the memory footprint of OLM, as it no longer maintains these additional resources. - -For more details about disabling Copied CSVs, see Additional resources. - -== Scale worker nodes to match workspace demand - -Although cluster autoscaling is a powerful {kubernetes} feature, you cannot always rely on it. Consider predictive scaling by analyzing load data to detect daily or weekly usage patterns. - -If your workloads follow a pattern with dramatic peaks throughout the day, provision worker nodes accordingly. For example, if workspaces increase during business hours and decrease during off-hours, predictive scaling adjusts the number of worker nodes. This ensures enough resources are available during peak load while minimizing costs during off-peak hours. - -You can also use open-source solutions such as Karpenter for configuration and lifecycle management of the worker nodes. Karpenter can dynamically provision and optimize worker nodes based on the specific requirements of the workloads. This helps improve resource utilization and reduce costs. - -== Distribute workloads across multiple clusters - -By design, {prod} is not multi-cluster aware. You can only have one instance per cluster. - -However, you can run {prod} in a multi-cluster environment by deploying {prod} in each cluster. Use a load balancer or Domain Name System (DNS)-based routing to direct traffic to the appropriate instance. This approach distributes the workload across clusters and provides redundancy in case of cluster failures. - -== How Developer Sandbox runs Dev Spaces across clusters - -You can test running {prod-short} in a multi-cluster environment by using the Developer Sandbox, a free trial environment by Red Hat. - -From an infrastructure perspective, the Developer Sandbox consists of multiple Red Hat OpenShift Service on AWS (ROSA) clusters. On each cluster, the productized version of {prod} is installed and configured using Argo CD. The workspaces.openshift.com URL is used as a single entry point to the {prod} instances across clusters. - -image::running-at-scale/snip_{project-context}-multi-cluster.png[Scheme of a multi-cluster environment] - -For implementation details about the multicluster redirector, see Additional resources. - -[IMPORTANT] -==== -The multi-cluster architecture of workspaces.openshift.com is part of the Developer Sandbox. It is a Developer Sandbox-specific solution that cannot be reused as-is in other environments. However, you can use it as a reference for implementing a similar solution well-tailored to your specific multicluster needs. -==== - -== Route developers to the correct cluster automatically - -Red Hat offers an open-source, Quarkus-based service that acts as a single gateway for developers. This service automatically redirects users to the correct {prod} instance on the appropriate cluster based on their {ocp} group membership. For the community-supported version, see Additional resources. - -== What the redirector requires - -A critical requirement for the multicluster redirector is that all users are provisioned to the host cluster where the redirector is deployed. Users authenticate through the OAuth flow of this cluster, even if they never run workloads there. The host cluster’s OpenShift Container Platform groups determine the routing logic. For deployment instructions, see Additional resources. - -== Map OpenShift groups to cluster URLs - -The routing configuration uses a `ConfigMap` that contains JSON to map OpenShift Container Platform groups to {prod} URLs. The redirector uses this file to update routing tables in real-time without requiring restarts. - -== How authentication and routing work - -The routing process follows these steps: - -. Authenticate by using OAuth through a proxy sidecar. -. Pass identity and group information through HTTP headers. -. Verify group memberships by using {ocp} API queries. -. Determine the appropriate {prod} URL by using a mapping lookup. -. Redirect the user to the designated cluster instance. - -If users belong to multiple {ocp} groups, they can choose their desired {prod} instance from a selection dashboard. - -[role="_additional-resources"] -.Additional resources - -* link:https://che.eclipseprojects.io/2025/04/29/@ilya.buziuk-running-at-scale.html[Running at scale] -* link:https://developers.redhat.com/articles/2026/01/23/enterprise-multi-cluster-scalability-openshift-dev-spaces[Enterprise multi-cluster scalability] -* link:https://kubernetes.io/[Kubernetes] -* link:https://github.com/microsoft/vscode[Visual Studio Code - Open Source ("Code - OSS")] -* link:https://www.jetbrains.com/remote-development/gateway/[JetBrains Gateway] -* link:https://kubernetes.io/docs/setup/best-practices/cluster-large/[Considerations for large clusters] -* link:https://www.redhat.com/en/technologies/cloud-computing/openshift[OpenShift Container Platform] -* link:https://docs.redhat.com/en/documentation/openshift_container_platform/4.18/html/scalability_and_performance/planning-your-environment-according-to-object-maximums#planning-your-environment-according-to-object-maximums[OpenShift Container Platform tested object maximums] -* xref:calculating-che-resource-requirements.adoc[] -* link:https://etcd.io/[etcd] -* link:https://olm.operatorframework.io/[Operator Lifecycle Manager (OLM)] -* link:https://github.com/operator-framework/enhancements/blob/master/enhancements/olm-toggle-copied-csvs.md[OLM toggle Copied CSVs enhancement proposal] -* link:https://olm.operatorframework.io/docs/advanced-tasks/configuring-olm/#disabling-copied-csvs[Disabling Copied CSVs in OLM] -* link:https://karpenter.sh/[Karpenter] -* link:https://developers.redhat.com/developer-sandbox[Developer Sandbox] -* link:https://www.redhat.com/en/technologies/cloud-computing/openshift/aws[Red Hat OpenShift Service on AWS (ROSA)] -* link:https://argo-cd.readthedocs.io/en/stable/[Argo CD] -* link:https://workspaces.openshift.com/[workspaces.openshift.com] -* link:https://github.com/codeready-toolchain/crw-multicluster-redirector[crw-multicluster-redirector GitHub repository] -* link:https://github.com/redhat-developer/devspaces-multicluster-redirector[devspaces-multicluster-redirector GitHub repository] diff --git a/modules/install/pages/using-chectl-to-configure-the-checluster-custom-resource-during-installation.adoc b/modules/install/pages/using-chectl-to-configure-the-checluster-custom-resource-during-installation.adoc index 04ab10d4cb..bd5fbf2e2a 100644 --- a/modules/install/pages/using-chectl-to-configure-the-checluster-custom-resource-during-installation.adoc +++ b/modules/install/pages/using-chectl-to-configure-the-checluster-custom-resource-during-installation.adoc @@ -17,7 +17,7 @@ include::partial$snip_persona-admin.adoc[] * An active `{orch-cli}` session with administrative permissions to the {orch-name} cluster. See {orch-cli-link}. -* `{prod-cli}`. See: xref:installing-the-chectl-management-tool.adoc[]. +* `{prod-cli}`. See: xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure * Create a `che-operator-cr-patch.yaml` YAML file that contains the subset of the `CheCluster` Custom Resource to configure: diff --git a/modules/install/partials/snip_preparing-images-for-a-restricted-environment.adoc b/modules/install/partials/snip_preparing-images-for-a-restricted-environment.adoc index dede7d6f1e..9eeb4d89cd 100644 --- a/modules/install/partials/snip_preparing-images-for-a-restricted-environment.adoc +++ b/modules/install/partials/snip_preparing-images-for-a-restricted-environment.adoc @@ -28,4 +28,4 @@ * An active `skopeo` session with administrative access to the private Docker registry. link:https://github.com/containers/skopeo#authenticating-to-a-registry[Authenticating to a registry], and link:https://docs.redhat.com/en/documentation/openshift_container_platform/{ocp4-ver}/html/disconnected_environments/installing-mirroring-disconnected-about[Mirroring images for a disconnected installation]. -* `{prod-cli}` for {prod-short} version {prod-ver}. See xref:install:installing-the-chectl-management-tool.adoc[]. +* `{prod-cli}` for {prod-short} version {prod-ver}. See xref:plan:installing-the-chectl-management-tool.adoc[]. diff --git a/modules/optimize/nav.adoc b/modules/optimize/nav.adoc index d8b474ec94..700b58e4b7 100644 --- a/modules/optimize/nav.adoc +++ b/modules/optimize/nav.adoc @@ -12,4 +12,8 @@ * xref:configuring-autoscaling.adoc[] ** xref:configuring-number-of-replicas.adoc[] ** xref:configuring-machine-autoscaling.adoc[] +* Reduce etcd pressure at scale +** xref:configuring-automatic-cleanup-of-inactive-workspaces.adoc[] +** xref:disabling-workspace-ca-bundle-mount.adoc[] +** xref:disabling-copied-csvs.adoc[] * xref:verify-optimization-impact.adoc[] diff --git a/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc b/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc new file mode 100644 index 0000000000..5f02908590 --- /dev/null +++ b/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc @@ -0,0 +1,61 @@ +:_content-type: PROCEDURE +:description: Configure the DevWorkspace Operator pruner to automatically delete workspaces that have not been started within a configurable retention period. +:keywords: optimize, etcd, cleanup, pruner, DevWorkspace, cron +:navtitle: Configure automatic cleanup of inactive workspaces + +[id="configuring-automatic-cleanup-of-inactive-workspaces"] += Configure automatic cleanup of inactive workspaces + +[role="_abstract"] +For large {kubernetes} deployments, particularly those involving a high number of custom resources such as `DevWorkspace` objects, which represent CDEs, etcd can become a significant performance bottleneck. Configure the {devworkspace} Operator pruner to automatically delete workspaces that have not been started within a configurable retention period. + +[IMPORTANT] +==== +Based on the load testing for 6,000 `DevWorkspace` objects, storage consumption for etcd was approximately 2.5GB. +==== + +.Prerequisites + +* An active `{orch-cli}` session with administrative permissions to the destination {orch-name} cluster. See {orch-cli-link}. + +* {devworkspace} Operator version 0.34.0 or later is installed on the cluster. + +.Procedure + +. Configure the `DevWorkspaceOperatorConfig` object to enable the cleanup cron job: ++ +[source,yaml,subs="+attributes"] +---- +apiVersion: controller.devfile.io/v1alpha1 +kind: DevWorkspaceOperatorConfig +metadata: + name: devworkspace-operator-config + namespace: {prod-namespace} +config: + workspace: + cleanupCronJob: + enabled: true + dryRun: false + retainTime: 2592000 + schedule: "0 0 1 * *" +---- ++ +retainTime:: The number of seconds a workspace must be inactive before it is marked for deletion. The default value `2592000` equals 30 days. ++ +schedule:: A cron expression that defines how often the pruner runs. The default value `"0 0 1 * *"` runs the pruner once per month. + +. Optional: To preview which workspaces would be deleted without removing them, set `dryRun` to `true`. + +.Verification + +* Verify that the cleanup cron job is scheduled: ++ +[source,bash,subs="+attributes"] +---- +$ {orch-cli} get cronjobs -n {prod-namespace} +---- + +[role="_additional-resources"] +.Additional resources + +* xref:plan:etcd-storage-limits.adoc[] diff --git a/modules/optimize/pages/disabling-copied-csvs.adoc b/modules/optimize/pages/disabling-copied-csvs.adoc new file mode 100644 index 0000000000..2deea7371a --- /dev/null +++ b/modules/optimize/pages/disabling-copied-csvs.adoc @@ -0,0 +1,49 @@ +:_content-type: PROCEDURE +:description: Disable Copied CSVs to reduce etcd storage usage and OLM memory consumption on large clusters. +:keywords: optimize, etcd, OLM, CSV, namespace +:navtitle: Disable Copied CSVs to reduce etcd usage + +[id="disabling-copied-csvs"] += Disable Copied CSVs to reduce etcd usage + +[role="_abstract"] +When an Operator is installed by the Operator Lifecycle Manager (OLM), a stripped-down copy of its ClusterServiceVersion (CSV) is created in every {namespace} the Operator watches. In clusters with many {namespace}s and cluster-wide Operators, Copied CSVs increase etcd storage usage, OLM memory consumption, and network bandwidth. Disabling Copied CSVs reduces the data stored in etcd and the memory footprint of OLM. + +.Prerequisites + +* An active `{orch-cli}` session with administrative permissions to the destination {orch-name} cluster. See {orch-cli-link}. + +.Procedure + +. Configure the `OLMConfig` object to disable Copied CSVs: ++ +[source,yaml] +---- +apiVersion: operators.coreos.com/v1 +kind: OLMConfig +metadata: + name: cluster +spec: + features: + disableCopiedCSVs: true +---- ++ +Disabling Copied CSVs reduces etcd storage usage and the memory footprint of OLM. + +.Verification + +* Verify that Copied CSVs are no longer created in user namespaces: ++ +[source,bash,subs="+attributes"] +---- +$ {orch-cli} get csv -n +---- ++ +The command should return no Copied CSVs for namespaces that do not have their own Operator subscriptions. + +[role="_additional-resources"] +.Additional resources + +* link:https://github.com/operator-framework/enhancements/blob/master/enhancements/olm-toggle-copied-csvs.md[OLM toggle Copied CSVs enhancement proposal] +* link:https://olm.operatorframework.io/docs/advanced-tasks/configuring-olm/#disabling-copied-csvs[Disabling Copied CSVs in OLM] +* xref:plan:etcd-storage-limits.adoc[] diff --git a/modules/optimize/pages/disabling-workspace-ca-bundle-mount.adoc b/modules/optimize/pages/disabling-workspace-ca-bundle-mount.adoc new file mode 100644 index 0000000000..6af46265d0 --- /dev/null +++ b/modules/optimize/pages/disabling-workspace-ca-bundle-mount.adoc @@ -0,0 +1,49 @@ +:_content-type: PROCEDURE +:description: Disable the full CA bundle mount to reduce etcd pressure from ca-certs-merged ConfigMaps in every user namespace. +:keywords: optimize, etcd, CA-bundle, ConfigMap, TLS +:navtitle: Disable the workspace CA bundle mount + +[id="disabling-workspace-ca-bundle-mount"] += Disable the workspace CA bundle mount + +[role="_abstract"] +By default, the Operator creates a `ca-certs-merged` ConfigMap in every user namespace. This ConfigMap contains the full Certificate Authorities (CAs) bundle. With a large number of Transport Layer Security (TLS) certificates in the cluster, these ConfigMaps consume significant etcd storage. Disable the full CA bundle mount to reduce etcd pressure. + +.Prerequisites + +* An active `{orch-cli}` session with administrative permissions to the destination {orch-name} cluster. See {orch-cli-link}. + +.Procedure + +. Configure the `CheCluster` Custom Resource to disable the workspace CA bundle mount: ++ +[source,bash,subs="+attributes"] +---- +$ {orch-cli} edit checluster/{prod-checluster} -n {prod-namespace} +---- ++ +[source,yaml] +---- +spec: + devEnvironments: + trustedCerts: + disableWorkspaceCaBundleMount: true +---- ++ +With this configuration, {prod-short} no longer mounts the full CA bundle under `/etc/pki/ca-trust/extracted/pem`. Only custom certificates are mounted under `/public-certs`. + +.Verification + +* Verify that new workspaces no longer mount the full CA bundle: ++ +[source,bash,subs="+attributes"] +---- +$ {orch-cli} get configmap ca-certs-merged -n +---- ++ +The ConfigMap should not exist in namespaces for newly created workspaces. + +[role="_additional-resources"] +.Additional resources + +* xref:plan:etcd-storage-limits.adoc[] diff --git a/modules/install/examples/proc_che-installing-the-chectl-management-tool-on-linux-or-macos.adoc b/modules/plan/examples/proc_che-installing-the-chectl-management-tool-on-linux-or-macos.adoc similarity index 100% rename from modules/install/examples/proc_che-installing-the-chectl-management-tool-on-linux-or-macos.adoc rename to modules/plan/examples/proc_che-installing-the-chectl-management-tool-on-linux-or-macos.adoc diff --git a/modules/install/examples/proc_che-installing-the-chectl-management-tool-on-windows.adoc b/modules/plan/examples/proc_che-installing-the-chectl-management-tool-on-windows.adoc similarity index 100% rename from modules/install/examples/proc_che-installing-the-chectl-management-tool-on-windows.adoc rename to modules/plan/examples/proc_che-installing-the-chectl-management-tool-on-windows.adoc diff --git a/modules/plan/examples/snip_che-example-devfile-disclaimer.adoc b/modules/plan/examples/snip_che-example-devfile-disclaimer.adoc new file mode 100644 index 0000000000..e69de29bb2 diff --git a/modules/plan/examples/snip_che-running-at-scale-calculate-resources.adoc b/modules/plan/examples/snip_che-running-at-scale-calculate-resources.adoc new file mode 100644 index 0000000000..5d59e06da3 --- /dev/null +++ b/modules/plan/examples/snip_che-running-at-scale-calculate-resources.adoc @@ -0,0 +1,6 @@ +:_content-type: SNIPPET + +[NOTE] +==== +For details about calculating resource requirements, see xref:calculating-che-resource-requirements.adoc[]. +==== diff --git a/modules/discover/examples/snip_che-supported-platforms.adoc b/modules/plan/examples/snip_che-supported-platforms.adoc similarity index 69% rename from modules/discover/examples/snip_che-supported-platforms.adoc rename to modules/plan/examples/snip_che-supported-platforms.adoc index 158f0a894b..eee61da56d 100644 --- a/modules/discover/examples/snip_che-supported-platforms.adoc +++ b/modules/plan/examples/snip_che-supported-platforms.adoc @@ -13,7 +13,10 @@ You can install {prod} on all major Public Clouds such as: * Microsoft Azure * Rancher -WARNING: Setting up link:https://kubernetes.io/docs/reference/access-authn-authz/authentication/[Users' Authentication] is required for deploying {prod-short} on {kubernetes} infrastructures. For {ocp} no additional setup is needed. +[WARNING] +==== +Setting up link:https://kubernetes.io/docs/reference/access-authn-authz/authentication/[Users' Authentication] is required for deploying {prod-short} on {kubernetes} infrastructures. For {ocp} no additional setup is needed. +==== The following options are available for the local installation: @@ -21,6 +24,7 @@ The following options are available for the local installation: * link:https://developers.redhat.com/products/openshift-local/overview[Red Hat OpenShift Local (formerly Red Hat CodeReady Containers)] +[role="_additional-resources"] .Additional resources * xref:install:con_installation-overview.adoc[] diff --git a/modules/install/images/running-at-scale/snip_che-multi-cluster.png b/modules/plan/images/running-at-scale/snip_che-multi-cluster.png similarity index 100% rename from modules/install/images/running-at-scale/snip_che-multi-cluster.png rename to modules/plan/images/running-at-scale/snip_che-multi-cluster.png diff --git a/modules/plan/nav.adoc b/modules/plan/nav.adoc new file mode 100644 index 0000000000..ac4b02474f --- /dev/null +++ b/modules/plan/nav.adoc @@ -0,0 +1,11 @@ +.Plan +* Plan deployment resources +** xref:supported-platforms.adoc[] +** xref:calculating-che-resource-requirements.adoc[] +** xref:cluster-capacity-limits.adoc[] +** xref:etcd-storage-limits.adoc[] +** xref:optimize:disabling-workspace-ca-bundle-mount.adoc[] +** xref:optimize:configuring-automatic-cleanup-of-inactive-workspaces.adoc[] +** xref:optimize:disabling-copied-csvs.adoc[] +** xref:multi-cluster-deployments.adoc[] +** xref:installing-the-chectl-management-tool.adoc[] diff --git a/modules/install/pages/calculating-che-resource-requirements.adoc b/modules/plan/pages/calculating-che-resource-requirements.adoc similarity index 93% rename from modules/install/pages/calculating-che-resource-requirements.adoc rename to modules/plan/pages/calculating-che-resource-requirements.adoc index 03f7b938cb..a9b37db9a2 100644 --- a/modules/install/pages/calculating-che-resource-requirements.adoc +++ b/modules/plan/pages/calculating-che-resource-requirements.adoc @@ -1,8 +1,8 @@ :_content-type: PROCEDURE :description: Size your cluster by calculating the CPU and memory requirements for the {prod-short} Operator, {devworkspace} Controller, and user workspaces so that your cluster can handle the expected number of concurrent users. -:keywords: install, calculating-che-resource-requirements +:keywords: plan, calculating-che-resource-requirements, sizing :navtitle: Size your cluster for {prod-short} -:page-aliases: administration-guide:calculating-che-resource-requirements.adoc, plan:calculating-che-resource-requirements.adoc +:page-aliases: administration-guide:calculating-che-resource-requirements.adoc, install:calculating-che-resource-requirements.adoc [id="calculating-{prod-id-short}-resource-requirements"] @@ -18,6 +18,12 @@ The pods contribute to the resource consumption in CPU and memory limits and req include::example$snip_{project-context}-example-devfile-disclaimer.adoc[] +.Prerequisites + +* You have a planned or existing {prod-short} deployment on {orch-name}. +* You have the devfiles that define the development environments for your users. +* You have an estimate of the number of concurrent workspaces that your users will run. + .Procedure . Identify the workspace resource requirements from the devfile `components` section. The following example uses the link:https://github.com/che-incubator/quarkus-api-example/blob/main/devfile.yaml[Quarkus API example devfile]. diff --git a/modules/plan/pages/cluster-capacity-limits.adoc b/modules/plan/pages/cluster-capacity-limits.adoc new file mode 100644 index 0000000000..e0bdf30117 --- /dev/null +++ b/modules/plan/pages/cluster-capacity-limits.adoc @@ -0,0 +1,58 @@ +:_content-type: CONCEPT +:description: Evaluate the tested cluster maximums, workspace resource requirements, and worker node capacity that constrain how many cloud development environments your cluster can support. +:keywords: plan, scale, cluster, maximums, capacity, workspace, worker-node +:navtitle: Tested maximums that constrain workspace scaling +:page-aliases: running-at-scale.adoc + +[id="cluster-capacity-limits"] += Tested maximums that constrain workspace scaling + +[role="_abstract"] +Evaluate the tested cluster maximums, workspace resource requirements, and worker node capacity that constrain how many cloud development environments your {orch-name} cluster can support. + +CDE workloads are particularly complex to scale. The underlying IDE solutions, such as Visual Studio Code - Open Source ("Code - OSS") or JetBrains Gateway, are designed as single-user applications, not as multitenant services. + +== Tested cluster maximums that constrain scaling + +While there is no strict limit on the number of resources in a {kubernetes} cluster, there are certain considerations for large clusters to remember. + +{ocp}, a certified distribution of {kubernetes}, provides a set of tested maximums for various resources. These maximums can serve as an initial guideline for planning your environment: + +.{ocp} tested cluster maximums +[%header,format=csv] +|=== +Resource type, Tested maximum +Number of nodes,2000 +Number of pods,150000 +Number of pods per node,2500 +Number of namespace,10000 +Number of services,10000 +Number of secrets,80000 +Number of config maps,90000 +|=== + +For example, it is generally not recommended to have more than 10,000 namespaces due to potential performance and management cost. In {prod}, each user is allocated a namespace. If you expect the user base to be large, consider spreading workloads across multiple "fit-for-purpose" clusters and potentially using solutions for multi-cluster orchestration. + +== Workspace size that determines cluster capacity + +When deploying {prod} on {kubernetes}, accurately calculate the resource requirements for each CDE, including memory and CPU or GPU needs. This determines the right sizing of the cluster. In general, the CDE size is limited by and cannot be bigger than the worker node size. + +The resource requirements for CDEs can vary significantly based on the specific workloads and configurations. A simple CDE might require only a few hundred megabytes of memory. A more complex one might need several gigabytes of memory and multiple CPU cores. + +For details about calculating resource requirements, see Additional resources. + +== Worker node capacity that matches workspace demand + +Although cluster autoscaling is a powerful {kubernetes} feature, you cannot always rely on it. Consider predictive scaling by analyzing load data to detect daily or weekly usage patterns. + +If your workloads follow a pattern with dramatic peaks throughout the day, provision worker nodes accordingly. For example, if workspaces increase during business hours and decrease during off-hours, predictive scaling adjusts the number of worker nodes. This ensures enough resources are available during peak load while minimizing costs during off-peak hours. + +You can also use open source solutions such as Karpenter for configuration and lifecycle management of the worker nodes. Karpenter can dynamically provision and optimize worker nodes based on the specific requirements of the workloads. This helps improve resource utilization and reduce costs. + +[role="_additional-resources"] +.Additional resources + +* xref:calculating-che-resource-requirements.adoc[] +* link:https://docs.redhat.com/en/documentation/openshift_container_platform/{ocp4-ver}/html/scalability_and_performance/planning-your-environment-according-to-object-maximums#planning-your-environment-according-to-object-maximums[{ocp} tested object maximums] +* link:https://kubernetes.io/docs/setup/best-practices/cluster-large/[Considerations for large clusters] +* link:https://karpenter.sh/[Karpenter] diff --git a/modules/plan/pages/etcd-storage-limits.adoc b/modules/plan/pages/etcd-storage-limits.adoc new file mode 100644 index 0000000000..2c514d7724 --- /dev/null +++ b/modules/plan/pages/etcd-storage-limits.adoc @@ -0,0 +1,26 @@ +:_content-type: CONCEPT +:description: Understand how {devworkspace} objects, CA certificate bundles, and Copied CSVs contribute to etcd growth so that you can take action before your cluster becomes unstable. +:keywords: plan, etcd, storage, DevWorkspace, scale, bottleneck +:navtitle: etcd pressure that limits cluster scale + +[id="etcd-storage-limits"] += etcd pressure that limits cluster scale + +[role="_abstract"] +Scaling {prod-short} to thousands of concurrent workspaces strains etcd, the primary datastore for {kubernetes} cluster state. Three categories of objects contribute most to etcd growth: CA certificate bundles, {devworkspace} custom resources, and Copied CSVs. The procedures that follow address each one. + +The primary datastore of {kubernetes} cluster configuration and state is etcd. It holds information about nodes, pods, services, and custom resources. + +As a distributed key-value store, etcd does not scale well past a certain threshold. As the size of etcd grows, so does the load on the cluster, risking its stability. + +[IMPORTANT] +==== +The default etcd size is 2 GB, and the recommended maximum is 8 GB. Exceeding the maximum limit can make the {kubernetes} cluster unstable and unresponsive. Even though the data stored in a `ConfigMap` cannot exceed 1 MiB by design, a few thousand relatively large `ConfigMap` objects can overload etcd storage. +==== + +The size of the objects stored in etcd is also a critical factor. Each object consumes space, and as the number of objects increases, the overall size of etcd grows. The larger the object, the more space it takes. For example, etcd can be overloaded with only a few thousand large {kubernetes} objects. + +[role="_additional-resources"] +.Additional resources + +* link:https://etcd.io/[etcd] diff --git a/modules/install/pages/installing-the-chectl-management-tool.adoc b/modules/plan/pages/installing-the-chectl-management-tool.adoc similarity index 76% rename from modules/install/pages/installing-the-chectl-management-tool.adoc rename to modules/plan/pages/installing-the-chectl-management-tool.adoc index 167c3de3f0..cecd50d63f 100644 --- a/modules/install/pages/installing-the-chectl-management-tool.adoc +++ b/modules/plan/pages/installing-the-chectl-management-tool.adoc @@ -1,8 +1,8 @@ :_content-type: ASSEMBLY :description: Set up {prod-cli} on Linux, macOS, or Windows so that you can deploy, update, and manage {prod-short} from the command line. -:keywords: install, installing-the-chectl-management-tool +:keywords: plan, installing-the-chectl-management-tool, cli :navtitle: Set up the {prod-cli} command-line tool -:page-aliases: administration-guide:installing-the-chectl-management-tool.adoc, plan:installing-the-chectl-management-tool.adoc, installation-guide:using-the-chectl-management-tool.adoc, overview:using-the-chectl-management-tool.adoc, using-the-chectl-management-tool.adoc +:page-aliases: administration-guide:installing-the-chectl-management-tool.adoc, install:installing-the-chectl-management-tool.adoc, installation-guide:using-the-chectl-management-tool.adoc, overview:using-the-chectl-management-tool.adoc, using-the-chectl-management-tool.adoc [id="installing-the-{prod-cli}-management-tool"] diff --git a/modules/plan/pages/multi-cluster-deployments.adoc b/modules/plan/pages/multi-cluster-deployments.adoc new file mode 100644 index 0000000000..ba1f110ebe --- /dev/null +++ b/modules/plan/pages/multi-cluster-deployments.adoc @@ -0,0 +1,66 @@ +:_content-type: CONCEPT +:description: Evaluate multi-cluster architectures for deployments that exceed single-cluster capacity, and review the Developer Sandbox reference implementation. +:keywords: plan, multi-cluster, scale, redirector, Developer-Sandbox +:navtitle: Multi-cluster architectures for large deployments + +[id="multi-cluster-deployments"] += Multi-cluster architectures for large deployments + +[role="_abstract"] +Evaluate multi-cluster architectures for {prod-short} deployments that exceed single-cluster capacity, and review the Developer Sandbox reference implementation that routes developers across clusters automatically. + +== Workloads distributed across multiple clusters + +By design, {prod} is not multi-cluster aware. You can only have one instance per cluster. + +However, you can run {prod} in a multi-cluster environment by deploying {prod} in each cluster. Use a load balancer or Domain Name System (DNS)-based routing to direct traffic to the appropriate instance. This approach distributes the workload across clusters and provides redundancy in case of cluster failures. + +== Developer Sandbox that runs Dev Spaces across clusters + +You can test running {prod-short} in a multi-cluster environment by using the Developer Sandbox, a free trial environment by Red Hat. + +From an infrastructure perspective, the Developer Sandbox consists of multiple Red Hat OpenShift Service on AWS (ROSA) clusters. On each cluster, the productized version of {prod} is installed and configured using Argo CD. The workspaces.openshift.com URL is used as a single entry point to the {prod} instances across clusters. + +.Developer Sandbox multi-cluster architecture +image::running-at-scale/snip_{project-context}-multi-cluster.png[Scheme of a multi-cluster environment] + +[IMPORTANT] +==== +The multi-cluster architecture of workspaces.openshift.com is part of the Developer Sandbox. It is a Developer Sandbox-specific solution that cannot be reused as-is in other environments. However, you can use it as a reference for implementing a similar solution well-tailored to your specific multicluster needs. +==== + +== Redirector that routes developers to the correct cluster + +Red Hat offers an open source, Quarkus-based service that acts as a single gateway for developers. This service automatically redirects users to the correct {prod} instance on the appropriate cluster based on their {ocp} group membership. For the community-supported version, see Additional resources. + +=== Prerequisites for the multicluster redirector + +A critical requirement for the multicluster redirector is that all users are provisioned to the host cluster where the redirector is deployed. Users authenticate through the OAuth flow of this cluster, even if they never run workloads there. The host cluster's {ocp} groups determine the routing logic. For deployment instructions, see Additional resources. + +=== OpenShift groups that map to cluster URLs + +The routing configuration uses a `ConfigMap` that contains JSON to map {ocp} groups to {prod} URLs. The redirector uses this file to update routing tables in real-time without requiring restarts. + +=== Authentication and routing steps + +The routing process follows these steps: + +. Authenticate by using OAuth through a proxy sidecar. +. Pass identity and group information through HTTP headers. +. Verify group memberships by using {ocp} API queries. +. Determine the appropriate {prod} URL by using a mapping lookup. +. Redirect the user to the designated cluster instance. + +If users belong to multiple {ocp} groups, they can choose the {prod} instance they need from a selection dashboard. + +[role="_additional-resources"] +.Additional resources + +* link:https://che.eclipseprojects.io/2025/04/29/@ilya.buziuk-running-at-scale.html[Running at scale] +* link:https://developers.redhat.com/articles/2026/01/23/enterprise-multi-cluster-scalability-openshift-dev-spaces[Enterprise multi-cluster scalability] +* link:https://developers.redhat.com/developer-sandbox[Developer Sandbox] +* link:https://www.redhat.com/en/technologies/cloud-computing/openshift/aws[Red Hat OpenShift Service on AWS (ROSA)] +* link:https://argo-cd.readthedocs.io/en/stable/[Argo CD] +* link:https://workspaces.openshift.com/[workspaces.openshift.com] +* link:https://github.com/codeready-toolchain/crw-multicluster-redirector[crw-multicluster-redirector GitHub repository] +* link:https://github.com/redhat-developer/devspaces-multicluster-redirector[devspaces-multicluster-redirector GitHub repository] diff --git a/modules/discover/pages/supported-platforms.adoc b/modules/plan/pages/supported-platforms.adoc similarity index 69% rename from modules/discover/pages/supported-platforms.adoc rename to modules/plan/pages/supported-platforms.adoc index 4a1baa35ce..023f9d7bb2 100644 --- a/modules/discover/pages/supported-platforms.adoc +++ b/modules/plan/pages/supported-platforms.adoc @@ -1,11 +1,11 @@ :_content-type: REFERENCE :description: {prod-short} runs on a specific range of {orch-name} versions and CPU architectures. -:keywords: platforms, compatibility, supported, discover -:navtitle: Supported platforms -:page-aliases: administration-guide:supported-platforms.adoc, installation-guide:supported-platforms.adoc +:keywords: platforms, compatibility, supported, plan +:navtitle: Supported platforms and architectures +:page-aliases: administration-guide:supported-platforms.adoc, installation-guide:supported-platforms.adoc, discover:supported-platforms.adoc [id="supported-platforms"] -= Supported platforms += Supported platforms and architectures [role="_abstract"] {prod-short} runs on a specific range of {orch-name} versions and CPU architectures. Confirm that your target cluster matches a supported combination before starting the installation. diff --git a/modules/upgrade/examples/snip_che-upgrading-the-chectl-management-tool.adoc b/modules/upgrade/examples/snip_che-upgrading-the-chectl-management-tool.adoc index 68b0e5992b..7f3ad72386 100644 --- a/modules/upgrade/examples/snip_che-upgrading-the-chectl-management-tool.adoc +++ b/modules/upgrade/examples/snip_che-upgrading-the-chectl-management-tool.adoc @@ -4,7 +4,7 @@ This section describes how to upgrade the `{prod-cli}` management tool. .Prerequisites -* `{prod-cli}`. See: xref:install:installing-the-chectl-management-tool.adoc[]. +* `{prod-cli}`. See: xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure diff --git a/modules/upgrade/pages/upgrading-che-using-the-cli-management-tool.adoc b/modules/upgrade/pages/upgrading-che-using-the-cli-management-tool.adoc index 265b353251..5ee932dec1 100644 --- a/modules/upgrade/pages/upgrading-che-using-the-cli-management-tool.adoc +++ b/modules/upgrade/pages/upgrading-che-using-the-cli-management-tool.adoc @@ -17,7 +17,7 @@ include::partial$snip_persona-admin.adoc[] * A running instance of a previous minor version of {prod-prev-short}, installed using the CLI management tool on the same instance of {platforms-name}, in the `{prod-namespace}` {platforms-namespace}. -* `{prod-cli}` for {prod-short} version {prod-ver}. See: xref:install:installing-the-chectl-management-tool.adoc[]. +* `{prod-cli}` for {prod-short} version {prod-ver}. See: xref:plan:installing-the-chectl-management-tool.adoc[]. .Procedure From 4cd0bdf0b66fbb09429e05f34c739b482e5813e5 Mon Sep 17 00:00:00 2001 From: Gaurav Trivedi Date: Thu, 17 Sep 2026 12:00:17 +0530 Subject: [PATCH 2/2] fix: resolve rebase conflict and clean up Plan module references Rebased jtbd-plan-category onto main to pick up 18 commits merged since this PR was last updated (GitOps install procedures, Install JTBD audit, Troubleshoot/Discover/Optimize promotions, dashboard Swagger route fix). Conflict resolution: - modules/install/pages/running-at-scale.adoc: modify/delete conflict. GitOps (#3141) added one line (include::partial$snip_persona-admin. adoc[]) to the old monolithic file after this branch had already split it into cluster-capacity-limits.adoc, etcd-storage-limits.adoc, and multi-cluster-deployments.adoc under modules/plan/. That single added line is itself a Parameter 4 violation (generic persona snippet include, forbidden by the JTBD checklist), so nothing needed to be ported forward. Resolved by keeping the deletion. Post-rebase fixes found while re-validating: - Removed 2 pre-existing include::partial$snip_persona-admin.adoc[] violations carried over from Install into calculating-che-resource-requirements.adoc and installing-the-chectl-management-tool.adoc (Parameter 4 FAIL) - Fixed 2 broken xrefs in already-merged content that the rebase pulled in, which still pointed at the pre-move paths for files this PR relocates to modules/plan/: - modules/troubleshoot/pages/troubleshooting-workspace-startup-failures.adoc: xref:install:calculating-che-resource-requirements.adoc[] -> xref:plan:calculating-che-resource-requirements.adoc[] - modules/discover/pages/what-is-che.adoc: xref:supported-platforms.adoc[] -> xref:plan:supported-platforms.adoc[] Verified: zero remaining snip_persona includes in modules/plan/, zero remaining bare/legacy xrefs anywhere in the tree pointing at the files this PR moves. Pre-existing Vale errors in troubleshooting-workspace-startup-failures.adoc (lines 44/46/55/69, CheDocs.Attributes) confirmed present on main before this rebase -- out of scope for this fix. Co-authored-by: Cursor --- .vale.ini | 6 ++++++ modules/discover/pages/what-is-che.adoc | 2 +- ...stalling-che-on-the-virtual-kubernetes-cluster.adoc | 4 ++-- ...uring-automatic-cleanup-of-inactive-workspaces.adoc | 4 ++-- .../pages/calculating-che-resource-requirements.adoc | 2 -- .../pages/installing-the-chectl-management-tool.adoc | 2 -- .../troubleshooting-workspace-startup-failures.adoc | 10 +++++----- 7 files changed, 16 insertions(+), 14 deletions(-) diff --git a/.vale.ini b/.vale.ini index 9ca66e0756..a25115ba7d 100644 --- a/.vale.ini +++ b/.vale.ini @@ -58,3 +58,9 @@ RedHat.ConfigMap = NO RedHat.Slash = NO RedHat.Spacing = NO RedHat.Spelling = NO + +# The DevWorkspaceOperatorConfig custom resource and its example YAML manifest +# must use the literal Kind and metadata name from the DevWorkspace Operator API, +# which legitimately contain the substring "devworkspace". +[modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc] +CheDocs.Attributes = NO diff --git a/modules/discover/pages/what-is-che.adoc b/modules/discover/pages/what-is-che.adoc index 223bf5d3d2..6a5f26df2a 100644 --- a/modules/discover/pages/what-is-che.adoc +++ b/modules/discover/pages/what-is-che.adoc @@ -72,4 +72,4 @@ See the development link:https://github.com/eclipse/che/wiki/Roadmap[roadmap] on .Additional resources * xref:architecture-overview.adoc[] -* xref:supported-platforms.adoc[] +* xref:plan:supported-platforms.adoc[] diff --git a/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc b/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc index d139cb8868..c6cbd705f1 100644 --- a/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc +++ b/modules/install/pages/proc_installing-che-on-the-virtual-kubernetes-cluster.adoc @@ -41,7 +41,7 @@ Check your {kubernetes} provider documentation on how to install it. [TIP] ==== Use the following command to install link:https://docs.nginx.com/nginx-ingress-controller/[NGINX Ingress Controller] -on Azure Kubernetes Service cluster: +on Azure {kubernetes} Service cluster: [source,bash,subs="attributes+"] ---- helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx @@ -566,7 +566,7 @@ get-token,\ [TIP] ==== Use the following command to install link:https://docs.nginx.com/nginx-ingress-controller/[NGINX Ingress Controller] -on Azure Kubernetes Service cluster: +on Azure {kubernetes} Service cluster: [source,bash,subs="attributes+"] ---- helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx diff --git a/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc b/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc index 5f02908590..3f6ed16866 100644 --- a/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc +++ b/modules/optimize/pages/configuring-automatic-cleanup-of-inactive-workspaces.adoc @@ -7,7 +7,7 @@ = Configure automatic cleanup of inactive workspaces [role="_abstract"] -For large {kubernetes} deployments, particularly those involving a high number of custom resources such as `DevWorkspace` objects, which represent CDEs, etcd can become a significant performance bottleneck. Configure the {devworkspace} Operator pruner to automatically delete workspaces that have not been started within a configurable retention period. +For large {kubernetes} deployments, particularly those with a high number of Cloud Development Environment (CDE) custom resources, etcd can become a significant performance bottleneck. Configure the pruner to automatically delete workspaces that have not been started within a configurable retention period. [IMPORTANT] ==== @@ -18,7 +18,7 @@ Based on the load testing for 6,000 `DevWorkspace` objects, storage consumption * An active `{orch-cli}` session with administrative permissions to the destination {orch-name} cluster. See {orch-cli-link}. -* {devworkspace} Operator version 0.34.0 or later is installed on the cluster. +* The `DevWorkspace` Operator, version 0.34.0 or later, is installed on the cluster. .Procedure diff --git a/modules/plan/pages/calculating-che-resource-requirements.adoc b/modules/plan/pages/calculating-che-resource-requirements.adoc index a9b37db9a2..36b6aa980c 100644 --- a/modules/plan/pages/calculating-che-resource-requirements.adoc +++ b/modules/plan/pages/calculating-che-resource-requirements.adoc @@ -11,8 +11,6 @@ [role="_abstract"] Size your cluster by calculating the CPU and memory requirements for the {prod-short} Operator, {devworkspace} Controller, and user workspaces so that your cluster can handle the expected number of concurrent users. -include::partial$snip_persona-admin.adoc[] - The {prod-short} Operator, {devworkspace} Controller, and user workspaces consist of a set of pods. The pods contribute to the resource consumption in CPU and memory limits and requests. diff --git a/modules/plan/pages/installing-the-chectl-management-tool.adoc b/modules/plan/pages/installing-the-chectl-management-tool.adoc index cecd50d63f..79d44a9aa6 100644 --- a/modules/plan/pages/installing-the-chectl-management-tool.adoc +++ b/modules/plan/pages/installing-the-chectl-management-tool.adoc @@ -11,8 +11,6 @@ [role="_abstract"] Set up `{prod-cli}` on Linux, macOS, or Windows so that you can deploy, update, and manage {prod-short} from the command line. -include::partial$snip_persona-admin.adoc[] - include::example$proc_{project-context}-installing-the-chectl-management-tool-on-windows.adoc[leveloffset=+1] include::example$proc_{project-context}-installing-the-chectl-management-tool-on-linux-or-macos.adoc[leveloffset=+1] diff --git a/modules/troubleshoot/pages/troubleshooting-workspace-startup-failures.adoc b/modules/troubleshoot/pages/troubleshooting-workspace-startup-failures.adoc index beb4fe10cf..a3bff05ad2 100644 --- a/modules/troubleshoot/pages/troubleshooting-workspace-startup-failures.adoc +++ b/modules/troubleshoot/pages/troubleshooting-workspace-startup-failures.adoc @@ -41,9 +41,9 @@ Diagnose and resolve common workspace startup failures based on error symptoms a | The container runtime does not trust the TLS certificate of the container registry. Import the registry Certificate Authority (CA) certificate into {prod-short}. |=== -== DevWorkspace errors +== {devworkspace} errors -.DevWorkspace error messages and resolutions +.{devworkspace} error messages and resolutions [cols="1,2",options="header"] |=== | Error message | Resolution @@ -52,7 +52,7 @@ Diagnose and resolve common workspace startup failures based on error symptoms a | The workspace did not reach the `Running` phase within the configured timeout. Increase `startTimeoutSeconds` in the `CheCluster` Custom Resource or investigate Pod events for resource or scheduling issues. | `Failed to create DevWorkspace: admission webhook denied the request` -| The {devworkspace} Operator webhook rejected the DevWorkspace. Verify that the {devworkspace} Operator is running and that CRDs are up to date. +| The {devworkspace} Operator webhook rejected the {devworkspace}. Verify that the {devworkspace} Operator is running and that CRDs are up to date. | `BadRequest` or `InfrastructureFailure` | An infrastructure-level error prevented workspace creation. Check the {devworkspace} Operator logs for details. @@ -66,7 +66,7 @@ Diagnose and resolve common workspace startup failures based on error symptoms a | Error message | Resolution | `exceeded quota` or `forbidden: exceeded quota` -| The user namespace has a ResourceQuota that prevents creating the workspace Pod or PVC. Increase the quota or reduce the workspace resource requests in the devfile. +| The user {namespace} has a ResourceQuota that prevents creating the workspace Pod or PVC. Increase the quota or reduce the workspace resource requests in the devfile. | `OOMKilled` | The workspace container exceeded its memory limit and was terminated. Increase the memory limit in the devfile `components` section or in the `CheCluster` Custom Resource defaults. @@ -75,4 +75,4 @@ Diagnose and resolve common workspace startup failures based on error symptoms a .Additional resources * xref:optimize:configuring-machine-autoscaling.adoc[] -* xref:install:calculating-che-resource-requirements.adoc[] +* xref:plan:calculating-che-resource-requirements.adoc[]