From 89f05358c62ffd5f55ecd0910a96b4fa33e141a7 Mon Sep 17 00:00:00 2001 From: Zita Dombi Date: Mon, 31 Aug 2026 15:11:24 +0200 Subject: [PATCH 1/3] HDDS-14777. Add documentation for admins to do ZDU to the website --- docs/01-overview.md | 2 +- .../03-operations/02-upgrade-and-downgrade.md | 213 ++++++++++++++---- .../03-feature-branches/02-merge-checklist.md | 2 +- 3 files changed, 172 insertions(+), 45 deletions(-) diff --git a/docs/01-overview.md b/docs/01-overview.md index 6ec52a9879..bc404c76ea 100644 --- a/docs/01-overview.md +++ b/docs/01-overview.md @@ -96,7 +96,7 @@ Aspects related to storage administration include: - **Unified Storage:** Can potentially serve as a common storage layer for different types of workloads. - **Management Tools:** Includes the Recon web UI for monitoring and CLI tools for administration. -- **Maintenance:** Supports features like rolling upgrades, node decommissioning, and data balancing. +- **Maintenance:** Supports features like [rolling upgrades](./administrator-guide/operations/upgrade-and-downgrade), node decommissioning, and data balancing. ### Hybrid Cloud Scenarios diff --git a/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md b/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md index 2f57cd38a8..62423be7ce 100644 --- a/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md +++ b/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md @@ -4,79 +4,206 @@ sidebar_label: Upgrade and Downgrade # Upgrade and Downgrade -Ozone supports non-rolling upgrades and downgrades, where all components are stopped first, and then restarted with the upgraded or downgraded versions. +Ozone supports two ways to move a cluster between software versions: -## Upgrade States +- **Rolling (zero downtime) upgrade (ZDU)**: components are restarted with the new + version one service at a time, in a fixed order, so the cluster remains fully + operational throughout the upgrade. This is the recommended approach for + clusters that are already running a ZDU-capable release. +- **Non-rolling upgrade**: all components are stopped first, and then restarted + with the upgraded (or downgraded) versions. This is simpler, but the cluster is + unavailable while the components are stopped. + +Both approaches share the same [upgrade states](#upgrade-states) and the same +[finalization](#finalization) commands. -After upgrading components, the upgrade process is divided into two states: +## Upgrade States -1. **Pre-finalized**: When the current components are stopped and the new versions are started, they will see that the data on disk was written by a previous version of Ozone and enter a pre-finalized state. In the pre-finalized state: - - The cluster can be downgraded at any time by stopping all components and restarting with the old versions. - - Backwards incompatible features introduced in the new version will be disallowed by the cluster. - - The cluster will remain fully operational with all functionality present in the old version still allowed. +After starting components with a newer version, the upgrade process is divided +into two states: + +1. **Pre-finalized**: When components are started with a new version, they see + that the data on disk was written by a previous version of Ozone and enter a + pre-finalized state. Internally, each component's *apparent version* (the + version persisted on disk, which determines the API it exposes and the format + it writes) is lower than its *software version* (the version of the running + bits). In the pre-finalized state: + - The cluster can be downgraded at any time by restarting components with the + old versions. + - Backwards incompatible features introduced in the new version will be + disallowed by the cluster. + - The cluster will remain fully operational with all functionality present in + the old version still allowed. - Any data created while pre-finalized will remain readable after downgrade. -2. **Finalized**: When a finalize command is given to OM or SCM, they will enter a finalized state. In the finalized state: +2. **Finalized**: When a finalize command is given to the cluster, the components + move their apparent version up to match their software version and enter a + finalized state. In the finalized state: - The cluster can no longer be downgraded. - All new features of the cluster introduced in the new version can be used. -### Querying finalization status +### Querying upgrade status -**OM**: `ozone admin om finalizationstatus`. If using OM HA, finalization status is checked for the quorum, not individual OMs. +Use a single command to query the finalization status of the whole cluster: -**SCM**: `ozone admin scm finalizationstatus`. SCM will report that finalization is complete once it has finalized and is aware of enough finalized Datanodes to form a write pipeline. The remaining Datanodes will finalize asynchronously and be incorporated into write pipelines after informing SCM that they have finalized. +```bash +ozone admin upgrade status +``` -**Datanodes**: `ozone admin datanode list` will list all Datanodes and their health state as seen by SCM. If SCM is finalized, then Datanodes whose health state is `HEALTHY` have informed SCM that they have finalized. Datanodes whose health state is `HEALTHY_READONLY` have not yet informed SCM that they have finished finalization. `HEALTHY_READONLY` (pre-finalized) Datanodes remain readable, so the cluster is operational even if some otherwise healthy Datanodes have not yet finalized. `STALE` or `DEAD` Datanodes will be told to finalize by SCM once they are reachable again. +This command is sent to OM, which gathers the HDDS (SCM and Datanode) status from +SCM and reports the combined result, for example: -## Steps to upgrade and downgrade +```text +Upgrade finalization status: + Cluster: FINALIZED + OM: FINALIZED + SCM: FINALIZED + Datanodes finalized: 3/3 +``` -Starting with your current version of Ozone, complete the following steps to upgrade to a newer version of Ozone. +- Add `-v` / `--verbose` to also show the apparent versions of OM, SCM, and the + Datanodes. +- Add `--json` to format the output as JSON. +- In OM HA, pass `--service-id=`; in a non-HA deployment, pass + `--service-host=`. -1. If using OM HA and currently running Ozone 1.2.0 or greater, prepare the Ozone Manager. If OM HA is not being used, this step can be skipped. +`ozone admin datanode list` lists all Datanodes and their health state (`HEALTHY`, `STALE`, or `DEAD`) as seen by SCM. Pre-finalized Datanodes are not placed in a separate read-only state: during a rolling upgrade, every reachable Datanode stays `HEALTHY` and fully participates in reads and writes, while SCM ensures a consistent write version is used across mixed-version Datanodes. To track how far Datanode finalization has progressed, use `ozone admin upgrade status` (the `Datanodes finalized: N/M` line, or `-v` for per-Datanode apparent versions) rather than the Datanode health state. `STALE` or `DEAD` Datanodes will be told to finalize by SCM once they are reachable again. - ```bash - ozone admin om prepare -id= - ``` +## Rolling upgrade (ZDU) - The `prepare` command will block the Ozone Managers from receiving all write requests. See [Ozone Manager Prepare For Upgrade](https://ozone.apache.org/docs/edge/design/omprepare.html) for more information +A rolling upgrade keeps the cluster available by restarting components with the +new version one at a time, relying on Ozone's existing fault tolerance so that +service continues while individual nodes restart. -2. Stop all components. +### Prerequisites -3. Replace artifacts of all components with the newer versions. +A rolling upgrade requires an HA deployment. Because components are restarted one +at a time while the rest continue to serve traffic, both OM and SCM must be +running in HA, so a quorum stays available throughout. A non-HA +deployment (a single OM or SCM) cannot be upgraded without downtime and must use +the [non-rolling upgrade](#non-rolling-upgrade) procedure instead. -4. Start the components - 1. Start the SCM and Datanodes as usual: +The cluster must also already be running a ZDU-capable Ozone release. To move from +an earlier release to the first ZDU-capable release, use the +[non-rolling upgrade](#non-rolling-upgrade) procedure. Once the cluster is on a +ZDU-capable release, all subsequent upgrades can be done as a rolling upgrade or as +a non-rolling upgrade. - ```bash - ozone --daemon start scm - ozone --daemon start datanode - ``` +### Component upgrade order - 2. Start the Ozone Manager using the `--upgrade` flag to take it out of prepare mode. +Components must be upgraded in the following order: - ```bash - ozone --daemon start om --upgrade - ``` +1. All SCMs +2. Recon +3. All Datanodes +4. All OMs +5. Client-facing processes: S3 Gateway, HttpFS, etc. - - There also exists a `--downgrade` flag which is an alias of `--upgrade`. The name used does not matter. - - **IMPORTANT**: All OMs must be started with the `--upgrade` or `--downgrade` flag in this step. If some of the OMs are not started with this flag by mistake, run `ozone admin om cancelprepare -id=` to make sure all OMs leave prepare mode. +This order ensures that for every internal client/server relationship inside +Ozone, the server is always running the same or a newer version than its clients +and is finalized first, so servers only need to remain backwards compatible. +Recon is upgraded together with SCM because it shares SCM's Datanode report +processing code. - At this point, the cluster is upgraded to a pre-finalized state and fully operational. The cluster can be downgraded by repeating the above steps, but restoring the older versions of components in step 3, instead of the newer versions. To finalize the cluster to use new features, continue on with the following steps. +### Steps - **Once the following steps are performed, downgrading will not be possible.** +1. Deploy the new software to all SCMs and restart them one at a time. +2. Deploy the new software to Recon and restart it. +3. Deploy the new software to the Datanodes and restart them. Restarting + Datanodes one at a time keeps data continuously available but takes the + longest. If temporary unavailability is acceptable, or data placement makes it + safe (for example, EC or rack-scatter placement), Datanodes can be restarted in + larger groups such as one rack at a time. +4. Deploy the new software to all OMs and restart them one at a time. Unlike + earlier versions, there is no "prepare for upgrade" step and no `--upgrade` + flag — OMs are started normally. +5. Deploy the new software and restart the client-facing processes (S3 Gateway, + HttpFS, etc.). -5. Finalize SCM +Throughout the upgrade, keep a quorum available: with three OMs or SCMs, at least +two must be up at all times. If a node fails to start on the new version, resolve +the issue or downgrade that node before restarting others. - ```bash - ozone admin scm finalizeupgrade - ``` +At this point the cluster is running the new software in a pre-finalized state and +is fully operational, but it is still "acting as" the old version — no data is +written in a new format and new features are unavailable. Regression testing can +be performed here, and a [downgrade](#downgrade) is still possible. + +To finalize the cluster and enable the new features, continue with +[Finalization](#finalization). + +## Non-rolling upgrade - At this point, SCM will tell all of the Datanodes to finalize. Once SCM has finalized enough Datanodes to form a write pipeline, it will return that finalization was successful. The remaining pre-finalized Datanodes will be in a read-only state until they indicate to SCM that they have finalized. Write requests will be directed to finalized Datanodes only. +In a non-rolling upgrade, all components are stopped and restarted together. The +cluster is unavailable while the components are stopped, but the procedure is +simpler than a rolling upgrade. -6. Finalize OM +1. Stop all components. + +2. Replace the artifacts of all components with the newer versions. + +3. Start the components: ```bash - ozone admin om finalizeupgrade -id= + ozone --daemon start scm + ozone --daemon start datanode + ozone --daemon start om ``` -At this point, the cluster is finalized and the upgrade is complete. + There is no longer a "prepare for upgrade" step, and the `--upgrade` / + `--downgrade` start flags are no longer required (they are deprecated no-ops). + Components are started normally. + +At this point, the cluster is upgraded to a pre-finalized state and fully +operational. It can still be [downgraded](#downgrade). To finalize the cluster +and use the new features, continue with [Finalization](#finalization). + +## Downgrade + +Before the cluster is finalized, it can be downgraded by restarting the +components with the old software. No data in a new format has been persisted, so +the old software will be able to read everything. The restart with the +downgraded software can be done either non-rolling (stop all components, restore +the old artifacts, and start them again) or rolling (restart the components in +the **reverse** of the upgrade order: client-facing processes, then OMs, then +Datanodes, then Recon, then SCMs). + +**Once the cluster is [finalized](#finalization), downgrading is not possible.** + +## Finalization + +Finalization is the same for both rolling and non-rolling upgrades. A single +command finalizes the whole cluster: + +```bash +ozone admin upgrade finalize +``` + +- The command is sent to OM, which orchestrates finalization across the cluster: + SCM finalizes first, then the Datanodes, and finally the OMs. Because + finalization is asynchronous, the command returns as soon as the process has + been started. +- Add `--wait` to have the command poll and block until the entire cluster is + finalized (interruptible with Ctrl-C). +- In OM HA, pass `--service-id=`; in a non-HA deployment, pass + `--service-host=`. + +Monitor progress with [`ozone admin upgrade status`](#querying-upgrade-status). +The command is idempotent, so finalization continues after OM restarts or leader +changes and can be safely re-issued. + +**After finalization, the cluster cannot be downgraded.** If a component sees a +version on disk that is newer than its own software version, it will refuse to +start. + +### Finalizing a cluster upgraded from a pre-ZDU release + +`ozone admin upgrade finalize` and `ozone admin upgrade status` require an OM that +supports ZDU. If they are run against a cluster whose OM predates ZDU support, they +will report that the OM does not support zero downtime upgrade and direct you to +the older commands, which are retained for this case: + +```bash +ozone admin scm finalizeupgrade +ozone admin om finalizeupgrade +``` diff --git a/docs/08-developer-guide/04-project/01-git/03-feature-branches/02-merge-checklist.md b/docs/08-developer-guide/04-project/01-git/03-feature-branches/02-merge-checklist.md index 53c4c51a56..c4efa75af8 100644 --- a/docs/08-developer-guide/04-project/01-git/03-feature-branches/02-merge-checklist.md +++ b/docs/08-developer-guide/04-project/01-git/03-feature-branches/02-merge-checklist.md @@ -48,7 +48,7 @@ All Github actions runs for a branch can be found at `github.com/apache/ozone/ac ## Incompatible Changes -Ozone currently supports non-rolling upgrades and downgrades even when backwards incompatible features are present. Backwards incompatible features should be added to the versioning framework so that they are not used until the Ozone upgrade is finalized, after which downgrading is not possible. +Ozone supports both rolling (zero downtime) and non-rolling upgrades and downgrades even when backwards incompatible features are present. Backwards incompatible features should be added to the versioning framework so that they are not used until the Ozone upgrade is finalized, after which downgrading is not possible. Because a rolling upgrade means components can run mixed versions at the same time, changes must be version gated so that the API surface and on-disk format stay consistent across components until the cluster is finalized. Client cross compatibility should also be maintained as much as possible, with sensible error messages provided when this is not possible. An old client should be able to talk to the new Ozone instance, and a new client should be able to talk to the old Ozone instance. From 79a9072b5849d5220484faea31cadd6c3e930b26 Mon Sep 17 00:00:00 2001 From: Zita Dombi Date: Tue, 1 Sep 2026 00:18:37 +0200 Subject: [PATCH 2/3] Some polishing --- .../03-operations/02-upgrade-and-downgrade.md | 12 ++---------- 1 file changed, 2 insertions(+), 10 deletions(-) diff --git a/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md b/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md index 62423be7ce..367a442e47 100644 --- a/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md +++ b/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md @@ -64,8 +64,6 @@ Upgrade finalization status: - Add `-v` / `--verbose` to also show the apparent versions of OM, SCM, and the Datanodes. - Add `--json` to format the output as JSON. -- In OM HA, pass `--service-id=`; in a non-HA deployment, pass - `--service-host=`. `ozone admin datanode list` lists all Datanodes and their health state (`HEALTHY`, `STALE`, or `DEAD`) as seen by SCM. Pre-finalized Datanodes are not placed in a separate read-only state: during a rolling upgrade, every reachable Datanode stays `HEALTHY` and fully participates in reads and writes, while SCM ensures a consistent write version is used across mixed-version Datanodes. To track how far Datanode finalization has progressed, use `ozone admin upgrade status` (the `Datanodes finalized: N/M` line, or `-v` for per-Datanode apparent versions) rather than the Datanode health state. `STALE` or `DEAD` Datanodes will be told to finalize by SCM once they are reachable again. @@ -185,8 +183,6 @@ ozone admin upgrade finalize been started. - Add `--wait` to have the command poll and block until the entire cluster is finalized (interruptible with Ctrl-C). -- In OM HA, pass `--service-id=`; in a non-HA deployment, pass - `--service-host=`. Monitor progress with [`ozone admin upgrade status`](#querying-upgrade-status). The command is idempotent, so finalization continues after OM restarts or leader @@ -201,9 +197,5 @@ start. `ozone admin upgrade finalize` and `ozone admin upgrade status` require an OM that supports ZDU. If they are run against a cluster whose OM predates ZDU support, they will report that the OM does not support zero downtime upgrade and direct you to -the older commands, which are retained for this case: - -```bash -ozone admin scm finalizeupgrade -ozone admin om finalizeupgrade -``` +the older, now-deprecated commands that are retained for this case +(`ozone admin scm finalizeupgrade` followed by `ozone admin om finalizeupgrade`). From 07279ed2beb2c7345fd794608783c57915d80b2b Mon Sep 17 00:00:00 2001 From: Zita Dombi Date: Tue, 1 Sep 2026 00:31:09 +0200 Subject: [PATCH 3/3] Fix spelling --- .../03-operations/02-upgrade-and-downgrade.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md b/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md index 367a442e47..571b6e1cb0 100644 --- a/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md +++ b/docs/05-administrator-guide/03-operations/02-upgrade-and-downgrade.md @@ -182,7 +182,7 @@ ozone admin upgrade finalize finalization is asynchronous, the command returns as soon as the process has been started. - Add `--wait` to have the command poll and block until the entire cluster is - finalized (interruptible with Ctrl-C). + finalized (interrupt with Ctrl-C). Monitor progress with [`ozone admin upgrade status`](#querying-upgrade-status). The command is idempotent, so finalization continues after OM restarts or leader