Summary
#467 moved open_store to after the execution engine is reachable, which closes the EL-retry window in #466. On current main (6e76402) the store is still opened before install_sigterm_handler runs.
A SIGTERM in that gap kills the process with the redb store open and without the quick-repair commit that savepoint writes. The next start then does a full repair. #466 measured that as 1h04 on a 103 GB testnet store. This report is the same failure mode on the window #467 left open. I have not re-run that store-size measurement.
This is not #360. That issue is the race after the handler is already installed (the handler task is aborted when the runtime drops). Here no handler is installed yet.
What the code does
In crates/malachite-app/src/node.rs, App::start opens the store and then awaits consensus-engine startup:
let (store, store_monitor) = self
.open_store(db_metrics, env_config.db_cache_size)
.await?;
// ...
let (channels, engine_handle) = self.start_consensus_engine(ctx, identity).await?;
start_consensus_engine awaits validator-proof creation and malachitebft_app_channel::start_engine.
Node::run installs the handler only after start() returns:
let mut handles = match self.start().await { /* ... */ };
install_sigterm_handler(&handles);
install_sigterm_handler is what calls store.savepoint() before exit. An unhandled SIGTERM does not run Drop, so that commit never happens. An orderly Err return from start() still drops the store and is fine. Only the signal in this window is the problem.
#467's commit message already calls this out: the patch is minimal, and a proper fix installs the signal handler before the store is opened.
Expected
SIGTERM at any point after open_store should take the same path as the installed handler: stop, savepoint, then exit. The next start should log Database opened, not Database repair in progress.
Proposed fix
Keep the change in node.rs. Install a SIGTERM handler that can savepoint before open_store, and keep it in place through start_consensus_engine. No consensus or Engine API behavior change.
Summary
#467 moved
open_storeto after the execution engine is reachable, which closes the EL-retry window in #466. On currentmain(6e76402) the store is still opened beforeinstall_sigterm_handlerruns.A SIGTERM in that gap kills the process with the redb store open and without the quick-repair commit that
savepointwrites. The next start then does a full repair. #466 measured that as 1h04 on a 103 GB testnet store. This report is the same failure mode on the window #467 left open. I have not re-run that store-size measurement.This is not #360. That issue is the race after the handler is already installed (the handler task is aborted when the runtime drops). Here no handler is installed yet.
What the code does
In
crates/malachite-app/src/node.rs,App::startopens the store and then awaits consensus-engine startup:start_consensus_engineawaits validator-proof creation andmalachitebft_app_channel::start_engine.Node::runinstalls the handler only afterstart()returns:install_sigterm_handleris what callsstore.savepoint()before exit. An unhandled SIGTERM does not runDrop, so that commit never happens. An orderlyErrreturn fromstart()still drops the store and is fine. Only the signal in this window is the problem.#467's commit message already calls this out: the patch is minimal, and a proper fix installs the signal handler before the store is opened.
Expected
SIGTERM at any point after
open_storeshould take the same path as the installed handler: stop,savepoint, then exit. The next start should logDatabase opened, notDatabase repair in progress.Proposed fix
Keep the change in
node.rs. Install a SIGTERM handler that cansavepointbeforeopen_store, and keep it in place throughstart_consensus_engine. No consensus or Engine API behavior change.