Repository navigation
StatefulSet restarter always restarts replica 0 immediately after initial rollout #111
Description
Activity
Actually this is how the restart controller is supposed to work. The Superset operator creates a stateful set labeled with
restarter.stackable.tech/enabled=true. The restart controller of the commons operator watches this labeled stateful set and adds the UIDs and resource versions of the referenced config maps and secrets as annotations to the pod template. This triggers a restart of the stateful set right after its first creation. The commons operator watches these config maps and secrets and updates the annotations whenever these resources change which in turn triggers a restart. This ensures that Superset always runs with the latest configuration.Even if the immediate restart is confusing, I do not see an elegant solution to prevent this, so I propose to close this issue.
If the outcome is "this is how it is supposed to work", then it might not be a bug, be we should still go back to the drawing board and look at what we've built, because this immediate restart is really something that shouldn't happen.
Maybe the title of the issue and also it's place in this repo is not correct anymore.
After talking to sigi, I'm moving it back to "Refinement: In Progress"
Taking a step back. Outcomes from the refinement discussion:
Observed Situation
- A clean system is easier to monitor, debug, and operate
- Random restarts are surprising in a bad way
- Logs are not mentioning reason for this behaviour - just the restart is visible
- We want our systems to be understandable, and observable (know why restart happens)
- The restart controller causes a restart, which might be avoided
- Two operators working on the same resource interact in sometimes hard to understand fashion (badly designed system?)
Investigation Worthy Threads
- Can we avoid this kind of initial restart? Are they strictly necessary?
- Mutating admission controller / webhook (?) link
- Port restarting logic into the operator-rs framework (it is a library!)? Reusable functions.
- Can we add logs to understand what's going on
- Add logging outputs to justify / announce restarts
- Root cause analysis: how was the decision made that we want to have this kind of pattern?
- Possible "design guideline" which could have been helpful: "only one operator is modifying a single resource"
Possible Next Steps
- Create spike investigating if restart is necessary?
- Address "how was restarting decision" in architecture meeting?
- Do we need somebody who feels responsible for higher-level architecture patterns?
@soenkeliebau @lfrancke unclear how to proceed here, some input would be great. Do we leave things as they are here, or do we pursue one of the investigation threads?
This ticket can be closed in favour of these other investigation threads.
That the commons-operator is reconciling doesn't really say much, that'll happen whenever the STS is modified externally too. If you want to check whether it has triggered a restart then you'll want to check the STS YAMLs for whether the restarter annotations change.
I'd also argue that the restart controller isn't the cause of the problem here, just a symptom. The restart controller will trigger when the app's config changes. Check for why that happens if you want to avoid this.
(Case in point, if the restart controller is at fault then this will happen for all product operators, not just Superset...)
Thanks Teo.
The restart controller of the commons operator watches this labeled stateful set and adds the UIDs and resource versions of the referenced config maps and secrets as annotations to the pod template. This triggers a restart of the stateful set right after its first creation.
I'm a bit confused now, though. This sounds like the answer that you were looking for, no?
I must be missing something still. I thought the restart happens because the restart controller changes the configmaps which in turn triggers a restart of the STS?The restart controller just watches the configmaps for changes, and triggers STS restarts based on that. The restart controller is just a messenger for the actual issue (that the CM changed). It never modifies the CMs itself.
Understood, thank you.
So, the described behavior has to come from somewhere else then.@siegfriedweber @fhennig @vsupalov - sorry, back to you for now?
If you want to check whether it has triggered a restart then you'll want to check the STS YAMLs for whether the restarter annotations change.
The restarter annotations are initially added to the STS YAMLs by the restart controller. This triggers the restart.
The restart controller will trigger when the app's config changes.
The restart controller also triggers when a new StatefulSet with the annotation
restarter.stackable.tech/enabled=truewas created.(Case in point, if the restart controller is at fault then this will happen for all product operators, not just Superset...)
It happens for all StatefulSets which are annotated with
restarter.stackable.tech/enabled=true. That is, it also happens for the Airflow operator. The other operators do not use this annotation yet.I think that I accurately described the issue in my first comment:
Actually this is how the restart controller is supposed to work. The Superset operator creates a stateful set labeled with
restarter.stackable.tech/enabled=true. The restart controller of the commons operator watches this labeled stateful set and adds the UIDs and resource versions of the referenced config maps and secrets as annotations to the pod template. This triggers a restart of the stateful set right after its first creation. The commons operator watches these config maps and secrets and updates the annotations whenever these resources change which in turn triggers a restart. This ensures that Superset always runs with the latest configuration.Please tell me if I am wrong.
@siegfriedweber and @teozkr will have a chat about this
Hm, the label "should" still be added quickly enough that we the pods shouldn't have been scheduled yet, but maybe my memory fails me on how much of a window we have, will have to do some testing tomorrow... If that is indeed the sad case then a fail-open (since we'd still need the controller as well, the worst that could happen would be falling back to the current extra restart) mutating webhook probably makes sense at some point, and should be completely transparent to the operators.
The downside is that doing a first webhook means doing a bunch of new scaffolding from scratch (certificates, kube API registration, etc)...
Ok, having tried it out now it does look like the first replica is restarted while all other replicas are created late enough that they are up to date. Avoiding that glitch will probably require us to do the webhook...
I have validated that a mutating webhook is able to avoid the initial restart for a simple STS... 0cbf72c
Moving this ticket to commons since this is a bug in the restarter, not in superset-op specifically.
Reacted by Felix Hennig31 remaining items
This is an old ticket and while it is very well written and detailed I'd like to see a refinement/plan before we start implementing anything.
The plan should basically be a summary of past and current discussions and it should explain what we plan on doing now and how.
Meta: Add a mutating webhook that adds restarter annotations for ConfigMaps + Secret instantly before the first Pod is created. This way the first Pod doesn't need to be updated immediately after it's creation.
We already have a working spike by @nightkrSteps:
- Extend stackable-webhook with supports for mutating webhooks (in addition to the current conversion webhooks)
- Take @nightkr spike branch and port it to the new stackable-webhook crate
a. The whole TLS server, cert creation and rotation is already solved by stackable-webhook, so we don't need to think about that
b. Things in kube-rs has changed since, needs some adopting
c. Use the types from https://docs.rs/kube/latest/kube/core/admission/index.html, so we have better typing and ergonomics in Rust code -> Basically the same as we have for CRD versioning with https://docs.rs/kube/latest/kube/core/conversion/index.html.
Small side note: Technically this only works in case the StatefulSet is created after the mounted Secrets and ConfigMaps, but after quickly checking trino-operator I think in practice this shouldn't be a problem.
I now deeply understand all parts and already have a working prototype for all of this. Happy to answer detail questions :)
Reacted by Lars Francke- moved this from Selected for Development to In Progress in Stackable End-to-End Coordination
on Nov 19, 2025 Thank you!
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Current behaviour
New restarter-enabled StatefulSet have their replica 0 restarted after the initial rollout is complete.
Expected Behaviour
The initial rollout should be completed "normally" with no extra restarts.
Why does this happen?
There is a race condition between Kubernetes' StatefulSet controller creating the first replica Pod and commons' StatefulSet restart controller adding the restart trigger labels. If the restarter loses the race then the first replica is created without the metadata, triggering a restart once it is added.
What can we do about it?
Add a mutating webhook (see the spike) that adds the relevant metadata. The webhook must not replace the existing controller, since webhook delivery is not reliable.
However, webhook delivery requires a bunch of extra infrastructure that we do not currently have, namely:
Tasks
Definition of done
podTemplateannotations as the controller currently doesSTS.metadata.generationshould stay1)failurePolicy: Ignore)failurePolicy)Original ticket
The StatefulSet of a Superset cluster is immediately restarted after its creation. This should not be necessary and should be prevented.After the restart, the Superset pods are annotated as follows:
This could be an indication that the restart controller of the commons-operator is involved.
The commons-operator is busy while the StatefulSet is restarted: