Created with: Astra (AI), under my supervision and review.
Candidate fix: moby/swarmkit#3297
Description
A valid unassigned task can remain Pending and desired Running after its eligible node returns. A service-level label update advances the service revision without replacing the task. If another scheduling pass occurs while the node is unavailable, the task is dropped from the scheduler's retry queue. An untouched control service recovers; force-updating the affected service creates a replacement that runs.
Related to moby/moby#36699, especially the 2023 node-restart report and 2024 follow-up. Those later reports have similar symptoms but do not establish the metadata-update/revision trigger. The original 2018 issue concerns tasks desired Shutdown; this report concerns tasks still desired Running. This is not a claim to resolve all of #36699.
Reproduce
The disposable lab harness creates labelled DinD daemons and cleans up only its own resources. It requires a local Linux-container Docker engine supporting privileged DinD. The original stock reproducer is pinned at be601541949e509f65b1096e4ca0c760202364b9:
git clone https://github.com/Mazyod/swarm-pending-task-repro.git
cd swarm-pending-task-repro
git checkout be601541949e509f65b1096e4ca0c760202364b9
./scripts/run.sh pause 29.8.1
./scripts/run.sh restart 29.1.5
The smaller scenario performs these steps inside a disposable single-node Swarm:
- Cache
busybox:1.37 and set the node's scheduling availability to Pause.
- Create services
control and updated, each running busybox:1.37 sleep 600; wait for both tasks to become Pending and record their IDs.
- Run
docker service update --detach --label-add repro.revision=1 updated. Its task ID remains unchanged.
- Run
docker node update --label-add repro.tick=1 <node> to trigger another scheduler pass while scheduling is still paused.
- Set node availability back to Active. The original control task runs; the updated service's original task stays Pending.
- Run
docker service update --detach --force updated. A new task ID reaches Running.
The worker-restart scenario first runs both services on a separate worker using a hostname constraint. It stops that worker's DinD container, waits for Down and replacement Pending tasks, updates only one service's label, triggers a scheduler pass by creating an unrelated service, and starts the same worker. Only the untouched control recovers on stock Engine.
Expected behavior
Both still-valid Pending tasks remain eligible for scheduling and run after node recovery.
Actual: the updated service's original Pending task remains unassigned and desired Running while the control runs. The lab observes five seconds after control recovery, and the source regression directly verifies loss of the retry queue entry.
docker version
Stock 29.8.1 disposable manager, Linux amd64:
Client:
Version: 29.8.1
API version: 1.56
Go version: go1.26.8
Git commit: 4a63305
Built: Tue Sep 15 16:24:42 2026
OS/Arch: linux/amd64
Context: default
Server: Docker Engine - Community
Engine:
Version: 29.8.1
API version: 1.56 (minimum version 1.40)
Go version: go1.26.8
Git commit: 464cd50
Built: Tue Sep 15 16:27:24 2026
OS/Arch: linux/amd64
Experimental: false
containerd:
Version: v2.3.5
GitCommit: 1294c24a7da8e5a793ed378161673abe94118892
runc:
Version: 1.5.1
GitCommit: v1.5.1-0-g8f2685a
docker-init:
Version: 0.19.0
GitCommit: de40ad0
docker info
Captured from that same disposable daemon before Swarm initialization:
Client:
Version: 29.8.1
Context: default
Debug Mode: false
Plugins:
buildx: Docker Buildx (Docker Inc.)
Version: v0.37.1
Path: /usr/local/libexec/docker/cli-plugins/docker-buildx
compose: Docker Compose (Docker Inc.)
Version: v5.5.1
Path: /usr/local/libexec/docker/cli-plugins/docker-compose
Server:
Containers: 0
Running: 0
Paused: 0
Stopped: 0
Images: 0
Server Version: 29.8.1
Storage Driver: vfs
Logging Driver: json-file
Cgroup Driver: cgroupfs
Cgroup Version: 2
Plugins:
Volume: local
Network: bridge host ipvlan macvlan null overlay
Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
CDI spec directories:
/etc/cdi
/var/run/cdi
Swarm: inactive
Runtimes: io.containerd.runc.v2 runc
Default Runtime: runc
Init Binary: docker-init
containerd version: 1294c24a7da8e5a793ed378161673abe94118892
runc version: v1.5.1-0-g8f2685a
init version: de40ad0
Security Options:
apparmor
Profile: default
seccomp
Profile: builtin
cgroupns
Kernel Version: 7.0.0-30-generic
Operating System: Alpine Linux v3.24 (containerized)
OSType: linux
Architecture: x86_64
CPUs: 12
Total Memory: 12.06GiB
Name: patch-manager
ID: 59fa6989-a661-431d-acd4-091e8c98f8ad
Docker Root Dir: /var/lib/docker
Debug Mode: true
File Descriptors: 26
Goroutines: 51
System Time: 2026-09-18T11:14:59.432676771Z
EventsListeners: 0
Experimental: false
Insecure Registries:
::1/128
127.0.0.0/8
Live Restore Enabled: false
Product License: Community Engine
Firewall Backend: iptables
EnableUserlandProxy: true
UserlandProxyPath: /usr/local/bin/docker-proxy
Additional Info
All evidence is from disposable labs. Stock Docker 29.1.5 and 29.8.1 reproduce the pause case; 29.1.5 reproduces worker stop/start. The Docker image digests were:
docker:29.1.5-dind: docker@sha256:3a33fc81fa4d38360f490f5b900e9846f725db45bb1d9b1fe02d849bd42a5cf2
docker:29.8.1-dind: docker@sha256:3f3c01aaaebf7cce837356b688b7c059a4749f10bd7660dec7c58fc454a283f0
busybox:1.37: busybox@sha256:9db7b59979c38555a39def84a31fb98b5296952f9e3afd4f6f11f05b07adfab0
One matching scheduler message from the 29.8.1 pause run:
time="2026-09-18T11:15:04.855048722Z" level=debug msg="task belongs to old revision of service" module=scheduler node.id=5zow56i4ep3rwzvrlg9x9dk4d task.id=sw8i16alhabowfjz5pc4tu0uf
In SwarmKit's scheduler, noSuitableNode takes the older-revision branch whenever service.SpecVersion.Index > task.SpecVersion.Index. For a task still desired Running, it neither transitions to Shutdown nor reaches enqueue in the other branch. The task remains in the store but disappears from later retries. The branch was introduced in SwarmKit #2998 to retire obsolete tasks desired Shutdown.
A candidate SwarmKit fix adds the desired-Shutdown condition to the older-revision branch. Its regression fails on stock because the task is missing from unassignedTasks, then passes with the fix and verifies assignment of the same task ID after node recovery. The full scheduler package, full scheduler race check, and full SwarmKit root-module suite pass on Go 1.27.1 linux/amd64, including the existing obsolete-task shutdown regression.
For runtime validation, built stock and patched managers from Moby 464cd50c3d9e92877d56940ea160de6fca7bea23 (29.8.1), vendoring SwarmKit ad0357aedfca72c288abb8395ec5603e1136a9e5. The scheduler file matches the tested SwarmKit HEAD byte-for-byte. The stock build reproduces both scenarios; the patched build recovers the same Pending task IDs in both, including with a stock worker. These are laboratory builds, not a released fix. A SwarmKit PR and a later Moby dependency update/backport are separate steps.
Created with: Astra (AI), under my supervision and review.
Candidate fix: moby/swarmkit#3297
Description
A valid unassigned task can remain Pending and desired Running after its eligible node returns. A service-level label update advances the service revision without replacing the task. If another scheduling pass occurs while the node is unavailable, the task is dropped from the scheduler's retry queue. An untouched control service recovers; force-updating the affected service creates a replacement that runs.
Related to moby/moby#36699, especially the 2023 node-restart report and 2024 follow-up. Those later reports have similar symptoms but do not establish the metadata-update/revision trigger. The original 2018 issue concerns tasks desired Shutdown; this report concerns tasks still desired Running. This is not a claim to resolve all of #36699.
Reproduce
The disposable lab harness creates labelled DinD daemons and cleans up only its own resources. It requires a local Linux-container Docker engine supporting privileged DinD. The original stock reproducer is pinned at
be601541949e509f65b1096e4ca0c760202364b9:git clone https://github.com/Mazyod/swarm-pending-task-repro.git cd swarm-pending-task-repro git checkout be601541949e509f65b1096e4ca0c760202364b9 ./scripts/run.sh pause 29.8.1 ./scripts/run.sh restart 29.1.5The smaller scenario performs these steps inside a disposable single-node Swarm:
busybox:1.37and set the node's scheduling availability to Pause.controlandupdated, each runningbusybox:1.37 sleep 600; wait for both tasks to become Pending and record their IDs.docker service update --detach --label-add repro.revision=1 updated. Its task ID remains unchanged.docker node update --label-add repro.tick=1 <node>to trigger another scheduler pass while scheduling is still paused.docker service update --detach --force updated. A new task ID reaches Running.The worker-restart scenario first runs both services on a separate worker using a hostname constraint. It stops that worker's DinD container, waits for Down and replacement Pending tasks, updates only one service's label, triggers a scheduler pass by creating an unrelated service, and starts the same worker. Only the untouched control recovers on stock Engine.
Expected behavior
Both still-valid Pending tasks remain eligible for scheduling and run after node recovery.
Actual: the updated service's original Pending task remains unassigned and desired Running while the control runs. The lab observes five seconds after control recovery, and the source regression directly verifies loss of the retry queue entry.
docker version
Stock 29.8.1 disposable manager, Linux amd64:
docker info
Captured from that same disposable daemon before Swarm initialization:
Additional Info
All evidence is from disposable labs. Stock Docker 29.1.5 and 29.8.1 reproduce the pause case; 29.1.5 reproduces worker stop/start. The Docker image digests were:
docker:29.1.5-dind:docker@sha256:3a33fc81fa4d38360f490f5b900e9846f725db45bb1d9b1fe02d849bd42a5cf2docker:29.8.1-dind:docker@sha256:3f3c01aaaebf7cce837356b688b7c059a4749f10bd7660dec7c58fc454a283f0busybox:1.37:busybox@sha256:9db7b59979c38555a39def84a31fb98b5296952f9e3afd4f6f11f05b07adfab0One matching scheduler message from the 29.8.1 pause run:
In SwarmKit's scheduler,
noSuitableNodetakes the older-revision branch wheneverservice.SpecVersion.Index > task.SpecVersion.Index. For a task still desired Running, it neither transitions to Shutdown nor reachesenqueuein the other branch. The task remains in the store but disappears from later retries. The branch was introduced in SwarmKit #2998 to retire obsolete tasks desired Shutdown.A candidate SwarmKit fix adds the desired-Shutdown condition to the older-revision branch. Its regression fails on stock because the task is missing from
unassignedTasks, then passes with the fix and verifies assignment of the same task ID after node recovery. The full scheduler package, full scheduler race check, and full SwarmKit root-module suite pass on Go 1.27.1 linux/amd64, including the existing obsolete-task shutdown regression.For runtime validation, built stock and patched managers from Moby
464cd50c3d9e92877d56940ea160de6fca7bea23(29.8.1), vendoring SwarmKitad0357aedfca72c288abb8395ec5603e1136a9e5. The scheduler file matches the tested SwarmKit HEAD byte-for-byte. The stock build reproduces both scenarios; the patched build recovers the same Pending task IDs in both, including with a stock worker. These are laboratory builds, not a released fix. A SwarmKit PR and a later Moby dependency update/backport are separate steps.