Skip to content

Swarm stops retrying a Pending task after a service-only label update #53722

Description

@Mazyod

Created with: Astra (AI), under my supervision and review.

Candidate fix: moby/swarmkit#3297

Description

A valid unassigned task can remain Pending and desired Running after its eligible node returns. A service-level label update advances the service revision without replacing the task. If another scheduling pass occurs while the node is unavailable, the task is dropped from the scheduler's retry queue. An untouched control service recovers; force-updating the affected service creates a replacement that runs.

Related to moby/moby#36699, especially the 2023 node-restart report and 2024 follow-up. Those later reports have similar symptoms but do not establish the metadata-update/revision trigger. The original 2018 issue concerns tasks desired Shutdown; this report concerns tasks still desired Running. This is not a claim to resolve all of #36699.

Reproduce

The disposable lab harness creates labelled DinD daemons and cleans up only its own resources. It requires a local Linux-container Docker engine supporting privileged DinD. The original stock reproducer is pinned at be601541949e509f65b1096e4ca0c760202364b9:

git clone https://github.com/Mazyod/swarm-pending-task-repro.git
cd swarm-pending-task-repro
git checkout be601541949e509f65b1096e4ca0c760202364b9
./scripts/run.sh pause 29.8.1
./scripts/run.sh restart 29.1.5

The smaller scenario performs these steps inside a disposable single-node Swarm:

  1. Cache busybox:1.37 and set the node's scheduling availability to Pause.
  2. Create services control and updated, each running busybox:1.37 sleep 600; wait for both tasks to become Pending and record their IDs.
  3. Run docker service update --detach --label-add repro.revision=1 updated. Its task ID remains unchanged.
  4. Run docker node update --label-add repro.tick=1 <node> to trigger another scheduler pass while scheduling is still paused.
  5. Set node availability back to Active. The original control task runs; the updated service's original task stays Pending.
  6. Run docker service update --detach --force updated. A new task ID reaches Running.

The worker-restart scenario first runs both services on a separate worker using a hostname constraint. It stops that worker's DinD container, waits for Down and replacement Pending tasks, updates only one service's label, triggers a scheduler pass by creating an unrelated service, and starts the same worker. Only the untouched control recovers on stock Engine.

Expected behavior

Both still-valid Pending tasks remain eligible for scheduling and run after node recovery.

Actual: the updated service's original Pending task remains unassigned and desired Running while the control runs. The lab observes five seconds after control recovery, and the source regression directly verifies loss of the retry queue entry.

docker version

Stock 29.8.1 disposable manager, Linux amd64:

Client:
 Version:           29.8.1
 API version:       1.56
 Go version:        go1.26.8
 Git commit:        4a63305
 Built:             Tue Sep 15 16:24:42 2026
 OS/Arch:           linux/amd64
 Context:           default

Server: Docker Engine - Community
 Engine:
  Version:          29.8.1
  API version:      1.56 (minimum version 1.40)
  Go version:       go1.26.8
  Git commit:       464cd50
  Built:            Tue Sep 15 16:27:24 2026
  OS/Arch:          linux/amd64
  Experimental:     false
 containerd:
  Version:          v2.3.5
  GitCommit:        1294c24a7da8e5a793ed378161673abe94118892
 runc:
  Version:          1.5.1
  GitCommit:        v1.5.1-0-g8f2685a
 docker-init:
  Version:          0.19.0
  GitCommit:        de40ad0

docker info

Captured from that same disposable daemon before Swarm initialization:

Client:
 Version:    29.8.1
 Context:    default
 Debug Mode: false
 Plugins:
  buildx: Docker Buildx (Docker Inc.)
    Version:  v0.37.1
    Path:     /usr/local/libexec/docker/cli-plugins/docker-buildx
  compose: Docker Compose (Docker Inc.)
    Version:  v5.5.1
    Path:     /usr/local/libexec/docker/cli-plugins/docker-compose

Server:
 Containers: 0
  Running: 0
  Paused: 0
  Stopped: 0
 Images: 0
 Server Version: 29.8.1
 Storage Driver: vfs
 Logging Driver: json-file
 Cgroup Driver: cgroupfs
 Cgroup Version: 2
 Plugins:
  Volume: local
  Network: bridge host ipvlan macvlan null overlay
  Log: awslogs fluentd gcplogs gelf journald json-file local splunk syslog
 CDI spec directories:
  /etc/cdi
  /var/run/cdi
 Swarm: inactive
 Runtimes: io.containerd.runc.v2 runc
 Default Runtime: runc
 Init Binary: docker-init
 containerd version: 1294c24a7da8e5a793ed378161673abe94118892
 runc version: v1.5.1-0-g8f2685a
 init version: de40ad0
 Security Options:
  apparmor
   Profile: default
  seccomp
   Profile: builtin
  cgroupns
 Kernel Version: 7.0.0-30-generic
 Operating System: Alpine Linux v3.24 (containerized)
 OSType: linux
 Architecture: x86_64
 CPUs: 12
 Total Memory: 12.06GiB
 Name: patch-manager
 ID: 59fa6989-a661-431d-acd4-091e8c98f8ad
 Docker Root Dir: /var/lib/docker
 Debug Mode: true
  File Descriptors: 26
  Goroutines: 51
  System Time: 2026-09-18T11:14:59.432676771Z
  EventsListeners: 0
 Experimental: false
 Insecure Registries:
  ::1/128
  127.0.0.0/8
 Live Restore Enabled: false
 Product License: Community Engine
 Firewall Backend: iptables
  EnableUserlandProxy: true
  UserlandProxyPath: /usr/local/bin/docker-proxy

Additional Info

All evidence is from disposable labs. Stock Docker 29.1.5 and 29.8.1 reproduce the pause case; 29.1.5 reproduces worker stop/start. The Docker image digests were:

  • docker:29.1.5-dind: docker@sha256:3a33fc81fa4d38360f490f5b900e9846f725db45bb1d9b1fe02d849bd42a5cf2
  • docker:29.8.1-dind: docker@sha256:3f3c01aaaebf7cce837356b688b7c059a4749f10bd7660dec7c58fc454a283f0
  • busybox:1.37: busybox@sha256:9db7b59979c38555a39def84a31fb98b5296952f9e3afd4f6f11f05b07adfab0

One matching scheduler message from the 29.8.1 pause run:

time="2026-09-18T11:15:04.855048722Z" level=debug msg="task belongs to old revision of service" module=scheduler node.id=5zow56i4ep3rwzvrlg9x9dk4d task.id=sw8i16alhabowfjz5pc4tu0uf

In SwarmKit's scheduler, noSuitableNode takes the older-revision branch whenever service.SpecVersion.Index > task.SpecVersion.Index. For a task still desired Running, it neither transitions to Shutdown nor reaches enqueue in the other branch. The task remains in the store but disappears from later retries. The branch was introduced in SwarmKit #2998 to retire obsolete tasks desired Shutdown.

A candidate SwarmKit fix adds the desired-Shutdown condition to the older-revision branch. Its regression fails on stock because the task is missing from unassignedTasks, then passes with the fix and verifies assignment of the same task ID after node recovery. The full scheduler package, full scheduler race check, and full SwarmKit root-module suite pass on Go 1.27.1 linux/amd64, including the existing obsolete-task shutdown regression.

For runtime validation, built stock and patched managers from Moby 464cd50c3d9e92877d56940ea160de6fca7bea23 (29.8.1), vendoring SwarmKit ad0357aedfca72c288abb8395ec5603e1136a9e5. The scheduler file matches the tested SwarmKit HEAD byte-for-byte. The stock build reproduces both scenarios; the patched build recovers the same Pending task IDs in both, including with a stock worker. These are laboratory builds, not a released fix. A SwarmKit PR and a later Moby dependency update/backport are separate steps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions