Skip to content

test(grpc-js): stabilize weighted round robin tests on slow runners - #3104

Merged
murgatroid99 merged 1 commit into
grpc:masterfrom
olavloite:wait-for-channel-readiness-in-tests
Oct 1, 2026
Merged

murgatroid99 merged 1 commit into
grpc:masterfrom
olavloite:wait-for-channel-readiness-in-tests

Conversation

@olavloite

@olavloite olavloite commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

In test-weighted-round-robin, test traffic was issued immediately against freshly
created clients with aggressive 100ms per-call deadlines. On slower or heavily
loaded CI environments (such as Windows Kokoro runners), cold-start connection
establishment, name resolution, and picker readiness frequently exceeded 100ms,
triggering intermittent DEADLINE_EXCEEDED ("Waiting for LB pick") failures.
Additionally, blind initial traffic could reach only the first backend before
the second finished connecting, leaving the second backend's metrics unrecorded.
Furthermore, executing 50 sequential RPC round-trips combined with arbitrary
sleep timers pushed total execution time beyond Mocha's default 2-second timeout
on Windows VMs (often failing at ~2.005s).

This change stabilizes the tests by:

  1. Ensuring the channel reaches readiness and that both backends have responded
    to initial traffic before measurement begins, guaranteeing that connections
    and initial metric tracking are established across all endpoints.
  2. Synchronizing deterministically against picker updates rather than relying on
    fixed sleep delays, and executing measurement RPCs concurrently over HTTP/2
    so the suite runs in ~1s while preserving the expected weight distribution.
  3. Increasing the test timeout to provide sufficient headroom against VM
    scheduling jitter on Windows CI.

In test-weighted-round-robin, test traffic was issued immediately against freshly
created clients with aggressive 100ms per-call deadlines. On slower or heavily
loaded CI environments (such as Windows Kokoro runners), cold-start connection
establishment, name resolution, and picker readiness frequently exceeded 100ms,
triggering intermittent DEADLINE_EXCEEDED ("Waiting for LB pick") failures.
Additionally, blind initial traffic could reach only the first backend before
the second finished connecting, leaving the second backend's metrics unrecorded.

Furthermore, executing 50 sequential RPC round-trips combined with arbitrary
sleep timers pushed total execution time beyond Mocha's default 2-second timeout
on Windows VMs (often failing at ~2.005s).

This change stabilizes the tests by:
1. Ensuring the channel reaches readiness and that both backends have responded
   to initial traffic before measurement begins, guaranteeing that connections
   and initial metric tracking are established across all endpoints.
2. Synchronizing deterministically against picker updates rather than relying on
   fixed sleep delays, and executing measurement RPCs concurrently over HTTP/2
   so the suite runs in ~1s while preserving the expected weight distribution.
3. Increasing the test timeout to provide sufficient headroom against VM
   scheduling jitter on Windows CI.
@olavloite
olavloite force-pushed the wait-for-channel-readiness-in-tests branch from 9e1e2f3 to 2675479 Compare October 1, 2026 10:58
@olavloite olavloite changed the title test(grpc-js): wait for channel readiness in weighted round robin tests test(grpc-js): stabilize weighted round robin tests on slow runners Oct 1, 2026
@murgatroid99
murgatroid99 merged commit c3101f9 into grpc:master Oct 1, 2026
10 checks passed
murgatroid99 pushed a commit that referenced this pull request Oct 7, 2026
…3104)

In test-weighted-round-robin, test traffic was issued immediately against freshly
created clients with aggressive 100ms per-call deadlines. On slower or heavily
loaded CI environments (such as Windows Kokoro runners), cold-start connection
establishment, name resolution, and picker readiness frequently exceeded 100ms,
triggering intermittent DEADLINE_EXCEEDED ("Waiting for LB pick") failures.
Additionally, blind initial traffic could reach only the first backend before
the second finished connecting, leaving the second backend's metrics unrecorded.
Furthermore, executing 50 sequential RPC round-trips combined with arbitrary
sleep timers pushed total execution time beyond Mocha's default 2-second timeout
on Windows VMs (often failing at ~2.005s).

This change stabilizes the tests by:
1. Ensuring the channel reaches readiness and that both backends have responded
   to initial traffic before measurement begins, guaranteeing that connections
   and initial metric tracking are established across all endpoints.
2. Synchronizing deterministically against picker updates rather than relying on
   fixed sleep delays, and executing measurement RPCs concurrently over HTTP/2
   so the suite runs in ~1s while preserving the expected weight distribution.
3. Increasing the test timeout to provide sufficient headroom against VM
   scheduling jitter on Windows CI.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants