Repository navigation
Conversation
This was referenced Oct 2, 2026
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…lls any more 30257df (Android device queries match the whole text and id) took out the last caller of this helper in the uiautomator2 driver. The unused linter has flagged it since, so Lint has failed on every push to main, and Build, which needs Lint, has been skipped each time. The same helper in the appium and devicelab drivers is untouched.
Maestro runs a runScript file as plain JavaScript. It reads the file and
evaluates ${...} only in the step's env, when: condition and label; the
script text goes to the engine as it is (YamlFluentCommand.kt:409-424,
Commands.kt:1029-1035, Orchestra.kt:723-737).
The runner expanded ${...} and $VAR across the whole file before running
it. A template literal that used the script's own variables was replaced
ahead of the script, against variables that did not exist yet:
`/v1/x?email=${encodeURIComponent(who)}&state=${state}` came out as
`/v1/x?email=undefined&state=`.
A script file now runs as written. Inline script text, which Maestro has no
equivalent of, keeps the expansion.
Maestro gives a sub-flow its own env scope: enterEnvScope saves the env and
leaveEnvScope puts that copy back, so a key the sub-flow added is gone when
it returns (GraalJsEngine.kt:223-238, around runSubFlow at
Orchestra.kt:1159-1197, which repeat, retry and runFlow all go through).
The runner's withEnvVars restored each key to the value it had before, and
a key that had none was set to "" rather than removed. After a runFlow,
retry or sub-flow with `env: {KEY: ...}`, KEY stayed defined: typeof KEY was
"string", `$KEY` expanded to nothing, and runShell saw KEY="" in its
environment.
withEnvVars now uses applyScopedEnv, which runScript's env already used:
a key is restored when it existed and removed when it did not.
Maestro builds its script http client with 5-minute read, write and call timeouts (GraalJsEngine.kt:28-34), and the CLI passes no client of its own (Orchestra.kt:138 and 161), so every call gets them. The runner's http.* gave a call 30 s unless the call set `timeout`, so a script that calls a slow endpoint (seeding test data, waiting on a server-side job) failed with "HTTP request failed" where Maestro waits. A local server that answered after 32 s failed the call at 30 s. A call without a timeout option now gets 5 minutes. The option still wins.
On a physical iPhone reached through usbmux and an SSH tunnel, WDA began dropping connections 90 minutes into a run: about one request in fourteen came back as EOF within about 10 ms, nearly all of them element reads sent in parallel (displayed, text, rect, name). Most were harmless, but a dropped text read made a copyTextFrom come back empty and failed the flow. A fresh WDA dropped none, in seven other runs. net/http does not help here. It sends a request again by itself only when it is a GET on a connection that had been used before (shouldRetryRequest), so a fresh connection that is hung up on, and any POST, come back as errors. A GET, or a POST that only finds elements, whose connection dies before any response (EOF, reset, broken pipe) is now sent once more, with a warning in the log. An action (a tap, typing, a swipe, launching an app) is never repeated: it may have reached WDA before the connection went. One more try, not a loop. The changelog entry for keeping idle connections said those requests are sent again. That is only true with this change, so the entry is reworded.
7 of 13 tasks
bulatgaleev
force-pushed
the
upstream-pr/b2-wda-resend-dropped-reads
branch
from
October 2, 2026 04:07
1bdd641 to
1bac95a
Compare
Contributor
|
Thanks @bulatgaleev, this is a nice one. Resending only reads and lookups after a dropped connection, and never actions, is exactly the right line. The "an action is not sent again" test makes it clear. Merging. And thanks for the whole series (#194–#197, #200). Each one was small, well explained and tested, which made them easy to review. They'll go out in the next release. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
WebDriverAgent sometimes closes a connection before it answers. On a physical iPhone reached through usbmux and an SSH tunnel, about one request in fourteen came back as EOF within about 10 ms, 90 minutes into a long run, nearly all of them element reads sent in parallel (displayed, text, rect, name). Most were harmless, but a dropped text read made
copyTextFromcopy an empty string and fail the flow. A fresh WDA dropped none in seven other runs. Go'snet/httpdoes not help here: it sends a request again by itself only when it is a GET on a connection that had been used before, so a fresh connection that is hung up on, and any POST, come back as errors. A GET, or a POST that only finds elements, whose connection dies before any response (EOF, reset, broken pipe) is now sent once more, with a warning in the log. An action (a tap, typing, a swipe, launching an app) is never repeated, because it may have reached WDA before the connection went. This one is about the link to WDA, not about a difference from Maestro.Type of Change
Changes Made
Client.getsends the request once more when the connection drops before a response (EOF, unexpected EOF, connection reset, broken pipe).Client.postdoes the same for a lookup, a path ending in/elementor/elements. It keeps the body as bytes so it can send it twice. Any other POST, and DELETE, are never sent again.pkg/driver/wda/dropped_connection_test.go: a read is sent again, a lookup is sent again, a tap is sent exactly once and its error comes back, and a request that fails twice fails (one more try, not a loop). Three of the four fail without the change.Related Issues
No existing issue found. #181 keeps idle connections open so fewer new ones are opened; this covers a connection that drops anyway.
Testing
go test ./pkg/driver/wda/go test -raceforpkg/executor,pkg/driver/wda,pkg/jsengine,pkg/flowandpkg/driver/uiautomator2passes at the top of the stack;go vet ./...,gofmt -land golangci-lint v2.13.2 (the version CI pins) report nothing on every branch of the stackmake testas a whole: not run, because its device tests drive whatever device is attached to the machinemake lint: the Makefile has nolinttarget, and the other lintersmake checkruns (staticcheck, revive, errcheck, nilaway, gosec) are not installed hereChecklist
Additional Notes
A resent request leaves a warning in the log:
WDA GET <path>: the connection dropped before a response (<error>), sending it again.Stack 5 of 7: #200 → #194 → #195 → #196 → #197 → #198 → #199. Merge in that order. This branch is built on #200, #194, #195 and #196, so it also contains their commits. This PR's own change is the top commit,
1bac95a. Once the PRs below it merge, the rest of the diff disappears, and nothing conflicts. All seven sit onmainatd3f738f(the merge of #185).