You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
nightly-performance.yml has failed on main every night since 2026-09-01 — ten consecutive runs, 09-01 through 09-10. dd-onprem-performance completes normally; ipi-onprem-performance exits 1.
There are two separable defects here:
A. The adapter hides the failure. Ten nights have produced no error message at all, because the script throws before it prints the output it captured.
B. The nightly depends on a data file that is committed monthly, and September's was never committed. The last green run is 2026-08-31; the first red one is the first nightly of September.
Draft PR #42 addresses B by fetching the file from Azure instead, and its runs are green. It does not address A.
A. The adapter swallows the example's output
common-ci/rust/run-performance-tests.ps1:
10$ErrorActionPreference="Stop"11$PSNativeCommandUseErrorActionPreference=$true
...
41Write-Host"Running performance example '$Bin'..."
...
44$output= cargo run --release --config source.toml -p $Package--bin $Bin2>&1|Out-String
...
48Write-Host$output
Line 11 turns a non-zero native exit code into a terminating error. So when cargo run fails at line 44, execution never reaches the Write-Host $output on line 48 — the output was captured into $output and then discarded. All CI can show is:
Running performance example 'ipi-onprem-performance'...
NativeCommandExitException: .../common-ci/rust/run-performance-tests.ps1:44
44 | $output = cargo run --release --config source.toml -p $Pa …
| Program "cargo" ended with non-zero exit code: 1 (0x01).
##[error]Process completed with exit code 1.
That is the entire diagnostic surface for ten nights of failure. The example's own error message — the thing that would identify the cause in seconds — is captured and thrown away.
This is worth fixing on its own merits, independently of B. Printing $output (or streaming rather than capturing) before rethrowing would have made this a one-night fix rather than a ten-night outage.
B. The ASN data file is stale
The workflow pins the data file explicitly:
.github/workflows/nightly-performance.yml, step Configure native source and data directories:
51DEGREES_IPI_PATH takes precedence over every tier fallback (examples/examples-shared/src/data_paths.rs:111-117), so the example loads exactly this file and cannot fall back to anything else.
That file is committed to the ip-intelligence-data submodule and refreshed monthly:
7d194429 2026-08-06 DATA: Add the Ipi V4 Asn file for August 2026
ad93ef35 2026-07-02 DATA: Add the Ipi V4 Asn file for July 2026
a9565a0a 2026-06-02 DATA: Add the Ipi V4 Asn file for June 2026
There is no September commit, and it is now 2026-09-10. The nightly went red on the first run of September:
2026-08-31 02:06 success
2026-09-01 02:07 failure <- first nightly of September
... every night since ...
2026-09-10 02:06 failure
Hypothesis (not yet proven): the August file stopped being loadable at the month boundary — an expiry or validity window is the natural reading of that correlation. I could not confirm the exact error because of defect A, and could not reproduce locally (no cargo on this machine). The correlation is strong but the mechanism should be confirmed once A is fixed.
What is confirmed is that replacing the committed file with a current one fixes the run — see below.
Fix A regardless of how B is resolved. Print the captured output before the error propagates. Cheap, and it is the reason this took ten nights to characterise.
Decide B deliberately — fetching at run time (PR Fix CI Performance Tests #42's approach) removes the monthly-commit dependency and is more robust, but makes the nightly depend on Azure availability and means the benchmark input changes under you, which is a real consideration for a performance trend graph. Pinning a committed file gives comparable numbers night to night; fetching gives a run that does not rot. Whichever is chosen, the failure mode should be loud.
If the committed-file approach is kept, commit the September ASN file and add a check that fails with a clear message when the file is unloadable.
Two things reviewers may want to look at, neither blocking:
It also relocates the adapter from common-ci/rust/ into rust/ci/, deliberately leaving the common-ci copy in place to be removed by a separate PR. Until that lands there are two identical copies, and only one is live.
Its get-lite-file-from-azure.ps1 call fetches both the Lite and ASN archives. The workflow's 51DEGREES_IPI_PATH still points at the ASN file, so the Lite download (~230 MB compressed, ~460 MB unpacked) appears to be unused by this job. Worth confirming it is needed.
Related
common-ci#215 — Standardise run-performance-tests.ps1: examples emit the results JSON, adapters just copy it. Overlaps with A; it reworks this script's output handling, but is scoped as a standardisation task and does not itself track this failure.
rust#35 — FEAT: Emit performance results JSON from the examples — the rust-side half of common-ci#215.
Summary
nightly-performance.ymlhas failed onmainevery night since 2026-09-01 — ten consecutive runs, 09-01 through 09-10.dd-onprem-performancecompletes normally;ipi-onprem-performanceexits 1.There are two separable defects here:
Draft PR #42 addresses B by fetching the file from Azure instead, and its runs are green. It does not address A.
A. The adapter swallows the example's output
common-ci/rust/run-performance-tests.ps1:Line 11 turns a non-zero native exit code into a terminating error. So when
cargo runfails at line 44, execution never reaches theWrite-Host $outputon line 48 — the output was captured into$outputand then discarded. All CI can show is:That is the entire diagnostic surface for ten nights of failure. The example's own error message — the thing that would identify the cause in seconds — is captured and thrown away.
This is worth fixing on its own merits, independently of B. Printing
$output(or streaming rather than capturing) before rethrowing would have made this a one-night fix rather than a ten-night outage.B. The ASN data file is stale
The workflow pins the data file explicitly:
.github/workflows/nightly-performance.yml, step Configure native source and data directories:51DEGREES_IPI_PATHtakes precedence over every tier fallback (examples/examples-shared/src/data_paths.rs:111-117), so the example loads exactly this file and cannot fall back to anything else.That file is committed to the
ip-intelligence-datasubmodule and refreshed monthly:There is no September commit, and it is now 2026-09-10. The nightly went red on the first run of September:
Hypothesis (not yet proven): the August file stopped being loadable at the month boundary — an expiry or validity window is the natural reading of that correlation. I could not confirm the exact error because of defect A, and could not reproduce locally (no cargo on this machine). The correlation is strong but the mechanism should be confirmed once A is fixed.
What is confirmed is that replacing the committed file with a current one fixes the run — see below.
Corroboration from PR #42
#42 (draft, @drasmart) deletes the committed ASN file and downloads a fresh copy from Azure before running the examples. Its runs are green:
The Azure blob's
Last-Modifiedis Mon, 07 Sep 2026 11:42:03 GMT — i.e. a September file exists upstream; it just never reached the submodule.Runs: 34451901570, 34455534461.
Suggested resolution
Notes on PR #42's scope
Two things reviewers may want to look at, neither blocking:
common-ci/rust/intorust/ci/, deliberately leaving the common-ci copy in place to be removed by a separate PR. Until that lands there are two identical copies, and only one is live.get-lite-file-from-azure.ps1call fetches both the Lite and ASN archives. The workflow's51DEGREES_IPI_PATHstill points at the ASN file, so the Lite download (~230 MB compressed, ~460 MB unpacked) appears to be unused by this job. Worth confirming it is needed.Related
run-performance-tests.ps1: examples emit the results JSON, adapters just copy it. Overlaps with A; it reworks this script's output handling, but is scoped as a standardisation task and does not itself track this failure.