Skip to content

ci: bencher v1, benches/perf/{lb,sharding,pooler}, detect perf regressions - #1617

Merged
jkaczman merged 7 commits into
mainfrom
jk-bencher-v1
Sep 28, 2026
Merged

jkaczman merged 7 commits into
mainfrom
jk-bencher-v1

Conversation

@jkaczman

@jkaczman jkaczman commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Automatically detect performance regressions in CI using Bencher

Runs our pre-existing load-balancer, sharding and pooler tests (loc. in benches/perf) in a workflow; measures mean TPS over 3 minutes, and looks at the history (from main) to see if the mean TPS dropped by more than 3%. If it does, it comments on the PR with the numbers, and emits an error in CI.

I tested Bencher many times, and the results were always <1% variance! Contrary to this, Blacksmith, GH Actions had variance of over 9%, and did not seem suitable for what we want. It's also free for open source :)

Implementation wise: this re-uses our pgdog-base-runtime image, re-builds with a built (via Blacksmith) pgdog binary off the tested PR's branch, copies over perf scripts, asks Bencher to run it, and Bencher handles everything from there.

Many future plans to add more tests on top of this

Something notable to mention:

  • This will not work when a fork is being merged in because of GitHub secrets. This is fixable, it will just require implementing Bencher's two-step workflow for that case. I'm going to follow up with another PR fixing that.

re #1585

@jkaczman
jkaczman marked this pull request as ready for review September 22, 2026 20:58
Comment thread benches/perf/lb/run.sh Outdated
@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

🐰 Bencher Report

ProjectPgDog
Branchjk-bencher-v1
TestbedIntel v1
Click to view all benchmark results
BenchmarkThroughputoperations / second (ops/s) x 1e3
lb📈 view plot
🚷 view threshold
29.35 ops/s x 1e3
pooler📈 view plot
🚷 view threshold
30.77 ops/s x 1e3
sharding📈 view plot
🚷 view threshold
19.91 ops/s x 1e3
🐰 View full continuous benchmarking report in Bencher

#
# Why --github-actions? This allows it to comment on the PR with the result.
# Why --adapter json? This allows us to use a 'custom harness'; basically nothing like rust's cargo test / criterion, we're running custom scripts
run: |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you want this to be fire and forget (don't wait for job to start/end). Otherwise, we're sitting in the action wasting minutes.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep. You're right! I missed that it edited the message.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

🐰 Bencher Report

ProjectPgDog
Branchjk-bencher-v1
TestbedIntel v1
Click to view all benchmark results
BenchmarkThroughputBenchmark Result
operations / second (ops/s) x 1e3
(Result Δ%)
Lower Boundary
operations / second (ops/s) x 1e3
(Limit %)
pooler📈 view plot
🚷 view threshold
30.88 ops/s x 1e3
(+0.47%)Baseline: 30.74 ops/s x 1e3
29.82 ops/s x 1e3
(96.55%)
🐰 View full continuous benchmarking report in Bencher

@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

🐰 Bencher Report

ProjectPgDog
Branchjk-bencher-v1
TestbedIntel v1
Click to view all benchmark results
BenchmarkThroughputBenchmark Result
operations / second (ops/s) x 1e3
(Result Δ%)
Lower Boundary
operations / second (ops/s) x 1e3
(Limit %)
lb📈 view plot
🚷 view threshold
29.61 ops/s x 1e3
(+0.13%)Baseline: 29.57 ops/s x 1e3
28.68 ops/s x 1e3
(96.87%)
🐰 View full continuous benchmarking report in Bencher

@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

🐰 Bencher Report

ProjectPgDog
Branchjk-bencher-v1
TestbedIntel v1
Click to view all benchmark results
BenchmarkThroughputBenchmark Result
operations / second (ops/s) x 1e3
(Result Δ%)
Lower Boundary
operations / second (ops/s) x 1e3
(Limit %)
sharding📈 view plot
🚷 view threshold
19.97 ops/s x 1e3
(+0.15%)Baseline: 19.94 ops/s x 1e3
19.34 ops/s x 1e3
(96.85%)
🐰 View full continuous benchmarking report in Bencher

@jkaczman

Copy link
Copy Markdown
Contributor Author

in addition to the fire-and-forget, also leaving a note here that I've contacted the author of Bencher via email to try to figure out why our runs aren't being parallelized

@jkaczman
jkaczman added this pull request to stack #1620 September 23, 2026 12:30
@jkaczman
jkaczman force-pushed the jk-bencher-v1 branch 2 times, most recently from 4b5d726 to 4285b8a Compare September 23, 2026 22:46
@jkaczman

jkaczman commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Ultimately:

  • we're now able to have 2 benchmark runs in parallel until they get more hardware on hand
    • it might be better for now to only run benchmarks when a tag is added to a PR to prevent long waiting times from non-sensitive PRs
  • fire-and-forget feature is being worked on by them, but I'm not sure of an ETA

@levkk

levkk commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

TBH doesn't look like this product is quite ready. Let's setup our own GH runner instead. If we use a C7 with like 4-8 cores, we should get pretty good results imo.

@codecov

codecov Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@jkaczman
jkaczman force-pushed the jk-bencher-v1 branch 3 times, most recently from ebda5bb to 245188e Compare September 28, 2026 17:32
@jkaczman

Copy link
Copy Markdown
Contributor Author

I believe this will have to be merged into main to be tested with the new fire-and-forget functionality. Bencher now internally fires repository_dispatch as a callback to trigger a secondary workflow, however, that's only ran off the default (in our case, main) branch.

@jkaczman
jkaczman added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit f353ba0 Sep 28, 2026
31 of 34 checks passed
@jkaczman
jkaczman deleted the jk-bencher-v1 branch September 28, 2026 22:33
pull Bot pushed a commit to TheTechOddBug/pgdog that referenced this pull request Sep 29, 2026
In addition to pgdogdev#1617, also run perf regression tests on fork PRs.

This uses [Bencher's recommended 2-step PR
workflow](https://bencher.dev/docs/how-to/github-actions/#pull-requests-from-forks)
(with some minor tweaks) to safely make sure that forks can't abuse
`GITHUB_TOKEN` / `BENCHER_API_KEY` / etc.

Steps:
- Workflow 1: runs on the PR branch, builds the Docker image with the
new pgdog binary, uploads image as an artifact (+ what PR it's from)
- Workflow 2: runs off main branch after Workflow 1 concludes, uses
secrets to upload the image, and run the Bencher CLI.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants