Skip to content

Testing to see blacksmith run variance - #1613

Closed
jkaczman wants to merge 13 commits into
mainfrom
jk-blacksmith-bench-testing
Closed

jkaczman wants to merge 13 commits into
mainfrom
jk-blacksmith-bench-testing

Conversation

@jkaczman

Copy link
Copy Markdown
Contributor

Before I try out Bencher, which is more complex to setup, curious to see how our current Blacksmith runners do in terms of variance. Not planning on merging this unless it works well.

Using 32vcpu runners (max) to try to get (at least mostly) dedicated machines, since each core is dedicated.

This just has a basic workflow right now to emit TPS from our current benches/perf/{lb,pooler,sharding} benches.

@jkaczman

Copy link
Copy Markdown
Contributor Author

Doesn't seem possible with Blacksmith.

  • With Linux (+ Windows) runners, we have no way of knowing what kind of gaming GPU we get, meaning we can't ensure the same kind every time, leading to variance
  • With macOS runners, even though they always have M4 CPUs, they use a big.little architecture with 8 performance and 4 efficiency cores; no way of knowing how the OS schedules out our tasks onto them... leading to variance.

Oh well! Worth a try.

@jkaczman jkaczman closed this Sep 22, 2026
@levkk

levkk commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

Yup!

@jkaczman jkaczman reopened this Sep 22, 2026
@jkaczman jkaczman closed this Sep 22, 2026
pull Bot pushed a commit to TheTechOddBug/pgdog that referenced this pull request Sep 29, 2026
…sions (pgdogdev#1617)

Automatically detect performance regressions in CI using
[Bencher](https://bencher.dev/)

Runs our pre-existing load-balancer, sharding and pooler tests (loc. in
benches/perf) in a workflow; measures mean TPS over 3 minutes, and looks
at the history (from main) to see if the mean TPS dropped by more than
3%. If it does, it comments on the PR with the numbers, and emits an
error in CI.

I tested Bencher many times, and the results were always <1% variance!
Contrary to this, [Blacksmith, GH Actions had variance of over
9%](pgdogdev#1613), and did not seem
suitable for what we want. It's also free for open source :)

Implementation wise: this re-uses our pgdog-base-runtime image,
re-builds with a built (via Blacksmith) pgdog binary off the tested PR's
branch, copies over perf scripts, asks Bencher to run it, and Bencher
handles everything from there.

Many future plans to add more tests on top of this

Something notable to mention:
- This will not work when a fork is being merged in because of GitHub
secrets. This is fixable, it will just require implementing Bencher's
two-step workflow for that case. I'm going to follow up with another PR
fixing that.

re pgdogdev#1585
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants