Repository navigation
ci: bencher v1, benches/perf/{lb,sharding,pooler}, detect perf regressions - #1617
Conversation
|
| Project | PgDog |
| Branch | jk-bencher-v1 |
| Testbed | Intel v1 |
Click to view all benchmark results
| Benchmark | Throughput | operations / second (ops/s) x 1e3 |
|---|---|---|
| lb | 📈 view plot 🚷 view threshold | 29.35 ops/s x 1e3 |
| pooler | 📈 view plot 🚷 view threshold | 30.77 ops/s x 1e3 |
| sharding | 📈 view plot 🚷 view threshold | 19.91 ops/s x 1e3 |
| # | ||
| # Why --github-actions? This allows it to comment on the PR with the result. | ||
| # Why --adapter json? This allows us to use a 'custom harness'; basically nothing like rust's cargo test / criterion, we're running custom scripts | ||
| run: | |
There was a problem hiding this comment.
I think you want this to be fire and forget (don't wait for job to start/end). Otherwise, we're sitting in the action wasting minutes.
There was a problem hiding this comment.
Yep. You're right! I missed that it edited the message.
There was a problem hiding this comment.
I forgot about this: GITHUB_TOKEN expires when the job finishes or after its effective maximum lifetime.
|
| Project | PgDog |
| Branch | jk-bencher-v1 |
| Testbed | Intel v1 |
Click to view all benchmark results
| Benchmark | Throughput | Benchmark Result operations / second (ops/s) x 1e3 (Result Δ%) | Lower Boundary operations / second (ops/s) x 1e3 (Limit %) |
|---|---|---|---|
| pooler | 📈 view plot 🚷 view threshold | 30.88 ops/s x 1e3(+0.47%)Baseline: 30.74 ops/s x 1e3 | 29.82 ops/s x 1e3 (96.55%) |
|
| Project | PgDog |
| Branch | jk-bencher-v1 |
| Testbed | Intel v1 |
Click to view all benchmark results
| Benchmark | Throughput | Benchmark Result operations / second (ops/s) x 1e3 (Result Δ%) | Lower Boundary operations / second (ops/s) x 1e3 (Limit %) |
|---|---|---|---|
| lb | 📈 view plot 🚷 view threshold | 29.61 ops/s x 1e3(+0.13%)Baseline: 29.57 ops/s x 1e3 | 28.68 ops/s x 1e3 (96.87%) |
|
| Project | PgDog |
| Branch | jk-bencher-v1 |
| Testbed | Intel v1 |
Click to view all benchmark results
| Benchmark | Throughput | Benchmark Result operations / second (ops/s) x 1e3 (Result Δ%) | Lower Boundary operations / second (ops/s) x 1e3 (Limit %) |
|---|---|---|---|
| sharding | 📈 view plot 🚷 view threshold | 19.97 ops/s x 1e3(+0.15%)Baseline: 19.94 ops/s x 1e3 | 19.34 ops/s x 1e3 (96.85%) |
|
in addition to the fire-and-forget, also leaving a note here that I've contacted the author of Bencher via email to try to figure out why our runs aren't being parallelized |
4b5d726 to
4285b8a
Compare
|
Ultimately:
|
|
TBH doesn't look like this product is quite ready. Let's setup our own GH runner instead. If we use a C7 with like 4-8 cores, we should get pretty good results imo. |
4285b8a to
28d085c
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
ebda5bb to
245188e
Compare
|
I believe this will have to be merged into main to be tested with the new fire-and-forget functionality. Bencher now internally fires |
c069afe to
8cfe197
Compare
In addition to pgdogdev#1617, also run perf regression tests on fork PRs. This uses [Bencher's recommended 2-step PR workflow](https://bencher.dev/docs/how-to/github-actions/#pull-requests-from-forks) (with some minor tweaks) to safely make sure that forks can't abuse `GITHUB_TOKEN` / `BENCHER_API_KEY` / etc. Steps: - Workflow 1: runs on the PR branch, builds the Docker image with the new pgdog binary, uploads image as an artifact (+ what PR it's from) - Workflow 2: runs off main branch after Workflow 1 concludes, uses secrets to upload the image, and run the Bencher CLI.
Automatically detect performance regressions in CI using Bencher
Runs our pre-existing load-balancer, sharding and pooler tests (loc. in benches/perf) in a workflow; measures mean TPS over 3 minutes, and looks at the history (from main) to see if the mean TPS dropped by more than 3%. If it does, it comments on the PR with the numbers, and emits an error in CI.
I tested Bencher many times, and the results were always <1% variance! Contrary to this, Blacksmith, GH Actions had variance of over 9%, and did not seem suitable for what we want. It's also free for open source :)
Implementation wise: this re-uses our pgdog-base-runtime image, re-builds with a built (via Blacksmith) pgdog binary off the tested PR's branch, copies over perf scripts, asks Bencher to run it, and Bencher handles everything from there.
Many future plans to add more tests on top of this
Something notable to mention:
re #1585