Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review. 📝 WalkthroughWalkthroughThe backup upload wraps its pipe input stream with ChangesBackup stream upload
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~10 minutes Change: Refactor Suggested reviewers: Merge Risk: ⚪ Minimal · up to The changed backup reader has no identified issue that needs resolution before merge. The planned full-backup rollout check remains appropriate. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This is a standalone efficiency optimization. It does not resolve the September 24 outage investigation and should not be deployed on the assumption that it prevents recurrence.
Backup uploads currently hand at most 8 KiB to the AWS SDK on each read, while SDK 2.46.2 allocates a 16 KiB backing array for every read. Multipart retries retain those arrays. A filled 32 MiB payload queue therefore retains roughly 64 MiB of backing arrays.
Fill the caller's buffer before returning from the backup pipe, except at EOF. The same queue then uses roughly 32 MiB of backing arrays. This changes only the backup reader; compression, object contents, multipart sizing, concurrency, retry behavior and publication order stay the same. Sparse namespaces can wait for more bytes, and finish when their writer closes.
Validation:
git diff --checkpassed. Independent review found no actionable regressions. Clojure CI also passed one1f513f61: full lint, uberjar build and all five test shards.This reduces Java heap allocations. With the fixed 90 GiB heap, the reduction need not translate into lower host RSS. The failed host exhausted root-volume read throughput while telemetry stopped, but retained evidence does not identify the initiating process or connect these buffers to the reads. Passing a full nightly backup would establish additional functional confidence; it would not by itself establish outage prevention. That requires attributing or reproducing the host stall, then showing that the proposed mitigation prevents it while the complete backup and concurrent request workload succeed.