Conversation
Creating a GZipCodec eagerly initialized compressor state and then immediately discarded it while initializing decompressor state. Callers using the streaming factories therefore paid for unrelated codec state during construction. Keep construction-time option validation, but leave compressor and decompressor allocation to the existing lazy initialization paths. Extend the option tests to cover compression-level and window-size validation independently.
|
|
Author
|
This is a nice perf win when reading gzip-compressed parquet files, btw. |
Member
Interesting, what kind of speedup are you observing? |
Author
|
No measurable speedup on files with a few dozen columns and reasonable row group sizes, but I have a workload on super wide files with tiny row groups (1600-ish columns, small file is 7.4MiB / 6 row groups, large file is 116MiB / 30 row groups). There, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
GZipCodec::Init()initializes compressor state and then immediately initializes decompressor state, which tears the compressor state down again. Compressor state is fairly expensive to initialize (thedeflateInit2call allocates ~256KiB), so this is rather wasteful. A caller that wants to compress has to do it twice (because decompressor initialization throws it away), and a caller that only wants to decompress still has to do the compressor setup. Whatever the caller wants to do with the GZipCodec, this is pure waste.What changes are included in this PR?
GZipCodec::Init()now validates the compression level and window size without allocating codec state. Since all the compression and decompression methods have to check the mode anyway before doing any work, this is safe.Are these changes tested?
Largely relying on existing test coverage; the interesting part is that
InitCompressor()validated the compression level internally, so this check has to be extracted. Split the unit test that covered this into a few separate cases to make it clearer.Are there any user-facing changes?
No.
Was AI used for this PR?
Yes, both to find the issue and to fix it (Codex).
PR code and description written by:
Reviewed before submission by: