Skip to content

GH-51308: [C++] Initialize gzip codec state lazily - #51483

Open
lorenzhs wants to merge 1 commit into
apache:mainfrom
firebolt-db:lorenz/gzip-lazy-init-upstream
Open

lorenzhs wants to merge 1 commit into
apache:mainfrom
firebolt-db:lorenz/gzip-lazy-init-upstream

Conversation

@lorenzhs

@lorenzhs lorenzhs commented Sep 24, 2026 •

Copy link
Copy Markdown

Rationale for this change

GZipCodec::Init() initializes compressor state and then immediately initializes decompressor state, which tears the compressor state down again. Compressor state is fairly expensive to initialize (the deflateInit2 call allocates ~256KiB), so this is rather wasteful. A caller that wants to compress has to do it twice (because decompressor initialization throws it away), and a caller that only wants to decompress still has to do the compressor setup. Whatever the caller wants to do with the GZipCodec, this is pure waste.

What changes are included in this PR?

GZipCodec::Init() now validates the compression level and window size without allocating codec state. Since all the compression and decompression methods have to check the mode anyway before doing any work, this is safe.

Are these changes tested?

Largely relying on existing test coverage; the interesting part is that InitCompressor() validated the compression level internally, so this check has to be extracted. Split the unit test that covered this into a few separate cases to make it clearer.

Are there any user-facing changes?

No.

Was AI used for this PR?

Yes, both to find the issue and to fix it (Codex).

PR code and description written by:

  • Human (PR desc)
  • AI (code)

Reviewed before submission by:

  • Human
  • AI
  • Not reviewed

Creating a GZipCodec eagerly initialized compressor state and then immediately discarded it while initializing decompressor state. Callers using the streaming factories therefore paid for unrelated codec state during construction.

Keep construction-time option validation, but leave compressor and decompressor allocation to the existing lazy initialization paths. Extend the option tests to cover compression-level and window-size validation independently.
@github-actions

Copy link
Copy Markdown

⚠️ GitHub issue #51308 has been automatically assigned in GitHub to PR creator.

@lorenzhs

Copy link
Copy Markdown
Author

This is a nice perf win when reading gzip-compressed parquet files, btw.

@pitrou

pitrou commented Sep 29, 2026

Copy link
Copy Markdown
Member

This is a nice perf win when reading gzip-compressed parquet files, btw.

Interesting, what kind of speedup are you observing?

@pitrou
pitrou requested a review from HuaHuaY September 29, 2026 09:06
@lorenzhs

lorenzhs commented Sep 29, 2026 •

Copy link
Copy Markdown
Author

No measurable speedup on files with a few dozen columns and reasonable row group sizes, but I have a workload on super wide files with tiny row groups (1600-ish columns, small file is 7.4MiB / 6 row groups, large file is 116MiB / 30 row groups). There, parquet-scan becomes ~3x faster on the small file (0.34 → 0.11s) and 1.8x faster on the large file (2.75 → 1.5s). I tried playing with --batch_size but it doesn't make much of a difference.

@pitrou pitrou changed the title GH-51308: [C++] initialize gzip codec state lazily GH-51308: [C++] Initialize gzip codec state lazily Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants