GH-51548: [Python] Fix silent corruption in non-native-endian dataframe imports - #51557
Open
tam3tamtam wants to merge 1 commit into
Open
tam3tamtam wants to merge 1 commit into
tam3tamtam wants to merge 1 commit into
Conversation
…dian columns in from_dataframe
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Fixes #51548
The dataframe interchange protocol includes byte order in each dtype description. The PyArrow importer currently wraps these buffers directly, so non-native-endian values can be interpreted incorrectly and silently corrupted.
What changes are included in this PR?
The importer now converts non-native-endian data and string offset buffers to native byte order when needed, using
array.array's C-levelbyteswap().If that conversion would require a copy and
allow_copy=False, it raises aRuntimeError.Regression tests cover numeric values, string offsets, sentinel nulls, 16-bit
and 64-bit values, and the
allow_copy=Falseand invalid-buffer error paths.Are these changes tested?
Yes. The focused tests for these changes passed:
Are there any user-facing changes?
Non-native-endian columns are now imported with their correct values. Importing
them may copy their buffers; with
allow_copy=False, conversion raises an error when a copy is required.This PR contains a "Critical Fix". Non-native-endian values could previously be interpreted with the wrong byte order and silently corrupted during dataframe interchange imports. This change preserves the producer's values.
Was AI used for this PR?
In accordance to the AI generation guidelines, please disclose below whether and how AI was used in this PR.
PR code and description written by:
Reviewed before submission by: