You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Closing an output stream always commits what was written, so a job that fails halfway through still leaves a partial object on S3. Arrow C++ already has Abort() for this, and the S3 stream implements it by aborting the multipart upload, but it wasn't reachable from Python.
What changes are included in this PR?
This adds NativeFile.abort(), which calls FileInterface::Abort() the same way close() calls Close(). The example from the issue now works:
It also changes the S3 output stream so that a failed abort still closes it. If the AbortMultipartUpload request failed, Abort() returned the error but left the stream open, so the close() from the with block (or the destructor) went on to complete the upload with the partial data. That's easy to run into, since s3:AbortMultipartUpload is a separate IAM permission: with a MinIO user that only had s3:PutObject, the partial object was committed anyway. It also didn't match the Abort() docs in interfaces.h, which say the stream is closed afterwards. The error is still returned, and the worst case is now an incomplete multipart upload, which a lifecycle rule can clean up.
Are these changes tested?
Yes. test_open_output_stream_abort runs on all the filesystem fixtures, with and without compression and buffering. It checks that nothing is written on S3, and that the mock filesystem sees an abort rather than a close. test_s3_output_stream_abort_after_part_upload aborts after a 10 MiB part has already been uploaded, which I don't think the C++ tests cover. test_s3_output_stream_failed_abort uses a MinIO user that isn't allowed to abort multipart uploads, and checks that the error is raised and that neither the with block nor the destructor completes the upload. There are also two small tests in test_io.py for aborting in-memory streams.
Backends other than S3 don't discard anything on abort today so the tests are pretty light for the other backends. I ran it against minio, azurite and the GCS testbench. Local, GCS and fsspec keep the written data, since their Abort() just closes. Azure leaves an empty blob, because the blob is created when the stream is opened. Those seem worth separate issues.
test_io.py and test_fs.py pass locally with S3, Azure and GCS enabled, and so does arrow-s3fs-test.
Are there any user-facing changes?
Yes, NativeFile.abort() is new.
Was AI used for this PR?
In accordance to the AI generation guidelines, please disclose below whether and how AI was used in this PR.
I used Claude Code to write the code, the tests and this description, and to run the tests locally.
Wait for pending uploads before aborting multipart upload
cpp/src/arrow/filesystem/s3fs.cc:1754
Abort() can race with outstanding multipart uploads. With the default background_writes=True, a 10 MiB write queues UploadPart and returns, but this immediately sends AbortMultipartUpload; an in-flight part may finish after the abort and leave multipart-upload storage behind. The new test avoids this path by calling flush(). Wait for upload_state_->pending_uploads_completed (without committing current_part_) before sending the abort, and still issue the abort if an upload failed.
Regarding Copilot's last comment: it seems to me that the background-upload race is Abort() pre-existing behaviour. Waiting for pending uploads would mean uploading everything already queued, which would be slow(er) and might defeat the point of aborting.
Skipping queued uploads when Abort() has been called could be a better fix, but I'm not sure this should be scoped in this PR?
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Closing an output stream always commits what was written, so a job that fails halfway through still leaves a partial object on S3. Arrow C++ already has
Abort()for this, and the S3 stream implements it by aborting the multipart upload, but it wasn't reachable from Python.What changes are included in this PR?
This adds
NativeFile.abort(), which callsFileInterface::Abort()the same wayclose()callsClose(). The example from the issue now works:It also changes the S3 output stream so that a failed abort still closes it. If the
AbortMultipartUploadrequest failed,Abort()returned the error but left the stream open, so theclose()from thewithblock (or the destructor) went on to complete the upload with the partial data. That's easy to run into, sinces3:AbortMultipartUploadis a separate IAM permission: with a MinIO user that only hads3:PutObject, the partial object was committed anyway. It also didn't match theAbort()docs ininterfaces.h, which say the stream is closed afterwards. The error is still returned, and the worst case is now an incomplete multipart upload, which a lifecycle rule can clean up.Are these changes tested?
Yes.
test_open_output_stream_abortruns on all the filesystem fixtures, with and without compression and buffering. It checks that nothing is written on S3, and that the mock filesystem sees an abort rather than a close.test_s3_output_stream_abort_after_part_uploadaborts after a 10 MiB part has already been uploaded, which I don't think the C++ tests cover.test_s3_output_stream_failed_abortuses a MinIO user that isn't allowed to abort multipart uploads, and checks that the error is raised and that neither thewithblock nor the destructor completes the upload. There are also two small tests intest_io.pyfor aborting in-memory streams.Backends other than S3 don't discard anything on abort today so the tests are pretty light for the other backends. I ran it against minio, azurite and the GCS testbench. Local, GCS and fsspec keep the written data, since their
Abort()just closes. Azure leaves an empty blob, because the blob is created when the stream is opened. Those seem worth separate issues.test_io.pyandtest_fs.pypass locally with S3, Azure and GCS enabled, and so doesarrow-s3fs-test.Are there any user-facing changes?
Yes,
NativeFile.abort()is new.Was AI used for this PR?
In accordance to the AI generation guidelines, please disclose below whether and how AI was used in this PR.
I used Claude Code to write the code, the tests and this description, and to run the tests locally.
PR code and description written by:
Reviewed before submission by: