Conversation
) ### Rationale for this change `Datum::null_count()` only counts physical nulls, so it gives the wrong answer for union, run-end encoded and dictionary data. `Array`, `ArrayData`, `ArraySpan` and `ChunkedArray` already have `ComputeLogicalNullCount()` for this, but `Datum` doesn't. ### What changes are included in this PR? - `Datum::ComputeLogicalNullCount()`, which delegates to the existing `ArrayData`/`ChunkedArray` implementations for array-like data. - `Scalar::IsLogicalNull()`, which is `!is_valid` for most types. `DictionaryScalar` overrides it because a valid index can still refer to a null dictionary value. Ill-formed scalars (e.g. an out-of-bounds index) are not handled defensively — the result is undefined for them, and validation rejects them. Like the array path, it doesn't recurse into nested values. ### Are these changes tested? Yes, in datum_test.cc and scalar_test.cc: union (sparse and dense), run-end encoded, dictionary and chunked inputs, scalars obtained from `Array::GetScalar()`, and a check that the scalar and array paths agree element by element. ### Are there any user-facing changes? The two new APIs above. The new virtual on `Scalar` changes the C++ ABI, so this shouldn't be backported to a patch release. * GitHub Issue: apache#50338 Authored-by: Rahul Goel <goel.rahul4200@gmail.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…e#50691) ### Rationale for this change This PR continues the simdjson migration by replacing the remaining `ObjectWriter` users with `JsonWriter`. With all usages migrated, the obsolete `ObjectWriter` implementation and its associated build configuration are removed. ### What changes are included in this PR? * Replace `ObjectWriter` with `JsonWriter` in: * `key_material.cc` * `key_metadata.cc` * `local_wrap_kms_client.cc` * `file_system_key_material_store.cc` * Remove the unused `ObjectWriter` implementation (`object_writer.cc` and `object_writer.h`). * Remove `object_writer.h` from the installed headers. * Update the Arrow build configuration to stop building `object_writer.cc`. * Link Parquet against `simdjson::simdjson` since it now includes `json_writer_internal.h`. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. Closes: apache#50690 * GitHub Issue: apache#50690 Lead-authored-by: Aaditya Srinivasan <aadityasri03@gmail.com> Co-authored-by: Antoine Pitrou <antoine@python.org> Signed-off-by: Antoine Pitrou <antoine@python.org>
…ter (apache#50708) ### Rationale for this change This PR continues the simdjson migration by replacing RapidJSON's `Writer` API with `JsonWriter` in extension type serialization. ### What changes are included in this PR? * Replace RapidJSON writer usage with `JsonWriter` in: * `FixedShapeTensorType::Serialize()` * `VariableShapeTensorType::Serialize()` * `OpaqueType::Serialize()` * Remove `rapidjson::Writer` usage from the migrated serializers. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. Closes: apache#50706 * GitHub Issue: apache#50706 Authored-by: Aaditya Srinivasan <aadityasri03@gmail.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…FixedSizeListTestCase (apache#50721) ### Rationale for this change Fix a valgrind error. ### What changes are included in this PR? Add `PrintTo` method to `FixedSizeListTestCase`. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. * GitHub Issue: apache#50718 Authored-by: Zehua Zou <zehuazou2000@gmail.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
### Rationale for this change This change enables building Arrow with `SIMDJSON_EXCEPTIONS=0` by updating the JSON object parser to use simdjson's non-throwing API. ### What changes are included? - Enable `SIMDJSON_EXCEPTIONS=0` for the bundled simdjson dependency. - Replace uses of `simdjson_result::value()` in `ObjectParser` with the non-throwing `get()` API. - Preserve existing error handling by returning Arrow `Status` values on simdjson errors. ### Are these changes tested? Yes. Fixes: apache#50654 * GitHub Issue: apache#50654 Authored-by: Aaditya Srinivasan <aadityasri03@gmail.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
…50717) ### Rationale for this change This should fix a number of C++ and Python CI builds now that simdjson is used for parsing geospatial and modular encryption metadata. ### Are these changes tested? By existing CI builds. ### Are there any user-facing changes? Just a bugfix in the build chain. * GitHub Issue: apache#50716 Authored-by: Antoine Pitrou <antoine@python.org> Signed-off-by: Antoine Pitrou <antoine@python.org>
apache#50705) ### Rationale for this change Fix apache#50636 - `test-r-macos-as-cran` nightly job fails compiling `visit({range_start, range_cur})` introduced in apache#50248. ```console /Users/runner/work/crossbow/crossbow/arrow/cpp/src/arrow/compute/kernels/vector_sort.cc:325:11: note: candidate function not viable: cannot convert initializer list argument to 'std::span<uint64_t>' (aka 'span<unsigned long long>') 325 | [&](std::span<uint64_t> indices) { SortNextColumn(indices, offset); }); ``` The job pins [macOS SDK 11.3](https://github.com/ursacomputing/crossbow/actions/runs/30420027426/job/90474737457#step:9:14), so libc++ there does not have C++20 iterator-pair span constructor available yet ([available with libc++ 14](https://libcxx.llvm.org/Status/Cxx20.html)). Similar problem as in recent apache#50295 ### What changes are included in this PR? Replace std::span iterator-pair constructor with subspan in `vector_sort.cc` Also replace std::ranges in `parquet/arrow/reader.cc` introduced in apache#50271 ### Are these changes tested? Yes, builds locally and crossbow `test-r-macos-as-cran` job succeeds. ### Are there any user-facing changes? No. * GitHub Issue: apache#50636 Authored-by: Tadeja Kadunc <tadeja.kadunc@gmail.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
…apache#50681) ### Rationale for this change We need this when we break our Yum repositories. ### What changes are included in this PR? We can recover Yum repositories too by `dev/release/binary-recover.sh`. ### Are these changes tested? Yes. I recovered our Yum repositories with this. ### Are there any user-facing changes? No. * GitHub Issue: apache#50674 Authored-by: Sutou Kouhei <kou@clear-code.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…nters (apache#50723) ### Rationale for this change The Arrow GDB pretty-printers support `Decimal128` and `Decimal256`, but do not support `Decimal32` or `Decimal64`. ### What changes are included in this PR? * Add GDB pretty-printer support and tests for `Decimal32` and `Decimal64`. * Remove the unused `max_type_id` variable. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. * GitHub Issue: apache#50660 Authored-by: fenfeng9 <fenfeng9@qq.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…rge list (view) and map values (apache#50534) ### Rationale for this change From apache#44183 - creating run-end encoded cols that contain list or struct values is unsupported. Adding support for both the encode and decode paths. ### What changes are included in this PR? - Adding support for `List`, `ListView`, `LargeList`, `LargeListView`, `Map`, `FixedSizeList` and `Struct` values in the `run_end_encode` and the `run_end_decode` paths. - Support list of lists as part of that as well. ### Are these changes tested? - Yes added covering tests ### Are there any user-facing changes? There are but they are not breaking. Currently the encode path throws that it is unsupported on one of these types. We are adding support. * GitHub Issue: apache#44183 Authored-by: Ben Magyar <ben.magyar@depop.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…e effect (apache#50719) ### Rationale for this change I noticed while editing a .pyx file that my changes weren't showing up after a rebuild — I had to build twice. Traced it to `BYPRODUCTS` being commented out in `UseCython.cmake`. ### What changes are included in this PR? I uncommented `BYPRODUCTS ${_generated_files}` in `cpp/cmake_modules/UseCython.cmake`. Without it, CMake doesn't realize the .cpp was updated in the same build pass, so it skips recompiling the .so until the next build. The line was commented out for older CMake compatibility — but the project requires CMake >= 3.25 now, and BYPRODUCTS has worked since 3.2, so that's no longer a concern. ### Are these changes tested? This is a build system fix so there's no unit test for it. The CI builds pyarrow from source and runs the full test suite, which will validate the build still works correctly. ### Are there any user-facing changes? No — this only improves the dev experience when iterating on .pyx files. One build instead of two. * GitHub Issue: apache#50702 Authored-by: Pratyush Adhikari <pratyushadk990@gmail.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
### Rationale for this change The simdjson API has a throwing and non-throwing subset. apache#50672 activated a compiler flag that disables the throwing subset of the API, which broke some CI builds. ### What changes are included in this PR? This changes `json_write_internal.cc` and `from_string.cc` to use the non-throwing simdjson api ### Are these changes tested? Yes ### Are there any user-facing changes? No * GitHub Issue: apache#50730 Authored-by: Alexander Taepper <alexander.taepper@gmail.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…che#50646) ### Rationale for this change Resolves [49305](apache#49305) `RecordBatchFileReader::CountRows` has existed in Arrow C++ (`cpp/src/arrow/ipc/reader.h`) but was never bound in Python, so the only way to get the total number of rows of an IPC file was: ```python num_rows = sum(reader.get_batch(i).num_rows for i in range(reader.num_record_batches)) ``` ### This findings are done in original opened Issue apache#49305 That deserializes every record batch just to read its length, which is wasteful and becomes expensive on remote filesystems. ### What changes are included in this PR? Adds `RecordBatchFileReader.count_rows()`: ```python with pa.ipc.open_file(source) as reader: reader.count_rows() ``` Three changes: * `python/pyarrow/includes/libarrow.pxd`: declare `CResult[int64_t] CountRows()` on `CRecordBatchFileReader`, which was the missing piece. * `python/pyarrow/ipc.pxi`: add `count_rows()` to `_RecordBatchFileReader`, released GIL around the call, with the same closed reader guard used by the existing `stats` property * `python/pyarrow/tests/test_ipc.py`: tests. On the naming: `count_rows()` follows the C++ method and is consistent with the existing `count_rows()` on `Dataset`, `Scanner` and `Fragment`. To be precise about the benefit, since the issue describes it as reading the count from the metadata: the C++ implementation still walks every block, but reads only each record batch's flatbuffer message header to pick up its length, and never touches the data buffers. So this is a reduction in bytes read rather than in the number of reads, and the gain shows up on remote filesystems and on files with large batches rather than in a local in memory benchmark. This is only added to the file reader. The stream reader has no footer and cannot count rows without consuming the stream. ### Are these changes tested? Yes, two tests in `python/pyarrow/tests/test_ipc.py`: * `test_file_count_rows`: count matches the sum of the written batch lengths, and counting does not consume the reader (count, `read_all()`, count again). * `test_file_count_rows_no_batches`: a file with a schema but no batches counts 0. Locally `test_ipc.py` passes (72 tests) and `test_feather.py` passes (83 passed, 8 skipped, 1 xfailed), the latter because the feather reader sits on the same file reader. The docstring example was run and produces the output shown. ### Are there any user-facing changes? Yes, a new public method `RecordBatchFileReader.count_rows()`. No existing behaviour changes. ### AI usage disclosure I used Claude Code to locate the unbound C++ method and the place where the declaration was missing, and to draft the binding, the docstring and the tests. I reviewed the result, rebuilt PyArrow locally, ran the test suites quoted above, and checked the C++ implementation of `CountRows` myself to confirm what it actually does before describing the benefit here. * GitHub Issue: apache#49305 Authored-by: Guja <127162872+GujaLomsadze@users.noreply.github.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…segfaults (apache#50734) ### Rationale for this change Fix apache#50688, per analysis of apache#50712, `arrow-s3fs-test` segfaults because of incompatible aws-sdk-cpp and aws-crt-cpp bottles on current homebrew-core (Homebrew/homebrew-core#295531 bumped aws-crt-cpp to 0.43.0 without rebuilding the aws-sdk-cpp bottle). `brew update` was added in apache#49491 and is no longer needed (grpc/protobuf v34 got fixed in homebrew-core in March Homebrew/homebrew-core@ 552efcae and current runner ships that Homebrew snapshot including that). ### What changes are included in this PR? Remove brew update to avoid incompatible aws-sdk-cpp and aws-crt-cpp bottles. ### Are these changes tested? Yes, fork succeeded on Python and cpp workflows, now CI here succeeds too. ### Are there any user-facing changes? No. * GitHub Issue: apache#50688 Authored-by: Tadeja Kadunc <tadeja.kadunc@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
### Rationale for this change The SVE128 code path has conflict with the SVE256 that we do not yet manage properly. - There was first the ODR violation in apacheGH-49921 - Now it seems that there may also be an issue with LTO Anyhow, after we fixed the inlining issue in Neon, the SVE128 had no clear advantages over Neon as expected, os this was due to be removed anyways. ### What changes are included in this PR? Remove SVE128 unpack ### Are these changes tested? In CI. ### Are there any user-facing changes? No * GitHub Issue: apache#50503 Lead-authored-by: AntoinePrv <AntoinePrv@users.noreply.github.com> Co-authored-by: Antoine Pitrou <antoine@python.org> Signed-off-by: Antoine Pitrou <antoine@python.org>
…th <type_traits> helpers (apache#50714) ### Rationale for this change This is the first part of simplifying functional helpers which are no longer required since the code-base supports more recent C++ versions. (See apache#50713 and apache#50250) ### What changes are included in this PR? This removes the `return_type` related helpers from `functional.h`. Also, the unused helpers `is_overloaded`, `enable_if_empty` and `enable_if_not_empty` are removed. ### Are these changes tested? Yes ### Are there any user-facing changes? No * GitHub Issue: apache#50713 Authored-by: Alexander Taepper <alexander.taepper@gmail.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
… GetMoreResults (apache#50700) ### Rationale for this change Fixes a bug in the implementation of ODBC `GetMoreResults` in the FlightSQL ODBC driver. According to https://learn.microsoft.com/en-us/sql/odbc/reference/appendixes/statement-transitions?view=sql-server-ver17#sqlmoreresults, we should return `SQL_NO_DATA` for some states we previously were throwing another error in. This appears to be exposed by a behavior of only the Windows ODBC driver manager: `GetMoreResults` always gets called even for metadata queries. ### What changes are included in this PR? - Changed implementation and test: `GetMoreResults` now always returns `SQL_NO_DATA`. ### Are these changes tested? Yes, in CI. ### Are there any user-facing changes? No. * GitHub Issue: apache#50578 Authored-by: Bryce Mecum <petridish@gmail.com> Signed-off-by: Bryce Mecum <petridish@gmail.com>
… the C++ FlightServerBase and the Python object to avoid leaking server (apache#50687) ### Rationale for this change PyFlightServer keeps a reference towards the Python server via `OwnedRefNoGIL server_`, the Python server also keeps a reference of the C++ `PyFlightServer` creating a cycle that is never freed during the process lifetime. ### What changes are included in this PR? Create a new `ReleasePythonServerRef` method that is called after any `server.Shutdown` (including at `__exit__`) allowing for the `OwnedRefNoGIL` to be cleared breaking the cycle. This lets normal reference counting free the previously leaked Python object. ### Are these changes tested? Yes, the newly added tests were failing leaking the references before the fix. Currently a test demonstrating a leak when not calling server.Shutdown is added for discussion purposes. ### Are there any user-facing changes? No * GitHub Issue: apache#50684 Authored-by: Raúl Cumplido <raulcumplido@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…yml reference index (apache#50745) ### Rationale for this change Errors due to functions missing from pkgdown docs ### What changes are included in this PR? Adds them ### Are these changes tested? Will run CI ### Are there any user-facing changes? No * GitHub Issue: apache#50744 Authored-by: Nic Crane <thisisnic@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…he#50741) ### Rationale for this change Unlike RapidJSON, simdjson is not header-only and needs to be linked to explicitly when linking against libarrow.a. This should fix the JNI builds on the C++ Extra workflow. ### Are these changes tested? Yes, by existing CI builds. ### Are there any user-facing changes? No. * GitHub Issue: apache#50739 Authored-by: Antoine Pitrou <antoine@python.org> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…_WITH_RE2 is disabled (apache#50754) ### Rationale for this change This change fixes the following build error: ``` /arrow/cpp/src/arrow/compute/kernels/scalar_string_ascii.cc:1753:16: error: unused variable ‘is_utf8’ [-Werror=unused-variable] 1753 | const bool is_utf8 = is_string_or_string_view(batch[0].type()->id()); | ^~~~~~~ ``` The `is_utf8` variable introduced by commit 374db36 is unused when `ARROW_WITH_RE2=OFF` is specified. ```diff static Status Exec(KernelContext* ctx, const ExecSpan& batch, ExecResult* out) { const MatchSubstringOptions& options = MatchSubstringState::Get(ctx); + const bool is_utf8 = is_string_or_string_view(batch[0].type()->id()); if (options.ignore_case) { ARROW_ASSIGN_OR_RAISE(auto matcher, - FindSubstringRegex::Make(options, InputType::is_utf8, true)); - applicator::ScalarUnaryNotNullStateful<OffsetType, InputType, FindSubstringRegex> + FindSubstringRegex::Make(options, is_utf8, true)); + applicator::ScalarUnaryNotNullStateful<OffsetType, InputPhysicalType, + FindSubstringRegex> kernel{std::move(matcher)}; return kernel.Exec(ctx, batch, out); return Status::NotImplemented("ignore_case requires RE2"); } - applicator::ScalarUnaryNotNullStateful<OffsetType, InputType, FindSubstring> kernel{ - FindSubstring(PlainSubstringMatcher(options))}; + applicator::ScalarUnaryNotNullStateful<OffsetType, InputPhysicalType, FindSubstring> + kernel{FindSubstring(PlainSubstringMatcher(options))}; return kernel.Exec(ctx, batch, out); } }; ``` Therefore, I move the declaration of `is_utf8` inside the `#ifdef ARROW_WITH_RE2` block to prevent this error. ### What changes are included in this PR? I move the declaration of `is_utf8` inside the `#ifdef ARROW_WITH_RE2` block to prevent this error. This PR does not includes breaking changes to public APIs. This PR does not contains a "Critical Fix". ### Are these changes tested? Yes. This change only moves the declaration of `is_utf8` and does not change any logic. Therefore, the existing tests introduced by 374db36 should continue to pass. These tests are already covered by CI, and CI passes successfully with this change. No new tests are added because this change only moves a variable declaration and does not affect behavior. I have confirmed that C++ CI checks pass on my fork. See: https://github.com/komainu8/arrow/actions/runs/30619400223/job/91120073626 ### Are there any user-facing changes? No. * GitHub Issue: apache#50752 Authored-by: Horimoto Yasuhiro <horimoto@clear-code.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…pache#50675) ### Rationale for this change The `ceil_temporal`, `floor_temporal`, and `round_temporal` functions currently support date, time, and timestamp inputs, but not duration inputs. ### What changes are included in this PR? - Add kernel registration for duration values with second, millisecond, microsecond, and nanosecond resolutions - Support rounding duration inputs using physical units through day - Treat week as seven physical days for duration inputs - Reject ambiguous calendar units such as month, quarter, and year - Reject `calendar_based_origin` for duration inputs - Add focused C++ tests covering all four duration resolutions, positive and negative values, null propagation, day and week rounding, and unsupported calendar behavior ### Are these changes tested? Yes. - `arrow-compute-scalar-temporal-test`: 55 tests passed - Applicable pre-commit C++ formatting and lint checks passed ### Are there any user-facing changes? Yes. Users can now pass duration values to `ceil_temporal`, `floor_temporal`, and `round_temporal` for supported physical units. ### AI assistance I used ChatGPT to help navigate the codebase and draft the initial implementation. I reviewed, revised, and tested the changes locally. * GitHub Issue: apache#50395 Authored-by: snigdhachoppac <snigchoppa@gmail.com> Signed-off-by: Rok Mihevc <rok@mihevc.org>
…eps.sh (apache#50610) ### Rationale for this change This is the sub issue apache#44748. * SC2027: The surrounding quotes actually unquote this. Remove or escape them. * SC2086: Double quote to prevent globbing and word splitting. * SC2223: This default assignment may cause DoS due to globbing. Quote it. ``` shellcheck ci/scripts/r_deps.sh In ci/scripts/r_deps.sh line 21: : ${R_BIN:=R} ^---------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_deps.sh line 23: : ${R_PRUNE_DEPS:=FALSE} ^--------------------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_deps.sh line 24: R_PRUNE_DEPS=`echo $R_PRUNE_DEPS | tr '[:upper:]' '[:lower:]'` ^-- SC2006 (style): Use $(...) notation instead of legacy backticks `...`. ^-----------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: R_PRUNE_DEPS=$(echo "$R_PRUNE_DEPS" | tr '[:upper:]' '[:lower:]') In ci/scripts/r_deps.sh line 26: : ${R_DUCKDB_DEV:=FALSE} ^--------------------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_deps.sh line 27: R_DUCKDB_DEV=`echo $R_DUCKDB_DEV | tr '[:upper:]' '[:lower:]'` ^-- SC2006 (style): Use $(...) notation instead of legacy backticks `...`. ^-----------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: R_DUCKDB_DEV=$(echo "$R_DUCKDB_DEV" | tr '[:upper:]' '[:lower:]') In ci/scripts/r_deps.sh line 31: pushd ${source_dir} ^-----------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: pushd "${source_dir}" In ci/scripts/r_deps.sh line 33: if [ ${R_PRUNE_DEPS} = "true" ]; then ^-------------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: if [ "${R_PRUNE_DEPS}" = "true" ]; then In ci/scripts/r_deps.sh line 46: ${R_BIN} -e "options(warn=2); install.packages('remotes'); remotes::install_cran(c('glue', 'rcmdcheck', 'sys')); remotes::install_deps(INSTALL_opts = '"${INSTALL_ARGS}"')" ^-------------^ SC2027 (warning): The surrounding quotes actually unquote this. Remove or escape them. ^-------------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: ${R_BIN} -e "options(warn=2); install.packages('remotes'); remotes::install_cran(c('glue', 'rcmdcheck', 'sys')); remotes::install_deps(INSTALL_opts = '""${INSTALL_ARGS}""')" In ci/scripts/r_deps.sh line 49: if [ ${R_DUCKDB_DEV} == "true" ]; then ^-------------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: if [ "${R_DUCKDB_DEV}" == "true" ]; then In ci/scripts/r_deps.sh line 55: ${R_BIN} -e "remotes::install_deps(dependencies = TRUE, INSTALL_opts = '"${INSTALL_ARGS}"')" ^-------------^ SC2027 (warning): The surrounding quotes actually unquote this. Remove or escape them. ^-------------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: ${R_BIN} -e "remotes::install_deps(dependencies = TRUE, INSTALL_opts = '""${INSTALL_ARGS}""')" For more information: https://www.shellcheck.net/wiki/SC2027 -- The surrounding quotes actually u... https://www.shellcheck.net/wiki/SC2086 -- Double quote to prevent globbing ... https://www.shellcheck.net/wiki/SC2223 -- This default assignment may cause... ``` ### What changes are included in this PR? * SC2027: Remove redundant quotes. * SC2086: Quote variable expansions. * SC2223: Quote default variable assignments. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. * GitHub Issue: apache#50609 Lead-authored-by: Hiroyuki Sato <hiroysato@gmail.com> Co-authored-by: Sutou Kouhei <kou@cozmixng.org> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
…50759) ### Rationale for this change Fix apache#50758 - `llvm-toolchain-20` package was [removed from Debian experimental on Jul 14](https://tracker.debian.org/news/1774937/removed-12018-1-from-experimental/), so the nightly `test-debian-experimental-cpp-gcc-15` fails with `E: Unable to locate package clang-20` / `llvm-20-dev`. ### What changes are included in this PR? Update LLVM 20 to 22 on Debian experimental in `dev/tasks/tasks.yml` ### Are these changes tested? Yes, built locally and verified via crossbow. ### Are there any user-facing changes? No. * GitHub Issue: apache#50758 Authored-by: Tadeja Kadunc <tadeja.kadunc@gmail.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
apache#50725) ### Rationale for this change This PR continues the simdjson migration by adding support for serializing `simdjson::ondemand::value` directly with `JsonWriter`. This provides a reusable API for future migration work and avoids requiring callers to implement their own recursive serialization logic. It also introduces a shared helper for dispatching `simdjson::ondemand::value` based on its JSON type, reducing duplicated type dispatch and extraction logic. ### What changes are included in this PR? * Add `JsonWriter::WriteValue(simdjson::ondemand::value)`. * Add `VisitJsonValue` to centralize JSON type dispatch and `simdjson` value extraction. * Recursively serialize: * objects * arrays * strings * booleans * null values * numeric values * Add unit tests covering: * simple objects * nested objects * objects containing arrays * complex nested values * empty objects * Use `simdjson::ondemand::document::get_value()` in tests to obtain the root `ondemand::value` before serialization. ### Are these changes tested? Yes. Added unit tests for `JsonWriter::WriteValue` covering the supported JSON value types and nested structures. ### Are there any user-facing changes? No. Closes: apache#50724 * GitHub Issue: apache#50724 Authored-by: Aaditya Srinivasan <aadityasri03@gmail.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
…ColumnDescriptor` as deprecated (apache#50738) ### Rationale for this change I offer two reasons for marking this method as deprecated: 1. It accepts a `distinct_count` parameter but lacks `has_distinct_count`, making it impossible to represent missing distinct counts, which could lead to misuse. 2. When I try to add `nan_count`, I found that it lacks a `ColumnDescriptor` parameter, making it impossible to obtain the logical type and determine the value of `has_nan_count` based on Parquet's logical type (FLOAT16). ### What changes are included in this PR? Mark `MakeStatistics` method without `ColumnDescriptor` as deprecated. ### Are these changes tested? Yes. ### Are there any user-facing changes? Yes. Mark `MakeStatistics` method without `ColumnDescriptor` as deprecated. * GitHub Issue: apache#50737 Authored-by: Zehua Zou <zehuazou2000@gmail.com> Signed-off-by: Gang Wu <ustcwg@gmail.com>
…ocker_configure.sh (apache#50767) ### Rationale for this change This is the sub issue apache#44748. * SC2006 (style): Use $(...) notation instead of legacy backticks * SC2046: Quote this to prevent word splitting. * SC2086: Double quote to prevent globbing and word splitting. * SC2223: This default assignment may cause DoS due to globbing. Quote it. ``` shellcheck ci/scripts/r_docker_configure.sh In ci/scripts/r_docker_configure.sh line 21: : ${R_BIN:=R} ^---------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_docker_configure.sh line 23: : ${ARROW_SOURCE_HOME:=/arrow} ^--------------------------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_docker_configure.sh line 29: cat ${ARROW_SOURCE_HOME}/ci/etc/rprofile >> $(${R_BIN} RHOME)/etc/Rprofile.site ^------------------^ SC2086 (info): Double quote to prevent globbing and word splitting. ^---------------^ SC2046 (warning): Quote this to prevent word splitting. Did you mean: cat "${ARROW_SOURCE_HOME}"/ci/etc/rprofile >> $(${R_BIN} RHOME)/etc/Rprofile.site In ci/scripts/r_docker_configure.sh line 33: echo "MAKEFLAGS=-j$(${R_BIN} -s -e 'cat(parallel::detectCores())')" >> $(R RHOME)/etc/Renviron.site ^--------^ SC2046 (warning): Quote this to prevent word splitting. In ci/scripts/r_docker_configure.sh line 36: if [ "`which dnf`" ]; then ^---------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: if [ "$(which dnf)" ]; then In ci/scripts/r_docker_configure.sh line 38: elif [ "`which yum`" ]; then ^---------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: elif [ "$(which yum)" ]; then In ci/scripts/r_docker_configure.sh line 40: elif [ "`which zypper`" ]; then ^------------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: elif [ "$(which zypper)" ]; then In ci/scripts/r_docker_configure.sh line 42: elif [ "`which apk`" ]; then ^---------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: elif [ "$(which apk)" ]; then In ci/scripts/r_docker_configure.sh line 50: : ${R_CUSTOM_CCACHE:=FALSE} ^-----------------------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_docker_configure.sh line 51: R_CUSTOM_CCACHE=`echo $R_CUSTOM_CCACHE | tr '[:upper:]' '[:lower:]'` ^-- SC2006 (style): Use $(...) notation instead of legacy backticks `...`. ^--------------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: R_CUSTOM_CCACHE=$(echo "$R_CUSTOM_CCACHE" | tr '[:upper:]' '[:lower:]') In ci/scripts/r_docker_configure.sh line 52: if [ ${R_CUSTOM_CCACHE} = "true" ]; then ^----------------^ SC2086 (info): Double quote to prevent globbing and word splitting. Did you mean: if [ "${R_CUSTOM_CCACHE}" = "true" ]; then For more information: https://www.shellcheck.net/wiki/SC2046 -- Quote this to prevent word splitt... https://www.shellcheck.net/wiki/SC2086 -- Double quote to prevent globbing ... https://www.shellcheck.net/wiki/SC2223 -- This default assignment may cause... ``` ### What changes are included in this PR? * SC2006 Use `$(...)` notation instead of legacy backticks * SC2046: Quote variable to prevent word splitting. * SC2086: Quote variable expansions. * SC2223: Quote default variable assignments. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. * GitHub Issue: apache#50766 Authored-by: Hiroyuki Sato <hiroysato@gmail.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
…nstall_system_dependencies.sh (apache#50772) ### Rationale for this change This is the sub issue apache#44748. * SC2223: This default assignment may cause DoS due to globbing. Quote it. * SC2006: Use $(...) notation instead of legacy backticked `...`. ``` shellcheck r_install_system_dependencies.sh r_install_system_dependencies.sh: r_install_system_dependencies.sh: openBinaryFile: does not exist (No such file or directory) palolovalley:arrow hsato$ shellcheck ci/scripts/r_install_system_dependencies.sh In ci/scripts/r_install_system_dependencies.sh line 22: : ${ARROW_SOURCE_HOME:=/arrow} ^--------------------------^ SC2223 (info): This default assignment may cause DoS due to globbing. Quote it. In ci/scripts/r_install_system_dependencies.sh line 25: if [ "`which dnf`" ]; then ^---------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: if [ "$(which dnf)" ]; then In ci/scripts/r_install_system_dependencies.sh line 27: elif [ "`which yum`" ]; then ^---------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: elif [ "$(which yum)" ]; then In ci/scripts/r_install_system_dependencies.sh line 29: elif [ "`which zypper`" ]; then ^------------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: elif [ "$(which zypper)" ]; then In ci/scripts/r_install_system_dependencies.sh line 31: elif [ "`which apk`" ]; then ^---------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: elif [ "$(which apk)" ]; then In ci/scripts/r_install_system_dependencies.sh line 59: if [ "$ARROW_S3" == "ON" ] && [ -f "${ARROW_SOURCE_HOME}/ci/scripts/install_minio.sh" ] && [ "`which wget`" ]; then ^----------^ SC2006 (style): Use $(...) notation instead of legacy backticks `...`. Did you mean: if [ "$ARROW_S3" == "ON" ] && [ -f "${ARROW_SOURCE_HOME}/ci/scripts/install_minio.sh" ] && [ "$(which wget)" ]; then For more information: https://www.shellcheck.net/wiki/SC2223 -- This default assignment may cause... https://www.shellcheck.net/wiki/SC2006 -- Use $(...) notation instead of le... ``` ### What changes are included in this PR? * SC2006 Use `$(...)` notation instead of legacy backticks `...` * SC2223: Quote default variable assignments. ### Are these changes tested? Yes. ### Are there any user-facing changes? No. * GitHub Issue: apache#50771 Authored-by: Hiroyuki Sato <hiroysato@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…pache#50761) ### Rationale for this change The nightly job `test-fedora-42-python-3` fails with `Cannot uninstall packaging 24.2` `╰─> The package's contents are unknown: no RECORD file was found for packaging.` ### What changes are included in this PR? venv in `linux-dnf-python-3.dockerfile` to install Python requirements there ### Are these changes tested? Yes, crossbow job passes. ### Are there any user-facing changes? No. * GitHub Issue: apache#50760 Authored-by: Tadeja Kadunc <tadeja.kadunc@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
### Rationale for this change The collaborator list defines triage role. It helps having triage people. https://github.com/apache/arrow/commits?author=Reranko05 ### What changes are included in this PR? Add to the list a collaborator that could benefit from triage role. ### Are these changes tested? Not relevant ### Are there any user-facing changes? No Authored-by: Rok Mihevc <rok@mihevc.org> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…XT (apache#51458) ### What changes are included in this PR? Allow applications compiled with `ARROW_EXTRA_ERROR_CONTEXT` enabled to link against libarrow compiled with `ARROW_EXTRA_ERROR_CONTEXT` disabled, and vice-versa. Also improve unit tests slightly. ### Are these changes tested? The original issue was tested manually. ### Are there any user-facing changes? No. ### Was AI used for this PR? In accordance to the [AI generation guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code), please disclose below whether and how AI was used in this PR. **PR code and description written by:** - [x] Human - [ ] AI **Reviewed before submission by:** - [x] Human - [ ] AI - [ ] Not reviewed * GitHub Issue: apache#51456 Authored-by: Antoine Pitrou <antoine@python.org> Signed-off-by: Antoine Pitrou <antoine@python.org>
… updates (apache#51462) Bumps the apache-infrastructure-actions-stash group with 2 updates: [apache/infrastructure-actions/stash/restore](https://github.com/apache/infrastructure-actions) and [apache/infrastructure-actions/stash/save](https://github.com/apache/infrastructure-actions). Updates `apache/infrastructure-actions/stash/restore` from ce952724eb5210790bd5d466d70d5d60ac3e6c21 to 547b55daa5d3c7e45383094ef738b7d16bc8ccc5 <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/apache/infrastructure-actions/commit/547b55daa5d3c7e45383094ef738b7d16bc8ccc5"><code>547b55d</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/170917ae51ffb0b9b500c477a1dac22b65f84ec0"><code>170917a</code></a> Add Google OSS Fuzz actions (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1319">#1319</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/ff5537ba6ba5a324b039c565ebdf2e245158f98f"><code>ff5537b</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/c69277c35974459da3ec9309741f1fdc7b8573ea"><code>c69277c</code></a> action-allowlist-review: bump astral-sh/setup-uv (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1318">#1318</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/5bec58c77de582da4fb4e00d36ddb8a012fd891a"><code>5bec58c</code></a> Remove Expired Refs</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/d0833c6bb44a8c522951ba9afa7ca1ac1fbed92d"><code>d0833c6</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/5bd4332e1cb503de93f84f30ccad378d312572f0"><code>5bd4332</code></a> allowlist: re-add astral-sh/setup-uv v7.0.0 as transitive dep (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1314">#1314</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/fc08b473a466384de1808003384bbaf87df75150"><code>fc08b47</code></a> check-for-transitive-failures: apply grace period before safety net (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1315">#1315</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/56d505ad951bd54a8dd691c90657a21af5cd1d3c"><code>56d505a</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/ceff3e786d2b0181a21f0884f9f2afd76b0cb377"><code>ceff3e7</code></a> action-allowlist-review: bump loadingalias/cargo-rail-action (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1313">#1313</a>)</li> <li>Additional commits viewable in <a href="https://github.com/apache/infrastructure-actions/compare/ce952724eb5210790bd5d466d70d5d60ac3e6c21...547b55daa5d3c7e45383094ef738b7d16bc8ccc5">compare view</a></li> </ul> </details> <br /> Updates `apache/infrastructure-actions/stash/save` from ce952724eb5210790bd5d466d70d5d60ac3e6c21 to 547b55daa5d3c7e45383094ef738b7d16bc8ccc5 <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/apache/infrastructure-actions/commit/547b55daa5d3c7e45383094ef738b7d16bc8ccc5"><code>547b55d</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/170917ae51ffb0b9b500c477a1dac22b65f84ec0"><code>170917a</code></a> Add Google OSS Fuzz actions (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1319">#1319</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/ff5537ba6ba5a324b039c565ebdf2e245158f98f"><code>ff5537b</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/c69277c35974459da3ec9309741f1fdc7b8573ea"><code>c69277c</code></a> action-allowlist-review: bump astral-sh/setup-uv (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1318">#1318</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/5bec58c77de582da4fb4e00d36ddb8a012fd891a"><code>5bec58c</code></a> Remove Expired Refs</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/d0833c6bb44a8c522951ba9afa7ca1ac1fbed92d"><code>d0833c6</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/5bd4332e1cb503de93f84f30ccad378d312572f0"><code>5bd4332</code></a> allowlist: re-add astral-sh/setup-uv v7.0.0 as transitive dep (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1314">#1314</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/fc08b473a466384de1808003384bbaf87df75150"><code>fc08b47</code></a> check-for-transitive-failures: apply grace period before safety net (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1315">#1315</a>)</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/56d505ad951bd54a8dd691c90657a21af5cd1d3c"><code>56d505a</code></a> Sync actions.yml, composite action, and approved_patterns.yml</li> <li><a href="https://github.com/apache/infrastructure-actions/commit/ceff3e786d2b0181a21f0884f9f2afd76b0cb377"><code>ceff3e7</code></a> action-allowlist-review: bump loadingalias/cargo-rail-action (<a href="https://redirect.github.com/apache/infrastructure-actions/issues/1313">#1313</a>)</li> <li>Additional commits viewable in <a href="https://github.com/apache/infrastructure-actions/compare/ce952724eb5210790bd5d466d70d5d60ac3e6c21...547b55daa5d3c7e45383094ef738b7d16bc8ccc5">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@ dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@ dependabot rebase` will rebase this PR - `@ dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@ dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@ dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@ dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@ dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@ dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@ dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
…pache#51468) ### Rationale for this change Add support for LLVM 23.1. My mistake. In the previous PR, I used the `ninja-debug-gandiva` CMake preset but didn't notice that some Parquet tests weren't being compiled; consequently, certain compilation errors within the Parquet component went undetected. This time I used the `ninja-debug` preset. I'm not sure if there are compilation errors in other presets—I don't use them often, and CI doesn't check them either—so we'll just fix them if we run into them. ### What changes are included in this PR? Remove static from some template functions in header files because LLVM 23 added `-Wunused-template` to `-Wall`. ### Are these changes tested? Yes. ### Are there any user-facing changes? Yes. ### Was AI used for this PR? In accordance to the [AI generation guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code), please disclose below whether and how AI was used in this PR. **PR code and description written by:** - [x] Human - [ ] AI **Reviewed before submission by:** - [x] Human - [ ] AI - [ ] Not reviewed * GitHub Issue: apache#51245 Authored-by: Zehua Zou <zehuazou2000@gmail.com> Signed-off-by: Zehua Zou <zehuazou2000@gmail.com>
… Datasets" using "date" expression rigth hand side of a filter (apache#51291) ### Rationale for this change Error when user tries to use local variable in dplyr pipeline in Arrow ### What changes are included in this PR? Make sure we attach them to the mask ### Are these changes tested? Yup ### Are there any user-facing changes? Yep * GitHub Issue: apache#39688 Authored-by: Nic Crane <thisisnic@gmail.com> Signed-off-by: Nic Crane <thisisnic@gmail.com>
### Rationale for this change Boolean values are bit-packed. When copying values from a sliced boolean array, the copy path did not include the array's slice offset. This caused `fill_null_forward`, `fill_null_backward`, and `replace_with_mask` to read values from earlier positions in the parent array. ### What changes are included in this PR? - Apply `ArraySpan::offset` to the bit index when copying boolean values from an `ArraySpan`. - Add regression coverage for sliced boolean inputs in `replace_with_mask`, `fill_null_forward`, and `fill_null_backward`. ### Are these changes tested? Yes. I ran `arrow-compute-vector-test` locally. ### Are there any user-facing changes? Only a bugfix. ### This PR contains a "Critical Fix". This fixes a bug that caused compute operations on sliced boolean arrays with a non-zero offset to produce incorrect values. * GitHub Issue: apache#51223 Authored-by: Abdul Azeem Makarim <114302821+A-makarim@users.noreply.github.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
…ls (apache#50269) ### Rationale for this change I was looking into apache#49889 and updated the handling of logical nulls in the `is_valid`, `is_null`, and `true_unless_null` kernels. It turns out there is some prior work here that I didn't see before I starting implementing my fix (apache#35058, apache#35036, apache#37642). For some reason, that work has stalled. ### What changes are included in this PR? This PR updates the `IsValidExec`, `IsNullExec`, and `TrueUnlessNullExec` kernels so that they call out to a new function `SetLogicalNullBits`. The `SetLogicalNullBits` function matches the logic in `ArraySpan::ComputeLogicalNullCount()`. The logic is similar enough that they may be some opportunities for code deduplication in `dict_util.cc` and `ree_util.cc`. ### Are these changes tested? Yes ### Are there any user-facing changes? No * GitHub Issue: apache#49889 Authored-by: Hadrian <hadrian@uchicago.edu> Signed-off-by: Antoine Pitrou <antoine@python.org>
… build (apache#51469) ### Rationale for this change After: - apache#49679 The nightly CI job for [test-r-macos-as-cran](https://github.com/ursacomputing/crossbow/actions/runs/35804144934/job/107001010915) failed due to ranges file not found: ``` [ 95%] Building CXX object src/arrow/CMakeFiles/arrow_compute_objlib.dir/compute/kernels/vector_search_sorted.cc.o /Users/runner/work/crossbow/crossbow/arrow/cpp/src/arrow/compute/kernels/vector_search_sorted.cc:23:10: fatal error: 'ranges' file not found 23 | #include <ranges> | ^~~~~~~~ 1 error generated. ``` ### What changes are included in this PR? Remove `#include <ranges>` as is never used on the file. ### Are these changes tested? Yes via CI ### Are there any user-facing changes? No ### Was AI used for this PR? In accordance to the [AI generation guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code), please disclose below whether and how AI was used in this PR. **PR code and description written by:** - [x] Human - [ ] AI **Reviewed before submission by:** - [x] Human - [ ] AI - [ ] Not reviewed * GitHub Issue: apache#51467 Authored-by: Raúl Cumplido <raulcumplido@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…che#51274) ### Rationale for this change apacheGH-50382 requests explicit value constructors for all existing `ArrowFormat::Type` subclasses before the broader builder API is exposed. Fixed-size lists still require callers to manually construct the parent validity buffer and flattened child array. This is the focused fixed-size-list prerequisite requested in the apacheGH-50382 maintainer discussion. It is distinct from variable-size `ListArray` support in apacheGH-51262. ### What changes are included in this PR? - Add `ArrowFormat::FixedSizeListArray.new(type, values)` while preserving the existing low-level four-argument constructor. - Build the parent validity bitmap and the child array from nested Ruby values. - Preserve exactly `type.size` child slots for null parent lists. - Reject non-null lists whose size differs from the declared fixed size. - Delegate child construction to the declared child field type. - Add focused coverage for typed construction, parent and child nulls, invalid sizes, empty input, and the low-level constructor. ### Are these changes tested? Yes. ```console RUBYLIB=/opt/homebrew/lib/ruby/gems/4.0.0/gems/red-arrow-25.0.1/lib GI_TYPELIB_PATH=/opt/homebrew/lib/girepository-1.0 bundle exec ruby test/run.rb test-fixed-size-list-array.rb # 6 tests, 12 assertions, 0 failures, 0 errors RUBYLIB=/opt/homebrew/lib/ruby/gems/4.0.0/gems/red-arrow-25.0.1/lib GI_TYPELIB_PATH=/opt/homebrew/lib/girepository-1.0 bundle exec rake test # 696 tests, 705 assertions, 0 failures, 0 errors ``` I also wrote an Arrow IPC file containing `[[1, 2], nil, [3, nil]]` with the new constructor and loaded it with the native Arrow reader; the values round-tripped unchanged. The tests use the current `red-arrow-format` sources with the locally installed Arrow 25.0.1 native extension. Upstream CI remains authoritative for the matching main-branch native runtime. ### Are there any user-facing changes? Yes. `ArrowFormat::FixedSizeListArray` gains a two-argument values constructor. The existing low-level constructor remains supported. ### AI assistance disclosure OpenAI Codex assisted with issue research, implementation, test generation, validation commands, and drafting this pull request. The submitted behavior is supported by the focused regression, full package suite, and IPC round-trip results above. * GitHub Issue: apache#51273 Lead-authored-by: Yifan Chen <emecii23@gmail.com> Co-authored-by: Yifan Chen <30335308+emecii@users.noreply.github.com> Signed-off-by: Sutou Kouhei <kou@clear-code.com>
apache#51160) ### Rationale for this change `Array.slice`, `ChunkedArray.slice`, `RecordBatch.slice`, and `Table.slice` accept Python and NumPy integer-like values but reject Arrow integer scalars, even though those scalars implement Python's integer-index protocol. This makes values produced by Arrow APIs unnecessarily unusable as slice offsets and lengths. ### What changes are included in this PR? Normalize non-`None` offsets and lengths with `operator.index()` at the four public Python wrapper boundaries before the existing validation and C++ calls. Add regression coverage for all signed and unsigned Arrow integer scalar types across the four wrappers, plus NumPy compatibility, invalid and null scalars, negative values, offset clamping, and int64 overflow behavior. ### Are these changes tested? Yes. Current Arrow C++ and editable PyArrow were built from source in an isolated Python 3.13 environment. The focused new tests and existing slice controls passed: 48 passed. Cython translation, Python test-file compilation, Flake8, and `git diff --check` also passed. ### Are there any user-facing changes? Yes. The four Python `.slice()` methods now accept non-null Arrow integer scalars for offsets and lengths. Existing Python/NumPy integer behavior and error behavior for unsupported values are preserved. ### AI assistance AI assisted analysis, implementation, test drafting, build work, and review. The account holder reviewed and approved the final four-file diff. The account holder did not personally run the commands reported above. * GitHub Issue: apache#51145 Authored-by: YusefSyed <211442445+YusefSyed@users.noreply.github.com> Signed-off-by: AlenkaF <frim.alenka@gmail.com>
… version from Homebrew (apache#51403) ### Rationale for this change Due to homebrew dropping support for building bottles for macOS Intel several CI jobs started failing. Some other arm64 related jobs also failed because old bottles are also not built. There's a ML discussion about dropping support for macOS Intel, see for more details: - https://lists.apache.org/thread/l3vmknbjhwf60crrfp3hwb7p5m4qjw7z ### What changes are included in this PR? Bump macOS arm runners to macos-26 and macOS Intel runners (that don't require homebrew bottles) to macos-26-intel. Drop the macOS Intel source verification jobs that require Homebrew bottles in order to run. ### Are these changes tested? Yes on CI. The only untested ones are the bumps on binary verification. ### Are there any user-facing changes? No ### Was AI used for this PR? In accordance to the [AI generation guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code), please disclose below whether and how AI was used in this PR. **PR code and description written by:** - [x] Human - [ ] AI **Reviewed before submission by:** - [x] Human - [ ] AI - [ ] Not reviewed * GitHub Issue: apache#51402 Authored-by: Raúl Cumplido <raulcumplido@gmail.com> Signed-off-by: Raúl Cumplido <raulcumplido@gmail.com>
…e#51357) ### Rationale for this change Parquet null-count statistics can be incorrect for fixed-width leaf columns nested below a repeated ancestor, such as `list<struct<...>>`. When a list is null or empty, its descendant leaf does not produce a value in the leaf values buffer. `DefLevelsToBitmap` intentionally excludes these repeated-ancestor nulls from its `null_count`. `MaybeCalculateValidityBits` was using that value directly for the Parquet column statistics, causing the null count to be undercounted. For example, a `list<struct<string, int32>>` column can report an incorrect `null_count` for the `int32` leaf even though the encoded definition levels correctly represent the null and empty list entries. ### What changes are included in this PR? * Update `MaybeCalculateValidityBits` to calculate the total null count from `batch_size - out_values_to_write`. * This includes nulls represented by repeated ancestors, such as null or empty lists, while preserving the existing validity bitmap behavior. * Add a regression test covering a fixed-width `int32` leaf under `list<struct<...>>`. ### Are these changes tested? Yes. Added a regression test that verifies the `int32` leaf reports: * `null_count = 3` * `num_values = 2` The focused regression test passes: `StatisticsTest.FixedWidthLeafUnderListStructNullCount` The existing `parquet-writer-test` also passes. ### Are there any user-facing changes? Yes. This fixes incorrect Parquet column statistics for affected nested fixed-width columns. The change does not alter the encoded data or public APIs. ### This PR contains a "Critical Fix". This fixes a bug that produces incorrect Parquet statistics. The underlying data remains correct, but the reported `null_count` for affected fixed-width leaf columns can be incorrect. ### Was AI used for this PR? **PR code and description written by:** * [x] Human * [ ] AI **Reviewed before submission by:** * [x] Human * [ ] AI * [ ] Not reviewed * GitHub Issue: apache#51097 Lead-authored-by: AnuragRaut08 <anuragtraut2003@gmail.com> Co-authored-by: Gang Wu <ustcwg@gmail.com> Signed-off-by: Gang Wu <ustcwg@gmail.com>
Add a canonical extension type for bounded ranges (mathematical intervals),
distinct from Arrow's calendar Interval (duration) type.
- Spec: docs/source/format/CanonicalExtensions.rst adds the Range section.
Storage is Struct<lower, upper> with both bounds nullable (null = +/-infinity,
treated as exclusive). A closed parameter (left/right/both/neither, pandas
vocabulary) is carried as JSON extension metadata; the subtype is read from
storage. Disambiguates from the calendar Interval type per DB convention
(INTERVAL = duration, RANGE/PERIOD = bounded set).
- C++ reference impl: cpp/src/arrow/extension/range.{h,cc} (RangeType/RangeArray)
with serialize/deserialize, storage validation, registration in the global
registry, tests, and CMake/meson wiring.
The closedness is no longer defaulted on the wire: empty metadata or a JSON object without a "closed" key is now rejected by Deserialize, so a serialized arrow.range is always unambiguous. The C++ convenience default argument for constructing a RangeType in code is left-closed ([lower, upper)), matching the PostgreSQL/Rust/Python range convention. Spec and tests updated.
Verified by building the arrow-canonical-extensions-test target (50/50 pass, 10/10 RangeType). Two fixes to the previously-uncompiled test: - include arrow/array/array_nested.h for the full StructArray definition (it is only forward-declared in type_fwd.h). - wrap the CheckDeserialize helper in an anonymous namespace to avoid a link-time collision with the identically named helper in opaque_test.cc.
Add a sibling canonical extension type to arrow.range that stores bound
inclusivity per value via non-nullable boolean lower_inc/upper_inc fields,
storage Struct<lower:T, upper:T, lower_inc:bool, upper_inc:bool>.
arrow.range carries a single type-level closed parameter, sufficient for
discrete ranges that canonicalize to one closedness (int4range, int8range,
daterange). Continuous ranges (numrange, tsrange, tstzrange) cannot be
canonicalized, so closedness must travel with each value. arrow.range_inc
mirrors PostgreSQL's internal range representation for that case; both types
coexist.
The type has no metadata parameters: inclusivity lives in storage, so
Serialize emits {} and Deserialize accepts empty/{}/extra keys. A null
(infinite) bound is always exclusive regardless of its flag.
Covers C++ (type, array, registration, tests), pyarrow bindings and tests,
and the format spec, status table, and C++/Python API docs.
Under CMAKE_UNITY_BUILD (Windows CI), range_test.cc and opaque_test.cc are merged into one translation unit. Both declared a CheckDeserialize helper (range's in an anonymous namespace, opaque's in namespace arrow), making the unqualified call ambiguous and failing the MSVC build with C2668. Rename the range helper to CheckRangeDeserialize to remove the collision.
Use JsonWriter for Serialize and the simdjson DOM helpers for Deserialize, matching the other canonical extension types.
Name the pair after arrow.fixed_shape_tensor and arrow.variable_shape_tensor: closedness is either one type parameter or stored per value. - arrow.range -> arrow.fixed_closedness_range - arrow.range_inc -> arrow.variable_closedness_range - C++: RangeType/RangeArray -> FixedClosednessRangeType/Array, RangeIncType/RangeIncArray -> VariableClosednessRangeType/Array, range()/range_inc() -> fixed_closedness_range()/variable_closedness_range() - pyarrow: range_/range_inc and the Range* classes follow the same names
Allow any orderable type as the range subtype and compare bounds with the order of that type; the spec defines only the storage layout. State that only null means unbounded and that all empty ranges denote the same set. Add string subtype cases to the C++ and Python tests.
__reduce__ rebuilt both range types from their parameters only, so a type with non-nullable bounds came back nullable after unpickling. Rebuild them from the storage type through the C++ Deserialize instead. This also keeps a type where only one bound is nullable, and tests cover both cases.
hoeze-minion
force-pushed
the
feat/arrow-range-extension
branch
from
September 26, 2026 21:38
5ff2abd to
ffcfd4d
Compare
The emptiness rule compared lower and upper without saying what happens with a null bound. A null bound means unbounded and cannot be compared, so the rule now applies only when both bounds are non-null.
_fixed_closedness_range_from_storage put closed into the JSON metadata without escaping. It now builds its prototype with closed, so fixed_closedness_range() rejects any value other than left, right, both and neither first.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Thanks for opening a pull request!
If this is your first pull request you can find detailed information on how to contribute here:
Please remove this line and the above text before creating your pull request.
Rationale for this change
What changes are included in this PR?
Are these changes tested?
Are there any user-facing changes?
This PR includes breaking changes to public APIs. (If there are any breaking changes to public APIs, please explain which changes are breaking. If not, you can remove this.)
This PR contains a "Critical Fix". (If the changes fix either (a) a security vulnerability, (b) a bug that caused incorrect or invalid data to be produced, or (c) a bug that causes a crash (even when the API contract is upheld), please provide explanation. If not, you can remove this.)