Skip to content

[Format] Add canonical extension types for bounded ranges #50027

Description

@Hoeze

Describe the enhancement requested

Arrow has no canonical way to represent a bounded range (a mathematical interval with a lower and an upper endpoint), e.g. a numeric range [0, 10), a date range, or a timestamp period. Today such data is modeled ad hoc with two separate columns or with system-specific extension types, which hurts interoperability. A canonical range type will be useful to libraries like Pandas, Polars/Polars-bio, IRanges/PyRanges, database connectors, ...

Note this is distinct from Arrow's existing calendar Interval type (INTERVAL_MONTHS / INTERVAL_DAY_TIME / INTERVAL_MONTH_DAY_NANO), which represents a duration (a signed amount of time), not a bounded set. Databases like PostgreSQL make the same distinction: SQL uses INTERVAL for durations and RANGE / PERIOD for bounded sets. This proposal follows that convention by naming the types arrow.fixed_closedness_range and arrow.variable_closedness_range.

Proposed design:

Two extension types, named after the existing arrow.fixed_shape_tensor / arrow.variable_shape_tensor pair:

  • arrow.fixed_closedness_range: the closedness is one type parameter shared by all values. This fits discrete ranges (e.g. PostgreSQL's int4range or daterange) and pandas intervals.
  • arrow.variable_closedness_range: the closedness is stored per value. Continuous ranges (e.g. PostgreSQL's numrange or tstzrange) need this, because one column can hold both [1, 5] and (1, 5).

Common to both types:

  • Both storage types start with the fields lower: T and upper: T. When a bound field is nullable, a null bound represents an unbounded (infinite) endpoint. Only null means unbounded.
    • Field names lower / upper (PostgreSQL convention) are chosen deliberately for ordering clarity. (Note that Pandas uses left / right for the field names)
    • The subtype T may be any orderable Arrow type (the numeric, temporal and decimal families, etc.). Nested or non-comparable types are out of scope. The spec defines only the storage layout; bounds are compared with the order of T.
  • A range with two non-null bounds is empty implicitly when lower > upper, or when lower == upper with at least one bound exclusive. A range with lower > upper is therefore valid (it denotes the empty set), not an error. All empty values denote the same empty set.

arrow.fixed_closedness_range:

  • Storage type: Struct<lower: T, upper: T>.
  • Metadata: a JSON object {"closed": "..."}.
    • Parameter closed: one of left, right, both, neither (pandas vocabulary; left = lower inclusive / upper exclusive, etc.).
    • closed is required on the wire so a serialized arrow.fixed_closedness_range is always unambiguous. Unknown JSON keys are ignored for forward compatibility.

arrow.variable_closedness_range:

  • Storage type: Struct<lower: T, upper: T, lower_inc: bool, upper_inc: bool>. The two flags are non-nullable and record per value whether each bound is inclusive.
  • Metadata: an empty JSON object {}, because the type has no parameters.

Relation to pandas

This mirrors pandas' interval support and deliberately reuses its vocabulary:

  • pandas.Interval is the scalar form: an immutable bounded interval whose closed parameter takes exactly left, right, both, or neither; the vocabulary adopted here for the closed metadata.
  • pandas.IntervalIndex / pandas.arrays.IntervalArray is the columnar form: it stores parallel left and right bound arrays, directly analogous to the proposed Struct<lower, upper> storage.
  • Crucially, closed is part of pandas' dtype itself (interval[T, left] and interval[T, right] are distinct dtypes), so a typed interval column carries exactly one closed: constructing an array from intervals with differing closedness raises ValueError, and concatenating columns of differing closedness falls back to untyped object dtype. This uniform, type-level closed maps one-to-one onto the closed parameter of arrow.fixed_closedness_range; no per-element closedness is required.

So arrow.fixed_closedness_range would give pandas' IntervalArray / IntervalIndex a natural, lossless Arrow representation for round-tripping.

Component(s)

Format

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions