Skip to content

ENH Use 3-way partitioning only on duplicate-heavy subarrays when sorting in trees - #17

Open
cakedev0 wants to merge 4 commits into
mainfrom
trees-opt/sort-adaptive-3way-main
Open

cakedev0 wants to merge 4 commits into
mainfrom
trees-opt/sort-adaptive-3way-main

Conversation

@cakedev0

@cakedev0 cakedev0 commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Reference Issues/PRs

Benchmarked with probabl-ai/scikit-learn-benchmarks#121.

What does this implement/fix? Explain your changes.

The trees sort with simultaneous_sort(..., use_three_way_partition=True): the feature values in the dense and sparse partitioners, y in the MAE criterion and the categories means. The 3-way partitioning is much faster when there are many duplicated values (up to 10-20x on constant subarrays), but ~20-25% slower than the 2-way partitioning when the values are distinct, which is the common case for continuous features.

This PR replaces the 3-way variant of simultaneous_sort by a mixed one: a 2-way introsort that switches to a 3-way partition at the recursion levels where the median-of-3 pivot sample contains duplicates, and uses it for all the sorts in trees. This check is local to each subarray: continuous features get the 2-way speed, while duplicate-heavy subarrays (constant, low-cardinality or zero-inflated features, deep nodes) still get the 3-way one.

The variant is selected with a SortPartitioning enum (TWO_WAY, MIXED) instead of the use_three_way_partition boolean. TWO_WAY stays the default, used by the neighbors and pairwise distances reductions as on main. The pure 3-way introsort is removed since it has no caller anymore, and both introsorts share the 2-way partition loop.

AI usage disclosure

I used AI assistance for:

  • Code generation
  • Test/benchmark generation
  • Research and understanding

Benchmarks

TLDR: RandomForest fit is ~10% faster on average, up to 1.16x on real datasets (fraud, year_prediction_msd) and up to 1.7x on synthetic data with continuous features. No slowdown beyond noise.

Single-threaded (n_estimators=4, n_jobs=1) fit speedup vs main, 45 RandomForest cases from the all_models config:

intel-laptop intel-gnr
geomean 1.11 1.07
min / max 0.994 / 1.72 0.944 / 1.46
fraud 1.17 1.13
year_prediction_msd 1.16 1.12
medical_charges_nominal 1.06 1.04
california_housing 1.05 1.06
kddcup09_churn 1.05 1.00
covtype 1.03 1.02
bank_marketing 1.01 1.00
amazon_employee_access 1.01 0.99
ames_housing 1.00 1.01
synthetic binary features 0.99 to 1.01 0.94 to 1.05

The binary features are the least favorable case, but it should be noise: with only 2 distinct values, the median-of-3 sample always contains duplicates, so this PR does the same 3-way partitioning as main (plus up to 3 swaps for the pivot selection). Locally, these cases are identical on both branches.

Multi-threaded (default config) results are noisier: geomean 1.10-1.12 over two runs on intel-laptop, 0.99-1.03 over two runs on intel-gnr (which runs 344 threads on these cases).

Other tree sorts switched to MIXED (single-threaded fit time vs main, best of 5):

  • MAE criterion (DecisionTreeRegressor(criterion="absolute_error"), 20k samples): 1.12 with a continuous y, 1.10 with 50% zeros in y, 1.06 with a 5-valued y
  • sparse partitioner (100k x 50, 10% density): 1.01-1.02 with continuous values, 0.99-1.00 with binary or small counts values

Micro-benchmarks of the sort alone, time relative to the fastest of 2-way/3-way for each case (lower is better):

data always 2-way always 3-way (main) this PR
distinct values 1.00 1.21-1.29 0.98-1.00
1 to 5 distinct values 1.9-23 1.00 0.97-1.12
80-95% zeros, rest continuous 1.3-6.9 1.00 0.79-1.07

Any other comments?

Least favorable cases found (sort alone, 1M values): MIXED is as slow as TWO_WAY when the introsort depth limit falls back to heapsort, which happens on some nearly sorted inputs (e.g. a sorted array with ~0.1-1% of random swaps). There, it can be up to ~1.7x slower than THREE_WAY. I think it's acceptable for trees: there's rarely more than one nearly sorted feature (e.g. a timestamp), and y or the categories means are unlikely to be nearly sorted. Picking the median-of-3 candidates away from the subarray ends (at n/4 and 3n/4) would avoid it, but it's out of scope here.

🤖 Generated with Claude Code

cakedev0 and others added 4 commits September 28, 2026 05:00
…sorts

The 2-way introsort now switches to a 3-way partition at levels where the
median of 3 pivot sample contains duplicates. The dense partitioner uses it
for numerical features instead of always using 3-way partitioning.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
simultaneous_sort takes a SortPartitioning enum (TWO_WAY, THREE_WAY or
MIXED) instead of a boolean. The 2-way introsort is restored as in main
and the one switching to 3-way partitioning on duplicate-heavy subarrays
becomes the MIXED one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sparse partitioner, the MAE criterion and the sort of the categories
means also use MIXED instead of THREE_WAY. The three introsorts share the
2-way and 3-way partition helpers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant