Conversation
…sorts The 2-way introsort now switches to a 3-way partition at levels where the median of 3 pivot sample contains duplicates. The dense partitioner uses it for numerical features instead of always using 3-way partitioning. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
simultaneous_sort takes a SortPartitioning enum (TWO_WAY, THREE_WAY or MIXED) instead of a boolean. The 2-way introsort is restored as in main and the one switching to 3-way partitioning on duplicate-heavy subarrays becomes the MIXED one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sparse partitioner, the MAE criterion and the sort of the categories means also use MIXED instead of THREE_WAY. The three introsorts share the 2-way and 3-way partition helpers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reference Issues/PRs
Benchmarked with probabl-ai/scikit-learn-benchmarks#121.
What does this implement/fix? Explain your changes.
The trees sort with
simultaneous_sort(..., use_three_way_partition=True): the feature values in the dense and sparse partitioners,yin the MAE criterion and the categories means. The 3-way partitioning is much faster when there are many duplicated values (up to 10-20x on constant subarrays), but ~20-25% slower than the 2-way partitioning when the values are distinct, which is the common case for continuous features.This PR replaces the 3-way variant of
simultaneous_sortby a mixed one: a 2-way introsort that switches to a 3-way partition at the recursion levels where the median-of-3 pivot sample contains duplicates, and uses it for all the sorts in trees. This check is local to each subarray: continuous features get the 2-way speed, while duplicate-heavy subarrays (constant, low-cardinality or zero-inflated features, deep nodes) still get the 3-way one.The variant is selected with a
SortPartitioningenum (TWO_WAY,MIXED) instead of theuse_three_way_partitionboolean.TWO_WAYstays the default, used by the neighbors and pairwise distances reductions as on main. The pure 3-way introsort is removed since it has no caller anymore, and both introsorts share the 2-way partition loop.AI usage disclosure
I used AI assistance for:
Benchmarks
TLDR: RandomForest fit is ~10% faster on average, up to 1.16x on real datasets (
fraud,year_prediction_msd) and up to 1.7x on synthetic data with continuous features. No slowdown beyond noise.Single-threaded (
n_estimators=4, n_jobs=1) fit speedup vs main, 45 RandomForest cases from theall_modelsconfig:fraudyear_prediction_msdmedical_charges_nominalcalifornia_housingkddcup09_churncovtypebank_marketingamazon_employee_accessames_housingThe binary features are the least favorable case, but it should be noise: with only 2 distinct values, the median-of-3 sample always contains duplicates, so this PR does the same 3-way partitioning as main (plus up to 3 swaps for the pivot selection). Locally, these cases are identical on both branches.
Multi-threaded (default config) results are noisier: geomean 1.10-1.12 over two runs on intel-laptop, 0.99-1.03 over two runs on intel-gnr (which runs 344 threads on these cases).
Other tree sorts switched to
MIXED(single-threaded fit time vs main, best of 5):DecisionTreeRegressor(criterion="absolute_error"), 20k samples): 1.12 with a continuousy, 1.10 with 50% zeros iny, 1.06 with a 5-valuedyMicro-benchmarks of the sort alone, time relative to the fastest of 2-way/3-way for each case (lower is better):
Any other comments?
Least favorable cases found (sort alone, 1M values):
MIXEDis as slow asTWO_WAYwhen the introsort depth limit falls back to heapsort, which happens on some nearly sorted inputs (e.g. a sorted array with ~0.1-1% of random swaps). There, it can be up to ~1.7x slower thanTHREE_WAY. I think it's acceptable for trees: there's rarely more than one nearly sorted feature (e.g. a timestamp), andyor the categories means are unlikely to be nearly sorted. Picking the median-of-3 candidates away from the subarray ends (at n/4 and 3n/4) would avoid it, but it's out of scope here.🤖 Generated with Claude Code