Skip to content

ENH Share X validation and categorical encoding across trees, forests, GB and HGB - #18

Open
cakedev0 wants to merge 9 commits into
feat/cat-forestfrom
trees-validation-refactor
Open

cakedev0 wants to merge 9 commits into
feat/cat-forestfrom
trees-validation-refactor

Conversation

@cakedev0

@cakedev0 cakedev0 commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Reference Issues/PRs

Built on top of feat/cat-forest (categorical support in forests, scikit-learn#34832). First step towards binning support for trees and towards unifying GB and HGB (scikit-learn#27873).

What does this implement/fix? Explain your changes.

Trees, forests, GradientBoosting* and HistGradientBoosting* now share the same X validation and categorical encoding, in the new sklearn/tree/_preprocessing.py.

  • Trees & forests: validation/encoding is done once in _validate_and_preprocess_X, then _fit_validated builds the tree. Forests call _fit_validated directly, so they don't need to know what kind of preprocessing trees do.
  • Encoding: categorical columns are ordinal-encoded straight into a F-ordered array, in the original column order. No ColumnTransformer anymore. This saves time, memory, and fixes bug in HGB (that was not re-ordering columns and had bugs because of that).
  • GB: uses the same helpers, which adds support for categorical_features and for missing values. GB trees are single-output squared_error regressors, so categorical splits work for all losses, including multi-class. predict_stages now uses the tree's goes_left: its manual X <= threshold traversal ignored categorical splits and missing_go_to_left.
  • HGB: uses the same helpers too, so categorical features are no longer moved first. This fixes 3 bugs when a categorical feature is not the first column: interaction_cst applied to the wrong features, partial_dependence(method="recursion") computed for the wrong feature, and the cardinality error reporting the wrong feature index.
  • HGB early stopping: the internal validation split now splits indices, bins X once (bin mapper fitted on the training samples only, via a new sample_indices parameter of _BinMapper.fit) and then splits the binned data column by column. This avoids a float64 copy of X. Fitted models are bit-identical.

Excluding tests and changelogs: +500 / −610 lines.

AI usage disclosure

I used AI assistance for:

  • Code generation (closely guided)
  • Test/benchmark generation
  • Documentation (changelogs)

Benchmarks

HGB fit, early stopping with the internal validation split, 20 iterations:

n × d fit time (before → after) peak RSS above X (before → after)
1M × 20 0.77 → 0.73 s 260 → 112 MB
1M × 100 2.49 → 2.24 s 1015 → 261 MB
3M × 50 4.12 → 3.90 s 1498 → 364 MB

Categorical encoding (500k rows, 16 numerical + 4 categorical pandas columns): trees 124 → 97 ms (the scatter copy after ColumnTransformer is gone), HGB 85 → 88 ms (the OrdinalEncoder dominates).

Any other comments?

🤖 Generated with Claude Code

cakedev0 and others added 9 commits September 27, 2026 14:59
Move validation and ordinal encoding of categorical features to
sklearn/tree/_preprocessing.py, used by trees, forests, GB and HGB.
Encoded columns are written directly into a Fortran-ordered array in the
original column order, without ColumnTransformer.

HGB no longer moves categorical features first, which fixes
interaction_cst, partial dependence with method="recursion" and the
cardinality error message when a categorical feature is not first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Split the sample indices instead of X, bin X once with thresholds fitted
on the training samples only (new sample_indices parameter of
_BinMapper.fit), then split the binned data column by column. This
avoids a float64 copy of X: peak memory is 2.3 to 4 times lower and fit
is slightly faster. Fitted models are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cakedev0

Copy link
Copy Markdown
Owner Author

Possible follow-ups:

  • HGB: validation don't have force dtype to float64. It could keep the dtype of the input (maybe just enforce float) and let the bin mapper handle both float32 and float64 (should be straightforward). This could save a copy of X.

@cakedev0

Copy link
Copy Markdown
Owner Author

@adam2392: heads-up about this PR. I'd like to open it on the official fork once yours (scikit-learn#34832) is merged.

I think it's really good: removes a lot a duplication, adds features, fixes bugs, moves things towards unifying GB and HGB.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant