Conversation
… to svd
use_no_center_cholesky matched on solver in ("auto", "cholesky") but didn't
check the array namespace. With array API dispatch to a non-numpy namespace,
solver="auto" silently resolves to "svd" instead of "cholesky" (see
resolve_solver), which needs X to actually be centered. This left X
uncentered while running svd, causing
test_cross_val_predict_array_api_compliance[...-Ridge] failures on
array_api_strict and torch in CI.
This reverts commit 6d7549d.
This reverts commit f5ed039.
…rays `_dense_mean_variance_axis0` reads each value of X once and accumulates in float64. Values are accumulated by tiles around the running mean and tiles are merged with Chan et al.'s pairwise update, which is stable for features with a large offset relative to their standard deviation. C- and F-contiguous inputs are processed without copies. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The algebraic centering of the cholesky solver computes X.T @ X - n * outer(X_offset, X_offset), which suffers from catastrophic cancellation for features with a large offset relative to their standard deviation. `_preprocess_data(skip_centering_if_safe=True)` now computes the mean and variance of X in a single pass and only skips centering (and the copy of X) when the rounding errors of working on uncentered data are small. Whether centering was skipped is returned with `return_centering_skipped=True`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Compare Ridge cholesky against a float64 reference fitted on the same data: for large offsets, the float32 mean computed by `solver="svd"` is not accurate enough to be a reference. The tolerance on the intercept accounts for the amplification of the errors on coef by the offset of the features. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ffsets With an unpenalized intercept, the conditioning of the problem degrades like 1 + mean**2 / var: lbfgs, newton-cg, newton-cholesky, sag and saga silently stop far from the optimum on features with a large offset relative to their standard deviation. Shifting X is an exact reparametrization, so these estimators now fit on centered X when the offset of a feature exceeds 10 standard deviations, and convert the intercept back (and the warm-started intercept to the centered problem). liblinear is excluded as it penalizes the intercept. The step size of sag/saga is computed again on the centered data. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-guarded-centering
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WIP: still experimental
Reference Issues/PRs
What does this implement/fix? Explain your changes.
Introduce yourself
AI usage disclosure
I used AI assistance for:
Any other comments?