Skip to content

Dense mean variance guarded centering - #16

Closed
cakedev0 wants to merge 28 commits into
mainfrom
dense-mean-variance-guarded-centering
Closed

cakedev0 wants to merge 28 commits into
mainfrom
dense-mean-variance-guarded-centering

Conversation

@cakedev0

Copy link
Copy Markdown
Owner

WIP: still experimental

Reference Issues/PRs

What does this implement/fix? Explain your changes.

Introduce yourself

AI usage disclosure

I used AI assistance for:

  • Code generation (e.g., when writing an implementation or fixing a bug)
  • Test/benchmark generation
  • Documentation (including examples)
  • Research and understanding

Any other comments?

cakedev0 and others added 28 commits August 21, 2026 22:21
… to svd

use_no_center_cholesky matched on solver in ("auto", "cholesky") but didn't
check the array namespace. With array API dispatch to a non-numpy namespace,
solver="auto" silently resolves to "svd" instead of "cholesky" (see
resolve_solver), which needs X to actually be centered. This left X
uncentered while running svd, causing
test_cross_val_predict_array_api_compliance[...-Ridge] failures on
array_api_strict and torch in CI.
This reverts commit 6d7549d.
This reverts commit f5ed039.
…rays

`_dense_mean_variance_axis0` reads each value of X once and accumulates in
float64. Values are accumulated by tiles around the running mean and tiles
are merged with Chan et al.'s pairwise update, which is stable for features
with a large offset relative to their standard deviation. C- and
F-contiguous inputs are processed without copies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The algebraic centering of the cholesky solver computes
X.T @ X - n * outer(X_offset, X_offset), which suffers from catastrophic
cancellation for features with a large offset relative to their standard
deviation.

`_preprocess_data(skip_centering_if_safe=True)` now computes the mean and
variance of X in a single pass and only skips centering (and the copy of X)
when the rounding errors of working on uncentered data are small. Whether
centering was skipped is returned with `return_centering_skipped=True`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Compare Ridge cholesky against a float64 reference fitted on the same data:
for large offsets, the float32 mean computed by `solver="svd"` is not
accurate enough to be a reference. The tolerance on the intercept accounts
for the amplification of the errors on coef by the offset of the features.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ffsets

With an unpenalized intercept, the conditioning of the problem degrades like
1 + mean**2 / var: lbfgs, newton-cg, newton-cholesky, sag and saga silently
stop far from the optimum on features with a large offset relative to their
standard deviation.

Shifting X is an exact reparametrization, so these estimators now fit on
centered X when the offset of a feature exceeds 10 standard deviations, and
convert the intercept back (and the warm-started intercept to the centered
problem). liblinear is excluded as it penalizes the intercept. The step size
of sag/saga is computed again on the centered data.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cakedev0 cakedev0 closed this Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants