perf(fit): copy NumPy fit data through PyArray - #1393
Merged
MilesCranmer merged 1 commit intoSep 28, 2026
Merged
Conversation
Wrap X, y and weights in `PythonCall.PyArray` and copy them into an owned Julia array. This skips PythonCall's generic array conversion, which compiled fresh code on every first fit. Co-authored-by: Miles Cranmer <miles.cranmer@gmail.com>
MilesCranmer
approved these changes
Sep 28, 2026
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
These PySR changes pair with the SymbolicRegression.jl PRs astroautomata/SymbolicRegression.jl#739 through #742, which compile PySR's search and PythonCall conversions at install time. Each PySR PR works on its own with the current SymbolicRegression.jl. The full speedup arrives once those PRs are released and PySR's version bound is raised.
Why
jl_arrayconverts NumPy arrays withpyconvert(Array{T}, x), which compiles PythonCall's generic conversion rules on every first fit. Wrapping the array inPythonCall.PyArrayand copying it gives the same owned JuliaArraythrough a narrow, typed path that SymbolicRegression.jl can precompile.Scope
pysr/julia_helpers.py: newjl_numpy_array(x), which returnsjl.copy(PythonCall.PyArray(x)).pysr/sr.py:X,y, andweightsusejl_numpy_arrayin_run. Type-spec inputs still go throughtype_spec_to_julia_array.Blast Radius
The data Julia receives is an owned copy, as before, so later changes to the NumPy array do not affect a running search. Fixed-seed fits give byte-identical halls of fame.
Verification
Zero-iteration first fit, measured with the matching SymbolicRegression.jl precompile statement: 3.68 s to 3.33 s.
With all seven PRs applied, on a 12-core Slurm node (Julia 1.13.1, Python 3.13), the README example's first fit went from 17.4 s to 6.95 s, and a zero-iteration first fit from 12.6 s to 2.71 s.
pysr/test/test_main.pyandpysr/test/test_type_specs.pypass (254 passed, 1 skipped). The pre-commit hooks pass.