Skip to content

Revert "Implement cpuinfo_deinitialize() to free heap-allocated globals" - #411

Merged
malfet merged 1 commit into
mainfrom
revert-387-vchelur/add-cpuinfo-deinitialize
Jul 7, 2026
Merged

Revert "Implement cpuinfo_deinitialize() to free heap-allocated globals"#411
malfet merged 1 commit into
mainfrom
revert-387-vchelur/add-cpuinfo-deinitialize

Conversation

@malfet

@malfet malfet commented Jul 7, 2026

Copy link
Copy Markdown
Collaborator

Reverts #387 because it breaks a lots of the existing clients, as both cpuinfo_initialize() and cpuinfo_deinitialize() should be reentrant and thread safe, but proposed implementation is not like that

@pytorch-bot pytorch-bot Bot added the ci-no-td label Jul 7, 2026
@meta-cla meta-cla Bot added the cla signed label Jul 7, 2026
@malfet
malfet merged commit b1a5d63 into main Jul 7, 2026
16 checks passed
crvineeth97 added a commit to microsoft/onnxruntime that referenced this pull request Sep 2, 2026
### Description
<!-- Describe your changes. -->

- Update `pytorch/cpuinfo` from
`4628dc060ce4e82345dc166bbac875609db4ff69` to
`66ee79c038d70dad9f08705b2c9b3e58f6d8f512`, the latest commit on cpuinfo
`main` as of August 27, 2026.
- Carry the thread-safe, reference-counted initialization and
deinitialization changes from
[pytorch/cpuinfo#400](pytorch/cpuinfo#400) as
one shared ORT patch used by both FetchContent and vcpkg.
- Reset Windows ARM64 cache-population state on each initialization so
repeated DLL load/unload cycles cannot reuse stale cache indices.
- Scope XNNPACK's cpuinfo references to hardware discovery so
XNNPACK-enabled ORT builds do not retain unmatched references during DLL
unload.

The cpuinfo patch can be removed after pytorch/cpuinfo#400, including
the ARM64 reinitialization fix, is merged and ORT updates to a revision
containing it. The XNNPACK compatibility patch can be removed after the
corresponding lifecycle fix is available upstream.

### Motivation and Context
<!-- - Why is this change required? What problem does it solve?
- If it fixes an open issue, please link to the issue here. -->


[#28245](#28245)
integrated `cpuinfo_deinitialize()` after
[pytorch/cpuinfo#387](pytorch/cpuinfo#387) added
it upstream. That implementation was later reverted by
[pytorch/cpuinfo#411](pytorch/cpuinfo#411)
because initialization and deinitialization were not safe for multiple
consumers.

Pinning the latest cpuinfo `main` without an ORT-side patch would
therefore make `cpuinfo_deinitialize()` a no-op again. Carrying the
corrected implementation keeps ORT independent of the pending upstream
review while preserving safe cleanup during dynamic DLL unload.

### Testing

- `onnxruntime_cpuinfo_refcount_test` covers sequential consumers,
concurrent consumers, and reinitialization after final release.
- Verified this test fails against ORT `main`'s pinned cpuinfo revision
(`4628dc060ce4e82345dc166bbac875609db4ff69`) and passes against the
patched revision (`66ee79c038d70dad9f08705b2c9b3e58f6d8f512`).
- In XNNPACK-enabled builds, the refcount test initializes XNNPACK and
verifies hardware discovery does not retain a cpuinfo reference.
- `onnxruntime_shared_lib_cpuinfo_dlopen_test` loads a small DLL
containing ORT's `CPUIDInfo` and cpuinfo, captures a cpuinfo
process-heap allocation, unloads the DLL, and verifies that allocation
was released.
- Verified the FetchContent patch sequences for the default, Linux, and
Windows ARM64/ARM64EC paths, plus the vcpkg patch sequence, apply with
zero rejected hunks.

---------

Co-authored-by: Vineeth Chelur <vchelur@microsoft.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant