Skip to content

Ship a linkable shared runtime and make the Python bindings use it - #21514

Open
shoumikhin wants to merge 63 commits into
mainfrom
gh/shoumikhin/78/head
Open

Ship a linkable shared runtime and make the Python bindings use it#21514
shoumikhin wants to merge 63 commits into
mainfrom
gh/shoumikhin/78/head

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Today pip install executorch gives you the Python side of ExecuTorch, but nothing a C++
program can link against. So if you want to write a small C++ app that loads and runs a
.pte model, you have to clone the repository, sync submodules, and build from source.

This change ships the runtime as a real shared library and exposes it through CMake, so the
wheel alone is enough.

What you get

Before, the wheel had everything welded into one Python extension file:

pip install executorch
  executorch/                  Python works
  _portable_lib.so   (~11 MB)  runtime + kernels + backends, all fused inside
  -> a C++ developer must build from source

After, the runtime is a separate library with a CMake package beside it:

pip install executorch
  executorch/lib/libexecutorch.so.1     the runtime, linkable
  executorch/include/                   the headers
  executorch/share/cmake/               so find_package works
  _portable_lib.so                      now depends on the runtime above

How you use it

Two lines in your build:

find_package(executorch REQUIRED)
target_link_libraries(my_app PRIVATE executorch::runtime)

And ordinary C++ in your program:

#include <executorch/extension/module/module.h>
#include <executorch/extension/tensor/tensor.h>
using namespace executorch::extension;

Module module("model.pte");
auto input = make_tensor_ptr({2, 4}, data.data());
auto result = module.forward(input);

Why the Python side changes too

executorch::runtime is an imported target whose location is worked out relative to the CMake
config file, so the package can be moved and no path from the build machine is baked in.

The Python extension has to move onto the same shared library in this change, not a later one.
Backends register themselves into one process-wide table when their library loads. If the
Python extension kept its own private copy of the runtime while a C++ app linked the new shared
one, a process using both would have two tables, and a backend registered in one would be
invisible to the other. Sharing the library is what keeps that single.

The shipped library is also the one the Python extension now depends on, which means the two
cannot drift apart: there is only one copy to load.

Tested

Built the wheel on Linux x86_64 and aarch64, installed it into a clean environment with no
source checkout reachable, then:

  • compiled and ran the C++ example above against the installed wheel;
  • confirmed exactly one backend registry across all shipped libraries, using symbol inspection;
  • confirmed the Python extension resolves the registry through the shared library rather than a
    private copy;
  • confirmed the CMake package still works after moving the installed directory, so nothing
    depends on an absolute path;
  • ran the existing Python tests, including .pte execution and custom-operator compilation
    through the older CMake contract, so existing users are unaffected.

[ghstack-poisoned]
@shoumikhin
shoumikhin requested a review from larryliu0820 as a code owner July 31, 2026 06:23
Copilot AI lite review requested due to automatic review settings July 31, 2026 06:23
@shoumikhin
shoumikhin requested a review from kirklandsign as a code owner July 31, 2026 06:23
@pytorch-bot

pytorch-bot Bot commented Jul 31, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21514

Note: Links to docs will display an error until the docs builds have been completed.

❌ 28 New Failures, 175 Pending, 2 Unrelated Failures

As of commit bb1959c with merge base 6bbb75a (image):

NEW FAILURES - The following jobs have failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 31, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Not ready to approve

The new shared-runtime linking logic uses GNU ld-specific flags without platform guards, which can break builds when EXECUTORCH_BUILD_SHARED is enabled on non-ELF toolchains.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

This review doesn't count toward merge requirements. Sign up for the private preview to control whether Copilot approvals count.

Pull request overview

This PR makes the ExecuTorch wheel usable as a C++ SDK by shipping a linkable shared runtime (libexecutorch.so) and exposing it via find_package(executorch) as executorch::runtime, while also updating the Python extensions to link against that shared runtime to preserve a single process-wide backend registry.

Changes:

  • Add a consolidated shared runtime target (executorch_shared / libexecutorch.so) and update multiple extension/tool targets to resolve runtime symbols from it.
  • Extend the wheel CMake package config to provide an imported target executorch::runtime resolved relative to the installed wheel for relocatable consumption.
  • Add CI wheel smoke coverage that verifies single-registry behavior and that a standalone C++ app can build/run via only find_package(executorch) + executorch::runtime.
File summaries
File Description
tools/cmake/Utils.cmake Adds helper to force linking against the shared runtime early on the link line.
tools/cmake/preset/pybind.cmake Enables EXECUTORCH_BUILD_SHARED for the Linux pybind preset to ship the shared runtime in wheels.
tools/cmake/executorch-wheel-config.cmake Exposes executorch::runtime imported target and keeps legacy _portable_lib discovery.
tools/cmake/Codegen.cmake Avoids linking executorch_core when shared runtime is enabled; forces shared runtime resolution.
setup.py Installs the shared runtime SONAME into the wheel and expands header install coverage for runtime-linked extensions.
kernels/quantized/CMakeLists.txt Adds wheel-runtime-relative RPATH when shared runtime is enabled.
extension/training/CMakeLists.txt Switches pybind training module to shared runtime path and fixes RPATH for wheel layout.
extension/llm/runner/CMakeLists.txt Forces shared runtime resolution and adjusts RPATH when shared runtime is enabled.
extension/llm/custom_ops/CMakeLists.txt Ensures wheel-built artifact has correct RPATH and resolves runtime from shared library.
devtools/etdump/CMakeLists.txt Avoids executorch whole-archive when shared runtime is present; links shared runtime instead.
devtools/bundled_program/CMakeLists.txt Avoids whole-archive duplication by linking shared runtime when enabled.
codegen/tools/CMakeLists.txt Makes selective_build resolve runtime via shared library and sets wheel-relative RPATH.
CMakeLists.txt Defines executorch_shared earlier and updates pybind/utility targets to avoid embedding duplicate registries.
.github/workflows/build-wheels-*.yml Expands workflow path triggers to include CMakeLists and tools/cmake/.
.ci/scripts/wheel/test_linux.py Runs the new C++ SDK wheel test as part of Linux wheel smoke tests.
.ci/scripts/wheel/test_linux_aarch64.py Runs the new C++ SDK wheel test as part of aarch64 Linux wheel smoke tests.
.ci/scripts/wheel/test_cpp_sdk.py New test that validates single-registry and that a standalone C++ app can link/run via the wheel.
Review details
  • Files reviewed: 20/20 changed files
  • Comments generated: 2
  • Review effort level: Lite

We're testing this review assessment. Please use 👍 or 👎 to tell us if it's correct.

Comment thread tools/cmake/Utils.cmake
Comment thread CMakeLists.txt Outdated
[ghstack-poisoned]
Copilot AI review requested due to automatic review settings July 31, 2026 06:51
shoumikhin added a commit that referenced this pull request Jul 31, 2026
Today `pip install executorch` gives you the Python half of ExecuTorch but
nothing a C++ application can link. The wheel installs a header subset and an
`executorch-config.cmake`, but that config only locates the Python extension so
custom-op builds can compile against it: it defines no runtime library and no
CMake targets. Writing a small C++ program that loads and runs a `.pte` means
cloning the repo, syncing submodules, and building from source.

This change ships the runtime as a real library and exposes it through
`find_package`, so a standalone application needs nothing but the wheel:

    find_package(executorch REQUIRED)
    target_link_libraries(my_app PRIVATE executorch::runtime)

`executorch::runtime` is an imported target whose location is resolved relative
to the config file itself, so the package stays relocatable and no path from the
machine that built the wheel is baked into it. A consumer picks up the wheel's
library directory in its RUNPATH, plus `$ORIGIN`-relative entries so an
application deployed next to a copy of the runtime keeps working without
`LD_LIBRARY_PATH`.

The Python extensions have to move to the shared runtime in the same change.
Backends register themselves into a single process-wide table owned by the
runtime, and that table is only process-wide if exactly one loaded library
defines it. The Python extension statically embeds the runtime today, so adding a
shared library beside it would give a process two independent registries, and a
backend could register into the one nobody reads. The extensions therefore link
the shared runtime instead of whole-archiving the static libraries, leaving
exactly one registry owner for both the Python and C++ paths.

Getting that right needs two linker details. The shared runtime is named through
a link option rather than an ordinary dependency, because CMake orders link
libraries so a static archive precedes what it depends on, which would let the
archive satisfy the runtime symbols first. It is also wrapped in
`--no-as-needed`, because a shared library with no already-referenced symbol at
the point it appears on the link line can be dropped, and a later static archive
would then supply the registry after all.

Registration still happens through static initializers exactly as before. No new
plugin or loader ABI is introduced.

The shared runtime is Linux-only for the wheel. macOS C++ consumers are served by
the existing Swift package distribution, and the runtime has no export
annotations for a Windows DLL. Every other build keeps linking the static
libraries, because the new behavior is gated on the existing
`EXECUTORCH_BUILD_SHARED` option, so iOS, Android, and embedded builds are
unaffected.

Test plan:

The wheel smoke test now covers this on Linux, so it is checked in CI rather than
only by hand. It asserts that exactly one shipped library defines the backend
registry, builds a standalone C++ program that only calls
`find_package(executorch)` and links `executorch::runtime`, runs it without
`LD_LIBRARY_PATH`, and checks the resulting binary depends on the shipped runtime
with a relocatable RUNPATH. The check was confirmed to fail when a second
registry definition is introduced deliberately. The wheel workflows now also run
when any `CMakeLists.txt` or anything under `tools/cmake` changes, since those
files decide what the wheel contains.

Built the wheel from a clean checkout on Linux x86_64 and on Linux aarch64, then
verified each against a fresh virtual environment with a normal
dependency-resolving install:

- `nm -DC` across every shipped shared object shows exactly one definition of the
  registry entry points; the Python extensions import them rather than defining
  their own.
- A C++ program whose CMake only calls `find_package(executorch)` and links
  `executorch::runtime` compiles against the installed wheel with no source
  checkout, then loads a `.pte` and lists its methods.
- `readelf -d` on that program lists the versioned runtime in `DT_NEEDED` with
  `$ORIGIN`-relative RUNPATH entries, and it runs without `LD_LIBRARY_PATH`.
- A C++ binary that links only `executorch::runtime` and loads the Python
  extension sees the backends that extension registered, which is the
  one-registry property stated above.
- `import executorch`, the registered backend list, and `.pte` execution through
  the Python bindings are unchanged, with outputs matching eager PyTorch.
- With `EXECUTORCH_BUILD_SHARED` off, the Python extension still builds
  self-contained and no shared runtime is produced, so the previous behavior is
  intact.
- Configured with CMake 3.28 as well as 3.31, since the project supports 3.24 and
  up and the two differ in how strictly they treat link features.

ghstack-source-id: ab9b01b
ghstack-comment-id: 5139976402
Pull-Request: #21514
@shoumikhin

Copy link
Copy Markdown
Contributor Author

Good catch, fixed in the latest push.

Both of these were real. The whole-archive loop used to go through
$<LINK_LIBRARY:WHOLE_ARCHIVE,...>, which CMake maps to the right flag per
platform, and switching to raw flags lost that. There is now a small
executorch_target_whole_archive helper next to the existing kernel-link helpers
that picks --whole-archive, -force_load, or /WHOLEARCHIVE based on the
platform, and the shared-runtime helper only adds --no-as-needed on ELF
platforms. No raw GNU flag is emitted outside those two helpers now.

Worth noting the option that reaches this code is only turned on for Linux today,
so this is about not breaking someone who enables it by hand rather than a live
failure. Verified by building the wheel and running the wheel checks on Linux
x86_64 and aarch64 after the change.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Not ready to approve

The new executorch::runtime CMake config has correctness/robustness issues (duplicate target definition risk and unconditional Python failure path) that can break consumers during configuration.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

This review doesn't count toward merge requirements. Sign up for the private preview to control whether Copilot approvals count.

Review details

Suppressed comments (3)

tools/cmake/executorch-wheel-config.cmake:121

  • With the new executorch::runtime target, a pure C++ consumer should be able to configure successfully even if python3 isn't available on PATH. Right now the unconditional execute_process(${PYTHON_EXECUTABLE} ...) followed by FATAL_ERROR prevents using the shared runtime without Python. Consider treating EXT_SUFFIX lookup as best-effort when the shared runtime was found, and skip _portable_lib discovery in that case.
find_library(
  _portable_lib_LIBRARY
  NAMES _portable_lib${EXT_SUFFIX}
  PATHS "${_executorch_package_root}/extension/pybindings/"
)

tools/cmake/executorch-wheel-config.cmake:121

  • find_library() will also search default system locations, which can accidentally pick up a different _portable_lib<EXT_SUFFIX> (e.g., from another install) and silently mix headers/libs. Since this config is meant to bind to the wheel it lives in, add NO_DEFAULT_PATH to restrict the search to the wheel directory.
find_library(
  _portable_lib_LIBRARY
  NAMES _portable_lib${EXT_SUFFIX}
  PATHS "${_executorch_package_root}/extension/pybindings/"
)

tools/cmake/executorch-wheel-config.cmake:68

  • executorch-config.cmake can be processed more than once in a single configure (e.g., multiple find_package(executorch) calls via subprojects). Unconditionally calling add_library(executorch::runtime ...) will then error with a duplicate target name. Guard the add_library call with if(NOT TARGET executorch::runtime) but still apply the target properties afterward.

This issue also appears in the following locations of the same file:

  • line 117
  • line 117
  add_library(executorch::runtime SHARED IMPORTED)
  set_target_properties(
    executorch::runtime
    PROPERTIES IMPORTED_LOCATION "${_executorch_runtime_library}"
               INTERFACE_INCLUDE_DIRECTORIES "${EXECUTORCH_INCLUDE_DIRS}"
  • Files reviewed: 20/20 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

We're testing this review assessment. Please use 👍 or 👎 to tell us if it's correct.

[ghstack-poisoned]
Copilot AI review requested due to automatic review settings July 31, 2026 14:53
shoumikhin added a commit that referenced this pull request Jul 31, 2026
Today `pip install executorch` gives you the Python half of ExecuTorch but
nothing a C++ application can link. The wheel installs a header subset and an
`executorch-config.cmake`, but that config only locates the Python extension so
custom-op builds can compile against it: it defines no runtime library and no
CMake targets. Writing a small C++ program that loads and runs a `.pte` means
cloning the repo, syncing submodules, and building from source.

This change ships the runtime as a real library and exposes it through
`find_package`, so a standalone application needs nothing but the wheel:

    find_package(executorch REQUIRED)
    target_link_libraries(my_app PRIVATE executorch::runtime)

`executorch::runtime` is an imported target whose location is resolved relative
to the config file itself, so the package stays relocatable and no path from the
machine that built the wheel is baked into it. A consumer picks up the wheel's
library directory in its RUNPATH, plus `$ORIGIN`-relative entries so an
application deployed next to a copy of the runtime keeps working without
`LD_LIBRARY_PATH`.

The Python extensions have to move to the shared runtime in the same change.
Backends register themselves into a single process-wide table owned by the
runtime, and that table is only process-wide if exactly one loaded library
defines it. The Python extension statically embeds the runtime today, so adding a
shared library beside it would give a process two independent registries, and a
backend could register into the one nobody reads. The extensions therefore link
the shared runtime instead of whole-archiving the static libraries, leaving
exactly one registry owner for both the Python and C++ paths.

Getting that right needs two linker details. The shared runtime is named through
a link option rather than an ordinary dependency, because CMake orders link
libraries so a static archive precedes what it depends on, which would let the
archive satisfy the runtime symbols first. It is also wrapped in
`--no-as-needed`, because a shared library with no already-referenced symbol at
the point it appears on the link line can be dropped, and a later static archive
would then supply the registry after all.

Registration still happens through static initializers exactly as before. No new
plugin or loader ABI is introduced.

The shared runtime is Linux-only for the wheel. macOS C++ consumers are served by
the existing Swift package distribution, and the runtime has no export
annotations for a Windows DLL. Every other build keeps linking the static
libraries, because the new behavior is gated on the existing
`EXECUTORCH_BUILD_SHARED` option, so iOS, Android, and embedded builds are
unaffected.

Test plan:

The wheel smoke test now covers this on Linux, so it is checked in CI rather than
only by hand. It asserts that exactly one shipped library defines the backend
registry, builds a standalone C++ program that only calls
`find_package(executorch)` and links `executorch::runtime`, runs it without
`LD_LIBRARY_PATH`, and checks the resulting binary depends on the shipped runtime
with a relocatable RUNPATH. The check was confirmed to fail when a second
registry definition is introduced deliberately. The wheel workflows now also run
when any `CMakeLists.txt` or anything under `tools/cmake` changes, since those
files decide what the wheel contains.

Built the wheel from a clean checkout on Linux x86_64 and on Linux aarch64, then
verified each against a fresh virtual environment with a normal
dependency-resolving install:

- `nm -DC` across every shipped shared object shows exactly one definition of the
  registry entry points; the Python extensions import them rather than defining
  their own.
- A C++ program whose CMake only calls `find_package(executorch)` and links
  `executorch::runtime` compiles against the installed wheel with no source
  checkout, then loads a `.pte` and lists its methods.
- `readelf -d` on that program lists the versioned runtime in `DT_NEEDED` with
  `$ORIGIN`-relative RUNPATH entries, and it runs without `LD_LIBRARY_PATH`.
- A C++ binary that links only `executorch::runtime` and loads the Python
  extension sees the backends that extension registered, which is the
  one-registry property stated above.
- `import executorch`, the registered backend list, and `.pte` execution through
  the Python bindings are unchanged, with outputs matching eager PyTorch.
- With `EXECUTORCH_BUILD_SHARED` off, the Python extension still builds
  self-contained and no shared runtime is produced, so the previous behavior is
  intact.
- Configured with CMake 3.28 as well as 3.31, since the project supports 3.24 and
  up and the two differ in how strictly they treat link features.

ghstack-source-id: 401229a
ghstack-comment-id: 5139976402
Pull-Request: #21514

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]
[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]
[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

[ghstack-poisoned]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/binaries/all Release PRs with this label will build wheels for all python versions ciflow/binaries ciflow/cuda ciflow/nightly ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants