Skip to content

Commit af6454e

Browse files
Bump the PyTorch pin to 2.14 (pytorch#22501)
## Summary PyTorch 2.14.0 is released, so move the pin from 2.13 to it. The tree is mid-cycle at 1.5.0 and the pin normally follows the current stable release. The pin itself is a few version strings plus a re-sync of the vendored c10 headers, which CI requires to match PyTorch's tree byte for byte. Everything else here is a change 2.14 forces, each as its own commit. ## What 2.14 forces **C++20 where ATen headers are compiled.** `intrusive_ptr.h` uses `operator<=>` and `TensorBase.h` uses `requires`, neither guarded, so a target that includes ATen cannot parse at C++17. This follows the pattern already in the tree: six CMake targets set `CXX_STANDARD 20` for exactly this reason. Added the Vulkan op tests to that list, and taught the Buck wrapper to make the same per-target distinction. Non-ATen targets, including everything the bare-metal presets build, still compile at C++17. **Two new AOTI entry points.** 2.14's generated wrapper calls `aoti_torch_is_defined` and `aoti_torch_empty_strided_pinned`, neither of which existed here. The first answers whether a tensor holds storage. The second allocates ordinary host memory, because there is no pinned allocator in this runtime and pinning only lets a copy to the device overlap other work. It refuses a device other than CPU, as PyTorch's own does. **Derived shim spellings.** Some custom operators now reach the fallback-kernel check under the name Inductor derives rather than the name they are registered under. Both spellings are accepted for the Metal ops and the CUDA int4 pack matmul. **Build and CI repairs.** `setup.py bdist_wheel` is gone, so the macOS source build uses the standard frontend. 2.14 is published for ROCm 7.2, not 7.1. PyTorch now compiles at C++20, which makes CMake scan for modules with scanners the images do not have. And a pip `cmake` in the build environment made the image build silently produce a wheel with no BLAS, so the image's own cmake is used and the result is now asserted rather than assumed. Two of these were already broken before this branch: the unpinned `katex` install, and the `cmake` interaction. ## Behaviour changes The re-sync is not cosmetic. `overflows()` answers differently for float to integer casts, in both directions. Filling an int8 tensor with 127.5 was refused and now gives 127. A value at 2^63 cast to int64 was accepted and wrapped, and is now refused. The portable kernels reach this through `check_overflow_cast`, so `full`, `full_like`, `fill`, `scalar_tensor`, `hardtanh`, `leaky_relu`, `scatter` and `constant_pad_nd` inherit it. The header is a faithful copy of upstream, so this is not ours to undo, but it should be visible rather than buried in a header diff. Neither direction had a test; both do now. Two smaller ones come with it. Building a delegate sorts its placeholders, and lifted tensor constants were sorted as if they were user inputs, which left a constant after one and made any later constant insert impossible. They now sort with the parameters and buffers. And dropout was missing from the list of operators that take their input's observer. It is an identity once a model is not training, which is how a quantized model is deployed, so measuring it separately gave the operator after it a different scale. XNNPACK then refused a reshape whose two sides disagreed. Both spellings are listed, as several other operators already are. ## Test plan Each commit carries its own. Covering the whole change: `compare_dirs.sh`, the header check CI runs, passes against a real `release/2.14` checkout and fails without the re-sync. All eight vendored headers are byte-identical to upstream; the one build-file edit follows a file upstream moved. The behaviour change was measured by compiling `overflows()` from both branches, not read off the diff. Both new cases fail on the old header, so neither is vacuous. The regenerated import library was checked member by member: 45 advertised names against 42 before, none lost, and its archive metadata is zeroed so the file is reproducible. Not covered locally: the image build, the CUDA, ROCm, Qualcomm and Metal jobs, the export suites, the Buck build, and anything needing Windows. CI has all of it. ## Known outstanding Three things this change does not repair. None of them is new here, and each needs its own change rather than being folded into a version bump. No job links the checked-in Windows link stub with the Microsoft linker. The cross build uses a GNU linker, the same family that produced the archive, so it cannot answer whether the Microsoft one accepts it. That needs a Windows lowering job, which does not exist yet. There is now a note beside the file describing how it is produced. The two Cortex-M model tests expect three quantize pairs where they used to expect one. Two of those three are round trips: the value is converted back to float and immediately converted again at the same scale, measured as 0.004997437 and 0.0000032501507 on both sides. So a pass that used to absorb them no longer matches. The pad operator those pairs sit around is created by this backend's own passes, after the shared pass that folds such pairs has already run, so absorbing them means changing the order or extending a pass another backend shares. Correct output, more work than needed. Asking for pinned host memory gives ordinary host memory. There is no pinned allocator in this runtime, and adding one needs a matching release path, so the entry point allocates ordinary memory and says so. A caller that wants a copy to the device to overlap other work will not get the overlap. cc @digantdesai @freddan80 @per @zingo @oscarandersson8218 @mansnils @Sebastian-Larsson @robell @rascani --------- Co-authored-by: PyTorch Bot <pytorchbot@users.noreply.github.com>
1 parent 5410b1a commit af6454e

43 files changed

Lines changed: 439 additions & 182 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.
Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1 +1 @@
1-
release/2.13
1+
release/2.14

‎.ci/docker/common/install_docs_reqs.sh‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,9 @@ if [ -n "$BUILD_DOCS" ]; then
2020

2121
apt-get update
2222
apt-get install -y --no-install-recommends yarn
23-
yarn global add katex --prefix /usr/local
23+
# katex 0.18.5 requires commander@15 / node >= 22.12; pin to the last
24+
# release compatible with the node 16 installed above
25+
yarn global add katex@0.18.4 --prefix /usr/local
2426

2527
sudo apt-get -y install doxygen
2628

‎.ci/docker/common/install_pytorch.sh‎

Lines changed: 15 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -76,21 +76,31 @@ install_pytorch_and_domains() {
7676
# the image compiler cannot satisfy. The venv inherits the image's
7777
# site-packages, so PyTorch still builds against the same numpy.
7878
#
79-
# Keep the list in sync with pytorch/pyproject.toml [build-system].requires.
79+
# Keep in sync with pytorch/pyproject.toml [build-system].requires.
8080
local build_venv=/tmp/pytorch-build-venv
8181
rm -rf "${build_venv}"
8282
conda_run python -m venv --system-site-packages "${build_venv}"
83+
# No pip cmake: scikit-build-core would prefer it over the image's, and it
84+
# searches site-packages, where MKL and libomp are not.
8385
conda_run "${build_venv}/bin/pip" install build "scikit-build-core>=1.0" \
84-
"setuptools>=77.0.0,<82" "cmake>=3.27,<4" ninja "packaging>=24.2" \
85-
"typing-extensions>=4.10.0" pyyaml six
86-
conda_run "${build_venv}/bin/python" -m build --wheel --no-isolation
86+
ninja "packaging>=24.2" "typing-extensions>=4.10.0" pyyaml six numpy
87+
# These images have no module scanner, and nothing here uses modules.
88+
conda_run env CMAKE_CXX_SCAN_FOR_MODULES=OFF \
89+
"${build_venv}/bin/python" -m build --wheel --no-isolation
8790
rm -rf "${build_venv}"
8891
pip_install "$(echo dist/*.whl)"
8992

93+
# A build with no BLAS succeeds silently. Run from / to import the wheel.
94+
(cd / && conda_run python -c "
95+
import torch
96+
assert torch._C.has_lapack, 'built without LAPACK'
97+
torch.linalg.qr(torch.randn(4, 4))
98+
")
99+
90100
# Grab the pinned audio and vision commits from PyTorch
91101
TORCHAUDIO_VERSION=release/2.11
92102
export TORCHAUDIO_VERSION
93-
TORCHVISION_VERSION=release/0.28
103+
TORCHVISION_VERSION=release/0.29
94104
export TORCHVISION_VERSION
95105

96106
install_domains

‎.ci/scripts/test-rocm-aoti.sh‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77

88
set -euo pipefail
99

10-
ROCM_VERSION="${ROCM_VERSION:-7.1}"
10+
ROCM_VERSION="${ROCM_VERSION:-7.2}"
1111
ROCM_PATH="${ROCM_PATH:-/opt/rocm}"
1212
PYTORCH_ROCM_INDEX="${PYTORCH_ROCM_INDEX:-https://download.pytorch.org/whl/test/rocm${ROCM_VERSION}}"
1313
TORCHAO_ROCM_WHEEL_BASE="${TORCHAO_ROCM_WHEEL_BASE:-https://download.pytorch.org/whl/nightly/rocm${ROCM_VERSION}}"

‎.ci/scripts/test-rocm-voxtral.sh‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77

88
set -euo pipefail
99

10-
ROCM_VERSION="${ROCM_VERSION:-7.1}"
10+
ROCM_VERSION="${ROCM_VERSION:-7.2}"
1111
ROCM_PATH="${ROCM_PATH:-/opt/rocm}"
1212
EXPECTED_ROCM_ARCH="${EXPECTED_ROCM_ARCH:-gfx950}"
1313
EXPECTED_WARP_SIZE="${EXPECTED_WARP_SIZE:-64}"

‎.ci/scripts/utils.sh‎

Lines changed: 23 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -106,8 +106,8 @@ install_pytorch_and_domains() {
106106
local python_version=$(python -c 'import platform; v=platform.python_version_tuple(); print(f"{v[0]}{v[1]}")')
107107
local torch_release=$(cat version.txt)
108108
# Download key must match the upload key below (basename of dist/*.whl,
109-
# which always carries setup.py's resolved +gitHASH). Branch-ref pins
110-
# like `release/2.13` would otherwise produce `+gitrelease` here and
109+
# which always carries the build's resolved +gitHASH). Branch-ref pins
110+
# like `release/2.14` would otherwise produce `+gitrelease` here and
111111
# never hit the cache.
112112
local torch_short_hash=$(git rev-parse --short=7 HEAD)
113113
local torch_wheel_path="cached_artifacts/pytorch/executorch/pytorch_wheels/${system_name}/${python_version}"
@@ -127,18 +127,31 @@ install_pytorch_and_domains() {
127127
if [[ "${torch_wheel_not_found}" == "1" ]]; then
128128
echo "No cached wheel found, continue with building PyTorch at ${TORCH_VERSION}"
129129

130-
# Install PyTorch's own build-time deps so the source build does not
131-
# silently inherit them from whatever else happens to be in the env
132-
# (e.g. executorch's requirements-ci.txt).
133-
pip install -r requirements-build.txt
134130
git submodule update --init --recursive
135131
if [[ "$(uname -m)" == "aarch64" ]]; then
136132
export BUILD_IGNORE_SVE_UNAVAILABLE=1
137133
fi
138-
USE_DISTRIBUTED=1 python setup.py bdist_wheel
134+
# PyTorch dropped setup.py. Build in a throwaway environment that can see the
135+
# active one, so its build requirements, which pin a cmake that would be
136+
# preferred over the one on PATH, cannot disturb what is installed here.
137+
#
138+
# Keep in sync with pytorch/pyproject.toml [build-system].requires.
139+
local build_venv=/tmp/pytorch-build-venv
140+
rm -rf "${build_venv}"
141+
python -m venv --system-site-packages "${build_venv}"
142+
"${build_venv}/bin/pip" install build "scikit-build-core>=1.0" ninja \
143+
"packaging>=24.2" "typing-extensions>=4.10.0" pyyaml six numpy
144+
USE_DISTRIBUTED=1 "${build_venv}/bin/python" -m build --wheel --no-isolation
145+
rm -rf "${build_venv}"
139146
pip install "$(echo dist/*.whl)"
140-
141-
# Invariant: the basename setup.py just produced must match the cache
147+
# A build with no BLAS succeeds silently, so check rather than assume.
148+
(cd / && python -c "
149+
import torch
150+
assert torch._C.has_lapack, 'built without LAPACK'
151+
torch.linalg.qr(torch.randn(4, 4))
152+
")
153+
154+
# Invariant: the basename the build just produced must match the cache
142155
# URL we'd reconstruct on the next run. If they diverge (someone edits
143156
# torch_wheel_name above, or PyTorch renames its wheels), the cache
144157
# will silently miss and every macOS run will fall back to a ~30-min
@@ -178,7 +191,7 @@ install_pytorch_and_domains() {
178191
# Grab the pinned audio and vision commits from PyTorch
179192
TORCHAUDIO_VERSION=release/2.11
180193
export TORCHAUDIO_VERSION
181-
TORCHVISION_VERSION=release/0.28
194+
TORCHVISION_VERSION=release/0.29
182195
export TORCHVISION_VERSION
183196

184197
install_domains

‎.ci/scripts/wheel/test_cpp_sdk.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -991,7 +991,7 @@ def test_every_shipped_header_compiles(work_dir: Path) -> None:
991991
# These say in their own text that they must not be included directly, and name the header to
992992
# include instead. Including one anyway is a use error rather than a packaging defect.
993993
"c10/util/complex_math.h",
994-
"c10/util/complex_utils.h",
994+
"torch/headeronly/util/complex_utils.h",
995995
)
996996

997997
source = work_dir / "header_probe.cpp"

‎.github/workflows/rocm.yml‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -169,7 +169,7 @@ jobs:
169169
strategy:
170170
fail-fast: false
171171
matrix:
172-
rocm-version: ["7.1"]
172+
rocm-version: ["7.2"]
173173
uses: pytorch/test-infra/.github/workflows/linux_job_v2.yml@main
174174
permissions:
175175
id-token: write
@@ -206,7 +206,7 @@ jobs:
206206
strategy:
207207
fail-fast: false
208208
matrix:
209-
rocm-version: ["7.1"]
209+
rocm-version: ["7.2"]
210210
with:
211211
timeout: 180
212212
no-sudo: true
@@ -248,7 +248,7 @@ jobs:
248248
strategy:
249249
fail-fast: false
250250
matrix:
251-
rocm-version: ["7.1"]
251+
rocm-version: ["7.2"]
252252
with:
253253
timeout: 180
254254
no-sudo: true

‎.github/workflows/windows-msvc.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -91,7 +91,7 @@ jobs:
9191
9292
- name: Install build dependencies
9393
shell: pwsh
94-
run: python -m pip install pyyaml torch==2.13.0 --extra-index-url https://download.pytorch.org/whl/test/cpu
94+
run: python -m pip install pyyaml torch==2.14.0 --extra-index-url https://download.pytorch.org/whl/test/cpu
9595

9696
- name: Build ExecuTorch
9797
shell: pwsh

‎backends/aoti/common_shims.cpp‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -159,6 +159,11 @@ AOTITorchError aoti_torch_get_numel(Tensor* tensor, int64_t* ret_numel) {
159159
return Error::Ok;
160160
}
161161

162+
AOTITorchError aoti_torch_is_defined(Tensor* tensor, bool* ret_is_defined) {
163+
*ret_is_defined = tensor != nullptr;
164+
return Error::Ok;
165+
}
166+
162167
// Device and layout utility functions
163168
int32_t aoti_torch_device_type_cpu() {
164169
// Let's say cpu is 0 for ET as well

0 commit comments

Comments
 (0)