Skip to content

About

tenferro benchmark suite

Resources

Stars

1 star

Watchers

1 watching

Forks

Repository files navigation

tenferro-benchmark

The M5 CPU refresh from tenferro-rs origin/main contains the latest CPU results, including the merged BLAS shared-session fix.

Short operations now use many operations in one session; see the M5 root-cause investigation.

Timing corrections are documented in timing policy revision 2. The M5 Max CPU refresh reran all CPU suites with revision 2 at 1 and 4 threads. Other targets retain earlier measurements and require affected-participant reruns before operation-performance comparisons.

Benchmark suite for tenferro-rs, comparing tenferro against PyTorch, JAX, Julia, HPTT, cuTENSOR, and other reference backends on CPU and GPU workloads.

Latest Results

The tracked reports under result/ are the source of truth for benchmark numbers (they are never duplicated into this README). Each file is the latest report for one target_profile × suite_id pair; older results live in git history only.

Target profile Suite Report
mac-cpu (Apple Silicon, native) cpu/einsum result/mac-cpu/cpu/einsum.md
mac-cpu cpu/cpu_ops result/mac-cpu/cpu/cpu_ops.md
mac-cpu cpu/linalg_jvp_vjp result/mac-cpu/cpu/linalg_jvp_vjp.md
mac-cpu cpu/fft result/mac-cpu/cpu/fft.md
mac-cpu cpu/public_api result/mac-cpu/cpu/public_api.md
mac-cpu cpu/session_matrix result/mac-cpu/cpu/session_matrix.md
mac-cpu cpu/small_work result/mac-cpu/cpu/small_work.md
mac-cpu cpu/permutation result/mac-cpu/cpu/permutation.md
linux-cpu (Linux devcontainer; collected as amd-cpu) cpu/einsum result/linux-cpu/cpu/einsum.md
linux-cpu cpu/cpu_ops result/linux-cpu/cpu/cpu_ops.md
linux-cpu cpu/linalg_jvp_vjp result/linux-cpu/cpu/linalg_jvp_vjp.md
linux-cpu cpu/permutation result/linux-cpu/cpu/permutation.md
amd-cpu (Linux devcontainer) cpu/small_work result/amd-cpu/cpu/small_work.md
amd-cpu cpu/large_ad result/amd-cpu/cpu/large_ad.md
amd-cpu (Linux devcontainer) cpu/perf_issues result/amd-cpu/cpu/perf_issues.md
linux-cpu linalg JVP/JVP repro result/linux-cpu/cpu/linalg_jvp_jvp.md
nvidia-gpu (CUDA devcontainer) gpu/dense result/nvidia-gpu/gpu/dense.md
nvidia-gpu gpu/einsum result/nvidia-gpu/gpu/einsum.md
nvidia-gpu gpu/sparse result/nvidia-gpu/gpu/sparse.md
nvidia-gpu gpu/tensornetwork result/nvidia-gpu/gpu/tensornetwork.md
nvidia-gpu gpu/linalg_jvp_vjp result/nvidia-gpu/gpu/linalg_jvp_vjp.md
nvidia-gpu gpu/permutation result/nvidia-gpu/gpu/permutation.md
nvidia-gpu gpu/perf_issues result/nvidia-gpu/gpu/perf_issues.md

Raw runs (per-timestamp run.yaml + machine-readable outputs) are written to data/results/<target_profile>/<suite_id>/<timestamp>/; the tracked report in result/ is regenerated from the newest run. See docs/results.md for the full layout.

Prerequisites

All profiles:

  • Rust toolchain (cargo).
  • uv for the Python environment (uv sync creates .venv from pyproject.toml).
  • External checkouts under extern/ (tenferro-rs, strided-rs, and problem data), fetched by ./scripts/setup_extern_deps.sh. Note that extern/strided-rs is required for any cargo build in this repository (Cargo resolves optional path-dependency manifests even when their feature is disabled), so run the setup script before building anything.

For mac-cpu, run the "All profiles" commands directly on the host. For the devcontainer profiles (amd-cpu / linux-cpu / nvidia-gpu), run them inside the corresponding devcontainer (devcontainer exec ... bash -lc 'uv sync && ./scripts/setup_extern_deps.sh') — the setup script expects the container's OPENBLAS_ROOT / MKLROOT environment, and the container's .venv must be built with the container's wheels (see the GPU note below).

Profile-specific:

  • mac-cpu: runs natively (no Docker); tenferro uses Accelerate.
  • amd-cpu / linux-cpu: the devcontainer CLI and Docker; tenferro defaults to OpenBLAS, oneMKL is optional.
  • nvidia-gpu: the CUDA devcontainer under .devcontainer/cuda/.
  • cpu/permutation, cpu/public_api, and cpu/einsum suites: Julia on PATH (e.g. juliaup or brew install julia) for the julia-base/strided-jl columns (cpu/permutation/cpu/public_api) and the omeinsum-jl column (cpu/einsum); the repo Project.toml pulls in JSON.jl, LinearAlgebra (stdlib), Strided.jl, and OMEinsum.jl via Pkg.instantiate. Without julia, those columns are skipped with a warning. For the HPTT column (present in the tracked latest cpu/permutation reports), also install cmake plus a C++ toolchain (macOS: brew install cmake) and pass PERMUTATION_EXTRA_FEATURES=hptt, because the hptt Cargo feature builds the vendored HPTT C++ library.

Workflow guides per platform: macOS CPU · Linux CPU devcontainer · NVIDIA GPU devcontainer.

Running the Benchmarks

CPU einsum + ops (macOS, native)

scripts/run_all.sh [NUM_THREADS] runs the CPU einsum suite (tenferro trace/eager vs PyTorch/JAX) plus the CPU ops microbenchmarks (primal linalg, JVP/VJP, eager backward), and regenerates result/<target_profile>/cpu/{einsum,cpu_ops,linalg_jvp_vjp}.md. Passing multiple thread counts runs those main suites once per thread count, then runs the FFT, public API, and permutation suites once over the same thread-count list, regenerating all tracked CPU reports.

uv sync
./scripts/setup_extern_deps.sh
BENCHMARK_TARGET_PROFILE=mac-cpu ./scripts/run_all.sh 1
BENCHMARK_TARGET_PROFILE=mac-cpu ./scripts/run_all.sh 4

To regenerate all tracked result/mac-cpu/cpu/*.md reports in one sequential orchestration, including FFT, public API, and permutation at 1T and 4T, run:

uv sync
./scripts/setup_extern_deps.sh
# Julia on PATH; for HPTT: brew install cmake (and a C++ toolchain)
PERMUTATION_EXTRA_FEATURES=hptt \
BENCHMARK_TARGET_PROFILE=mac-cpu ./scripts/run_all.sh 1 4

Quick smoke (single small instance, one run, no warmup):

BENCHMARK_TARGET_PROFILE=mac-cpu \
BENCH_INSTANCE=bin_matmul_256 \
BENCH_RUNS=1 \
BENCH_WARMUPS=0 \
PUBLICATION_GATE_SUITE=small \
  ./scripts/run_all.sh 1

Useful environment variables: BENCH_INSTANCE (restrict to one einsum instance), BENCH_RUNS / BENCH_WARMUPS (iteration counts), TENFERRO_CPU_FEATURES (native or a BLAS provider: blas-openblas, blas-mkl, blas-accelerate; Linux defaults to blas-openblas, macOS defaults to blas-accelerate), RUN_FFT_SUITE=0, RUN_PUBLIC_API_SUITE=0, and RUN_PERMUTATION_SUITE=0 (skip one of the follow-up suites in a multi-thread-count run_all.sh invocation; HPTT still needs PERMUTATION_EXTRA_FEATURES=hptt).

Note: a full or smoke run_all.sh invocation overwrites the tracked latest reports under result/<target_profile>/. If you only ran a smoke subset, restore them before committing (git checkout -- result/).

CPU einsum + ops (Linux devcontainer)

Same runner, executed inside the devcontainer from the host:

devcontainer up --workspace-folder .
devcontainer exec --workspace-folder . bash -lc '
  BENCHMARK_TARGET_PROFILE=amd-cpu ./scripts/run_all.sh 1'

Small-work public API suite (cpu/small_work)

Reuses the 154 cases from feat/95-small-work (38b9a83): F64 add/einsum/solve/gather/reduce_sum and C64 einsum. Twelve further f64 einsum cases carry explicit subscripts and operand_shapes and guard tenferro-rs CPU fixes: abcd,dbef->acef at extent 4 (#1897, canonical-fallback fixed cost), ax,asb->xsb at D=4, s=2 (#1899, prepared einsum dispatch) and ij,jk->ik with 95x95 times 95x1 (#1904, in-session tiny GEMM). Each runs as prepared execute (prepared-repeat), prepared execute_into into a preallocated destination (prepared-into-repeat), the plain in-session dot_general_read_into (dot-general-into-shared) and, as a diagnostic, ordinary einsum (concrete-shared). Shared/prepared operation routes use one clock interval for at least 1024 operations. Fresh-session, eager and compiled per-call routes are opt-in diagnostics. No cross-library equivalents or automatic performance gates are added.

devcontainer up --workspace-folder .
devcontainer exec --workspace-folder . bash -lc '
  TENFERRO_CPU_FEATURES=blas-mkl \
  BENCHMARK_TARGET_PROFILE=amd-cpu ./scripts/run_small_work.sh 1 4'

# Selected case (also overwrites the latest report; run the full suite last):
devcontainer exec --workspace-folder . bash -lc '
  TENFERRO_CPU_FEATURES=blas-mkl \
  BENCH_INSTANCE=add_f64_concrete_fresh ./scripts/run_small_work.sh 1'

Follow the existing checkout-freshness policy in AGENTS.md. For a worktree whose Git directory is not mounted in the container, pass --remote-env BENCHMARK_COMMIT="$(git rev-parse HEAD)" to devcontainer exec. The CLI can also be invoked as npx --yes @devcontainers/cli.

BENCH_INSTANCE accepts comma-separated case IDs; list them using python scripts/suite_instances.py --suite-file benchmarks/cpu/small_work.yaml --format lines. The suite is included in multi-thread-count run_all.sh calls (disable with RUN_SMALL_WORK_SUITE=0), or opt in with RUN_SMALL_WORK_SUITE=1 for a single-thread-count call. Standalone collection avoids rerunning other suites.

Numerical values, solve residuals, and applicable AD gradients are checked outside timing. The report records batch-normalized median/IQR in ns, CoV, chain totals versus per-operation normalization, and timing boundaries. Preparation is excluded from prepared execution rows and measured separately. Noisy rows remain visible. Failed cases have no latency and the runner exits nonzero. Raw samples and metadata are retained under data/results/<target_profile>/cpu/small_work/<timestamp>/; the generated latest report is result/<target_profile>/cpu/small_work.md.

Performance-issue workloads (cpu/perf_issues, gpu/perf_issues)

Every open tenferro-rs performance issue gets its workload added as cases when the issue is opened, with the issue number recorded next to each case (issues in data/instances/perf_issues.json / data/instances/gpu_perf_issues.json, or the suite description / instance intent where an existing suite already covers it). A performance-fix PR reports numbers from those cases. There is no automated audit, gate or scheduled run. Each family carries its reference arm(s) (faer direct, host sgemm, PyTorch, cudarc memcpy, the eager or batched counterpart, ...); reference_for names the compared case. Cases are built only from public tenferro-rs APIs (CpuBackend::new() / with_threads(n), with_backend_session, public concrete/eager/traced ops). Steady-state rows, per-call session-entry diagnostics, eager AD workflows, counter runs and the GPU first-call diagnostic are reported in separate sections.

Edit scripts/generate_perf_issue_cases.py (it also writes the #1865/#1863 cpu/einsum instances), rerun it and bump the manifest version. Collection:

devcontainer exec --workspace-folder . bash -lc '
  TENFERRO_CPU_FEATURES=blas-mkl BENCHMARK_TARGET_PROFILE=amd-cpu \
  ./scripts/run_perf_issues.sh 1 4'
devcontainer exec --workspace-folder . --config .devcontainer/cuda/devcontainer.json \
  bash -lc 'BENCHMARK_TARGET_PROFILE=nvidia-gpu ./scripts/run_gpu_perf_issues.sh'

BENCH_COVERAGE=full adds the diagnostics, BENCH_EFFORT=scan is a low-repetition screen, BENCH_INSTANCE filters case IDs and PERF_ISSUES_CORRECTNESS_ONLY=1 executes and validates every case without timing. Raw runs go to data/results/<target_profile>/{cpu,gpu}/perf_issues/<timestamp>/.

tenferro-rs issue Cases Suite
#1897, #1899, #1904 generic einsum cases (*_abcd-dbef-acef_d4, *_ax-asb-xsb_d4s2, *_ij-jk-ik_m95k95n1) cpu/small_work
#2000, #1884 batched_lu_factor / batched_lu_solve 1024x{8,16}, cpu/linalg_batch_families cpu/public_api
#1803 B/C batched_solve, *_backward rows cpu/cpu_ops
#1865 bin_omeinsum_{matmul_10x10,batched_matmul_8x8_batch_4,high_d_12x12_contract_4_batch_4} cpu/einsum
#1863 bin_permuted_r4_abcd_dbef_acef_d32 (1/4(/8)-thread runs) cpu/einsum
#1992, #2003, #1995 decode_proj_f32_*, decode_block_copy_volume_f32_d1024_len8 cpu/perf_issues
#1900 gemm_{c64,f64}_mm256_t1 cpu/perf_issues
#1615 conj_dot_{f64,c64}_* cpu/perf_issues
#2007 small_solve_f64_* cpu/perf_issues
#1803 A eager_backward_matmul2x2_f64_leaves* cpu/perf_issues
#1990 tanh_chain_f32_* cpu/perf_issues
#2006 (+ #2010 PR-B1 single call) {layer_norm,rms_norm}_f32_* cpu/perf_issues
#1975 (#2010 PR-B1) activation_{erf,sigmoid,silu,softplus,gelu,gelu_tanh}_f32_* cpu/perf_issues
#1976 (#2010 PR-B1) {softmax,log_softmax,masked_softmax}_f32_*, reduce_mean_f32_* cpu/perf_issues
#2008 (#2010 PR-B1) take_along_axis_rows_f64_* cpu/perf_issues
#1885 §4.1 small_contraction_abcd-dbef-acef_f64_*_d4 cpu/perf_issues
#2009 transfer_{up,down}_f64_* gpu/perf_issues
#1887 alloc_zero_f64_* gpu/perf_issues
#1885 §3.1-§3.4 small_blocks_gemm_*, batched_{qr,svd}_*, qr_live_buffers_*, first_call_transpose_* gpu/perf_issues

CPU permutation suite (cpu/permutation)

A standalone materialize/copy-kernel benchmark comparing tenferro-rs to_contiguous against strided-rs, HPTT, Julia Base, and Strided.jl. An internal untimed odometer implementation provides the correctness reference. Every timed call includes fresh destination allocation for a common end-to-end materialization comparison. Spec: docs/permutation-suite.md.

uv sync
./scripts/setup_extern_deps.sh
# Multiple thread counts are measured sequentially in one run.
# Without PERMUTATION_EXTRA_FEATURES=hptt the HPTT column is omitted (`-`);
# without `julia` on PATH the Julia columns are omitted. Tracked latest
# reports include both.
PERMUTATION_EXTRA_FEATURES=hptt \
BENCHMARK_TARGET_PROFILE=mac-cpu ./scripts/run_permutation.sh 1 4

For a quick trial run, the suite honors PATTERN_ID (restrict to one pattern), BENCH_RUNS, and BENCH_WARMUPS:

BENCHMARK_TARGET_PROFILE=mac-cpu \
PATTERN_ID=transpose_2d_2048 \
BENCH_RUNS=1 \
BENCH_WARMUPS=0 \
  ./scripts/run_permutation.sh 1

This writes result/<target_profile>/cpu/permutation.md. Pattern definitions live in data/instances/permutation_patterns.json and are read by both the Rust and Julia runners; result records are validated against schemas/permutation-result.schema.json. The suite is also included by ./scripts/run_all.sh 1 4 unless RUN_PERMUTATION_SUITE=0 is set.

GPU suites (CUDA devcontainer)

The repo .venv is shared between the CPU and CUDA devcontainers (it lives in the bind-mounted workspace). A CPU-side uv sync — including the CPU devcontainer's own post-create hook — replaces the CUDA wheels with CPU ones. The GPU Python runners then silently skip pytorch-cuda / jax-cuda (they exit 0 and the report is simply missing those columns) rather than failing. Before collecting GPU results, install the CUDA Python backends inside the container and verify they see the GPU:

devcontainer up --workspace-folder . --config .devcontainer/cuda/devcontainer.json

# One-time, and again after ANY CPU-side `uv sync`:
devcontainer exec --workspace-folder . --config .devcontainer/cuda/devcontainer.json \
  bash -lc '
    (uv sync --frozen || uv sync)
    uv pip install "torch>=2.12.0" --extra-index-url https://download.pytorch.org/whl/cu126
    uv pip install "jax[cuda12]"
    ./scripts/setup_extern_deps.sh
    uv run python -c "import torch, jax; assert torch.cuda.is_available(); jax.devices(\"cuda\"); print(\"CUDA OK:\", torch.__version__)"'

If nvidia-smi fails inside a previously created container ("Failed to initialize NVML" / CUDA_ERROR_NO_DEVICE) while the host GPU is fine, docker restart <container> usually restores GPU access — no rebuild needed.

# gpu/dense, gpu/einsum, gpu/sparse, gpu/tensornetwork:
devcontainer exec --workspace-folder . --config .devcontainer/cuda/devcontainer.json \
  bash -lc 'BENCHMARK_TARGET_PROFILE=nvidia-gpu ./scripts/run_gpu_suite.sh'

# gpu/linalg_jvp_vjp (separate from the standard GPU suite):
devcontainer exec --workspace-folder . --config .devcontainer/cuda/devcontainer.json \
  bash -lc 'BENCHMARK_TARGET_PROFILE=nvidia-gpu ./scripts/run_gpu_linalg_jvp_vjp.sh'

# gpu/permutation (standalone, like cpu/permutation):
devcontainer exec --workspace-folder . --config .devcontainer/cuda/devcontainer.json \
  bash -lc 'BENCHMARK_TARGET_PROFILE=nvidia-gpu ./scripts/run_gpu_permutation.sh'

gpu/permutation honors the same quick-trial variables as the CPU suite (PATTERN_ID, BENCH_RUNS, BENCH_WARMUPS), plus GPU_BENCH_DEVICE for the CUDA ordinal:

devcontainer exec --workspace-folder . --config .devcontainer/cuda/devcontainer.json \
  bash -lc 'BENCHMARK_TARGET_PROFILE=nvidia-gpu \
    PATTERN_ID=transpose_2d_2048 BENCH_RUNS=1 BENCH_WARMUPS=0 \
    ./scripts/run_gpu_permutation.sh'

After any GPU run, check the generated report for the full backend set (the Comparison Backends section lists what each suite compares; for gpu/permutation that is two tenferro columns, cuTENSOR, PyTorch CUDA, JAX CUDA, and memcpy-d2d). A column that is - on every row means that backend's runner skipped — usually the CUDA-wheels issue above. Like run_all.sh, these scripts overwrite the tracked latest reports under result/nvidia-gpu/; after a trial or partial run, restore them before committing (git checkout -- result/).

The GPU tensor network benchmark uses problem data from extern/TensorNetworkBenchmarks/, based on the upstream TensorNetworkBenchmarks repository; see docs/tensornetwork-gpu.md.

Linux linalg AD repro (OpenBLAS / oneMKL)

Reproduces result/linux-cpu/cpu/linalg_jvp_jvp.md with the devcontainer's source-built OpenBLAS (/opt/openblas):

devcontainer up --workspace-folder . --remove-existing-container
devcontainer exec --workspace-folder . bash -lc '
  python3 - <<PY
import ctypes
lib = ctypes.CDLL("/opt/openblas/lib/libopenblas.so")
lib.openblas_get_config.restype = ctypes.c_char_p
lib.openblas_get_parallel.restype = ctypes.c_int
print(lib.openblas_get_config().decode())
print(f"parallel={lib.openblas_get_parallel()}")
PY'
devcontainer exec --workspace-folder . bash -lc '
  export TENFERRO_CPU_FEATURES=blas-openblas
  export PUBLICATION_GATE_FEATURES=blas-openblas
  ./scripts/reproduce_linux_cpu_linalg_jvp_jvp.sh'

For the oneMKL variant (/opt/intel/oneapi/mkl/latest), replace both blas-openblas values with blas-mkl. Verify the OpenBLAS build through the runtime API above instead of relying on strings.

Measurement Policy

  • Benchmarks are always run sequentially — never multiple benchmark processes at once, including different thread-count variants of the same suite (see AGENTS.md).
  • Thread counts are controlled via RAYON_NUM_THREADS / OMP_NUM_THREADS / JULIA_NUM_THREADS. CPU pinning is unavailable on macOS; on Linux the devcontainer convention applies no taskset/numactl pinning either. The effective thread environment is recorded in each run's run.yaml.
  • Every run records provenance (target profile, suite, tenferro-rs commit, CPU/GPU info) in data/results/.../run.yaml.
  • A provider's thread count does not prove which CPUs its threads may use: BLAS/LAPACK provider threads inherit the CPU mask of the thread that creates them. At least one run per lane records the observed per-thread Cpus_allowed_list groups with scripts/cpu_provider_affinity_check.py (--output FILE.json -- COMMAND), which also reports whether the run was signalled. A 4-thread provider whose threads are all confined to one CPU is a configuration failure, not a result.

Comparison Backends

CPU einsum reports compare tenferro-trace, tenferro-eager, pytorch-cpu, jax-cpu, and omeinsum-jl (Julia; see docs/einsum-suite.md). GPU reports compare tenferro-cuda-trace, tenferro-cuda-eager, pytorch-cuda, jax-cuda, and vendor-specific CUDA backends where meaningful. The cpu/permutation suite has its own backend set (tenferro-rs to_contiguous, HPTT, strided-rs, Julia Base, Strided.jl, memcpy); gpu/permutation compares tenferro CUDA transpose paths against cuTENSOR, PyTorch/JAX CUDA, and a device-to-device memcpy baseline.

The cpu/public_api suite additionally compares two Julia columns, julia-base (natural Base/LinearAlgebra spellings, e.g. permutedims!, cholesky, eigen, gather/slice indexing, reshape/transpose/@view metadata-only views, broadcast-into/mul!/copyto! output reuse, and complex Base/LinearAlgebra spellings) and strided-jl (natural Strided.jl @strided fused-broadcast spellings), populated only where each spelling naturally applies: strided-jl covers the elementwise/chain/transpose rows, the elementwise cpu/output_reuse _into rows, and the elementwise cpu/complex rows (conj/mul/div/exp/log); reductions and dense linalg have no natural Strided.jl spelling, so those stay julia-base-only. julia-base also now covers cpu/indexing_layout, cpu/view_metadata, cpu/output_reuse, and cpu/complex, with a few rows left missing where Julia Base has no natural spelling: pad (edge padding), broadcast_in_dim_view (a zero-stride broadcast array view), and extract_diagonal (batched diagonal extraction). Julia is column-major like tenferro-rs, so these columns need no PyTorch/JAX-style layout reconstruction to preserve the same logical fixture values. Attribution: Strided.jl is prior art for tenferro-rs' strided-rs kernel layer (both implement strided-array views and fused, cache-blocked elementwise/permutation kernels); the strided-jl column exists to make that lineage visible in the comparison, not merely to add another backend.

C++ Torch/LibTorch runners are intentionally removed; PyTorch Python is the ATen comparison backend. The PyTorch CPU provider is detected at run time and recorded in run.yaml and generated reports. The default Linux image uses a binary PyTorch wheel; the separate provider-matched OpenBLAS image source-builds PyTorch against the same OpenBLAS as tenferro.

CPU provider pairs

A CPU result compares two provider stacks, and each side is fixed by a different mechanism:

Side Selection Linux default
tenferro TENFERRO_CPU_FEATURES (native, blas-openblas, blas-mkl) OpenBLAS from OPENBLAS_ROOT (/opt/openblas)
PyTorch CPU Build-time choice inside the wheel wheel-bundled Intel MKL

Consequences to state in a report instead of assuming provider identity:

  • A blas-openblas lane is tenferro=OpenBLAS vs PyTorch=bundled MKL, not OpenBLAS on both sides. The provider-matched image above is the only supported way to align them, and it applies to OpenBLAS only.
  • A blas-mkl lane compares different MKL builds: the system oneAPI MKL that tenferro links against, and the older MKL inside the wheel (reported by torch.__config__). Both sides say "MKL" but they are not the same library or OpenMP runtime, and replacing the wheel's MKL by preloading the system one is not supported.
  • Record both sides with run.yaml (tenferro features and BLAS implementation) plus the reference's own provider report (torch.__config__ BLAS/MKL/OpenMP lines), so a reader can see the pair that was actually measured.

Documentation

Development Checks

Run these after changing benchmark scripts or schemas:

uv run python scripts/validate_benchmark_suite.py benchmarks/cpu/einsum.yaml benchmarks/cpu/permutation.yaml
uv run python scripts/validate_benchmark_suite.py benchmarks/gpu/dense.yaml benchmarks/gpu/einsum.yaml benchmarks/gpu/sparse.yaml benchmarks/gpu/tensornetwork.yaml
bash tests/test_suite_result_layout.sh
bash tests/test_run_all_docs_outputs.sh
bash tests/test_clean_extern_deps.sh
bash tests/test_setup_extern_tenferro_checkout.sh
bash tests/test_permutation_result_schema.sh
cmake -S cpp -B build/cpp-plan-test
cmake --build build/cpp-plan-test --target einsum_plan_test
ctest --test-dir build/cpp-plan-test --output-on-failure

License

MIT

About

tenferro benchmark suite

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages