Skip to content

Release v0.1 - #28

Open
ThrudPrimrose wants to merge 776 commits into
mainfrom
release-v0.1
Open

ThrudPrimrose wants to merge 776 commits into
mainfrom
release-v0.1

Conversation

@ThrudPrimrose

Copy link
Copy Markdown
Collaborator

No description provided.

Yakup Koray Budanaz added 30 commits September 29, 2026 15:11
Measured leaders at the new XL: spgemm_hash 262144^3 with 4.2M nnz (numba 10.9 s), vexx_k ngrid 80 /
nbnd 36 (numba 6.6 s), quatrex_rgf NE 48 (numba 3.9 s), mg_vcycle ncycles 3 (C 14.6 s),
mixed_precision_ir N 9000 (numba ~16 s from 19.7 s at 9600), rk4_ensemble NSYS 568117 (C 14.9 s),
sgs_pcg 152^3 (C 10.2 s), jfnk_bratu N 704 (C ~14 s from 18.6 s at 768), householder_qr M 800000
and lanczos_reorth 135^3 (new numba ~16 s from 20.6 / 20.4 s at the old XL), nfa_frontier T
51772817 (new numba ~12 s from 5.4 s at L). ls3df_scf takes the old L as XL and a smaller L; its
numba leader is estimated (~17 s) until the rebuilt image stops the scipy-openblas crash. Fuzz upper
bounds follow the new XL.

Hand parallel numba references replace the generated ones for householder_qr (row-wise passes;
1063 s -> 20.6 s at the old XL), lanczos_reorth (two classical Gram-Schmidt passes over rows of Q;
> 2 h -> 20.4 s) and nfa_frontier (prange over components; > 2 h at XL).
# Conflicts:
#	hpcagent_bench/benchmarks/scientific_computing/sparse_linear_algebra/spgemm_hash/spgemm_hash.yaml
…ng's dtype (torch agent, interrupted; tests not re-run)
Schema v1 (layout, layout_request, size_scale, scale_axes on grades; race record per input),
same-configuration arms folded, sourceless finals dropped. Conflicts: the dbv1 side of
scoring/service/regrade/test_cpu_refuses_gpu; Score.layout_request now carries the submission's
sparse_config. The v11 GPU waves (gpuv2-/gpuv4-) recorded device cpu and the first qwen triton
wave language c: migrate_db.recorded_wrong corrects them, so they fold into their arms.
ML precision: bf16 references compile and grade through their binding's dtype, bf16/fp16 tolerances
bounded by the reduction depth, background warm-up of the ML denominator behind agent requests,
golden outputs kept in the track's dtype, and the preparation job. Conflicts: emit_bridge keeps the
sparse layout docs and gains the abi emit; the old XL scale-every-dimension sizing (track_scaled,
xl_size_scale) stays removed for the constant-bytes rule; Benchmark keeps the track of the loaded
manifest.
…iants are gone and Benchmark.get_data no longer takes one (every run-framework call failed with a TypeError after the sparse merge)
…on_krylov, householder_qr, jfnk_bratu, lanczos_reorth, mg_vcycle, mixed_precision_ir, rk45_ensemble, sgs_pcg, sparse_cholesky) under the scicomp budget
…installed versions (the wheels' bundled scipy-openblas crashed under numba prange), numba on its OpenMP layer, gated by 2 x nproc concurrent BLAS callers
…oot: containers/ holds images only, and the agent payload stays outside the package so no image carries tool scripts (bound from the checkout at launch). The staged-launch test runs the driver the way run_cluster.sh does (its own directory first on sys.path), and the hand numba references added since the ignore list was written get their negations.
… image's targets from its <image>/image.sh (targets and roles, pinned base, build args, mirror and cache inputs); build_and_verify.sbatch <image> verifies every role it wrote through verify_image.sbatch, which renders the role's production EDF template (ce_render_edf, shared with install_edfs.sh) and runs the GPU-arch, MPI and GH200 device checks where they apply; registry.sh promote|push|pull (+ registry.sbatch) replaces promote_image.sh, push_image.sh, pull_image.sh, push_images.sbatch and pull_images.sbatch, and promotion re-renders only the EDFs of the images it moved. The per-image build.sh and build.sbatch are gone (-17 scripts, +9 files).
…ged is stale (harness/grading_cuts.yaml: tsvc_2_s3112 and warpx_boris_push tolerances, channel_flow and warpx_field_gather numba references, quatrex_rgf / vexx_k / ls3df_scf XL), and the owed worklist also takes every sourced submission a since-fixed grading failed, so the correct blocked prefix sums the one-ulp s3112 atol rejected are graded again (+38 owed items, none lost)
…ortran and HIP submissions rebuild with ASan+UBSan (HIP device code for xnack+, HSA_XNACK=1) and run once at preset S in a sealed, freshly exec'd child with the sanitizer runtime preloaded, so the numpy arrays carry redzones; CUDA runs under compute-sanitizer memcheck. A memory error rejects (reason "sanitizer: ..."), undefined behaviour alone flags the grade suspect, a missing runtime is recorded as not applied. docs/anti_cheat.md lists every gate, its verdict and where it lives.
… reports the spack prefix the view links to, which failed the CPU image build) and look for bundled copies where wheels put them (<pkg>.libs)
…ner of best-of(numba, c) for 53 scicomp and solver kernels (sweeps 655298, 655303, 655611); sparse_cholesky XL is EDGE 36 (its fastest baseline took 24.6 s at 40)
…profv3, so the module docstring, the rocprof skill and one test comment stop describing a v1 route
…OutputCounter, SUSPECT_RATIO) and its agent_driver pass-through
…ment

gmres binds y to the length-N matvec result and later to the length-m lstsq solution; the deferred malloc used the last shape (m), so the matvec copy overflowed the heap and the C baseline died with SIGSEGV at preset S. The reassign marker carries its own shape, so it now sizes the buffer.
Yakup Koray Budanaz added 30 commits October 2, 2026 21:30
…i200 setups pass RECORD_STUDY naming mi200), clean test removed, legacy tag spelling in make_problems_select
Neither is an NPBench kernel. Their hand-written JAX references stay, and the reference test checks them beside the tag.
Validated on an MI300 node (ROCm torch and triton, preset S, float64) against the NumPy oracle: 29 of 56 failed.
Autotuned kernels that update a buffer in place now name it in restore_value, so the tuning trials stop accumulating
into it. Scalars reach fp64 kernels through pointers instead of as fp32 arguments. Entry parameters follow the manifest
names. Per kernel: stockham_fft ended in the wrong buffer for odd K and built its twiddles in fp32; doitgen only
handled NP up to its block size; scattering_self_energies used 5-D tiles in the complex matmul helper, a 6-D offset
helper for the 7-D D array, and one w too few; cavity_flow applied the pressure boundary conditions to the previous
sweep, left the result in the wrong buffer for odd nit, and its cooperative launch no longer depends on the first
autotune configs; channel_flow ping-ponged into uninitialized buffers and returned a step count the harness bound to p;
jacobi_1d ran one sweep too many; trisolv and deriche read stores of other threads without a barrier.
…dware> off it (apply_hardware pins it; a given RECORD_STUDY must name the hardware); legacy NULL datatype handling, the mw4x5-final stamp and the tag.sh references deleted; baseline_sweep and docs use the llr40 tag
… that stays rebuilt

llvm builds with SPACK_ALWAYS_CFLAGS and SPACK_ALWAYS_CXXFLAGS set to -pipe on the install command instead of cflags and cxxflags on the spec, which made spack fail in python's build hook (no compiler-wrapper in the llvm spec, build 660463); the llvm spec and its hash are the ones the buildcache already holds. In the cpu build 660464 the scipy sync reinstalled numpy from its wheel, because uv reinstalls a package whose build settings differ and that sync carried none, and it also swapped torch's CPU wheel and rich for other releases because it named no extra. Every sync of the image environment now carries the install's extra and judge-proxy group, the scipy sync and every later one skip numpy (and scipy), so the source-built numpy keeps its OpenBLAS.
compute kept 8-element blocks under the optimizer budget, which needs more programs than the launch grid allows over its 87M
elements, so it now tries the large blocks first. seidel_2d added the seven neighbours in a different order than the
reference. mandelbrot1 took its bounds as fp32 scalars and a None dtype, and both mandelbrot kernels computed the
coordinates with other arithmetic than np.linspace.
… vary_inputs_pool_size, REDUCTIONS_FINAL, pooled_seeds, grade_under pool env
…ard left from before the kernel rename cannot fail them
… new floor), and a 256-colour README overview figure under the 500 KiB limit
… finds the tool registry in hpcagent_agent

Build 660586: numpy_on_openblas.sh syncs the image's extra, which on amdgpu includes cupy, so the sync tried the HIP source build before the layer that sets its environment; the Dockerfile now passes --no-install-package cupy to it, as the first sync does. Build 660583 verified everything but the tools check, which looked for tools/mcp_server.py under the bound agent directory, where the package move put it at hpcagent_agent/tools/mcp_server.py.
…sions, ratchet counts, user-path scan sample, jax ignore test, registry toctree
… one final grade per submission (credited stamp first, then newest, under the oldest id); mwd-final and mw4x5-final registered as retired stamps; the v3 migration checks grades per timing_reduction
…1-final2) is migrated and verified, its v2 copy archived
…_map table resolves each port's upstream model for the agreement tests), scripts/check_inputs_finite.py and the kimi SGLang smoke sbatch (serve-only.sbatch smokes every serving candidate)
…city_tendencies C++ and its oracle test (numpy vs the original Fortran stays in tests/ports); the LULESH port test compiles the corpus lulesh_reference.f90 instead of its tests/ports duplicate
…ummary are the analysis), the stdlib signed-rank copy (EXACT_MAX_N and use_exact move into stats.summary; the audits test scipy's exact and corrected paths) and the GPU profiler smoke scripts; test_iteration_counts keeps the iteration_counts half
… a serve-private endpoint through alps-endpoint.sh); seissol_tensor_contraction points at seissol_batched_gemm's identical REFERENCES.md
…out= and astype(copy=) rules (each with its file scope and fix) in place of two near-identical scripts
…ned.py (setups_figure, paired_figure and their readers, drawn by nothing since the llr40 figures) with its tests and registry pins
…t drove it; the 219 loop_level_reasoning _reference.c files stay as frozen sources (headers say edit by hand), checked by the ABI and numeric-oracle tests
…arguments (--experiment, --setups, --include-incomplete, --repeats) and population.select_setups replace three copies; plot_scaling's --setup is --setups like the others
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant