Skip to content

NVRTC compile error: could not open source file vector_types.h in TsodyksMarkram synapse test - #122

Open
AlekSimpson wants to merge 7 commits into
nightlyfrom
SC-119_FixNvrtcIncludePathForVectorTypes
Open

AlekSimpson wants to merge 7 commits into
nightlyfrom
SC-119_FixNvrtcIncludePathForVectorTypes

Conversation

@AlekSimpson

Copy link
Copy Markdown
Owner

compile_kernel in src/core/backend.cpp invoked nvrtcCompileProgram with zero options; NVRTC
(unlike nvcc) has no default include search path, so any generated CUDA source that #includes a
CUDA header — e.g. <vector_types.h>, emitted by gpu_source.cpp:1411 for any codegen path that
walks shared-basis U/V edge storage via float4 (needs_edge_walk_anywhere) — failed with "could
not open source file". This is a systemic gap, not specific to the one test it was first noticed
on: no generated kernel ever got a CUDA include path, other paths just didn't need one.

Fix passes -I<cuda>/include on every compile_kernel call, baked in at compile time via a new
SPIKECOREC_CUDA_INCLUDE_DIR define derived from the Makefile's existing CUDA_PATH (with a
matching in-source fallback default), mirroring the existing SPIKECOREC_NML_STD_LIB_DIR pattern.

This ticket turned out to be entangled with #116 (already merged as PR #120) — AssembledModel
throws on the first kernel-compile failure, so before #116 landed, ~41 of #116's tests were also
silently blocked behind this bug once #116's extern "C" issue was fixed. Verified: standalone,
this fix eliminates all vector_types.h/NVRTC-include errors with pass count holding at 322/59
(the entangled test still fails via #116's error, expected since #116 isn't in this branch's
ancestry). Combined with #116, the suite goes to 363 passed / 18 failed, and
TsodyksMarkram.per_edge_state_evolves_correctly_across_repeated_presynaptic_spikes passes. The
remaining 18 failures are #117 (xcrun/MSL, macOS-only) and #118 (WeightMatrix numeric), untouched.

Branched off #114 (merged with #108/#109), so this diff shows those fixes too until they merge first.

Acceptance criteria

Reviewer summary: PASS. Independently reproduced both the standalone (322/59, zero
vector_types errors) and combined-with-#116 (363/18, TsodyksMarkram passing) results in a throwaway
scratch worktree, since removed. Code review: sound CUDA_PATH-derived include path, sane
fallback default, correct argument count into nvrtcCompileProgram, exactly one call site in the
codebase (no missed spots), no scope creep.

Closes #119

AlekSimpson and others added 7 commits July 17, 2026 11:54
…gcc CUDA build

types.h defines Vector<T> and UnorderedMap<K,V> aliases but relied on transitive
includes from <string>/<any> that only hold on macOS clang/libc++, not Linux
gcc/libstdc++. Add the two missing includes directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JWoo8NdTRVcnAzxdiXJbJ
… build ambiguity

CUDA's <vector_types.h> (pulled in via cuda_runtime.h) also declares a global
::float4, which collided with spikecorec::float4 wherever this file's bare
`float4` was looked up outside the spikecorec namespace (the anonymous-
namespace refit helpers, plus every other unqualified use for consistency).
Pure name-resolution fix, no behavior change.
… header (ticket #108)

backend.h only included <cuda.h> (the CUDA Driver API header) but called
Runtime API functions/constants (cudaMemPrefetchAsync, cudaCpuDeviceId,
cudaMemAdvise, cudaMemAdviseSetReadMostly) declared in <cuda_runtime.h>.
Add <cuda_runtime.h> alongside <cuda.h>, which is still required for the
Driver API types CUfunction/CUmodule used by KernelHandle.
…109)

backend.cpp's CUDA branches tried to launch raw __global__ kernels directly
with nvcc-only triple-chevron syntax from a file compiled by host g++, and
referenced an undeclared `stream` identifier at every site. Implements all
10 launch_* functions declared in kernels.cuh (plus launch_step_no_active_
optimization, needed for gpu_step's active_set_optimization_enabled==false
path, which kernels.cuh had no wrapper for) inside kernels.cu, each doing
the config computation + <<<...>>> launch of its corresponding kernel, and
rewires backend.cpp's 10 gpu_* CUDA branches to call them with stream=nullptr,
keeping the existing post-launch synchronize_gpu_work() (required due to
concurrentManagedAccess=0 on this Jetson). Removes the duplicate `s64
total_pairs` redeclarations in gpu_neighbor_weights/gpu_k2tree_get_neighbors_batch
and the local spikecorec::LaunchConfig variables that collided with
spikecorec::cuda::LaunchConfig. Also adds the missing `coefficients`
parameter to kernels.cuh's launch_neighbor_weights declaration (already
present on gpu_neighbor_weights/the kernel itself, just missing from the
header) and const_casts gpu_step's const float4* U/V to the non-const
pointers the mutating step kernels require.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JWoo8NdTRVcnAzxdiXJbJ
….cpp to fix CUDA build ambiguity

Same root cause and fix pattern as #107 (weight_matrix.cpp): bare float4
collides with CUDA's built-in ::float4 from <vector_types.h> wherever both
types are visible via this file's `using namespace spikecorec;`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JWoo8NdTRVcnAzxdiXJbJ
…pers' into SC-119_FixNvrtcIncludePathForVectorTypes
…clude CUDA headers

nvrtcCompileProgram was invoked with zero compile options, so NVRTC (unlike nvcc) had no
include search path at all -- any generated CUDA source that #includes a CUDA header (e.g.
gpu_source.cpp's `#include <vector_types.h>`, emitted whenever a codegen path needs float4
to walk the shared-basis U/V edge storage) failed with "could not open source file". Now
passes -I<cuda>/include, baked in at compile time from the same CUDA_PATH the Makefile
already uses for the host compiler, fixing this generally rather than special-casing one
kernel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JWoo8NdTRVcnAzxdiXJbJ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant