Conversation
KernelRuntime enumerates a cubin's symbols by shelling out to `$CUDA_HOME/bin/cuobjdump -symbols` and asserting `exit_code == 0`. cuobjdump is packaged separately from nvcc (conda `cuda-cuobjdump`, `cuda-command-line-tools`), so a JIT-only CUDA install that has nvcc but not cuobjdump compiles kernels fine and then dies here with an opaque `Assertion error (...: exit_code == 0)` — the captured command output, which call_external_command already collects via `2>&1`, was discarded. Check for cuobjdump up front with an actionable message, and on non-zero exit surface the command, exit code, and its captured output instead of the bare assertion. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gyu6PMVkDQZkoMkGKQC63M
| // install that has `nvcc` but not `cuobjdump` compiles kernels fine and then fails | ||
| // here, so check for it explicitly with an actionable message instead of a bare | ||
| // `exit_code == 0` assertion. | ||
| if (not std::filesystem::exists(cuobjdump_path)) |
There was a problem hiding this comment.
🔵 suggestion: std::filesystem::exists(cuobjdump_path) uses the throwing overload; on exotic failures (e.g. permission-denied on a parent directory) it can throw std::filesystem::filesystem_error instead of the friendly DGException. Consider the noexcept overload std::filesystem::exists(cuobjdump_path, ec) so any filesystem hiccup still funnels into the actionable message. Low impact since the codebase already uses the throwing overload elsewhere (e.g. check_validity).
🤖 v5
| "`cuda-command-line-tools`) or point `CUDA_HOME` at a full CUDA toolkit.", | ||
| cuobjdump_path.c_str())); | ||
|
|
||
| const auto command = fmt::format("{} -symbols {}", cuobjdump_path.c_str(), cubin_path.c_str()); |
There was a problem hiding this comment.
🔵 suggestion: Pre-existing (not introduced here): the command interpolates cuobjdump_path and cubin_path unquoted, so a CUDA_HOME or DG_JIT_CACHE_DIR containing spaces breaks the popen'd shell command. The new error message that echoes the full command actually makes this easier to diagnose, but wrapping both paths in quotes would fix it outright.
🤖 v5
| // install that has `nvcc` but not `cuobjdump` compiles kernels fine and then fails | ||
| // here, so check for it explicitly with an actionable message instead of a bare | ||
| // `exit_code == 0` assertion. | ||
| if (not std::filesystem::exists(cuobjdump_path)) |
There was a problem hiding this comment.
🔵 suggestion: The existence pre-check only tests that a file exists at $CUDA_HOME/bin/cuobjdump; a present-but-non-executable or mismatched-architecture binary still falls through to the (now well-diagnosed) exit-code path. If you want the very first failure to be maximally actionable, consider also checking std::filesystem::status(...).permissions() for the execute bit in the same pre-flight block — optional, since the improved non-zero-exit reporting already surfaces the underlying error text.
🤖 v4f
🤖 ds-review-bot Code Reviewv6变更仅改进 cuobjdump 缺失或执行失败时的诊断信息,成功路径未受影响,未发现会破坏现有行为的缺陷。 v5Commit b75f054 ('jit: actionable error when cuobjdump is missing or fails') changes exactly one file, csrc/jit/kernel_runtime.hpp, and fully matches the issue's intent: (1) an up-front std::filesystem::exists check on v4fThe MR improves the diagnostics around Files reviewed: 1 📍 未定位到 diff 的评论🔵 suggestion 🔵 suggestion |
Problem
KernelRuntimefinds a cubin's kernel symbol by shelling out to$CUDA_HOME/bin/cuobjdump -symbols kernel.cubinand then assertingDG_HOST_ASSERT(exit_code == 0).cuobjdumpis part of the CUDA toolkit but is packaged separately fromnvcc— e.g. conda'scuda-cuobjdump(orcuda-command-line-tools).A JIT-only CUDA install that has
nvccbut notcuobjdumpcompiles kernelsfine and then dies here with:
which gives the user nothing to go on.
call_external_commandalreadycaptures the command's combined stdout/stderr (it appends
2>&1), but thatoutput — which contains the actual
cuobjdump: not found/ error text — isdiscarded on the failure path.
This is easy to hit from the
kernels/Hugging Face flow and from minimalCUDA containers.
Change
Diagnostics only, no behavior change on the success path:
cuobjdumpexists up front and, if not, report its expected path andhow to fix it (install
cuda-cuobjdump/cuda-command-line-tools, or pointCUDA_HOMEat a full toolkit).instead of the bare
exit_code == 0assertion.Reproduce
A CUDA env with
nvccbut nocuobjdump(e.g.conda create -c nvidia cuda-nvccwithoutcuda-cuobjdump) — any JIT kernelload reaches this path.
🤖 Generated with Claude Code