You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 36b209a
Browse filesBrowse the repository at this point in the historyBrowse files
GPU: take the work-item indices from the kernel parameters
Metal has no ambient work-item builtins: the thread and threadgroup positions
exist only as attributes on the kernel entry point, so get_global_id() and its
siblings cannot be expressed in a device function the way CUDA's threadIdx or
OpenCL's get_global_id() can.
Every one of these call sites is inside a function that already receives
nBlocks, nThreads, iBlock and iThread, or is one call away from one, so they now
use those directly. The substitution is exact on every backend: CUDA, HIP and
OpenCL pass precisely these values into Thread(), and the CPU backend passes
nThreads = 1 and iThread = 0, which is what the CPU definitions of the macros
already assumed.
Four helpers had no index in scope and gain one parameter: sortInBlock,
GPUTPCCFClusterizer::buildCluster, GPUTPCCFNoiseSuppression::findMinimaAndPeaks
and GPUTPCCFPeakFinder::isPeak. GPUCA_SHARED_CACHE and GPUCA_TBB_KERNEL_LOOP
likewise take the indices as arguments rather than capturing them from the
expansion context.
The get_*() macros remain for the kernel entry point in
GPUReconstructionKernelMacros.h, which is the one place where Metal does provide
them.
constfloat lowPtThresh = Param().rec.tpc.rejectQPtB5 * 1.1f; // Might need to merge tracks above the threshold with parts below the rejection threshold
1940
-
for (uint32_t i = get_global_id(0); i < mMemory->nMergedTracks; i += get_global_size(0)) {
1940
+
for (uint32_t i = (iBlock * nThreads + iThread); i < mMemory->nMergedTracks; i += (nBlocks * nThreads)) {
0 commit comments