Skip to content

Port thrust extrema algs to use CUB - #4970

Closed
bernhardmgruber wants to merge 1 commit into
NVIDIA:mainfrom
bernhardmgruber:port_extrema_to_cub
Closed

Port thrust extrema algs to use CUB#4970
bernhardmgruber wants to merge 1 commit into
NVIDIA:mainfrom
bernhardmgruber:port_extrema_to_cub

Conversation

@bernhardmgruber

@bernhardmgruber bernhardmgruber commented Jun 12, 2025

Copy link
Copy Markdown
Contributor

Fixes #1626

thrust.bench.reduce.basic.base on RTX 5090

# base

## [0] NVIDIA GeForce RTX 5090

|  T{ct}  |  Elements  |   Ref Time |   Ref Noise |   Cmp Time |   Cmp Noise |      Diff |   %Diff |  Status  |
|---------|------------|------------|-------------|------------|-------------|-----------|---------|----------|
|   I8    |    2^16    |  18.996 us |       9.69% |  18.976 us |       7.08% | -0.020 us |  -0.10% |   SAME   |
|   I8    |    2^20    |  21.213 us |      10.46% |  20.692 us |       6.91% | -0.521 us |  -2.46% |   SAME   |
|   I8    |    2^24    |  38.706 us |       4.53% |  38.294 us |       4.14% | -0.412 us |  -1.06% |   SAME   |
|   I8    |    2^28    | 235.330 us |       0.90% | 234.404 us |       0.76% | -0.926 us |  -0.39% |   SAME   |
|   I16   |    2^16    |  18.491 us |       9.88% |  18.110 us |       7.43% | -0.381 us |  -2.06% |   SAME   |
|   I16   |    2^20    |  21.037 us |      10.49% |  20.769 us |       8.16% | -0.268 us |  -1.27% |   SAME   |
|   I16   |    2^24    |  43.850 us |       2.66% |  43.891 us |       2.63% |  0.041 us |   0.09% |   SAME   |
|   I16   |    2^28    | 368.571 us |       0.89% | 370.061 us |       3.76% |  1.490 us |   0.40% |   SAME   |
|   I32   |    2^16    |  19.128 us |      14.85% |  19.000 us |       9.60% | -0.128 us |  -0.67% |   SAME   |
|   I32   |    2^20    |  22.113 us |       7.64% |  22.448 us |       9.29% |  0.335 us |   1.52% |   SAME   |
|   I32   |    2^24    |  65.135 us |       3.37% |  66.471 us |       7.21% |  1.336 us |   2.05% |   SAME   |
|   I32   |    2^28    | 691.652 us |       0.52% | 691.844 us |       0.42% |  0.192 us |   0.03% |   SAME   |
|   I64   |    2^16    |  20.190 us |       7.42% |  20.795 us |       7.30% |  0.605 us |   2.99% |   SAME   |
|   I64   |    2^20    |  26.892 us |       3.10% |  27.530 us |       5.01% |  0.638 us |   2.37% |   SAME   |
|   I64   |    2^24    | 108.388 us |       1.93% | 108.820 us |       1.53% |  0.431 us |   0.40% |   SAME   |
|   I64   |    2^28    |   1.328 ms |       0.19% |   1.327 ms |       0.19% | -0.612 us |  -0.05% |   SAME   |
|  I128   |    2^16    |  20.705 us |      25.83% |  20.381 us |       7.45% | -0.324 us |  -1.56% |   SAME   |
|  I128   |    2^20    |  32.301 us |       6.55% |  32.129 us |       4.53% | -0.172 us |  -0.53% |   SAME   |
|  I128   |    2^24    | 199.346 us |       1.07% | 200.923 us |       1.24% |  1.577 us |   0.79% |   SAME   |
|  I128   |    2^28    |   2.597 ms |       0.13% |   2.599 ms |       0.15% |  1.824 us |   0.07% |   SAME   |
|   F32   |    2^16    |  19.672 us |       9.66% |  19.637 us |       7.84% | -0.035 us |  -0.18% |   SAME   |
|   F32   |    2^20    |  21.911 us |       6.59% |  21.860 us |       6.58% | -0.051 us |  -0.23% |   SAME   |
|   F32   |    2^24    |  64.753 us |       2.92% |  64.713 us |       2.61% | -0.040 us |  -0.06% |   SAME   |
|   F32   |    2^28    | 687.966 us |       0.30% | 689.294 us |       0.33% |  1.328 us |   0.19% |   SAME   |
|   F64   |    2^16    |  22.964 us |       4.34% |  23.073 us |       5.25% |  0.109 us |   0.48% |   SAME   |
|   F64   |    2^20    |  28.493 us |       4.50% |  28.488 us |       4.24% | -0.005 us |  -0.02% |   SAME   |
|   F64   |    2^24    | 109.701 us |       1.61% | 109.534 us |       1.49% | -0.167 us |  -0.15% |   SAME   |
|   F64   |    2^28    |   1.333 ms |       0.37% |   1.332 ms |       0.30% | -0.987 us |  -0.07% |   SAME   |

@bernhardmgruber
bernhardmgruber requested a review from a team as a code owner June 12, 2025 16:15
@bernhardmgruber
bernhardmgruber requested a review from elstehle June 12, 2025 16:15
@github-project-automation github-project-automation Bot moved this to Todo in CCCL Jun 12, 2025
@cccl-authenticator-app cccl-authenticator-app Bot moved this from Todo to In Review in CCCL Jun 12, 2025
Comment on lines +160 to +166
// TODO(bgruber): the previous thrust implementation avoided creating an initial value. Should we bring this back?
// There is no reduction in CUB without init.
const auto offset = thrust::cuda_cub::detail::reduce_n_impl(
policy,
zip_first,
num_items,
tuple_t{cub::FutureValue(first), offset_t{0}},

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I picked the first element of the input range as initial value, since I cannot come up with an identity element for any user-defined binary predicate. I am not worried about the performance impact of loading that first element twice. However, it still feels a bit hackish. Does anyone have any better ideas?

Comment on lines +168 to +172
[](execution_policy<Derived>& policy, const tuple_t* result_ptr) {
// TODO(bgruber): I only want to download the offset (element 1) and not the value (element 0), but how can I
// legally form a pointer to that tuple element?
// return get_value(policy, &thrust::get<1>(*result_ptr));
return thrust::get<1>(get_value(policy, result_ptr));

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@miscco If I have a tuple<A, B>* pointing to device memory, can I legally get a B* to the second tuple element? get<1> requires to dereference the pointer, which ... may be UB?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am wondering about the situation here.

If there is a tuple<A, B> at the memory location the pointer points at, then it is perfectly fine to use get even if the element in the A slot is not properly initialized

@elstehle

Copy link
Copy Markdown
Contributor

Thrust's version previously supported large custom types using virtual shared memory. That's currently not supported by cub::Reduce. For other algorithms, this is something we made sure to tackle first. Are we consciously not doing this for the extrema computation?

@bernhardmgruber
bernhardmgruber marked this pull request as draft June 12, 2025 16:27
@copy-pr-bot

copy-pr-bot Bot commented Jun 12, 2025

Copy link
Copy Markdown
Contributor

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@cccl-authenticator-app cccl-authenticator-app Bot moved this from In Review to In Progress in CCCL Jun 12, 2025
@github-actions

Copy link
Copy Markdown
Contributor
🟨 CI finished in 1h 50m: Pass: 63%/130 | Total: 2d 14h | Avg: 29m 03s | Max: 1h 16m | Hits: 78%/73181
  • 🟥 thrust: Pass: 0%/47 | Total: 16h 25m | Avg: 20m 57s | Max: 1h 14m

    🟥 cmake_options
      🟥 -DTHRUST_DISPATCH_TYPE=Force32bit Pass:   0%/2   | Total: 18m 54s | Avg:  9m 27s | Max: 18m 54s
    🟥 cpu
      🟥 amd64              Pass:   0%/45  | Total: 15h 45m | Avg: 21m 00s | Max:  1h 14m
      🟥 arm64              Pass:   0%/2   | Total: 39m 43s | Avg: 19m 51s | Max: 20m 42s
    🟥 ctk
      🟥 12.0               Pass:   0%/5   | Total:  2h 22m | Avg: 28m 24s | Max:  1h 03m
      🟥 12.9               Pass:   0%/42  | Total: 14h 02m | Avg: 20m 04s | Max:  1h 14m
    🟥 cudacxx
      🟥 ClangCUDA19        Pass:   0%/2   | Total: 30m 14s | Avg: 15m 07s | Max: 16m 01s
      🟥 nvcc12.0           Pass:   0%/5   | Total:  2h 22m | Avg: 28m 24s | Max:  1h 03m
      🟥 nvcc12.9           Pass:   0%/40  | Total: 13h 32m | Avg: 20m 19s | Max:  1h 14m
    🟥 cudacxx_family
      🟥 ClangCUDA          Pass:   0%/2   | Total: 30m 14s | Avg: 15m 07s | Max: 16m 01s
      🟥 nvcc               Pass:   0%/45  | Total: 15h 54m | Avg: 21m 13s | Max:  1h 14m
    🟥 cxx
      🟥 Clang14            Pass:   0%/4   | Total:  1h 12m | Avg: 18m 02s | Max: 19m 01s
      🟥 Clang15            Pass:   0%/2   | Total: 41m 29s | Avg: 20m 44s | Max: 22m 05s
      🟥 Clang16            Pass:   0%/2   | Total: 35m 50s | Avg: 17m 55s | Max: 18m 04s
      🟥 Clang17            Pass:   0%/2   | Total: 38m 28s | Avg: 19m 14s | Max: 20m 00s
      🟥 Clang18            Pass:   0%/2   | Total: 35m 01s | Avg: 17m 30s | Max: 17m 40s
      🟥 Clang19            Pass:   0%/7   | Total:  1h 26m | Avg: 12m 22s | Max: 19m 01s
      🟥 GCC7               Pass:   0%/2   | Total: 39m 13s | Avg: 19m 36s | Max: 19m 53s
      🟥 GCC8               Pass:   0%/1   | Total: 21m 40s | Avg: 21m 40s | Max: 21m 40s
      🟥 GCC9               Pass:   0%/2   | Total: 42m 20s | Avg: 21m 10s | Max: 22m 03s
      🟥 GCC10              Pass:   0%/2   | Total: 40m 35s | Avg: 20m 17s | Max: 20m 31s
      🟥 GCC11              Pass:   0%/2   | Total: 40m 44s | Avg: 20m 22s | Max: 21m 52s
      🟥 GCC12              Pass:   0%/2   | Total: 39m 27s | Avg: 19m 43s | Max: 19m 47s
      🟥 GCC13              Pass:   0%/10  | Total:  1h 57m | Avg: 11m 46s | Max: 22m 02s
      🟥 MSVC14.29          Pass:   0%/2   | Total:  2h 07m | Avg:  1h 03m | Max:  1h 03m
      🟥 MSVC14.43          Pass:   0%/3   | Total:  2h 26m | Avg: 48m 50s | Max:  1h 14m
      🟥 NVHPC25.5          Pass:   0%/2   | Total: 59m 10s | Avg: 29m 35s | Max: 31m 26s
    🟥 cxx_family
      🟥 Clang              Pass:   0%/19  | Total:  5h 09m | Avg: 16m 17s | Max: 22m 05s
      🟥 GCC                Pass:   0%/21  | Total:  5h 41m | Avg: 16m 16s | Max: 22m 03s
      🟥 MSVC               Pass:   0%/5   | Total:  4h 34m | Avg: 54m 53s | Max:  1h 14m
      🟥 NVHPC              Pass:   0%/2   | Total: 59m 10s | Avg: 29m 35s | Max: 31m 26s
    🟥 gpu
      🟥 h100               Pass:   0%/2   | Total: 13m 18s | Avg:  6m 39s | Max: 13m 18s
      🟥 rtx2080            Pass:   0%/35  | Total: 14h 01m | Avg: 24m 02s | Max:  1h 14m
      🟥 rtx4090            Pass:   0%/10  | Total:  2h 10m | Avg: 13m 03s | Max:  1h 11m
    🟥 jobs
      🟥 Build              Pass:   0%/40  | Total: 16h 25m | Avg: 24m 37s | Max:  1h 14m
      🟥 TestCPU            Pass:   0%/3  
      🟥 TestGPU            Pass:   0%/4  
    🟥 sm
      🟥 90                 Pass:   0%/2   | Total: 13m 18s | Avg:  6m 39s | Max: 13m 18s
      🟥 90;90a;100         Pass:   0%/1   | Total: 22m 02s | Avg: 22m 02s | Max: 22m 02s
    🟥 std
      🟥 17                 Pass:   0%/21  | Total:  9h 30m | Avg: 27m 09s | Max:  1h 14m
      🟥 20                 Pass:   0%/24  | Total:  6h 35m | Avg: 16m 29s | Max:  1h 11m
    
  • 🟩 cub: Pass: 100%/47 | Total: 1d 18h | Avg: 54m 29s | Max: 1h 16m | Hits: 74%/57723

    🟩 cpu
      🟩 amd64              Pass: 100%/45  | Total:  1d 16h | Avg: 54m 27s | Max:  1h 16m | Hits:  75%/55211 
      🟩 arm64              Pass: 100%/2   | Total:  1h 50m | Avg: 55m 09s | Max: 57m 52s | Hits:  68%/2512  
    🟩 ctk
      🟩 12.0               Pass: 100%/5   | Total:  5h 07m | Avg:  1h 01m | Max:  1h 13m | Hits:  69%/6095  
      🟩 12.9               Pass: 100%/42  | Total:  1d 13h | Avg: 53m 40s | Max:  1h 16m | Hits:  75%/51628 
    🟩 cudacxx
      🟩 ClangCUDA19        Pass: 100%/2   | Total:  1h 03m | Avg: 31m 37s | Max: 32m 17s | Hits:  75%/2161  
      🟩 nvcc12.0           Pass: 100%/5   | Total:  5h 07m | Avg:  1h 01m | Max:  1h 13m | Hits:  69%/6095  
      🟩 nvcc12.9           Pass: 100%/40  | Total:  1d 12h | Avg: 54m 46s | Max:  1h 16m | Hits:  75%/49467 
    🟩 cudacxx_family
      🟩 ClangCUDA          Pass: 100%/2   | Total:  1h 03m | Avg: 31m 37s | Max: 32m 17s | Hits:  75%/2161  
      🟩 nvcc               Pass: 100%/45  | Total:  1d 17h | Avg: 55m 30s | Max:  1h 16m | Hits:  74%/55562 
    🟩 cxx
      🟩 Clang14            Pass: 100%/4   | Total:  3h 55m | Avg: 58m 46s | Max:  1h 00m | Hits:  69%/5026  
      🟩 Clang15            Pass: 100%/2   | Total:  1h 51m | Avg: 55m 36s | Max: 55m 40s | Hits:  69%/2509  
      🟩 Clang16            Pass: 100%/2   | Total:  1h 53m | Avg: 56m 59s | Max: 57m 01s | Hits:  69%/2509  
      🟩 Clang17            Pass: 100%/2   | Total:  1h 49m | Avg: 54m 59s | Max: 56m 48s | Hits:  69%/2509  
      🟩 Clang18            Pass: 100%/2   | Total:  1h 47m | Avg: 53m 59s | Max: 54m 31s | Hits:  69%/2509  
      🟩 Clang19            Pass: 100%/7   | Total:  5h 03m | Avg: 43m 21s | Max:  1h 05m | Hits:  79%/8435  
      🟩 GCC7               Pass: 100%/2   | Total:  1h 57m | Avg: 58m 47s | Max: 59m 48s | Hits:  68%/2512  
      🟩 GCC8               Pass: 100%/1   | Total:  1h 00m | Avg:  1h 00m | Max:  1h 00m | Hits:  68%/1256  
      🟩 GCC9               Pass: 100%/2   | Total:  2h 00m | Avg:  1h 00m | Max:  1h 01m | Hits:  68%/2512  
      🟩 GCC10              Pass: 100%/2   | Total:  1h 57m | Avg: 58m 37s | Max: 59m 07s | Hits:  68%/2513  
      🟩 GCC11              Pass: 100%/2   | Total:  2h 00m | Avg:  1h 00m | Max:  1h 02m | Hits:  68%/2509  
      🟩 GCC12              Pass: 100%/2   | Total:  2h 07m | Avg:  1h 03m | Max:  1h 06m | Hits:  68%/2509  
      🟩 GCC13              Pass: 100%/11  | Total:  7h 58m | Avg: 43m 27s | Max:  1h 07m | Hits:  85%/13824 
      🟩 MSVC14.29          Pass: 100%/2   | Total:  2h 27m | Avg:  1h 13m | Max:  1h 14m | Hits:  74%/2140  
      🟩 MSVC14.43          Pass: 100%/2   | Total:  2h 32m | Avg:  1h 16m | Max:  1h 16m | Hits:  74%/2140  
      🟩 NVHPC25.5          Pass: 100%/2   | Total:  2h 17m | Avg:  1h 08m | Max:  1h 09m | Hits:  68%/2311  
    🟩 cxx_family
      🟩 Clang              Pass: 100%/19  | Total: 16h 21m | Avg: 51m 40s | Max:  1h 05m | Hits:  72%/23497 
      🟩 GCC                Pass: 100%/22  | Total: 19h 02m | Avg: 51m 55s | Max:  1h 07m | Hits:  77%/27635 
      🟩 MSVC               Pass: 100%/4   | Total:  4h 59m | Avg:  1h 14m | Max:  1h 16m | Hits:  74%/4280  
      🟩 NVHPC              Pass: 100%/2   | Total:  2h 17m | Avg:  1h 08m | Max:  1h 09m | Hits:  68%/2311  
    🟩 gpu
      🟩 h100               Pass: 100%/3   | Total:  1h 30m | Avg: 30m 18s | Max: 35m 48s | Hits:  89%/3771  
      🟩 rtx2080            Pass: 100%/36  | Total:  1d 11h | Avg: 59m 08s | Max:  1h 16m | Hits:  69%/43902 
      🟩 rtxa6000           Pass: 100%/8   | Total:  5h 41m | Avg: 42m 40s | Max:  1h 07m | Hits:  92%/10050 
    🟩 jobs
      🟩 Build              Pass: 100%/39  | Total:  1d 14h | Avg: 58m 45s | Max:  1h 16m | Hits:  69%/47671 
      🟩 DeviceLaunch       Pass: 100%/1   | Total: 39m 07s | Avg: 39m 07s | Max: 39m 07s | Hits:  99%/1257  
      🟩 GraphCapture       Pass: 100%/1   | Total: 31m 33s | Avg: 31m 33s | Max: 31m 33s | Hits:  99%/1257  
      🟩 HostLaunch         Pass: 100%/3   | Total:  1h 47m | Avg: 35m 55s | Max: 37m 07s | Hits:  99%/3769  
      🟩 TestGPU            Pass: 100%/3   | Total:  1h 30m | Avg: 30m 17s | Max: 34m 27s | Hits:  99%/3769  
    🟩 sm
      🟩 90                 Pass: 100%/3   | Total:  1h 30m | Avg: 30m 18s | Max: 35m 48s | Hits:  89%/3771  
      🟩 90;90a;100         Pass: 100%/1   | Total: 59m 25s | Avg: 59m 25s | Max: 59m 25s | Hits:  68%/1257  
    🟩 std
      🟩 17                 Pass: 100%/21  | Total: 21h 06m | Avg:  1h 00m | Max:  1h 16m | Hits:  69%/25525 
      🟩 20                 Pass: 100%/26  | Total: 21h 34m | Avg: 49m 47s | Max:  1h 15m | Hits:  78%/32198 
    
  • 🟩 cudax: Pass: 100%/26 | Total: 3h 04m | Avg: 7m 06s | Max: 15m 01s | Hits: 91%/15130

    🟩 cpu
      🟩 amd64              Pass: 100%/22  | Total:  2h 43m | Avg:  7m 25s | Max: 15m 01s | Hits:  91%/12710 
      🟩 arm64              Pass: 100%/4   | Total: 21m 33s | Avg:  5m 23s | Max:  5m 42s | Hits:  90%/2420  
    🟩 ctk
      🟩 12.0               Pass: 100%/3   | Total: 24m 59s | Avg:  8m 19s | Max: 15m 01s | Hits:  88%/1514  
      🟩 12.9               Pass: 100%/23  | Total:  2h 39m | Avg:  6m 57s | Max: 13m 50s | Hits:  91%/13616 
    🟩 cudacxx
      🟩 nvcc12.0           Pass: 100%/3   | Total: 24m 59s | Avg:  8m 19s | Max: 15m 01s | Hits:  88%/1514  
      🟩 nvcc12.9           Pass: 100%/23  | Total:  2h 39m | Avg:  6m 57s | Max: 13m 50s | Hits:  91%/13616 
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/26  | Total:  3h 04m | Avg:  7m 06s | Max: 15m 01s | Hits:  91%/15130 
    🟩 cxx
      🟩 Clang14            Pass: 100%/2   | Total:  9m 52s | Avg:  4m 56s | Max:  5m 25s | Hits:  91%/1212  
      🟩 Clang15            Pass: 100%/1   | Total:  5m 41s | Avg:  5m 41s | Max:  5m 41s | Hits:  90%/605   
      🟩 Clang16            Pass: 100%/1   | Total:  6m 01s | Avg:  6m 01s | Max:  6m 01s | Hits:  90%/605   
      🟩 Clang17            Pass: 100%/1   | Total:  5m 49s | Avg:  5m 49s | Max:  5m 49s | Hits:  90%/605   
      🟩 Clang18            Pass: 100%/1   | Total:  5m 53s | Avg:  5m 53s | Max:  5m 53s | Hits:  90%/605   
      🟩 Clang19            Pass: 100%/4   | Total: 24m 31s | Avg:  6m 07s | Max:  8m 21s | Hits:  93%/2420  
      🟩 GCC10              Pass: 100%/2   | Total: 11m 20s | Avg:  5m 40s | Max:  5m 49s | Hits:  90%/1212  
      🟩 GCC11              Pass: 100%/1   | Total:  6m 50s | Avg:  6m 50s | Max:  6m 50s | Hits:  90%/605   
      🟩 GCC12              Pass: 100%/1   | Total:  6m 18s | Avg:  6m 18s | Max:  6m 18s | Hits:  90%/605   
      🟩 GCC13              Pass: 100%/8   | Total: 53m 42s | Avg:  6m 42s | Max: 12m 23s | Hits:  92%/4840  
      🟩 MSVC14.39          Pass: 100%/1   | Total: 15m 01s | Avg: 15m 01s | Max: 15m 01s | Hits:  78%/304   
      🟩 MSVC14.43          Pass: 100%/1   | Total: 13m 50s | Avg: 13m 50s | Max: 13m 50s | Hits:  78%/306   
      🟩 NVHPC25.5          Pass: 100%/2   | Total: 20m 10s | Avg: 10m 05s | Max: 10m 32s | Hits:  88%/1206  
    🟩 cxx_family
      🟩 Clang              Pass: 100%/10  | Total: 57m 47s | Avg:  5m 46s | Max:  8m 21s | Hits:  91%/6052  
      🟩 GCC                Pass: 100%/12  | Total:  1h 18m | Avg:  6m 30s | Max: 12m 23s | Hits:  92%/7262  
      🟩 MSVC               Pass: 100%/2   | Total: 28m 51s | Avg: 14m 25s | Max: 15m 01s | Hits:  78%/610   
      🟩 NVHPC              Pass: 100%/2   | Total: 20m 10s | Avg: 10m 05s | Max: 10m 32s | Hits:  88%/1206  
    🟩 gpu
      🟩 h100               Pass: 100%/2   | Total: 14m 23s | Avg:  7m 11s | Max:  9m 43s | Hits:  95%/1210  
      🟩 rtx2080            Pass: 100%/24  | Total:  2h 50m | Avg:  7m 06s | Max: 15m 01s | Hits:  90%/13920 
    🟩 jobs
      🟩 Build              Pass: 100%/23  | Total:  2h 34m | Avg:  6m 43s | Max: 15m 01s | Hits:  90%/13315 
      🟩 Test               Pass: 100%/3   | Total: 30m 27s | Avg: 10m 09s | Max: 12m 23s | Hits:  99%/1815  
    🟩 sm
      🟩 90                 Pass: 100%/3   | Total: 18m 54s | Avg:  6m 18s | Max:  9m 43s | Hits:  93%/1815  
      🟩 90a                Pass: 100%/1   | Total:  4m 47s | Avg:  4m 47s | Max:  4m 47s | Hits:  90%/605   
    🟩 std
      🟩 17                 Pass: 100%/4   | Total: 24m 59s | Avg:  6m 14s | Max:  9m 38s | Hits:  90%/2418  
      🟩 20                 Pass: 100%/22  | Total:  2h 39m | Avg:  7m 16s | Max: 15m 01s | Hits:  91%/12712 
    
  • 🟩 packaging: Pass: 100%/4 | Total: 15m 18s | Avg: 3m 49s | Max: 4m 07s

    🟩 cpu
      🟩 amd64              Pass: 100%/4   | Total: 15m 18s | Avg:  3m 49s | Max:  4m 07s
    🟩 ctk
      🟩 12.0               Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  3m 47s
      🟩 12.9               Pass: 100%/2   | Total:  7m 51s | Avg:  3m 55s | Max:  4m 07s
    🟩 cudacxx
      🟩 nvcc12.0           Pass: 100%/2   | Total:  7m 27s | Avg:  3m 43s | Max:  3m 47s
      🟩 nvcc12.9           Pass: 100%/2   | Total:  7m 51s | Avg:  3m 55s | Max:  4m 07s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/4   | Total: 15m 18s | Avg:  3m 49s | Max:  4m 07s
    🟩 cxx
      🟩 Clang14            Pass: 100%/1   | Total:  3m 40s | Avg:  3m 40s | Max:  3m 40s
      🟩 Clang19            Pass: 100%/1   | Total:  3m 44s | Avg:  3m 44s | Max:  3m 44s
      🟩 GCC12              Pass: 100%/1   | Total:  3m 47s | Avg:  3m 47s | Max:  3m 47s
      🟩 GCC13              Pass: 100%/1   | Total:  4m 07s | Avg:  4m 07s | Max:  4m 07s
    🟩 cxx_family
      🟩 Clang              Pass: 100%/2   | Total:  7m 24s | Avg:  3m 42s | Max:  3m 44s
      🟩 GCC                Pass: 100%/2   | Total:  7m 54s | Avg:  3m 57s | Max:  4m 07s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/4   | Total: 15m 18s | Avg:  3m 49s | Max:  4m 07s
    🟩 jobs
      🟩 Test               Pass: 100%/4   | Total: 15m 18s | Avg:  3m 49s | Max:  4m 07s
    
  • 🟩 stdpar: Pass: 100%/4 | Total: 17m 16s | Avg: 4m 19s | Max: 4m 37s

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total:  8m 56s | Avg:  4m 28s | Max:  4m 37s
      🟩 arm64              Pass: 100%/2   | Total:  8m 20s | Avg:  4m 10s | Max:  4m 12s
    🟩 ctk
      🟩 12.9               Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 cudacxx
      🟩 nvcc12.9           Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 cxx
      🟩 NVHPC25.5          Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 cxx_family
      🟩 NVHPC              Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 gpu
      🟩 rtx2080            Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 jobs
      🟩 Build              Pass: 100%/4   | Total: 17m 16s | Avg:  4m 19s | Max:  4m 37s
    🟩 std
      🟩 17                 Pass: 100%/2   | Total:  8m 31s | Avg:  4m 15s | Max:  4m 19s
      🟩 20                 Pass: 100%/2   | Total:  8m 45s | Avg:  4m 22s | Max:  4m 37s
    
  • 🟩 cccl_c_parallel: Pass: 100%/2 | Total: 13m 37s | Avg: 6m 48s | Max: 10m 20s | Hits: 98%/328

    🟩 cpu
      🟩 amd64              Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 ctk
      🟩 12.9               Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 cudacxx
      🟩 nvcc12.9           Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 cudacxx_family
      🟩 nvcc               Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 cxx
      🟩 GCC13              Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 cxx_family
      🟩 GCC                Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 gpu
      🟩 rtx2080            Pass: 100%/2   | Total: 13m 37s | Avg:  6m 48s | Max: 10m 20s | Hits:  98%/328   
    🟩 jobs
      🟩 Build              Pass: 100%/1   | Total:  3m 17s | Avg:  3m 17s | Max:  3m 17s | Hits:  98%/164   
      🟩 Test               Pass: 100%/1   | Total: 10m 20s | Avg: 10m 20s | Max: 10m 20s | Hits:  98%/164   
    

👃 Inspect Changes

Modifications in project?

Project
CCCL Infrastructure
CCCL Packaging
libcu++
CUB
+/- Thrust
CUDA Experimental
stdpar
python
CCCL C Parallel Library
Catch2Helper

Modifications in project or dependencies?

Project
CCCL Infrastructure
+/- CCCL Packaging
libcu++
+/- CUB
+/- Thrust
+/- CUDA Experimental
+/- stdpar
python
+/- CCCL C Parallel Library
+/- Catch2Helper

🏃‍ Runner counts (total jobs: 130)

# Runner
89 linux-amd64-cpu16
11 windows-amd64-cpu16
10 linux-arm64-cpu16
7 linux-amd64-gpu-rtx2080-latest-1
6 linux-amd64-gpu-rtxa6000-latest-1
4 linux-amd64-gpu-h100-latest-1
3 linux-amd64-gpu-rtx4090-latest-1

@bernhardmgruber

Copy link
Copy Markdown
Contributor Author

Thrust's version previously supported large custom types using virtual shared memory. That's currently not supported by cub::Reduce. For other algorithms, this is something we made sure to tackle first. Are we consciously not doing this for the extrema computation?

Yes. Running out of SMEM is extremely unlikely so we just provide a static assert to users instead, see #6062.

@bernhardmgruber
bernhardmgruber marked this pull request as ready for review September 29, 2025 18:23
@cccl-authenticator-app cccl-authenticator-app Bot moved this from In Progress to In Review in CCCL Sep 29, 2025
@github-actions

Copy link
Copy Markdown
Contributor

😬 CI Workflow Results

🟥 Finished in 4h 38m: Pass: 97%/159 | Total: 6d 13h | Max: 4h 24m | Hits: 74%/189478

See results here.

@bernhardmgruber
bernhardmgruber marked this pull request as draft September 30, 2025 08:12
@cccl-authenticator-app cccl-authenticator-app Bot moved this from In Review to In Progress in CCCL Sep 30, 2025
@bernhardmgruber

bernhardmgruber commented Sep 30, 2025

Copy link
Copy Markdown
Contributor Author

I found some problems with the current approach. It's somewhat of a hack anyway. I think to properly implement this, we would need some extensions to cub::DeviceReduce, documented in #1626

@bernhardmgruber

Copy link
Copy Markdown
Contributor Author

Superseded by #8292

@github-project-automation github-project-automation Bot moved this from In Progress to Done in CCCL Aug 9, 2026
@bernhardmgruber
bernhardmgruber deleted the port_extrema_to_cub branch August 9, 2026 23:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Refactor thrust/extrema.h to use cub::DeviceReduce

3 participants