When scheduling .gpu_tile(x, y, xo, yo, xi, yi, ...).split(xi, xio, xii, 2) you don't want the inner xii loop to inherit the gpu_threads type. Same for splitting GPU-block loops I think, although that's less common when typing out schedules.
What do you think?
When scheduling
.gpu_tile(x, y, xo, yo, xi, yi, ...).split(xi, xio, xii, 2)you don't want the innerxiiloop to inherit thegpu_threadstype. Same for splitting GPU-block loops I think, although that's less common when typing out schedules.What do you think?