Skip to content

Commit 9234be6

Browse files
authored
Stop emitting an all-zero weight_zero_point (pytorch#23216)
Summary: A symmetric weight has a zero point of zero in every channel, so the per-channel path was materializing one int32 per output channel to say nothing. On wakeword stage 2 that is 4 of the 12 bytes per channel of qparam overhead, and the HiFi kernels do not even read it: the nnlib entry points behind conv take a `sym8s` weight operand, which has no weight zero point at all. Two halves, both here because the first is inert without the second. Schema. `Tensor weight_zero_point` becomes `Tensor? weight_zero_point` on the tensor-qparam overloads of conv1d (ncl/nlc and depthwise), conv2d (nchw/nhwc), linear and fully_connected. Absent means zero. The asymmetric path is untouched and still supplies a value, so this is additive rather than a narrowing and existing callers keep working. Only the tensor-qparam overloads move - 14 of the 68 declarations. The scalar (`int`/`SymInt`) forms serialize inline in the instruction rather than as a constant tensor, so they cost nothing and are left alone. `quantized_transposed_conv` is left alone too: it is not on the per-channel emission path, so changing it would be churn without payoff. It is deliberately not `Tensor? weight_zero_point=None`. The argument sits mid-signature, ahead of `bias_scale`/`out_scale`, and a positional schema cannot default an argument before non-defaulted ones. Callers pass `None` explicitly, which emits no constant. `Tensor? offset` in these same schemas is the existing precedent for that shape. Emitter. `weight_zero_point_arg` resolves the zero point and returns `None` when the vector is all zeros, so fusion omits the operand instead of lifting a constant. The test decides on the resolved values rather than on whether `zero_points` was set, so an affine dequantize that happens to be all zeros is also omitted, and a genuinely nonzero one is still carried through. Two shapes of kernel needed different handling. Linear and fully_connected read a single scalar, so they share `resolve_weight_zero_point` in `generic/operators/quantized_linear.h`. Conv reads a pointer and a stride, so those kernels fall back to a static zero with a stride of zero, which makes every channel read that same zero and leaves the inner loops untouched. Reviewed By: DrJessop Differential Revision: D118197851 Pull Request resolved: pytorch#23216
1 parent 06dc587 commit 9234be6

32 files changed

Lines changed: 264 additions & 123 deletions

‎backends/cadence/aot/functions.yaml‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -265,12 +265,12 @@
265265
- arg_meta: null
266266
kernel_name: impl::generic::dequantize_per_tensor_asym32s_out
267267

268-
- func: cadence::quantized_conv2d_nchw.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
268+
- func: cadence::quantized_conv2d_nchw.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
269269
kernels:
270270
- arg_meta: null
271271
kernel_name: impl::generic::quantized_conv2d_nchw_out
272272

273-
- func: cadence::quantized_conv2d_nhwc.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
273+
- func: cadence::quantized_conv2d_nhwc.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
274274
kernels:
275275
- arg_meta: null
276276
kernel_name: impl::generic::quantized_conv2d_nhwc_out
@@ -284,7 +284,7 @@
284284
- arg_meta: null
285285
kernel_name: impl::generic::quantized_layer_norm_per_tensor_out
286286

287-
- func: cadence::quantized_linear.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
287+
- func: cadence::quantized_linear.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor? weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
288288
kernels:
289289
- arg_meta: null
290290
kernel_name: impl::generic::quantized_linear_out
@@ -474,7 +474,7 @@
474474
- arg_meta: null
475475
kernel_name: impl::generic::quantized_conv2d_nhwc_depthwise_asym8uxsym8u_asym8u_per_tensor_out
476476

477-
- func: cadence::quantized_fully_connected.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
477+
- func: cadence::quantized_fully_connected.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor? weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
478478
kernels:
479479
- arg_meta: null
480480
kernel_name: impl::generic::quantized_fully_connected_out

‎backends/cadence/aot/functions_hifi.yaml‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -380,12 +380,12 @@
380380
- arg_meta: null
381381
kernel_name: impl::HiFi::dequantize_per_tensor_asym16s_out
382382

383-
- func: cadence::quantized_conv2d_nchw.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
383+
- func: cadence::quantized_conv2d_nchw.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
384384
kernels:
385385
- arg_meta: null
386386
kernel_name: impl::HiFi::quantized_conv2d_nchw_out
387387

388-
- func: cadence::quantized_conv2d_nhwc.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
388+
- func: cadence::quantized_conv2d_nhwc.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
389389
kernels:
390390
- arg_meta: null
391391
kernel_name: impl::HiFi::quantized_conv2d_nhwc_out
@@ -469,7 +469,7 @@
469469
- arg_meta: null
470470
kernel_name: impl::HiFi::quantized_layer_norm_per_tensor_out
471471

472-
- func: cadence::quantized_linear.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
472+
- func: cadence::quantized_linear.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor? weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
473473
kernels:
474474
- arg_meta: null
475475
kernel_name: impl::HiFi::quantized_linear_out
@@ -549,7 +549,7 @@
549549
- arg_meta: null
550550
kernel_name: impl::HiFi::quantized_matmul_asym8uxasym8u_asym8u_out
551551

552-
- func: cadence::quantized_fully_connected.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
552+
- func: cadence::quantized_fully_connected.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor? weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
553553
kernels:
554554
- arg_meta: null
555555
kernel_name: impl::HiFi::quantized_fully_connected_out

‎backends/cadence/aot/functions_vision.yaml‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -190,17 +190,17 @@
190190
- arg_meta: null
191191
kernel_name: impl::vision::dequantize_per_tensor_out
192192

193-
- func: cadence::quantized_conv.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, bool channel_last=False, *, Tensor(a!) out) -> Tensor(a!)
193+
- func: cadence::quantized_conv.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, bool channel_last=False, *, Tensor(a!) out) -> Tensor(a!)
194194
kernels:
195195
- arg_meta: null
196196
kernel_name: impl::vision::quantized_conv_out
197197

198-
- func: cadence::quantized_conv2d_nchw.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
198+
- func: cadence::quantized_conv2d_nchw.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
199199
kernels:
200200
- arg_meta: null
201201
kernel_name: impl::vision::quantized_conv2d_nchw_out
202202

203-
- func: cadence::quantized_conv2d_nhwc.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
203+
- func: cadence::quantized_conv2d_nhwc.out(Tensor input, Tensor weight, Tensor bias, int[] stride, SymInt[] padding, int[] dilation, int groups, int input_zero_point, Tensor? weight_zero_point, Tensor bias_scale, float out_scale, int out_zero_point, Tensor out_multiplier, Tensor out_shift, *, Tensor(a!) out) -> Tensor(a!)
204204
kernels:
205205
- arg_meta: null
206206
kernel_name: impl::vision::quantized_conv2d_nhwc_out
@@ -214,7 +214,7 @@
214214
- arg_meta: null
215215
kernel_name: impl::vision::quantized_layer_norm_per_tensor_out
216216

217-
- func: cadence::quantized_linear.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
217+
- func: cadence::quantized_linear.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor? weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
218218
kernels:
219219
- arg_meta: null
220220
kernel_name: impl::vision::quantized_linear_out
@@ -254,7 +254,7 @@
254254
- arg_meta: null
255255
kernel_name: impl::vision::quantized_conv_per_tensor_out
256256

257-
- func: cadence::quantized_fully_connected.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
257+
- func: cadence::quantized_fully_connected.out(Tensor src, Tensor weight, Tensor bias, int src_zero_point, Tensor? weight_zero_point, Tensor out_multiplier, Tensor out_shift, int out_zero_point, Tensor? offset, *, Tensor(a!) out) -> Tensor(a!)
258258
kernels:
259259
- arg_meta: null
260260
kernel_name: impl::vision::quantized_fully_connected_out

0 commit comments

Comments
 (0)