You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
XNNPACK: Route tanh GELU to approxgelu (pytorch#22851)
### Summary
The XNNPACK backend currently ignores GELU's `approximate` argument and
serializes both modes as `XNNGelu`. As a result,
`nn.GELU(approximate="tanh")` executes exact GELU after delegation. For
FP32 input `[-2.7]`, the baseline differs from the tanh reference by
approximately `4.73e-4`, failing comparison at `atol=rtol=1e-5`.
This PR serializes tanh GELU as `XNNApproxGelu` and dispatches it to
XNNPACK's existing `xnn_unary_approxgelu` operation. Default/exact GELU
and FP16 fallback remain unchanged. The new node is appended to both
FlatBuffer unions to preserve existing node IDs.
Affected XNNPACK models need to be re-exported, and the new node
requires an updated runtime.
### Test plan
I tested the fix on macOS arm64 using rebuilt runtimes based on
`500849ba5b`. Baseline validation reproduced the FP32 tanh GELU mismatch
at input `[-2.7]`, failing comparison at `atol=rtol=1e-5`. The current
parameterized regression tests retain this input and tolerance.
All 7 GELU tests passed, covering default GELU, both approximation
modes, static/dynamic shapes, and FP16 fallback. Local lintrunner
passed.
Authored with assistance from OpenAI Codex.
cc @GregoryComer@digantdesai@cbilgin@JakeStevens
0 commit comments