Skip to content

fix(fpu): round fused multiply-add only once - #11

Merged
RossComputerGuy merged 1 commit into
LilithSemi:masterfrom
murdoa:fix/fma-single-rounding
Oct 6, 2026
Merged

RossComputerGuy merged 1 commit into
LilithSemi:masterfrom
murdoa:fix/fma-single-rounding

Conversation

@murdoa

@murdoa murdoa commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Failure

fmadd.d currently returns zero for

(1 + 2^-52) * (1 - 2^-52) - 1

The exact answer is -2^-104 (0xb970000000000000). The multiplier rounds the product to 1 before the addend is applied, so cancellation destroys a representable result.

RISC-V's fused operations compute the product and sum before rounding once. The intermediate rounding also changes range behavior: maxFinite * 2 - maxFinite becomes infinity rather than maxFinite, and minSubnormal * 0.5 + minSubnormal returns one rather than two minimum subnormals under RNE. These errors affect numerical results and would make exception reporting based on the intermediate multiply incorrect.

Fix

Keep the existing multiplier, retain its full product, and align/add the third operand before reducing the result to the existing round/pack path. Cancellation therefore preserves the product's low bits, and intermediate product overflow/underflow is not mistaken for overflow/underflow of the fused result. Three-operand special cases and all four product/addend sign combinations are handled directly.

The ordinary add, multiply, divide and conversion paths remain. The FMA reference tests now use exact integer significands/exponents rather than Dart's separately rounded a * b + c expression.

Boundaries

This fixes single rounding under the unit's existing RNE behavior. Architectural rounding-mode selection and exception-flag reporting remain separate follow-up integration; this PR does not claim to implement them. It adds full-product alignment/addition state and changes FMA latency, but introduces no timing pipeline or board changes. Area/timing, synthesis and hardware have not been validated.

Validation

Against f403d31, the same 55 named tests run in fresh processes give 34 pass / 21 fail → 53 pass / 2 fail, with no lost passes or missing cases.

  • All 24 new full-core cases pass: both in-order executors, all four FMA forms, cancellation and intermediate overflow/underflow.
  • Three new unit cases cover 2,364 vectors across native F32, F64 and F32 on the F64 unit, including targeted cancellation and random bit patterns.
  • Existing arithmetic/conversion and core FP tests are included. Changed files analyze cleanly.
  • The exact oracle independently agrees with libm fma/fmaf on 21,664 cases, with NaNs canonicalized.

The two remaining failures are identical baseline F32 move/boxing failures: the load/store round-trip expects 123 but gets -4294967173; the sign/min/max/class/move case expects 1065353216 but gets -3229614080. Neither is changed or counted as an FMA regression. This is targeted Dart simulation, not a full-suite or external-RTL validation claim.

@RossComputerGuy
RossComputerGuy merged commit 60b1b2a into LilithSemi:master Oct 6, 2026
1 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants