Skip to content

fix(fpu): integrate FCSR rounding and retired exception flags - #12

Merged
RossComputerGuy merged 1 commit into
LilithSemi:masterfrom
murdoa:fix/fcsr-retirement
Oct 8, 2026
Merged

RossComputerGuy merged 1 commit into
LilithSemi:masterfrom
murdoa:fix/fcsr-retirement

Conversation

@murdoa

@murdoa murdoa commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

fflags, frm and fcsr currently trap as unimplemented, and the iterative arithmetic units round to nearest-even regardless of the requested mode. Software can neither select dynamic rounding nor read accumulated FP exceptions.

For example, adding 1.0 + 2^-53 with round-up should return 0x3ff0000000000001 and set NX. Returning 1.0 instead changes the numerical result; failing to record NX also hides that rounding occurred.

This wires FCSR through both in-order executors:

  • Implement the three CSR aliases over shared state, including field masking and sticky exception flags.
  • Select static rounding or frm, latch the selected mode, and reject reserved modes before execution.
  • Generate exceptions in the existing arithmetic, sqrt and conversion paths, then carry them with the instruction to successful retirement.
  • Reject FP instructions and FP CSR accesses with FS Off, retaining the existing FS/SD dirty-state machinery.

FMA uses the single-rounding implementation from #11. Its flags describe the fused result, not an intermediate multiply.

Integration also exposed a few related problems: zero-source CSRRS/CSRRC still asserted the write port, so reading fflags would dirty FS; narrow FP results and raw moves needed consistent boxing/sign extension; and RV32 microcoded FP reads forwarded 64 bits into 32-bit latches. These are fixed here.

Scope

Scalar FP on static and microcoded in-order cores, tested with RV32 F and RV64 F/D. H FCSR remains unavailable, and OoO explicitly leaves it disabled until precise FP retirement is implemented. Vector FP status is not included.

No replacement FPU, timing pipeline or board changes.

Testing

Compared against 60b1b2a, using identical tests in fresh processes:

  • All 201 baseline passes remain passing.
  • 186 behavioral assertion failures are fixed.
  • Ten additional RV32 microcoded cases fail elaboration on baseline and pass here. These are counted separately, not as behavioral reproductions.
  • 397 cases pass on this branch.

Coverage includes 106 new full-core cases, producer/CSR tests, existing FP controls and 52 straddling-fetch regressions. Baseline component adapters only ignore new inputs and supply zero for missing flag outputs.

Both trees use the same fixture corrections: explicit FS initialization, byte-enable-aware memory for word stores, and the correct RV64 sign-extension expectation for FMV.X.W.

Integrate non-H in-order FCSR aliases, FS legality, static and dynamic rounding, and sticky exception flags at successful retirement. Preserve the existing FS dirty mechanism and leave H and OoO FCSR disabled.

Normalize narrow computational operands and results while preserving raw move semantics. Suppress zero-source CSR set/clear writes so status reads do not dirty FP state. Cover both executors and repair RV32 microcoded FP read-width elaboration.
@RossComputerGuy
RossComputerGuy merged commit 40eab7f into LilithSemi:master Oct 8, 2026
1 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants