Skip to content

Loading-path shim fixes, host setjmp/longjmp, recompiler fixes, CI - #3

Merged
aaronateataco merged 6 commits into
mainfrom
claude/repository-review-dzfwn1
Sep 24, 2026
Merged

aaronateataco merged 6 commits into
mainfrom
claude/repository-review-dzfwn1

Conversation

@aaronateataco

Copy link
Copy Markdown
Member

Summary

Loading stall (each fix has a sync_harness case that fails on the old code):

  • The FS completion queue had no lock. Under contention the old queue delivered a completion twice, then wrapped its count to 4,294,967,295, so it read as full forever. The queue now has a mutex, and both pumps claim atomically. This is the first certain bug found with the shape of the stall. A hardware run will say whether it is the cause.
  • AUTO events released every waiter on a single signal, and a quick second signal was lost. Replaced with wake credits (EVCREDIT). Cemu's coreinit_Synchronization.cpp agrees with the new semantics.
  • OSWaitSemaphore parked forever on a completion only it could pump. It now waits in 1ms slices and pumps between them, like OSWaitEvent.

Recompiler:

  • setjmp/longjmp are done on the host (XenonRecomp's approach, adapted to C). The default names are setjmp/_setjmp/__setjmp; add others with --setjmp-name/--longjmp-name.
  • --extern-globals works again. It had refused every object since 2026-08-29.
  • Relocations into non-SHF_ALLOC sections are skipped and reported per section.

CI (.github/workflows/check.yml): builds recomp, runs the shim harness normally and under ThreadSanitizer, and runs blaster's verify.sh. Nothing reads the RPX.

Docs: docs/prior-art-2026-09.md.

Verified

  • sync_harness: 15 cases, 20 of 20 runs
  • TSan: races only on the deliberately unlocked counters
  • verify.sh: 25 of 25
  • CI green on the head commit

Not yet verified

Nothing here has run on a Switch. The shim headers are included by every generated translation unit, so the next jouster build is a full rebuild. On regenerate, check two things:

  • setjmp is converted
  • only debug sections are listed under "skipped relocation(s)"

🤖 Generated with Claude Code

https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg


Generated by Claude Code

AUTO events released every waiter. Each parked waiter compared one
per-event epoch that every signal bumped, so a single OSSignalEvent on an
AUTO event released all of them. A second signal landing before the first
woken waiter ran still saw waiting_count > 0, bumped the epoch again and
latched nothing -- lost. EVSTATE had measured two waiters parked on the
loading event, so this was live, and it was ours rather than the engine's.

Replaced with a count of wake credits granted to parked waiters, never
more than are waiting. An AUTO signal grants one if some waiter has none
and otherwise latches; a waiter leaves on a credit, then the latch. The
credit lives in the struct, so a signal that lands while its waiter is
off running the pump survives. epoch stays, as EVSTATE's diagnostic only;
init_epoch is what OSInitEvent moves to release waiters on a re-init.
OSWaitEventWithTimeout now returns TRUE only for a credit or the latch --
it used to tell a timed-out waiter TRUE whenever another waiter had been
signalled. OSResetEvent clears the latch and leaves granted credits,
which on Cafe OS are threads already off the queue.

The FS completion queue had no lock. It is filled by whichever thread
issues a read and drained by whichever thread next pumps, and the pump's
guard was a plain int two threads could both see clear. With four
producers and four pumpers the old queue delivered a completion twice,
then wrapped its count to 4294967295 -- after which it reads as full
forever and every later completion is dropped. A completion that never
runs is a work item that never reaches status 2 and a block never
released: the shape of the loading stall. Whether it is the cause is for
a hardware run to say. The queue now has a mutex held only to push or pop
one entry, never across a callback, and both pumps claim atomically.

sync_harness.c gains eight cases. Four fail on the old code (one signal
releasing two waiters, a lost second signal, two timed waiters both told
TRUE, and the queue stress); all fourteen pass on this, 20 runs out of
20. Under ThreadSanitizer the only races left are the deliberately
unlocked diagnostic counters. Headers compile clean under gcc and clang
with -Wall -Wextra. Not yet built for the Switch: that needs the runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg
It pumped FS completions once on entry and then parked on the condvar for
good. If the release it waits for comes from a completion queued after it
parked, and no other thread happens to pump, nothing delivers it -- the
same starvation OSWaitEvent was cured of on 2026-09-12, on a primitive
the loading path also uses (a count-1 semaphore, SEMINIT 0x4503598).

Now 1ms slices with the completion pump between them, lock dropped across
the pump because the callback may signal this same semaphore. Contract
unchanged: it returns only after taking a count. Only the completion
pump, not the archive pump, which is a workaround for one specific wait.

New sync_harness case deadlocks on the old code and passes on this; the
full harness passes 15 runs out of 15.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg
…sections

blaster's verify.sh has been red at its multi-object case since
2026-08-29. That change put .text into global_section_base, so it keeps
its real address -- and main.cpp's --extern-globals guard refused any
object whose global_section_base was non-empty, which after that was
every object. The guard now ignores .text: it is code, not data anyone
has to initialize.

Behind it, a second problem the same object showed. Every .rela.*
section was read, including .rela.debug_info, .rela.debug_frame and
.rela.debug_line, so debug sections were given synthetic bases in guest
memory and relocated pointers were written into them. Relocations are
now read only for targets with SHF_ALLOC. Skipped relocations are
printed per section, because this could not be checked against the
retail RPX from here: a regenerate that lists anything but debug
sections has a loaded section flagged in a way this did not expect.

verify.sh: all 23 pipelines pass, host and ARM64 under QEMU, including
the multi-object case that was failing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg
conquertron had no CI; jouster's runs on the owner's PC because it needs
the game. Everything here needs neither:

  recompiler  recomp and both verification tools against the pinned
              Capstone 5.0.3 (PowerPC only, cached)
  shims       the shim headers compile clean with -Wall -Wextra -Werror
              under gcc and clang; sync_harness five times; and once more
              under ThreadSanitizer, where every report must be one of the
              deliberately unlocked diagnostic counters -- total reports
              are compared with allowed ones, so a race on heap memory,
              which has no "Location is global" line, still fails
  verify      blaster's verify.sh, 23 recompile-and-compare pipelines,
              host and ARM64 under QEMU

Nothing reads the RPX and nothing is uploaded. verify checks out
Arkchemy/blaster, which needs a BLASTER_TOKEN secret if blaster goes
private. Each step was run locally before this was pushed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg
A follow-up to the 2026-09-12 prior-art note, covering only what is new
or was missing: three public Wii U static recompilers (one running Wind
Waker HD's startup), XenonRecomp as the MIT-licensed reference for
setjmp/longjmp and jump tables, Cemu read against today's event fix (it
agrees) and for where real FS callbacks run (dedicated AppIO threads --
the architecture that would retire the pumps), RPX tooling, the LLVM
paired-singles work, and the Switch shader compilers. Every licence was
checked at the source; the ones not stated are marked as such.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg
A recompiled longjmp restores guest registers and then carries on through
whatever host C frames are live: it cannot unwind the host call stack the
recompiled code runs on. ROADMAP.md listed it; XMLREAD's note flagged it
as worth suspecting on the rapidxml error path.

codegen now recognises calls to the guest's setjmp and longjmp by name
(setjmp/_setjmp/__setjmp and the longjmp equivalents by default; more
with --setjmp-name/--longjmp-name) and emits:

  PPC_HOST_SETJMP(ctx, ppc_setjmp(ctx));
      the recompiled setjmp still runs, so the guest jmp_buf is written
      as before; then the guest register file is saved to a host slot
      keyed by the jmp_buf's guest address and the HOST setjmp is called
      inline, in the caller's own C function -- which is why it is a
      macro. On the second return the guest registers are restored and
      r3 carries the longjmp value.

  ppc_host_longjmp(ctx);
      looks up the slot and host-longjmps to it; 0 becomes 1. A jmp_buf
      with no recorded setjmp aborts loudly rather than guessing.

The approach is XenonRecomp's for the Xbox 360 (MIT), adapted to C. The
register file is PpcContext up to `tb`; the reservation is dropped as a
context switch would. Indirect calls to setjmp are not intercepted.

Proved by a new blaster verify.sh pipeline (testdata/setjmp.c): longjmp
out of six frames, longjmp with 0, and two jmp_bufs live at once. It
returns -1 -1 -1 before this change and matches the native build after,
at guest -O0, -O1 and -O2, on the host at -O0 and -O2 and on ARM64 under
QEMU. verify.sh: 25 of 25.

Not yet checked against the retail RPX: whether its setjmp is called one
of the default names is the first thing a regenerate should confirm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qKyztZy5fFoipSXHhDnSg
Copilot AI lite review requested due to automatic review settings September 24, 2026 18:12
@aaronateataco
aaronateataco merged commit 9daa4d8 into main Sep 24, 2026
5 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical unresolved issues remain in host jump handling, filesystem pump cleanup, imported-call interception, and external-global initialization.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 4 High severity · 2 Medium severity · 1 Low severity

Open (7)
What changed in this PR

This PR improves synchronization shims, host setjmp/longjmp handling, recompiler behavior, relocation filtering, and CI validation.

Changes:

  • Adds locked FS queues, wake credits, and cooperative semaphore pumping.
  • Adds configurable host jump handling and restores external-global support.
  • Adds relocation filtering, harness coverage, documentation, and CI checks.
File Reviewed change
src/​main.cpp Adds jump options and external-global validation.
src/​elf_loader.cpp Skips relocations targeting non-allocated sections.
src/​codegen.h Declares configurable jump-function names.
src/​codegen.cpp Routes recognized calls through host jump handling.
ROADMAP.md Records completed synchronization and recompiler work.
include/​ppc_runtime.h Implements host jump storage and register restoration.
include/​cafeos_state.c Updates event fallback initialization.
include/​cafeos_coreinit_sync.h Adds wake credits and cooperative waiting.
include/​cafeos_coreinit_fs.h Adds locked completion queues and pump claims.
hosttest/​sync_harness.c Adds synchronization and queue stress tests.
hosttest/​README.md Documents harness and TSan usage.
docs/​prior-art-2026-09.md Documents related recompiler and runtime prior art.
.github/​workflows/​check.yml Adds build, harness, TSan, and verification jobs.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

for (int i = 0; i < 10; i++) ctx->r[3 + i] = saved_r[i];
ctx->lr = saved_lr;
g_arkchemy_fs_pumping = 0;
__atomic_store_n(&g_arkchemy_fs_pumping, 0, __ATOMIC_RELEASE);
for (i = 0; i < 10; i++) ctx->r[3 + i] = saved_r[i];
ctx->lr = saved_lr;
g_arkchemy_archive_pumping = 0;
__atomic_store_n(&g_arkchemy_archive_pumping, 0u, __ATOMIC_RELEASE);
Comment thread src/codegen.cpp
auto it = img.call_relocs.find(insn_addr);
if (it != img.call_relocs.end()) {
return "ppc_" + it->second + "(ctx);";
return direct_call_stmt(it->second);
Comment thread src/main.cpp
Comment on lines +84 to +87
bool has_own_data = false;
for (const auto &kv : img.global_section_base)
if (kv.first != ".text") { has_own_data = true; break; }
if (extern_globals && has_own_data) {
Comment on lines +95 to +97
total=$(grep -c 'SUMMARY: ThreadSanitizer' /tmp/tsan.txt || true)
allowed=$(grep 'Location is global' /tmp/tsan.txt \
| grep -cE "'(g_ark_sev_|g_ark_wev_|g_arkchemy_event_|g_case_done|limit)" || true)
Comment thread src/elf_loader.cpp
// ever lists anything here other than debug sections, a loaded
// section is flagged in a way this did not expect.
if (!(rd_be32(shdr(target_idx) + 8) & 0x2u /* SHF_ALLOC */)) {
skipped_nonalloc[target_name] += rela_sh_size_entries(rd_be32(rela_sh + 20), rela_entsize);
Comment thread docs/prior-art-2026-09.md
Comment on lines +140 to +141
2. **Port XenonRecomp's `setjmp`/`longjmp` approach** (MIT, with
attribution). It is a ROADMAP item with a known-good answer.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants