Skip to content

chore: port-to-v6 rolling port - #25565

Merged
fcarreiro merged 1 commit into
v6from
cb/port-to-v6
Oct 1, 2026
Merged

fcarreiro merged 1 commit into
v6from
cb/port-to-v6

Conversation

@AztecBot

@AztecBot AztecBot commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Rolling integration PR for changes labelled port-to-v6 on next. Each merged source PR is cherry-picked onto cb/port-to-v6 as its own commit, keeping (#N) in the title.

Ported

  • #25548 fix(bb.js): report a dead bb process as retryable, and replace it on request (clean pick)

Notes


Created by claudebox · group: slackbot · Slack thread

…request (#25548)

Addresses
AztecProtocol/barretenberg-claude#4427 from
the bb.js side. The aztec-node side is
[aztec-node#354](aztec-labs-eng/aztec-node#354).

A pooled bb verifier whose process dies is handed back to the pool and
handed out again on every later borrow. The pool cannot tell, because
the socket backend marks itself permanently unusable and every later
call fails with a bare `Socket not connected`. The node therefore
reports a dead helper process as an invalid transaction proof,
persistently, and books it into the metric that means users are
submitting bad proofs.

This gives an owner the two things it needs, in the shape the AVM
simulator pool already relies on.

### A failed call says whether retrying can help

An environmental failure now carries `retry: true`: the bb process died,
its connection broke, or it could not be started. Only a binary that
cannot be executed stays non-retryable, because retrying cannot fix it.

The bare property is the whole contract, feature-detected rather than
imported:

```ts
if (err instanceof Error && (err as Error & { retry?: unknown }).retry === true) { ... }
```

That is deliberately the same convention `ipc-runtime`'s transports use,
so a caller written against one works unchanged against the other. It is
also what lets a verifier distinguish a bad proof from a dead helper,
which is the misattribution half of the issue.

### The socket backend can replace a dead process, on request

`respawn` is a new `BackendOptions` flag, off by default. With it on,
the bb process and its connection are one unit swapped as a whole: the
next call starts a replacement, while calls already in flight still fail
retryably. Concurrent callers that find the connection down share one
replacement, so a single death costs a single process, and a replacement
that arrives after `destroy()` is not left running.

It is opt-in because a replacement has none of the state a command
sequence establishes: no SRS loaded over the connection, no Chonk
accumulation, no batch-verifier session with its registered keys. An
owner that runs any such sequence must leave it off, so a death fails
loudly instead of the next call quietly running against a process that
has forgotten everything. The verifier pool qualifies because every
verification carries its proof and key in the call.

With this, a pool needs no liveness check and no maintenance loop:
returning an instance unconditionally becomes correct, exactly as it
already is for the AVM pool.

### Also

`BarretenbergSync.initSingleton()` cached a failed initialization for
the life of the process, so one bad spawn was permanent. It now clears
the failure, as the asynchronous singleton already did.

### Testing

`native_socket.test.ts`, against the fake bb the existing tests use: a
killed bb fails the call retryably and keeps failing without the option;
with it on the next call is served by a replacement process, a different
pid; four concurrent calls that find the connection down start exactly
one replacement; `destroy()` during a replacement leaves no process
running; a replacement that cannot start fails retryably too; a
connection that breaks while the process keeps running leaves no bb
behind.

`singleton.test.ts` covers the initialization fix and fails without it.

Checked against a real bb as well: with the option off a killed bb gives
a retryable error and keeps doing so, and with it on the next call
transparently returns the same hash from a fresh process.

### What this does not cover

`BarretenbergSync` runs on the shared-memory backend, which has no
respawn and no way to report a death mid-call: the NAPI receive loop
retries without a deadline, and because the call blocks the event loop
the process exit is never even observed, so the caller wedges rather
than fails. That path is untouched here and is deliberately left alone:
the synchronous bb runs one thread doing hashes and signatures, so it is
the least likely process on the machine to be killed, and the fix would
mean threading a liveness check into the C++ client for a case that may
never happen. #25546 covers the idle-death half of it by replacing a
dead singleton, so the two PRs cover different backends rather than one
superseding the other.

### Note on direction

bb.js's hand-written backends are replaced by `ipc-runtime`'s in the
codegen migration (#25362). Nothing above is lost in that move:
`ipc-runtime`'s spawned backend already carries the same `retry`
contract and the same opt-in respawn, so the callers written against
this keep working and the implementation here is deleted. That is why
this is expressed as the retry contract rather than as a liveness query,
which would have to become part of the generated client's interface.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit d5b4b8a)
@AztecBot AztecBot added ci-draft Run CI on draft PRs. ci-no-fail-fast Sets NO_FAIL_FAST in the CI so the run is not aborted on the first failure claudebox Owned by claudebox. it can push to this PR. labels Oct 1, 2026
@fcarreiro
fcarreiro marked this pull request as ready for review October 1, 2026 09:58
@fcarreiro
fcarreiro enabled auto-merge October 1, 2026 09:58
@fcarreiro
fcarreiro merged commit cf7ea33 into v6 Oct 1, 2026
21 of 25 checks passed
@fcarreiro
fcarreiro deleted the cb/port-to-v6 branch October 1, 2026 10:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-draft Run CI on draft PRs. ci-no-fail-fast Sets NO_FAIL_FAST in the CI so the run is not aborted on the first failure claudebox Owned by claudebox. it can push to this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants