Skip to content

CL-6457: Provider connect must not deploy anything - #187

Merged
TheGreatAxios merged 3 commits into
mainfrom
cl-6457-async-provisioning
Aug 21, 2026
Merged

TheGreatAxios merged 3 commits into
mainfrom
cl-6457-async-provisioning

Conversation

@TheGreatAxios

Copy link
Copy Markdown
Contributor

Pasting an Anthropic key on onboarding sat on a static "Connecting…" for over two minutes. The connect request was deploying five default workflows — around twenty seconds each — before it answered. Long enough to read as frozen.

Connecting now deploys nothing.

What connect does now

Connect keeps only the durable, fast half: persist the credential, seed its model catalog, answer. Deploying a bench's default workflows moves to a background drain, and no HTTP route deploys anything any more — POST /complete and POST /complete-setup both report where the agents stand and hand the work off, and a new GET /api/onboarding/provisioning-status serves anything that needs to poll.

completeCredentialSetup still composes both halves for harnesses that need a fully deployed bench before a scenario begins (@workbench/evals' real target), and its comments now say plainly that a route calling it re-creates the freeze.

Why the drain converges

It is not a job queue. It is a pass over onboarding.pending_seed, where the row is the work item and means "this bench has a credential and does not yet have its agents". The properties fall out of that framing:

  • Idempotent — every pass re-reads the bench's real asset and deployment state first, and seedTenant underneath is ensure-then-create. A pass over a finished bench deploys nothing and clears the row.
  • Convergent — a pass that gets partway leaves the row, so the next one picks up exactly what is still missing.
  • Restart-safe — no outstanding work lives in the process. A hub that dies mid-deploy leaves the row behind and the next boot's first tick finishes it. In-memory state is only ever a dedupe guard or a retry backoff.

A failing bench backs off (15s, doubling to a 10-minute ceiling) rather than hammering a down sidecar, and is never dropped for failing.

Sessions

A background loop has no request to borrow cookies from. The provisioner takes sessionFor as a seam and the hub fills it by minting the bench owner's own session in process. It has to be the owner's: the hub resolves a tenant by looking up a principal for that user, and rights do not flow from a parent org to a child bench, so an administrator would simply be refused. These sessions carry a workbench-bench-provisioner user agent so they are never mistaken for a human sign-in, and are cached per user rather than minted per tick.

What someone sees

Connecting hands the person forward immediately. If they land before their agents exist they get the warm loading state — "Getting your agents ready…" with live progress, N of M ready — which resolves into their workbench on its own. Whether the bench is still provisioning is asked of the bench rather than read off an error message: still deploying is someone waiting, not someone broken. This closes a real hole, since the zero-workbench land path previously threw "No default setup agent found" and showed an error card.

Tests

The headline test is a clock, not a mock: the deploy step is a seam that takes five seconds, and connect has to answer in under one without that seam ever starting. A route that awaits a deploy again fails on elapsed time.

Alongside it: drain idempotency, convergence of a half-provisioned bench, restart-resume by a fresh provisioner with empty memory, no double-deploy under overlapping passes, backoff, and the client's connected/pending rendering.

Scoped green — packages/onboarding 149 pass / 0 fail, apps/web 748 pass / 0 fail, apps/hub 137 pass / 0 fail, typecheck and lint clean across all three.

Fixes CL-6457

Connecting a provider must deploy nothing (CL-6457). The live repro:
pasting a key sat on "Connecting…" for over two minutes because the
connect request deployed five default workflows — around twenty seconds
each — before it answered.

The headline test is a clock, not a mock: the deploy step is handed in
as a seam that takes five seconds, and the connect route has to answer
in well under a second without that seam ever starting. A route that
awaits a deploy again fails here on elapsed time.

Around it, the properties a background drain has to hold and no HTTP
request can hold for it: idempotence (a pass over an already-seeded
bench deploys nothing), convergence (a half-provisioned bench finishes
on a later pass), restart-resume (a fresh process with empty memory
picks up what a crashed one left behind), no double-deploy under
overlapping passes, and a pending-seed row that survives precisely so
the drain can finish the job.
Connect now does the durable, fast half only — persist the credential,
seed its catalog — and answers in seconds. Deploying a bench's default
workflows moves to a background drain, so no HTTP route deploys
anything any more.

The drain treats the existing pending-seed row as its work item: the row
means "this bench has a credential and does not yet have its agents".
That framing is what makes the properties fall out rather than have to
be engineered. Every pass re-reads the bench's real asset and deployment
state before acting, so a pass over a finished bench deploys nothing. A
pass that gets partway leaves the row, so the next one finishes it. And
because the row is in Postgres rather than in this process, a hub that
dies mid-deploy resumes the bench on its next boot — in-memory state is
only ever a dedupe guard or a retry backoff, never a fact the system
needs to be correct.

Sessions are the one thing a background loop cannot inherit, since there
is no request to borrow cookies from. The provisioner takes that as a
seam and the hub fills it by minting the bench owner's own session in
process. It has to be the owner's: the hub resolves a tenant by looking
up a principal for that user, and rights do not flow from a parent org
to a child bench, so an administrator would simply be refused. These
sessions are tagged by user agent so they are never mistaken for a
human sign-in.

On the client, connecting hands the person forward immediately. If they
land before their agents exist, they get the warm loading state with
live progress — how many agents are ready out of how many — which
resolves into their workbench on its own. Whether the bench is still
provisioning is asked of the bench rather than read off an error
message: a bench still deploying is someone waiting, not someone broken.
Records the rule connecting now follows — connecting deploys nothing —
along with the fast/slow split it comes from, why the pending-seed row
makes the drain idempotent, convergent and restart-safe, why the
provisioning session has to be the bench owner's own, and what someone
actually sees when they land before their agents exist.

Corrects the model-seeding note that still described onboarding as
deploying the default workflow set inline.
@TheGreatAxios
TheGreatAxios merged commit 23879e4 into main Aug 21, 2026
0 of 2 checks passed
@TheGreatAxios
TheGreatAxios deleted the cl-6457-async-provisioning branch August 25, 2026 15:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant