CL-6457: Provider connect must not deploy anything - #187
Merged
Merged
Conversation
Connecting a provider must deploy nothing (CL-6457). The live repro: pasting a key sat on "Connecting…" for over two minutes because the connect request deployed five default workflows — around twenty seconds each — before it answered. The headline test is a clock, not a mock: the deploy step is handed in as a seam that takes five seconds, and the connect route has to answer in well under a second without that seam ever starting. A route that awaits a deploy again fails here on elapsed time. Around it, the properties a background drain has to hold and no HTTP request can hold for it: idempotence (a pass over an already-seeded bench deploys nothing), convergence (a half-provisioned bench finishes on a later pass), restart-resume (a fresh process with empty memory picks up what a crashed one left behind), no double-deploy under overlapping passes, and a pending-seed row that survives precisely so the drain can finish the job.
Connect now does the durable, fast half only — persist the credential, seed its catalog — and answers in seconds. Deploying a bench's default workflows moves to a background drain, so no HTTP route deploys anything any more. The drain treats the existing pending-seed row as its work item: the row means "this bench has a credential and does not yet have its agents". That framing is what makes the properties fall out rather than have to be engineered. Every pass re-reads the bench's real asset and deployment state before acting, so a pass over a finished bench deploys nothing. A pass that gets partway leaves the row, so the next one finishes it. And because the row is in Postgres rather than in this process, a hub that dies mid-deploy resumes the bench on its next boot — in-memory state is only ever a dedupe guard or a retry backoff, never a fact the system needs to be correct. Sessions are the one thing a background loop cannot inherit, since there is no request to borrow cookies from. The provisioner takes that as a seam and the hub fills it by minting the bench owner's own session in process. It has to be the owner's: the hub resolves a tenant by looking up a principal for that user, and rights do not flow from a parent org to a child bench, so an administrator would simply be refused. These sessions are tagged by user agent so they are never mistaken for a human sign-in. On the client, connecting hands the person forward immediately. If they land before their agents exist, they get the warm loading state with live progress — how many agents are ready out of how many — which resolves into their workbench on its own. Whether the bench is still provisioning is asked of the bench rather than read off an error message: a bench still deploying is someone waiting, not someone broken.
Records the rule connecting now follows — connecting deploys nothing — along with the fast/slow split it comes from, why the pending-seed row makes the drain idempotent, convergent and restart-safe, why the provisioning session has to be the bench owner's own, and what someone actually sees when they land before their agents exist. Corrects the model-seeding note that still described onboarding as deploying the default workflow set inline.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pasting an Anthropic key on onboarding sat on a static "Connecting…" for over two minutes. The connect request was deploying five default workflows — around twenty seconds each — before it answered. Long enough to read as frozen.
Connecting now deploys nothing.
What connect does now
Connect keeps only the durable, fast half: persist the credential, seed its model catalog, answer. Deploying a bench's default workflows moves to a background drain, and no HTTP route deploys anything any more —
POST /completeandPOST /complete-setupboth report where the agents stand and hand the work off, and a newGET /api/onboarding/provisioning-statusserves anything that needs to poll.completeCredentialSetupstill composes both halves for harnesses that need a fully deployed bench before a scenario begins (@workbench/evals' real target), and its comments now say plainly that a route calling it re-creates the freeze.Why the drain converges
It is not a job queue. It is a pass over
onboarding.pending_seed, where the row is the work item and means "this bench has a credential and does not yet have its agents". The properties fall out of that framing:seedTenantunderneath is ensure-then-create. A pass over a finished bench deploys nothing and clears the row.A failing bench backs off (15s, doubling to a 10-minute ceiling) rather than hammering a down sidecar, and is never dropped for failing.
Sessions
A background loop has no request to borrow cookies from. The provisioner takes
sessionForas a seam and the hub fills it by minting the bench owner's own session in process. It has to be the owner's: the hub resolves a tenant by looking up a principal for that user, and rights do not flow from a parent org to a child bench, so an administrator would simply be refused. These sessions carry aworkbench-bench-provisioneruser agent so they are never mistaken for a human sign-in, and are cached per user rather than minted per tick.What someone sees
Connecting hands the person forward immediately. If they land before their agents exist they get the warm loading state — "Getting your agents ready…" with live progress, N of M ready — which resolves into their workbench on its own. Whether the bench is still provisioning is asked of the bench rather than read off an error message: still deploying is someone waiting, not someone broken. This closes a real hole, since the zero-workbench land path previously threw "No default setup agent found" and showed an error card.
Tests
The headline test is a clock, not a mock: the deploy step is a seam that takes five seconds, and connect has to answer in under one without that seam ever starting. A route that awaits a deploy again fails on elapsed time.
Alongside it: drain idempotency, convergence of a half-provisioned bench, restart-resume by a fresh provisioner with empty memory, no double-deploy under overlapping passes, backoff, and the client's connected/pending rendering.
Scoped green —
packages/onboarding149 pass / 0 fail,apps/web748 pass / 0 fail,apps/hub137 pass / 0 fail, typecheck and lint clean across all three.Fixes CL-6457