Skip to content

Split walking-skeleton CI and cut repeated install tax - #471

Merged
TheGreatAxios merged 4 commits into
mainfrom
cl-7187-split-the-walking-skeleton-ci-job-and-cut-repeated-install
Aug 30, 2026
Merged

TheGreatAxios merged 4 commits into
mainfrom
cl-7187-split-the-walking-skeleton-ci-job-and-cut-repeated-install

Conversation

@TheGreatAxios

Copy link
Copy Markdown
Contributor

Fixes https://linear.app/abklabs/issue/CL-7187/split-the-walking-skeleton-ci-job-and-cut-repeated-install-tax

Walking-skeleton is no longer one blob: e2e, isolation, and database-backed package+hub suites each get their own job, Postgres, and timeout.

Setup is a composite that installs Bun, restores caches, and runs frozen install. Checkout stays in the workflow so we do not clone twice. Jobs that do not need merge-base stay shallow; typecheck, build-test, and structural take a full clone.

CI package fan-out uses all runner cores. Lint result caches are keyed on Bun, lockfile, eslint config, and prettier config. node_modules restore is the optional experiment — revert if workspace linking goes stale.

Database-backed suites still run every PR with E2E_REQUIRED=1.

The db-suites job pointed DATABASE_URL at a bare Postgres service
container and ran the suites straight away. Nothing had created or
migrated the database, so the hub's mount tests failed with
"database \"workbench\" does not exist". The package suites resolve
their URL through e2eDatabaseUrl(), which appends an _e2e suffix, so
they need a second database created the same way.

Both are now created and migrated by scripts/db-setup.ts before the
suites run. e2e and isolation are unaffected: they provision their
own schema through the harness.
@TheGreatAxios

Copy link
Copy Markdown
Contributor Author

Reviewed against CL-7187's acceptance criteria, the CI YAML diff, and the two open questions about test coverage and flakiness.

Verified against acceptance criteria — all satisfied:

  • e2e, isolation, and db-suites (hub + package DB-backed suites) now run as separate jobs, each with its own Postgres and timeout. scripts/ci-jobs.test.ts pins the split so a rename can't silently revert it.
  • Setup (checkout stays inline, Bun + cache restore + frozen install) is shared via .github/actions/setup-workbench; typecheck/build-test/structural keep fetch-depth: 0 for merge-base, lint/e2e/isolation/db-suites go shallow.
  • resolveConcurrency uses every core under GITHUB_ACTIONS=true and leaves two free locally; test stays sequential either way (SEQUENTIAL_SCRIPTS), unaffected.
  • Lint cache key is now hashFiles('bun.lock', 'eslint.config.ts', '.prettierrc.json', '.bun-version') instead of github.sha, so it can't restore a stale "clean" against a bumped tool.
  • node_modules restore is the additional experimental cache layer, keyed the same way as the bun install cache.
  • db-suites creates and migrates both workbench and workbench_e2e before running the hub and package DB-backed suites, with E2E_REQUIRED=1 on both.

Same-tests-still-run verdict: confirmed by diffing the old walking-skeleton job against the new e2e/isolation/db-suites jobs step by step. All four original suites (walking skeleton e2e, two-org isolation, hub DB-backed suites, package DB-backed suites) still run, with the same E2E_REQUIRED=1 hard-failure-on-skip guard. Isolation and e2e each provision their own schema through the harness independently (confirmed in test/isolation/setup.ts, which imports scripts/db-setup.ts itself) — the split job doesn't depend on the old job's incidental ordering for schema setup, which the old single-job version implicitly did. Nothing dropped.

CI flakiness ("stolen state cookie" test) verdict: that test lives in packages/connections/packages/onboarding and runs under build-test's bun run test, which fans out per-package at concurrency 1 (SEQUENTIAL_SCRIPTS) both before and after this PR — this PR's resolveConcurrency change only affects non-sequential scripts (typecheck, lint), not test. build-test's own job body is unchanged apart from the composite-action refactor. What this PR does change is the total number of concurrent CI jobs per PR run, from 5 to 7 (walking-skeleton split into three). That's a plausible, if indirect, contributor to the "CI resource contention" a separate investigation already attributed the flake to — more concurrent jobs compete for the same account's Actions capacity — but it isn't a change to the flaky test's own job internals, and the split's stated goal (a flake naming the suite that failed, wall-clock as the slowest job not the sum) is the intended tradeoff.

bunx prettier --check, bun run lint, and bun run check:structural all clean on this branch. scripts/ci-jobs.test.ts and scripts/run-all.test.ts (18 tests) pass. Commit messages: no prefixes, no tracker refs, within 72 chars.

No code changes needed.

@TheGreatAxios
TheGreatAxios merged commit 609cbda into main Aug 30, 2026
13 of 14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant