fix(gateway): enforce JWT auth and give every caller credentials first - #12
Merged
Merged
Conversation
added 14 commits
September 11, 2026 11:30
…edentials The gateway never validated a JWT. `server.ts` called `app.register(authPlugin)`, which encapsulates the plugin, so its `onRequest` hook applied only to routes registered inside that child context - and there were none. healthRoutes, v2Routes, v1Routes and multipassRoutes are siblings on the parent, so the hook never ran: with AUTH_PUBLIC_KEY configured, GET /api/v2/ontologies still answered 200 with real data, `request.claims` and `request.orgRid` stayed empty, and nothing downstream could enforce anything. Credentials land before enforcement does, so no commit leaves a caller broken. Callers first: - scripts/lib/auth.sh (`of_curl`) and scripts/lib/auth.ts (`authHeaders`) resolve a bearer token from OPENFOUNDRY_TOKEN, or from OPENFOUNDRY_CLIENT_ID/_CLIENT_SECRET via the client_credentials grant, read from the environment or the gitignored .env. Nothing is baked into the repo. With neither set they send no header, which is exactly today's behaviour. - Wired into the Kelava sync scripts, every seed script, seed-dev-data.ts, seed-kelava-alerts.sh, test-integration.sh, the CI readiness probe and tests/integration/kelava-sync.test.ts. - The console has ~80 bare `fetch` call sites, so one interceptor in lib/authFetch.ts attaches the stored token to gateway requests. - bootstrap-admin.sh keeps no credential on purpose: it runs before one exists, against the signup route the middleware already exempts. - .env.example registers AUTH_PUBLIC_KEY, JWT_PRIVATE_KEY/JWT_PUBLIC_KEY and the OPENFOUNDRY_* credential keys, all empty. Then enforcement: - Call authPlugin directly, like rateLimitPlugin, with a comment explaining why so the next person does not re-register it. - Import AUTH_PUBLIC_KEY as a key object via the new `importVerificationKey`. The PEM's raw bytes were being handed to jose as an HMAC secret, which cannot verify an ES256 token, so enforcement alone would have rejected every caller. A malformed key now fails at startup. - Default AUTH_ISSUER to "openfoundry-multipass", the only issuer this platform mints. The old "openfoundry" matched nothing. - Plugin state moved into the closure; module-level state leaked between server instances in one process. Trusted identity for downstream services: - proxy.ts always drops client-supplied X-User-Id / X-User-Roles and re-asserts them from the verified claims, so role-based enforcement finally has an identity source a client cannot forge. - Both headers are set together or not at all: the permission middleware denies a request with a user id but no roles, so a half-filled identity would 403 every downstream route. - Roles come from a new optional `roles` claim minted by svc-multipass. A token without it conveys no roles and downstream behaves exactly as before. - The proxy also stops forwarding the incoming Content-Length, which no longer describes the re-serialised body; a pretty-printed payload was failing with a 502 (this is why seed-kelava-alerts.sh never actually created its monitors). - `local` used outside a function aborted sync-kelava-incremental.sh. Verified against a gateway with AUTH_PUBLIC_KEY set: unauthenticated /api/v2/ontologies is 401 and the same request with a token is 200; the seed scripts, seed-dev-data.ts, the alert seeder and all 12 integration suites pass; the console logs in and every page loads. The Kelava sync was exercised against a local stand-in for Sanocare's Postgres - only its auth path is proven, not a sync against the real source.
PR #9 landed the client-header strip and documented, accurately at the time, that role enforcement still did not work: nothing set x-user-id / x-user-roles from a verified token, `request.claims` was never populated, and issued tokens carried no roles claim. This change is the trusted hop those notes were waiting for, so AGENTS.md, deploy/.env.production and deploy/docker-compose.yml said the opposite of what the code now does. ENFORCE_PERMISSIONS=true no longer means "every gateway-proxied route answers 403"; it means the roles the caller's token carries are enforced, given AUTH_PUBLIC_KEY and a token minted with a roles claim.
…e; harden console token handling
…resh; fix token restore
…fresh lint counts
Przyval
force-pushed
the
fm/openfoundry-auth-tidak-menegakkan
branch
from
September 11, 2026 04:56
011a44c to
16140e2
Compare
sdk-compat's Admin.Users.getCurrent() reads identity out of the bearer token, so it 401s whenever the caller has none. The suite used to hardcode admin/admin123; now that every caller resolves credentials from the environment through scripts/lib/auth.*, CI has to supply them. Set the seeded dev client_id/secret as job-level env so the readiness probe, the seed script and the vitest suites all authenticate the same way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Kapten mengembangkan OpenFoundry - emulator Palantir Foundry buatannya sendiri - menuju kesetaraan dengan Foundry.
Pemeriksaan mutu menemukan bahwa gateway OpenFoundry TIDAK PERNAH memvalidasi JWT sama sekali. Peninjau membuktikannya dengan reproduksi Fastify minimal: services/svc-gateway/src/server.ts memanggil app.register(authPlugin), yang mengenkapsulasi plugin itu, sehingga hook onRequest di dalamnya hanya berlaku untuk rute yang didaftarkan DI DALAM konteks anak itu - dan tidak ada. healthRoutes, v2Routes, dan multipassRoutes adalah saudara yang terdaftar di induk, jadi hook-nya tidak pernah berjalan. Akibat konkret: dengan AUTH_PUBLIC_KEY terpasang sekalipun, GET /api/v2/ontologies tanpa otentikasi tetap dijawab 200, request.claims dan request.orgRid tidak pernah terisi, dan karena itu penegakan izin apa pun di hilir tidak berfungsi.
Perbaikan kodenya kecil - panggil authPlugin langsung seperti rateLimitPlugin sudah dilakukan di baris 91, dengan komentar yang menjelaskan persis alasannya. Yang TIDAK kecil adalah akibatnya, dan itu sebabnya ini menunggu keputusan kapten sampai sekarang: begitu otentikasi benar-benar menegakkan, scripts/sync-kelava.sh, skrip penyemai, dan tests/integration/kelava-sync.test.ts semuanya mulai menerima 401 karena tidak satu pun mengirim header Authorization. Sinkronisasi data operasional Safe and Care - usaha kapten sendiri - berhenti bekerja.
KEPUTUSAN KAPTEN, 2026-09-11: siapkan kredensial dulu, baru nyalakan otentikasi, DALAM PR YANG SAMA. Alasannya: sinkronisasi Safe and Care tidak boleh putus di tengah jalan, jadi tidak boleh ada jendela di mana otentikasi menegakkan sementara pemanggilnya belum punya kredensial.
What Changed
services/svc-gateway/src/server.tsnow callsauthPlugin(app, { config })directly instead ofapp.register(...), so itsonRequesthook actually reaches the sibling route plugins (health, v1/v2, multipass) - previously registration encapsulated the plugin and every request was served unauthenticated even withAUTH_PUBLIC_KEYset. All route registration moved below that call, the key is now imported as a real key object via newimportVerificationKey/importSigningKeyhelpers in@openfoundry/auth-tokens(raw PEM bytes were being read as an HMAC secret), the defaultAUTH_ISSUERis corrected toopenfoundry-multipass, and empty env vars are treated as unset rather than as "check nothing".scripts/lib/auth.sh(of_curl) andscripts/lib/auth.ts(authHeaders) source a bearer token fromOPENFOUNDRY_TOKENor a client-credentials exchange, and the demo seeds, Kelava sync scripts and integration tests were converted to use them; the console installs onewindow.fetchinterceptor (apps/app-console/src/lib/authFetch.ts) frommain.tsx, with session restore/expiry decisions extracted intoapps/app-console/src/lib/session.ts.NODE_ENV=production, sources tokenrolesfrom the authenticated client/user and mints none from the unauthenticated/authorizeor refresh grants; the proxy drops the staleContent-Length, and.env.example,deploy/,AGENTS.mdandGUIDE.mdwere updated to match. New tests cover gateway auth and header stripping, console auth fetch and session restore, plus expanded auth-tokens, sdk-oauth and multipass suites.Risk Assessment
Testing
I ran the targeted unit suites for every package this change touches (svc-gateway, app-console, auth-tokens, svc-multipass, sdk-oauth) and they are green, then stood up a real three-service stack with AUTH_PUBLIC_KEY configured and exercised it as an end user would: unauthenticated and tampered-token calls to /api/v2/ontologies are rejected 401, /metrics is guarded while /status/health stays open, a client_credentials token unlocks the same call at 200, and the shipped scripts/lib/auth.sh helper authenticates a shell caller unchanged. To prove the regression rather than just the current state, I temporarily restored the pre-fix encapsulated
app.register(authPlugin)line on a second gateway and confirmed it still served the same request 200, then restored the file and verified the worktree is clean. On the UI side I drove the console in a real browser against that enforcing gateway and captured screenshots of the login screen, the authenticated Ontology Explorer rendering live data, and confirmed the authFetch interceptor attaches the bearer token; a half-written session (token without user) now clears and lands on the login screen with no dead token left behind. No failures and nothing actionable found.Evidence: Gateway auth end-to-end CLI transcript (401 unauthenticated, 200 with token, pre-fix 200 control, console + session checks)
~/.no-mistakes/evidence/01M275S8KTMCZ7HMFYQSCB4R2C/console-01-login-unauthenticated.png)~/.no-mistakes/evidence/01M275S8KTMCZ7HMFYQSCB4R2C/console-03-ontology-authenticated.png)Evidence: Before/after control
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
services/svc-multipass/src/store/client-store.ts:49- The seeded confidential dev clients now carry platform roles (admin/admin123-> ADMIN,developer/dev123-> EDITOR) andClientStore's constructor seeds them unconditionally, unlikeDEV_USERSin services/svc-multipass/src/routes/auth.ts:49 which is emptied when NODE_ENV=production. Concrete path: deploy/docker-compose.yml runs NODE_ENV=production with ENFORCE_PERMISSIONS=true; once AUTH_PUBLIC_KEY is set (which this change's docs now present as the working posture), anyone who can reach the public gateway can POST /multipass/api/oauth2/token with grant_type=client_credentials and the repo-published admin/admin123, receive a token whose signedrolesclaim is ADMIN, and the gateway proxy (services/svc-gateway/src/proxy.ts) re-asserts X-User-Roles: ADMIN on all 118 guarded downstream routes. Before this change therolesclaim did not exist and no identity header was ever set, so these credentials conveyed nothing. Smallest remedy is to gate the roles (or the confidential dev clients themselves) on NODE_ENV the same way DEV_USERS is - but that changes which credentials the captain's scripts can use in a production deployment, so it needs a decision rather than a silent fix.services/svc-multipass/src/routes/oauth.ts:185- The authorization_code flow storesclient.roleson the auth code, so a user token's roles come from the OAuth client rather than the authenticated user. The console's production login path (loginWithOAuth, PKCE against clientopenfoundry-console/openfoundry-dev, which declare no roles) therefore yields a token with norolesclaim; the proxy then sends neither X-User-Id nor X-User-Roles, and with ENFORCE_PERMISSIONS=true every guarded route answers 403 for a legitimately logged-in operator. The dev password login (services/svc-multipass/src/routes/auth.ts:133) does carry user roles, which is why this is invisible locally. It also contradicts the new deploy/.env.production:57 comment thattruenow "enforces the roles a token actually carries". The remedy - sourcing roles from the approved user record at authorize time - adds state beyond this change's stated intent, so it needs authorization.tests/integration/api-endpoints.test.ts:11- Only tests/integration/kelava-sync.test.ts was routed throughauthHeaders. tests/integration/api-endpoints.test.ts, tests/integration/sdk-compat.test.ts and tests/integration/tenant-isolation.test.ts still call the gateway at localhost:8080 with a barefetchand no Authorization header, so they 401 wholesale the moment AUTH_PUBLIC_KEY is configured - the same gap the intent required closing before enforcement is switched on. Mechanical fix: give each suite theapiFetchwrapper kelava-sync.test.ts now uses.apps/app-console/src/lib/authFetch.ts:62-isApiRequesttestsurl.startsWith(API_BASE_URL). API_BASE_URL isimport.meta.env.VITE_API_URL ?? "http://localhost:8080", and??does not replace an empty string - a same-origin build that setsVITE_API_URL=makes API_BASE_URL"", every URL matches the prefix, and the interceptor attaches the bearer token to every outbound fetch including cross-origin ones (the existing "leaves third-party requests alone" test would flip to sending the token). Guard the absolute-URL branch on a non-empty API_BASE_URL.services/svc-multipass/src/store/token-store.ts:12- Roles are snapshotted onto the stored refresh token at issue time and replayed verbatim on refresh, so a role revocation does not take effect until the refresh token expires or is revoked. Acceptable for the in-memory demo store; noting the tradeoff rather than asking for a change.🔧 Fix: gate dev OAuth clients on NODE_ENV; source token roles from user
6 issues (1 error, 3 warnings, 2 infos) still open:
services/svc-multipass/src/routes/oauth.ts:178- GET /multipass/api/oauth2/authorize authenticates nobody: it auto-registers any unknown client_id (line 138), accepts any redirect_uri, and auto-approves as the hardcoded user "admin". The fix round now sources the auth code's roles from that hardcoded user (devUserRoles("admin") = ["ADMIN"]) instead of the client. Concrete path on any deployment where NODE_ENV is not "production" (start.sh, docker-compose.dev, any staging box) but AUTH_PUBLIC_KEY is set: an anonymous caller hits GET /multipass/api/oauth2/authorize?response_type=code&client_id=anything&redirect_uri=http://attacker/cb (gateway-exempt via SKIP_AUTH_PREFIXES), omits code_challenge so no PKCE verifier is required, POSTs the code to /oauth2/token, and receives a signed token whose roles claim is ADMIN. proxy.ts then asserts x-user-id/x-user-roles: ADMIN downstream, so the caller has full write authority on every guarded route. Before this commit that same token carried client.roles (undefined for an auto-registered client), so it conveyed no roles. This defeats the authentication the change exists to add. The remedy is a product decision - either stop minting roles from a flow that authenticates no one (only the password login at routes/auth.ts:130 actually verifies a user), or gate the auto-register/auto-approve behaviour - so it needs authorization rather than an automatic patch.apps/app-console/src/lib/authFetch.ts:63- isApiRequest treats any URL that merely starts with API_BASE_URL as the gateway. With API_BASE_URL="https://api.example.com", a fetch to "https://api.example.com.attacker.test/x" matches the prefix (path becomes ".attacker.test/x", which no CREDENTIAL_ENDPOINTS prefix excludes) and the bearer token is attached and sent to the attacker's host. Compare origins, or require the character after the prefix to be "/" or "?" (or the URL to end there), instead of a bare startsWith.apps/app-console/src/lib/authFetch.ts:30- The interceptor reads the raw token out of localStorage and never consults TokenManager; nothing outside AuthContext.tsx calls tokenManager.getAccessToken (grep across apps/app-console/src returns no other caller), so the refresh_token grant is never exercised. Concrete sequence: an operator logs in, works for longer than TOKEN_EXPIRY_SECONDS (3600), and from then on every page's fetch carries an expired token - with AUTH_PUBLIC_KEY set the gateway answers 401 on every screen and the console shows the user as still logged in with no recovery short of a manual re-login. The intent requires that enabling enforcement leaves every legitimate caller working; the remedy (routing the interceptor through TokenManager and handling a 401 retry) adds machinery beyond what the change currently does, so it needs authorization.deploy/.env.production:60- deploy/docker-compose.yml sets NODE_ENV=production for every service, and the fix round made ClientStore seed nothing there while DEV_USERS is already empty there. In that configuration no flow mints a token with a roles claim: client_credentials has no registered client (no HTTP registration endpoint exists), /authorize resolves devUserRoles("admin") to undefined, and the signup route (routes/auth.ts:215) builds its token with no roles field. proxy.ts therefore forwards no identity headers, and with ENFORCE_PERMISSIONS=true every guarded downstream route 403s - including the token the bootstrap-admin flow issues. The comment this change added at deploy/.env.production:57-59 states thattruenow "enforces the roles a token actually carries", which reads as working enforcement rather than a deployment that denies everything. Either the comment should say no production path yet mints roles, or a real production role source is needed - both are product calls.scripts/lib/auth.sh:73- of_resolve_token runs once when the library is sourced and the token is reused for the whole script run. A full Kelava sync that runs longer than TOKEN_EXPIRY_SECONDS (3600 by default) will start receiving 401s partway through with no re-mint, which is the exact caller the intent says must not break mid-run. Noting the limit rather than asking for a refresh loop now; it only bites on long runs against an enforcing gateway.services/svc-gateway/src/middleware/auth.ts:30- SKIP_AUTH_PREFIXES exempts /status/* but not GET /metrics (registered by metricsPlugin from @openfoundry/health). Once AUTH_PUBLIC_KEY is set, a Prometheus scraper pointed at the gateway gets 401. No scrape config lives in this repo, so nothing in-tree breaks - flagging it so an external scraper is not surprised.🔧 Fix: stop minting roles from unauthenticated authorize; harden console token handling
5 issues (3 warnings, 2 infos) still open:
services/svc-gateway/src/proxy.ts:103- Forwarding x-user-roles turns role enforcement on even where ENFORCE_PERMISSIONS is false. createHeaderBasedHook (packages/permissions/src/middleware.ts:147-168) short-circuits to 'granted' only when the role set contains the permission; when a role header IS present but lacks it, control falls through every enforcePermissions escape hatch (hasClaims is false downstream, userId is set) and reachesthrow permissionDenied(...). Concrete: gateway with AUTH_PUBLIC_KEY set, ENFORCE_PERMISSIONS unset (the demo/deploy default after this change), caller uses the seededdeveloperclient_credentials pair -> token carries roles=EDITOR -> any route guarded by admin:manage returns 403, where before this change the same call succeeded. Same for a console login as analyst (VIEWER) on any write route. That contradicts the intent's requirement that turning authentication on must not break existing callers, and it makes the newly written 'ENFORCE_PERMISSIONS=false' comments in deploy/ inaccurate. Remedy is a product call (gate the header assertion on the enforcement flag, or fix the hook's fall-through), so it needs authorization.services/svc-multipass/src/routes/oauth.ts:344- handleRefreshTokenGrant validates nothing but the refresh token itself - no client_id, no client_secret - and this change added arolesfield to the stored refresh token (line 408) that it now replays into the new access token (line 365). Concrete: a client_credentials exchange with the seeded admin/admin123 pair returns a refresh token; anyone holding only that value (it is not a secret in the client-authentication sense, and scripts/lib/auth.sh keeps it in-process but the grant is open to any caller) can POST grant_type=refresh_token with no client credentials and receive a fresh roles=ADMIN token, repeatedly via rotation. Before this change a refreshed token carried no roles, so the same leak conveyed no authority. Remedy (requiring client authentication on the refresh grant) changes the token API's accepted requests, so it needs authorization rather than a silent fix.apps/app-console/src/context/AuthContext.tsx:115- The interceptor now goes through TokenManager, but TokenManager.setToken computes #expiresAt = Date.now() + expiresIn*1000 from the relative expiresIn persisted in localStorage. The restore-on-mount path replays a token stored an arbitrary time ago, so a token that expired an hour back is treated as fresh for another full hour: #isExpiringSoon() is false, no refresh is attempted, and every request carries an expired token that the gateway 401s. Concrete: OAuth login at T (expiresIn 3600), reload the console at T+2h -> restore sets expiresAt=T+2h+1h -> getToken() returns the dead access token even though a valid refresh token is present. The claimed fix therefore does not hold on the most common path (page reload). The smallest honest remedy is persisting an absolute expiry alongside the token, i.e. new persisted state, so the remedy - not the defect - is what needs authorization.apps/app-console/src/lib/authFetch.ts:97- The new boundary check rejects a URL whose remainder after API_BASE_URL does not start with /, ? or #. If VITE_API_URL is configured with a trailing slash ("http://gateway:8080/"), then "http://gateway:8080/api/v2/ontologies" leaves rest="api/v2/ontologies", isApiRequest returns false, and no bearer token is attached - every console request 401s under enforcement with no diagnostic. Normalise by stripping trailing slashes from API_BASE_URL before the comparison.services/svc-multipass/src/store/client-store.ts:72- The comment added with the production gate says "in production, clients are registered at runtime", but no client-registration endpoint exists; deploy/.env.production (added in the same branch) correctly states that none can be registered. Two comments in the same change describe opposite realities; the client-store one should say that production seeds no client and that the only runtime registration is the implicit one /authorize performs.🔧 Fix: drop gateway identity assertion and roleless refresh; fix token restore
5 issues (1 error, 3 warnings, 1 info) still open:
deploy/.env.production:25- The gateway's default issuer was corrected to "openfoundry-multipass" (services/svc-gateway/src/config.ts:91) because that is the onlyissthe platform signs (services/svc-multipass/src/routes/oauth.ts:383, routes/auth.ts:127), but the deployment still pins the old value: deploy/.env.production:25 sets AUTH_ISSUER=openfoundry and deploy/docker-compose.yml:191 defaults to the same. Concrete failure: set AUTH_PUBLIC_KEY in that deployment (the whole point of this change), log in, and validateToken rejects the issuer - every authenticated request answers 401 InvalidAuthToken, including the Safe and Care sync callers the intent requires not to break. Fix: set both to openfoundry-multipass.services/svc-multipass/src/routes/oauth.ts:126- Enforcement is now real at the gateway, but the token-minting path it exempts is unauthenticated: GET /multipass/api/oauth2/authorize accepts any client_id (auto-registering it as a public client, line 138-150), accepts any redirect_uri, and approves the hardcoded user "admin" (line 173). handleAuthorizationCodeGrant then skips the secret check for public clients (line 277). So any caller who can reach the gateway does authorize -> token and obtains a signed token that passes the new middleware; with AUTH_PUBLIC_KEY set the gateway rejects anonymous requests but nothing stops an anonymous caller from minting credentials first. This is pre-existing code, but the change makes it the single load-bearing hole in the feature it claims to deliver. The remedy - authenticating /authorize (a real login/consent step) - extends this PR's scope, so it needs your decision, not an automatic patch; at minimum the claim that the gateway now enforces authentication should be qualified in AGENTS.md and deploy notes.apps/app-console/src/context/AuthContext.tsx:126- The restore fix is undone in storage by its own persistence hook. tokenManager.setToken(token, expiresAt) stores the correct absolute deadline in memory, but setToken also calls onTokenChange, and the handler at line 103-108 re-persists expiresAt = Date.now() + token.expiresIn * 1000 - a brand-new full hour. Concrete: OAuth login at T (expiresIn 3600); reload at T+2h -> in-memory expiry is correct and a refresh is attempted, but storage now claims T+3h. If that refresh fails (the dev password login mints no refresh token, so getToken throws and authFetch falls back to the stored value), or the user simply reloads again at T+2.5h, the restored session is treated as fresh and the dead access token is sent - exactly the bug this round was meant to close. Fix within the existing stored state: have the restore path re-write the entry with the original expiresAt (or suppress the persist for that setToken call).services/svc-multipass/src/store/client-store.ts:76- The newif (process.env.NODE_ENV === "production") return;leaves the ClientStore empty in production, and deploy/docker-compose.yml sets NODE_ENV: production for every service. There is no client-registration endpoint, so handleClientCredentialsGrant always throws "unknown client" there. The credential path this PR builds for scripts and tests (OPENFOUNDRY_CLIENT_ID/_CLIENT_SECRET exchanged via client_credentials, scripts/lib/auth.sh and auth.ts) therefore cannot work against the production deployment at all: mintAccessToken logs a warning and returns null, and the caller proceeds unauthenticated into 401s. Only a pre-minted OPENFOUNDRY_TOKEN works, and nothing in the deployment can mint one for a service. That interacts directly with the requirement that credentials be ready before enforcement is switched on. The remedy (a registration path or a seeded production client from env) adds machinery beyond this change, so it is your call - at minimum .env.example and the deploy notes should state that only OPENFOUNDRY_TOKEN is usable in production.packages/auth-tokens/src/claims.ts:48- Two comments still describe the behavior the final commit deleted. claims.ts:46-49 says "The downstream X-User-Roles header is set by the gateway from this claim and from nothing else - see services/svc-gateway/src/proxy.ts", and services/svc-multipass/src/store/client-store.ts:17-21 says "The gateway turns them into the downstream X-User-Roles header". proxy.ts now sets no identity headers at all. Reword both to match oauth.ts:386 ("minted, but with no consumer yet").🔧 Fix: fix deploy issuer drift, persist absolute token expiry
2 issues (1 warning, 1 info) still open:
deploy/docker-compose.yml:237- Turning on the feature this PR delivers is impossible in the shipped deployment. The gateway verifies with AUTH_PUBLIC_KEY (deploy/docker-compose.yml:190, deploy/.env.production:24), but the svc-multipass service block passes no JWT_PRIVATE_KEY/JWT_PUBLIC_KEY and deploy/.env.production defines neither, so services/svc-multipass/src/routes/oauth.ts:84 and routes/auth.ts:90 each fall back to generating their OWN ephemeral ES256 pair at startup. Concrete failure: operator follows the new .env.example guidance, sets AUTH_PUBLIC_KEY in deploy/.env.production, restarts - every token multipass mints is signed by a key nobody configured, so validateToken fails and every authenticated request answers 401 InvalidAuthToken (and password-login tokens and OAuth tokens are signed by two different ephemeral keys even then). Fix in the files this change already edits: add JWT_PRIVATE_KEY/JWT_PUBLIC_KEY to deploy/.env.production and pass them through to svc-multipass in docker-compose.yml, alongside AUTH_PUBLIC_KEY.services/svc-multipass/tests/multipass.test.ts:216- The comment sweep for stale identity claims missed the test file added in the same round. multipass.test.ts:216 says "the gateway turns a token's roles claim into X-User-Roles for all guarded routes" and the docblock at ~line 404 says the gateway "forwards no X-User-Id/X-User-Roles for a token with no roles claim", implying it forwards them when roles are present. services/svc-gateway/src/proxy.ts sets no identity headers at all in any case. Reword both to match the wording now used in claims.ts and oauth.ts (minted, no consumer yet).🔧 Fix: plumb JWT signing keys into deploy; fix stale test comments
3 issues (2 warnings, 1 info) still open:
services/svc-multipass/src/routes/oauth.ts:80- The newly documented way to supply the signing keys cannot actually work. The gateway's new importVerificationKey (packages/auth-tokens/src/token-validator.ts:74) normalises literal\nescapes before importSPKI, but svc-multipass imports the same env-sourced PEMs raw: routes/oauth.ts:80-81importPKCS8(config.jwtPrivateKey, "ES256")/importSPKI(config.jwtPublicKey)and routes/auth.ts:88. Concrete failure: an operator follows the new deploy/.env.production:23-32 and .env.example:27-33 instructions, and since a docker-compose.envfile cannot carry a multi-line value, puts the PEM in with\nescapes. The gateway imports its key fine; svc-multipass throws inside authRoutes/oauthRoutes at startup and the service never comes up, so every login and token request fails. If instead the operator sets only one of the two vars, theconfig.jwtPrivateKey && config.jwtPublicKeyguard silently falls through to an ephemeral key pair and every token is then rejected by the gateway with 401 InvalidAuthToken - the exact failure the previous fix round set out to close. Smallest honest fix: run the multipass-side PEMs through the same normalisation (reuse the helper, or.replace(/\\n/g, "\n").trim()), and warn when exactly one of the pair is set instead of silently generating ephemeral keys.scripts/lib/auth.ts:99- scripts/lib/auth.ts:99 usesprocess.env.OPENFOUNDRY_AUTH_URL ?? baseUrl, which only falls back when the variable is absent, not when it is empty. .env.example:49-50 shipsOPENFOUNDRY_AUTH_URL=empty, and scripts/sync-kelava.sh:24 / sync-kelava-incremental.sh:21 doset -a && source "$ROOT_DIR/.env", exporting it as an empty string to any tsx caller in that shell. Trace: OPENFOUNDRY_CLIENT_ID/_SECRET set, OPENFOUNDRY_AUTH_URL="" -> authUrl "" ->fetch("/multipass/api/oauth2/token")is not a valid absolute URL in Node, throws, is caught at line 113, logs "could not reach" and returns null -> the caller proceeds unauthenticated and gets 401s from an enforcing gateway, with the cause reported as unreachability rather than misconfiguration. The shell twin gets this right (${OPENFOUNDRY_AUTH_URL:-...}in scripts/lib/auth.sh:69), so the two are out of step. Use|| baseUrl.packages/auth-tokens/src/claims.ts:64-parseRolesis added and exported from packages/auth-tokens (claims.ts:64, index.ts:6) but has no caller or test anywhere in the repo - the roles claim is deliberately left with no consumer until the permission-model work lands. Noting it so it is understood as a placed-ahead API rather than live code; no action needed unless you prefer it to arrive with its consumer.🔧 Fix: share PEM normalisation, fix empty env fallback, drop parseRoles
2 warnings still open:
services/svc-gateway/src/server.ts:160- Per-tenant rate limiting can never engage, and this change is what makes that visible.rateLimitPluginis called on the root instance at server.ts:93,authPluginat server.ts:160, and Fastify runs same-contextonRequesthooks in registration order - so the rate-limit hook always runs before the auth hook.extractKey(services/svc-gateway/src/rate-limiter.ts:275) readsrequest.orgRidfirst, which is still undefined at that moment, and falls through to the Authorization-header hash. Concrete trace: a request with a valid token carryingorg: "org-1"is bucketed asauth:<hash-of-that-token>, notorg:org-1; an hour later the console refreshes its token and the same tenant gets a brand-new empty bucket, so the documented "per-tenant isolation" never holds and the limit is per-token instead. Before this changeorgRidwas never populated at all, so the dead branch was invisible; now middleware/auth.ts:130 comments that it extracts the org claim "for RLS and per-tenant rate limiting", which is not what happens. The remedy is a decision, not a mechanical edit: moving theauthPlugincall aboverateLimitPluginmakes the org bucket work but also means unauthenticated floods are rejected before they are counted, which is a deliberate trade-off about whether the 401 path stays rate-limited. Either reorder deliberately, or correct the claim in the auth comment and inextractKey's documented precedence.services/svc-gateway/src/config.ts:94- An emptyAUTH_ISSUERorAUTH_AUDIENCEsilently turns the corresponding check off instead of falling back to the default.env()(config.ts:80) usesprocess.env[key] ?? fallback, so an exported empty string survives as "", and middleware/auth.ts:103-104 then maps it throughconfig.authIssuer || undefined- which jose reads as "do not verify this claim". Concrete path created by this change: the diff deletes theAUTH_ISSUER=openfoundryline from deploy/.env.production and adds a comment saying it is "deliberately not set". An operator following that comment by blanking the value they already have (AUTH_ISSUER=) rather than deleting the line gets a gateway that accepts anyissand anyaud, with no warning and every test still green - the opposite of the intent of pinning the issuer to what multipass signs. This is the same empty-string-environment class as theOPENFOUNDRY_AUTH_URLfix in this branch, but here empty has a security meaning. Two honest readings exist - treat empty as absent (use the default, so the check always runs) or keep empty as an intentional escape hatch and document it - so which one is wanted is the author's call. Note that apps/app-console/src/config.ts was correctly left on??: there an emptyVITE_API_URLis a supported same-origin build that lib/authFetch.ts explicitly handles.🔧 Fix: treat empty env vars as unset in gateway config
3 issues (1 warning, 2 infos) still open:
apps/app-console/src/context/AuthContext.tsx:129- Once the gateway enforces, an expired console session leaves the UI authenticated but every request 401ing, with no path back to the login screen. Concrete trace: log in with the dev password flow (login()at AuthContext.tsx:167 callssetTokenwith a response that carries norefreshToken), leave the tab for overTOKEN_EXPIRY_SECONDS(3600), reload. The restore path at line 129 callssetToken(token, expiresAt)with a past deadline and unconditionallysetCurrentUser(user)- it never checks the deadline it just restored - soisAuthenticatedis true. authFetch'scurrentToken()then asks the manager,#isExpiringSoon()is true,#refresh()throws for want of a refresh token, and the catch falls back toreadStoredToken(), which returns the dead token. Every page therefore renders useApi's "Authentication required (HTTP 401)" banner while the navbar still shows the user logged in, and nothing clears the session or redirects. This is reachable only because this change makes the gateway enforce, so the branch creates it. Flagged as ask-user because the remedy is new behavior rather than a correction: either treat a restored-and-expired session as logged out, or clear the session on a 401 from the interceptor - both are product decisions about what the user sees.services/svc-multipass/src/config.ts:39- Theenv()docstring added by the last fix round was copied verbatim from services/svc-gateway/src/config.ts and justifies itself with AUTH_ISSUER, which svc-multipass never reads - thisenv()serves JWT_PRIVATE_KEY, JWT_PUBLIC_KEY, DATABASE_URL and the expiry values. The rule it states is right; the example points at the wrong service. Rewrite the example in terms of a key this config actually reads (e.g. JWT_PRIVATE_KEY, where empty means "generate an ephemeral pair").services/svc-gateway/src/middleware/auth.ts:29-SKIP_AUTH_PREFIXESexempts /status/, /multipass/api/oauth2/ and the login endpoint, but not /metrics (registered via metricsPlugin at server.ts:150). With AUTH_PUBLIC_KEY set, a Prometheus scrape of the URL scripts/start-production.sh:220 advertises now gets 401. Nothing in deploy/ scrapes it today and the container healthcheck uses /status/health, so nothing breaks - noting it as a consequence to expect when a scraper is added, not a defect. Requiring a token on /metrics is a defensible posture.🔧 Fix: clear unrecoverable expired console session on restore
1 info still open:
apps/app-console/src/context/AuthContext.tsx:124- The last fix round was asked for a provider round-trip test (persist an expired, refresh-less session, mount the provider, assert not authenticated and storage cleared). What landed is apps/app-console/src/lib/session.test.ts, a pure unit test of theisRestorablepredicate only. The predicate is correct, but the wiring that actually fixes the defect - theif (isRestorable(stored))branch at AuthContext.tsx:124 callingclearStoredSession()instead ofsetCurrentUser(user)- has no test. Inverting that condition, or dropping theelsearm, would leave the console authenticated-but-401ing again and every test still passes. Reported as ask-user because the remedy, not the defect, needs authorization: apps/app-console declares no jsdom, no @testing-library/react and no vitest.config.ts (it falls back to vite.config.ts, which has notestblock), so a mount test means adding test dependencies and an environment - new machinery beyond this change's scope. The alternative is to accept the predicate-level coverage and say so.🔧 Fix: extract and test console session restore decision
1 info still open:
ℹ️
apps/app-console/src/lib/session.ts:59- When exactly one of the two persisted keys is present,restoreSessionreturns null at line 59 without clearing storage, so the unusable half-session survives the reload that found it - contrary to the function's own docstring. Concrete path:openfoundry_useris removed (evicted, cleared by a user, or written by a build that only set the token key) whileopenfoundry_tokenholds an hour-old access token. The provider shows the login screen, butreadStoredToken()(apps/app-console/src/lib/authFetch.ts:65) still finds the token key and the interceptor attachesAuthorization: Bearer <dead token>to every API request, which the gateway now rejects with 401 instead of the request simply going unauthenticated. Remedy is mechanical and stays inside this change: callclearStoredSession()on that!storedToken || !storedUserpath when either key is present, and cover it with the same stubbed-storage test style already in session.test.ts.apps/app-console/src/lib/session.ts:59-restoreSessionreturns null at theif (!storedToken || !storedUser) return null;guard without clearing storage, contradicting its own docstring ("in which case the stored keys are cleared"). Concrete path:openfoundry_useris missing (evicted, cleared, or never written after a failed login) whileopenfoundry_tokenstill holds an hour-old access token. The provider shows the login screen, butreadStoredToken()(apps/app-console/src/lib/authFetch.ts:65) still finds the token key, so the interceptor attachesAuthorization: Bearer <dead token>to every API request and the now-enforcing gateway answers 401 instead of the request simply going out unauthenticated. Remedy is mechanical and inside the change's scope: callclearStoredSession()on that path when either key is present, and cover it with the same stubbed-storage style already in session.test.ts.ℹ️
services/svc-gateway/src/proxy.ts:91- The header loop now has two consecutive skips for the same header: the pre-existingif (lowerKey === "content-length") continue;and the newif (BODY_DESCRIBING_HEADERS.has(lowerKey)) continue;. The second is unreachable for its only member, so the documented set is dead. Keep one: either drop the literal check and let the set own the rule (matching the new docstring), or drop the set.🔧 Fix: clear orphaned half-sessions; drop duplicate content-length skip
2 issues (1 warning, 1 info) still open:
services/svc-gateway/src/server.ts:150-await app.register(metricsPlugin)runs at line 150, beforeauthPlugin(app, { config })at line 160. Fastify snapshots the parent's hooks when it creates an encapsulated child context, andmetricsPluginis a plain async plugin (not wrapped in fastify-plugin), so theGET /metricsroute it registers is created before the authonRequesthook exists on the root instance and never inherits it. Concrete path: with AUTH_PUBLIC_KEY set,GET /metricswith no Authorization header returns 200 Prometheus text, whileGET /metrics/:service(registered by healthRoutes after line 160) correctly returns 401. This directly contradicts the note this change added at AGENTS.md:122 ("SKIP_AUTH_PREFIXES does not exempt /metrics, so with AUTH_PUBLIC_KEY set a Prometheus scrape gets 401 - a deliberate posture"), so the one route escaping the new enforcement is the one documented as covered. The same encapsulation applies to the plugin's own request-timing hooks, which therefore record nothing for any real route - the same bug class this change exists to fix. Remedy is a posture decision, not a mechanical correction: either move the metricsPlugin registration below authPlugin (scrapes then need a token, as AGENTS claims) or add/metricsto SKIP_AUTH_PREFIXES and correct the AGENTS note; separately, making the metrics hooks actually cover routes needs fastify-plugin, which is beyond this change's scope.deploy/.env.production:74- ENFORCE_PERMISSIONS goes true -> false in deploy/.env.production and deploy/docker-compose.yml. The reasoning checks out against packages/permissions/src/middleware.ts: with the gateway deliberately asserting no X-User-* headers,truemakes every guarded route 403 rather than enforce anything, so the previous value was deny-all, not protection. Noting the posture change explicitly: production guarded routes are now open to any authenticated caller, and the change documents that in three places (deploy files, GUIDE.md, AGENTS.md) as pending the permission-model unification. No action needed unless the captain wants deny-all retained until that work lands.🔧 Fix: register gateway routes after auth guard; pin /metrics 401
1 info still open:
services/svc-gateway/tests/auth.test.ts:195- The new "guards /metrics like every other route" test asserts only that an unauthenticated GET /metrics is 401. Fastify's 404 context inherits root onRequest hooks, so any unrouted path returns 401 too - the assertion holds whether or not /metrics exists, and would still pass if metricsPlugin were removed or its route renamed. It also cannot fail in the ordering it was written to pin, since _addHook recurses into kChildren and a root hook reaches earlier-registered plugins either way. Add the authenticated half: GET /metrics with a valid signed token must return 200 with Prometheus text.✅ **Test** - passed
✅ No issues found.
npx turbo run test --filter=@openfoundry/svc-gateway --filter=@openfoundry/svc-multipass --filter=@openfoundry/auth-tokens --filter=@openfoundry/sdk-oauth --filter=@openfoundry/app-console(all green)Manual E2E: booted svc-gateway/svc-multipass/svc-ontology/svc-objects/svc-actions with a shared ES256 keypair andAUTH_PUBLIC_KEYset; probedGET /api/v2/ontologiesunauthenticated (401), with a garbage bearer (401), with client-assertedx-user-id/x-user-roles(401), with a realclient_credentialstoken (200), andGET /status/health(200)Before/after regression: temporarily restored the pre-fixawait app.register(authPlugin, { config })on the same running stack - unauthenticated request served HTTP 200; restored the fix - same request 401bash scripts/demo/seed-pest-control.shagainst the enforcing gateway withOPENFOUNDRY_CLIENT_ID/_SECRETset (full dataset seeded viaof_curl)npx tsx scripts/seed-dev-data.tswith and without credentials (401 without; 65 authenticated ontology/object writes with)npx vitest run --config tests/vitest.config.ts tests/integration/kelava-sync.test.tsagainst the enforcing gateway, with credentials (9/9 pass) and without (fails with an explicit 401 credential message rather than skipping)Browser (Playwright, 1440x900): dev login through the enforcing gateway, ontology data rendered, then backdated the persistedexpiresAton a refresh-token-less session and reloaded - landed on the login screen with both localStorage keys cleared✅ No issues found.
npx turbo run test --filter=@openfoundry/svc-gateway --filter=@openfoundry/auth-tokens --filter=@openfoundry/svc-multipass --filter=@openfoundry/sdk-oauth --filter=@openfoundry/app-console(16/16 tasks, gateway 55 tests, console 51 tests)Live stack: svc-ontology :18081, svc-multipass :18084, svc-gateway :18080 booted with a generated ES256 key pair in AUTH_PUBLIC_KEY / JWT_PRIVATE_KEYcurl $G/api/v2/ontologiesunauthenticated -> 401 MissingAuthToken;-H 'Authorization: Bearer not.a.token'-> 401 InvalidAuthTokencurl $G/metrics-> 401 andcurl $G/status/health-> 200 (guarded vs skipped prefixes)OAuth2 client_credentials token exchange through the gateway, thencurl -H "Authorization: Bearer $TOKEN" $G/api/v2/ontologies-> 200 with ontology datasource scripts/lib/auth.sh && of_curl $G/api/v2/ontologieswith OPENFOUNDRY_CLIENT_ID/_SECRET -> 200 (shell caller credential path)Before/after control: temporarily restored the pre-fixapp.register(authPlugin, { config })line, booted a second gateway on :18090 -> unauthenticated GET /api/v2/ontologies returned 200 (bug reproduced); file restored, worktree cleanBrowser (chrome-devtools-axi) against console on :3000 with VITE_API_URL pointed at the enforcing gateway: unauthenticated /ontology -> /login; dev login admin/admin123 -> 200; /ontology renders live ontologies; inspected request headers showing the authFetch-attached Bearer tokenBrowser: removed theopenfoundry_userkey leavingopenfoundry_tokenbehind, reloaded /ontology -> redirected to /login and localStorage token cleared🔧 **Document** - 1 issue found → auto-fixed ✅
AGENTS.md:42- AGENTS.md recordspnpm run testas green at 97/97, but this change adds four new test files (apps/app-console/src/lib/authFetch.test.ts, apps/app-console/src/lib/session.test.ts, services/svc-gateway/tests/auth.test.ts, services/svc-gateway/tests/trusted-identity.test.ts) and extends three existing suites, so that per-package task count is likely stale. I could not correct it because this phase is forbidden from running tests; the validation phase should re-runpnpm run testand update the number in AGENTS.md to the observed value.🔧 Fix: verify AGENTS.md task and lint counts; no changes
✅ Re-checked - no issues remain.
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.