Skip to content

fix(gateway): enforce JWT auth and give every caller credentials first - #12

Merged
Przyval merged 15 commits into
masterfrom
fm/openfoundry-auth-tidak-menegakkan
Sep 11, 2026
Merged

Przyval merged 15 commits into
masterfrom
fm/openfoundry-auth-tidak-menegakkan

Conversation

@Przyval

@Przyval Przyval commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Intent

Kapten mengembangkan OpenFoundry - emulator Palantir Foundry buatannya sendiri - menuju kesetaraan dengan Foundry.

Pemeriksaan mutu menemukan bahwa gateway OpenFoundry TIDAK PERNAH memvalidasi JWT sama sekali. Peninjau membuktikannya dengan reproduksi Fastify minimal: services/svc-gateway/src/server.ts memanggil app.register(authPlugin), yang mengenkapsulasi plugin itu, sehingga hook onRequest di dalamnya hanya berlaku untuk rute yang didaftarkan DI DALAM konteks anak itu - dan tidak ada. healthRoutes, v2Routes, dan multipassRoutes adalah saudara yang terdaftar di induk, jadi hook-nya tidak pernah berjalan. Akibat konkret: dengan AUTH_PUBLIC_KEY terpasang sekalipun, GET /api/v2/ontologies tanpa otentikasi tetap dijawab 200, request.claims dan request.orgRid tidak pernah terisi, dan karena itu penegakan izin apa pun di hilir tidak berfungsi.

Perbaikan kodenya kecil - panggil authPlugin langsung seperti rateLimitPlugin sudah dilakukan di baris 91, dengan komentar yang menjelaskan persis alasannya. Yang TIDAK kecil adalah akibatnya, dan itu sebabnya ini menunggu keputusan kapten sampai sekarang: begitu otentikasi benar-benar menegakkan, scripts/sync-kelava.sh, skrip penyemai, dan tests/integration/kelava-sync.test.ts semuanya mulai menerima 401 karena tidak satu pun mengirim header Authorization. Sinkronisasi data operasional Safe and Care - usaha kapten sendiri - berhenti bekerja.

KEPUTUSAN KAPTEN, 2026-09-11: siapkan kredensial dulu, baru nyalakan otentikasi, DALAM PR YANG SAMA. Alasannya: sinkronisasi Safe and Care tidak boleh putus di tengah jalan, jadi tidak boleh ada jendela di mana otentikasi menegakkan sementara pemanggilnya belum punya kredensial.

What Changed

  • services/svc-gateway/src/server.ts now calls authPlugin(app, { config }) directly instead of app.register(...), so its onRequest hook actually reaches the sibling route plugins (health, v1/v2, multipass) - previously registration encapsulated the plugin and every request was served unauthenticated even with AUTH_PUBLIC_KEY set. All route registration moved below that call, the key is now imported as a real key object via new importVerificationKey/importSigningKey helpers in @openfoundry/auth-tokens (raw PEM bytes were being read as an HMAC secret), the default AUTH_ISSUER is corrected to openfoundry-multipass, and empty env vars are treated as unset rather than as "check nothing".
  • Every caller gained credentials in the same change: new scripts/lib/auth.sh (of_curl) and scripts/lib/auth.ts (authHeaders) source a bearer token from OPENFOUNDRY_TOKEN or a client-credentials exchange, and the demo seeds, Kelava sync scripts and integration tests were converted to use them; the console installs one window.fetch interceptor (apps/app-console/src/lib/authFetch.ts) from main.tsx, with session restore/expiry decisions extracted into apps/app-console/src/lib/session.ts.
  • svc-multipass now seeds its dev OAuth clients only outside NODE_ENV=production, sources token roles from the authenticated client/user and mints none from the unauthenticated /authorize or refresh grants; the proxy drops the stale Content-Length, and .env.example, deploy/, AGENTS.md and GUIDE.md were updated to match. New tests cover gateway auth and header stripping, console auth fetch and session restore, plus expanded auth-tokens, sdk-oauth and multipass suites.

Risk Assessment

⚠️ Medium: A large, security-sensitive change that finally enforces gateway JWT validation, but it is credential-first as the captain required, leaves the demo default (empty AUTH_PUBLIC_KEY) behaving exactly as before, and is backed by behavioural tests on every new path; the only outstanding item is one test assertion that cannot fail.

Testing

I ran the targeted unit suites for every package this change touches (svc-gateway, app-console, auth-tokens, svc-multipass, sdk-oauth) and they are green, then stood up a real three-service stack with AUTH_PUBLIC_KEY configured and exercised it as an end user would: unauthenticated and tampered-token calls to /api/v2/ontologies are rejected 401, /metrics is guarded while /status/health stays open, a client_credentials token unlocks the same call at 200, and the shipped scripts/lib/auth.sh helper authenticates a shell caller unchanged. To prove the regression rather than just the current state, I temporarily restored the pre-fix encapsulated app.register(authPlugin) line on a second gateway and confirmed it still served the same request 200, then restored the file and verified the worktree is clean. On the UI side I drove the console in a real browser against that enforcing gateway and captured screenshots of the login screen, the authenticated Ontology Explorer rendering live data, and confirmed the authFetch interceptor attaches the bearer token; a half-written session (token without user) now clears and lands on the login screen with no dead token left behind. No failures and nothing actionable found.

Evidence: Gateway auth end-to-end CLI transcript (401 unauthenticated, 200 with token, pre-fix 200 control, console + session checks)
\### 1. Unauthenticated request to a protected API (AUTH_PUBLIC_KEY set)
$ curl -s -o /dev/null -w '%{http_code}' http://localhost:18080/api/v2/ontologies
{"errorCode":"CUSTOM_CLIENT","errorName":"MissingAuthToken","errorInstanceId":"cb7831bc-e389-4a13-841c-caad9b951b3b","parameters":{},"statusCode":401}
HTTP 401

\### 2. Unauthenticated /metrics (pinned as guarded)
HTTP 401

\### 3. Health endpoints stay open (SKIP_AUTH_PREFIXES)
/status/health -> HTTP 200

\### 4. Garbage / tampered bearer token
{"errorCode":"CUSTOM_CLIENT","errorName":"InvalidAuthToken","errorInstanceId":"4881cabd-6747-4083-82aa-f04e7a0b1e2c","parameters":{},"statusCode":401}
HTTP 401

\### 5. Obtain real credentials via OAuth2 client_credentials (through the gateway)
$ curl -s http://localhost:18080/multipass/api/oauth2/token -d grant_type=client_credentials -d client_id=admin -d client_secret=***
access_token (first 40 chars): eyJhbGciOiJFUzI1NiIsInR5cCI6IkpXVCJ9.eyJ...

\### 6. Same protected API, now with the token
{"data":[{"rid":"ri.ontology.main.ontology.8a6f21ac-60ed-4a80-a3a0-468daa618171","apiName":"sample-ontology","displayName":"Sample Ontology","description":"A sample ontology for development","version":1},{"rid":"ri.ontology.main.ontology.cd3c1d37-43fb-4610-9af6-defa878ac7cd","apiName":"test-kelava-sync","displayName":"Test Kelava Sync","description":"Integration test ontology","version":1}]}
HTTP 200

\### 7. The shipped shell credential helper (scripts/lib/auth.sh) against the same gateway
{"data":[{"rid":"ri.ontology.main.ontology.8a6f21ac-60ed-4a80-a3a0-468daa618171","apiName":"sample-ontology","displayName":"Sample Ontology","description":"A sample ontology for development","version":1},{"rid":"ri.ontology.main.ontology.cd3c1d37-43fb-4610-9af6-defa878ac7cd","apiName":"test-kelava-sync","displayName":"Test Kelava Sync","description":"Integration test ontology","version":1}]}
HTTP 200

\### 8. Before/after control: same binary with the pre-fix `app.register(authPlugin)` line restored
pre-fix  gateway :18090  unauthenticated GET /api/v2/ontologies -> HTTP 200  (served unauthenticated, the reported bug)
post-fix gateway :18080  unauthenticated GET /api/v2/ontologies -> HTTP 401 MissingAuthToken

\### 9. Console (real browser, Vite dev server on :3000, VITE_API_URL -> enforcing gateway)
- unauthenticated visit to /ontology  -> redirected to /login (screenshot console-01-login-unauthenticated.png)
- dev login admin/admin123           -> POST /multipass/api/auth/login 200, token + user stored
- /ontology after login              -> GET /api/v2/ontologies 200, "Sample Ontology" / "Test Kelava Sync" rendered
                                        (screenshot console-03-ontology-authenticated.png)
- request headers on that call, attached by apps/app-console/src/lib/authFetch.ts:
    authorization: Bearer eyJhbGciOiJFUzI1NiIsInR5cCI6IkpXVCJ9...  (iss=openfoundry-multipass, aud=openfoundry-api, roles=ADMIN)

\### 10. Half-written session (token present, user key missing), then reload
- /ontology -> redirected to /login AND localStorage openfoundry_token == null
  (no dead token left for the interceptor to attach)
  • Evidence: Console login screen when unauthenticated against the enforcing gateway (local file: ~/.no-mistakes/evidence/01M275S8KTMCZ7HMFYQSCB4R2C/console-01-login-unauthenticated.png)
  • Evidence: Console Ontology Explorer rendering live data after login, requests carrying the bearer token (local file: ~/.no-mistakes/evidence/01M275S8KTMCZ7HMFYQSCB4R2C/console-03-ontology-authenticated.png)
Evidence: Before/after control
pre-fix gateway :18090 unauthenticated GET /api/v2/ontologies -> HTTP 200 (bug reproduced)
post-fix gateway :18080 unauthenticated GET /api/v2/ontologies -> HTTP 401 MissingAuthToken
post-fix gateway :18080 with client_credentials token -> HTTP 200 + ontology data

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • 🚨 services/svc-multipass/src/store/client-store.ts:49 - The seeded confidential dev clients now carry platform roles (admin/admin123 -> ADMIN, developer/dev123 -> EDITOR) and ClientStore's constructor seeds them unconditionally, unlike DEV_USERS in services/svc-multipass/src/routes/auth.ts:49 which is emptied when NODE_ENV=production. Concrete path: deploy/docker-compose.yml runs NODE_ENV=production with ENFORCE_PERMISSIONS=true; once AUTH_PUBLIC_KEY is set (which this change's docs now present as the working posture), anyone who can reach the public gateway can POST /multipass/api/oauth2/token with grant_type=client_credentials and the repo-published admin/admin123, receive a token whose signed roles claim is ADMIN, and the gateway proxy (services/svc-gateway/src/proxy.ts) re-asserts X-User-Roles: ADMIN on all 118 guarded downstream routes. Before this change the roles claim did not exist and no identity header was ever set, so these credentials conveyed nothing. Smallest remedy is to gate the roles (or the confidential dev clients themselves) on NODE_ENV the same way DEV_USERS is - but that changes which credentials the captain's scripts can use in a production deployment, so it needs a decision rather than a silent fix.
  • ⚠️ services/svc-multipass/src/routes/oauth.ts:185 - The authorization_code flow stores client.roles on the auth code, so a user token's roles come from the OAuth client rather than the authenticated user. The console's production login path (loginWithOAuth, PKCE against client openfoundry-console/openfoundry-dev, which declare no roles) therefore yields a token with no roles claim; the proxy then sends neither X-User-Id nor X-User-Roles, and with ENFORCE_PERMISSIONS=true every guarded route answers 403 for a legitimately logged-in operator. The dev password login (services/svc-multipass/src/routes/auth.ts:133) does carry user roles, which is why this is invisible locally. It also contradicts the new deploy/.env.production:57 comment that true now "enforces the roles a token actually carries". The remedy - sourcing roles from the approved user record at authorize time - adds state beyond this change's stated intent, so it needs authorization.
  • ⚠️ tests/integration/api-endpoints.test.ts:11 - Only tests/integration/kelava-sync.test.ts was routed through authHeaders. tests/integration/api-endpoints.test.ts, tests/integration/sdk-compat.test.ts and tests/integration/tenant-isolation.test.ts still call the gateway at localhost:8080 with a bare fetch and no Authorization header, so they 401 wholesale the moment AUTH_PUBLIC_KEY is configured - the same gap the intent required closing before enforcement is switched on. Mechanical fix: give each suite the apiFetch wrapper kelava-sync.test.ts now uses.
  • ⚠️ apps/app-console/src/lib/authFetch.ts:62 - isApiRequest tests url.startsWith(API_BASE_URL). API_BASE_URL is import.meta.env.VITE_API_URL ?? "http://localhost:8080", and ?? does not replace an empty string - a same-origin build that sets VITE_API_URL= makes API_BASE_URL "", every URL matches the prefix, and the interceptor attaches the bearer token to every outbound fetch including cross-origin ones (the existing "leaves third-party requests alone" test would flip to sending the token). Guard the absolute-URL branch on a non-empty API_BASE_URL.
  • ℹ️ services/svc-multipass/src/store/token-store.ts:12 - Roles are snapshotted onto the stored refresh token at issue time and replayed verbatim on refresh, so a role revocation does not take effect until the refresh token expires or is revoked. Acceptable for the in-memory demo store; noting the tradeoff rather than asking for a change.

🔧 Fix: gate dev OAuth clients on NODE_ENV; source token roles from user
6 issues (1 error, 3 warnings, 2 infos) still open:

  • 🚨 services/svc-multipass/src/routes/oauth.ts:178 - GET /multipass/api/oauth2/authorize authenticates nobody: it auto-registers any unknown client_id (line 138), accepts any redirect_uri, and auto-approves as the hardcoded user "admin". The fix round now sources the auth code's roles from that hardcoded user (devUserRoles("admin") = ["ADMIN"]) instead of the client. Concrete path on any deployment where NODE_ENV is not "production" (start.sh, docker-compose.dev, any staging box) but AUTH_PUBLIC_KEY is set: an anonymous caller hits GET /multipass/api/oauth2/authorize?response_type=code&client_id=anything&redirect_uri=http://attacker/cb (gateway-exempt via SKIP_AUTH_PREFIXES), omits code_challenge so no PKCE verifier is required, POSTs the code to /oauth2/token, and receives a signed token whose roles claim is ADMIN. proxy.ts then asserts x-user-id/x-user-roles: ADMIN downstream, so the caller has full write authority on every guarded route. Before this commit that same token carried client.roles (undefined for an auto-registered client), so it conveyed no roles. This defeats the authentication the change exists to add. The remedy is a product decision - either stop minting roles from a flow that authenticates no one (only the password login at routes/auth.ts:130 actually verifies a user), or gate the auto-register/auto-approve behaviour - so it needs authorization rather than an automatic patch.
  • ⚠️ apps/app-console/src/lib/authFetch.ts:63 - isApiRequest treats any URL that merely starts with API_BASE_URL as the gateway. With API_BASE_URL="https://api.example.com", a fetch to "https://api.example.com.attacker.test/x" matches the prefix (path becomes ".attacker.test/x", which no CREDENTIAL_ENDPOINTS prefix excludes) and the bearer token is attached and sent to the attacker's host. Compare origins, or require the character after the prefix to be "/" or "?" (or the URL to end there), instead of a bare startsWith.
  • ⚠️ apps/app-console/src/lib/authFetch.ts:30 - The interceptor reads the raw token out of localStorage and never consults TokenManager; nothing outside AuthContext.tsx calls tokenManager.getAccessToken (grep across apps/app-console/src returns no other caller), so the refresh_token grant is never exercised. Concrete sequence: an operator logs in, works for longer than TOKEN_EXPIRY_SECONDS (3600), and from then on every page's fetch carries an expired token - with AUTH_PUBLIC_KEY set the gateway answers 401 on every screen and the console shows the user as still logged in with no recovery short of a manual re-login. The intent requires that enabling enforcement leaves every legitimate caller working; the remedy (routing the interceptor through TokenManager and handling a 401 retry) adds machinery beyond what the change currently does, so it needs authorization.
  • ⚠️ deploy/.env.production:60 - deploy/docker-compose.yml sets NODE_ENV=production for every service, and the fix round made ClientStore seed nothing there while DEV_USERS is already empty there. In that configuration no flow mints a token with a roles claim: client_credentials has no registered client (no HTTP registration endpoint exists), /authorize resolves devUserRoles("admin") to undefined, and the signup route (routes/auth.ts:215) builds its token with no roles field. proxy.ts therefore forwards no identity headers, and with ENFORCE_PERMISSIONS=true every guarded downstream route 403s - including the token the bootstrap-admin flow issues. The comment this change added at deploy/.env.production:57-59 states that true now "enforces the roles a token actually carries", which reads as working enforcement rather than a deployment that denies everything. Either the comment should say no production path yet mints roles, or a real production role source is needed - both are product calls.
  • ℹ️ scripts/lib/auth.sh:73 - of_resolve_token runs once when the library is sourced and the token is reused for the whole script run. A full Kelava sync that runs longer than TOKEN_EXPIRY_SECONDS (3600 by default) will start receiving 401s partway through with no re-mint, which is the exact caller the intent says must not break mid-run. Noting the limit rather than asking for a refresh loop now; it only bites on long runs against an enforcing gateway.
  • ℹ️ services/svc-gateway/src/middleware/auth.ts:30 - SKIP_AUTH_PREFIXES exempts /status/* but not GET /metrics (registered by metricsPlugin from @openfoundry/health). Once AUTH_PUBLIC_KEY is set, a Prometheus scraper pointed at the gateway gets 401. No scrape config lives in this repo, so nothing in-tree breaks - flagging it so an external scraper is not surprised.

🔧 Fix: stop minting roles from unauthenticated authorize; harden console token handling
5 issues (3 warnings, 2 infos) still open:

  • ⚠️ services/svc-gateway/src/proxy.ts:103 - Forwarding x-user-roles turns role enforcement on even where ENFORCE_PERMISSIONS is false. createHeaderBasedHook (packages/permissions/src/middleware.ts:147-168) short-circuits to 'granted' only when the role set contains the permission; when a role header IS present but lacks it, control falls through every enforcePermissions escape hatch (hasClaims is false downstream, userId is set) and reaches throw permissionDenied(...). Concrete: gateway with AUTH_PUBLIC_KEY set, ENFORCE_PERMISSIONS unset (the demo/deploy default after this change), caller uses the seeded developer client_credentials pair -> token carries roles=EDITOR -> any route guarded by admin:manage returns 403, where before this change the same call succeeded. Same for a console login as analyst (VIEWER) on any write route. That contradicts the intent's requirement that turning authentication on must not break existing callers, and it makes the newly written 'ENFORCE_PERMISSIONS=false' comments in deploy/ inaccurate. Remedy is a product call (gate the header assertion on the enforcement flag, or fix the hook's fall-through), so it needs authorization.
  • ⚠️ services/svc-multipass/src/routes/oauth.ts:344 - handleRefreshTokenGrant validates nothing but the refresh token itself - no client_id, no client_secret - and this change added a roles field to the stored refresh token (line 408) that it now replays into the new access token (line 365). Concrete: a client_credentials exchange with the seeded admin/admin123 pair returns a refresh token; anyone holding only that value (it is not a secret in the client-authentication sense, and scripts/lib/auth.sh keeps it in-process but the grant is open to any caller) can POST grant_type=refresh_token with no client credentials and receive a fresh roles=ADMIN token, repeatedly via rotation. Before this change a refreshed token carried no roles, so the same leak conveyed no authority. Remedy (requiring client authentication on the refresh grant) changes the token API's accepted requests, so it needs authorization rather than a silent fix.
  • ⚠️ apps/app-console/src/context/AuthContext.tsx:115 - The interceptor now goes through TokenManager, but TokenManager.setToken computes #expiresAt = Date.now() + expiresIn*1000 from the relative expiresIn persisted in localStorage. The restore-on-mount path replays a token stored an arbitrary time ago, so a token that expired an hour back is treated as fresh for another full hour: #isExpiringSoon() is false, no refresh is attempted, and every request carries an expired token that the gateway 401s. Concrete: OAuth login at T (expiresIn 3600), reload the console at T+2h -> restore sets expiresAt=T+2h+1h -> getToken() returns the dead access token even though a valid refresh token is present. The claimed fix therefore does not hold on the most common path (page reload). The smallest honest remedy is persisting an absolute expiry alongside the token, i.e. new persisted state, so the remedy - not the defect - is what needs authorization.
  • ℹ️ apps/app-console/src/lib/authFetch.ts:97 - The new boundary check rejects a URL whose remainder after API_BASE_URL does not start with /, ? or #. If VITE_API_URL is configured with a trailing slash ("http://gateway:8080/"), then "http://gateway:8080/api/v2/ontologies" leaves rest="api/v2/ontologies", isApiRequest returns false, and no bearer token is attached - every console request 401s under enforcement with no diagnostic. Normalise by stripping trailing slashes from API_BASE_URL before the comparison.
  • ℹ️ services/svc-multipass/src/store/client-store.ts:72 - The comment added with the production gate says "in production, clients are registered at runtime", but no client-registration endpoint exists; deploy/.env.production (added in the same branch) correctly states that none can be registered. Two comments in the same change describe opposite realities; the client-store one should say that production seeds no client and that the only runtime registration is the implicit one /authorize performs.

🔧 Fix: drop gateway identity assertion and roleless refresh; fix token restore
5 issues (1 error, 3 warnings, 1 info) still open:

  • 🚨 deploy/.env.production:25 - The gateway's default issuer was corrected to "openfoundry-multipass" (services/svc-gateway/src/config.ts:91) because that is the only iss the platform signs (services/svc-multipass/src/routes/oauth.ts:383, routes/auth.ts:127), but the deployment still pins the old value: deploy/.env.production:25 sets AUTH_ISSUER=openfoundry and deploy/docker-compose.yml:191 defaults to the same. Concrete failure: set AUTH_PUBLIC_KEY in that deployment (the whole point of this change), log in, and validateToken rejects the issuer - every authenticated request answers 401 InvalidAuthToken, including the Safe and Care sync callers the intent requires not to break. Fix: set both to openfoundry-multipass.
  • ⚠️ services/svc-multipass/src/routes/oauth.ts:126 - Enforcement is now real at the gateway, but the token-minting path it exempts is unauthenticated: GET /multipass/api/oauth2/authorize accepts any client_id (auto-registering it as a public client, line 138-150), accepts any redirect_uri, and approves the hardcoded user "admin" (line 173). handleAuthorizationCodeGrant then skips the secret check for public clients (line 277). So any caller who can reach the gateway does authorize -> token and obtains a signed token that passes the new middleware; with AUTH_PUBLIC_KEY set the gateway rejects anonymous requests but nothing stops an anonymous caller from minting credentials first. This is pre-existing code, but the change makes it the single load-bearing hole in the feature it claims to deliver. The remedy - authenticating /authorize (a real login/consent step) - extends this PR's scope, so it needs your decision, not an automatic patch; at minimum the claim that the gateway now enforces authentication should be qualified in AGENTS.md and deploy notes.
  • ⚠️ apps/app-console/src/context/AuthContext.tsx:126 - The restore fix is undone in storage by its own persistence hook. tokenManager.setToken(token, expiresAt) stores the correct absolute deadline in memory, but setToken also calls onTokenChange, and the handler at line 103-108 re-persists expiresAt = Date.now() + token.expiresIn * 1000 - a brand-new full hour. Concrete: OAuth login at T (expiresIn 3600); reload at T+2h -> in-memory expiry is correct and a refresh is attempted, but storage now claims T+3h. If that refresh fails (the dev password login mints no refresh token, so getToken throws and authFetch falls back to the stored value), or the user simply reloads again at T+2.5h, the restored session is treated as fresh and the dead access token is sent - exactly the bug this round was meant to close. Fix within the existing stored state: have the restore path re-write the entry with the original expiresAt (or suppress the persist for that setToken call).
  • ⚠️ services/svc-multipass/src/store/client-store.ts:76 - The new if (process.env.NODE_ENV === "production") return; leaves the ClientStore empty in production, and deploy/docker-compose.yml sets NODE_ENV: production for every service. There is no client-registration endpoint, so handleClientCredentialsGrant always throws "unknown client" there. The credential path this PR builds for scripts and tests (OPENFOUNDRY_CLIENT_ID/_CLIENT_SECRET exchanged via client_credentials, scripts/lib/auth.sh and auth.ts) therefore cannot work against the production deployment at all: mintAccessToken logs a warning and returns null, and the caller proceeds unauthenticated into 401s. Only a pre-minted OPENFOUNDRY_TOKEN works, and nothing in the deployment can mint one for a service. That interacts directly with the requirement that credentials be ready before enforcement is switched on. The remedy (a registration path or a seeded production client from env) adds machinery beyond this change, so it is your call - at minimum .env.example and the deploy notes should state that only OPENFOUNDRY_TOKEN is usable in production.
  • ℹ️ packages/auth-tokens/src/claims.ts:48 - Two comments still describe the behavior the final commit deleted. claims.ts:46-49 says "The downstream X-User-Roles header is set by the gateway from this claim and from nothing else - see services/svc-gateway/src/proxy.ts", and services/svc-multipass/src/store/client-store.ts:17-21 says "The gateway turns them into the downstream X-User-Roles header". proxy.ts now sets no identity headers at all. Reword both to match oauth.ts:386 ("minted, but with no consumer yet").

🔧 Fix: fix deploy issuer drift, persist absolute token expiry
2 issues (1 warning, 1 info) still open:

  • ⚠️ deploy/docker-compose.yml:237 - Turning on the feature this PR delivers is impossible in the shipped deployment. The gateway verifies with AUTH_PUBLIC_KEY (deploy/docker-compose.yml:190, deploy/.env.production:24), but the svc-multipass service block passes no JWT_PRIVATE_KEY/JWT_PUBLIC_KEY and deploy/.env.production defines neither, so services/svc-multipass/src/routes/oauth.ts:84 and routes/auth.ts:90 each fall back to generating their OWN ephemeral ES256 pair at startup. Concrete failure: operator follows the new .env.example guidance, sets AUTH_PUBLIC_KEY in deploy/.env.production, restarts - every token multipass mints is signed by a key nobody configured, so validateToken fails and every authenticated request answers 401 InvalidAuthToken (and password-login tokens and OAuth tokens are signed by two different ephemeral keys even then). Fix in the files this change already edits: add JWT_PRIVATE_KEY/JWT_PUBLIC_KEY to deploy/.env.production and pass them through to svc-multipass in docker-compose.yml, alongside AUTH_PUBLIC_KEY.
  • ℹ️ services/svc-multipass/tests/multipass.test.ts:216 - The comment sweep for stale identity claims missed the test file added in the same round. multipass.test.ts:216 says "the gateway turns a token's roles claim into X-User-Roles for all guarded routes" and the docblock at ~line 404 says the gateway "forwards no X-User-Id/X-User-Roles for a token with no roles claim", implying it forwards them when roles are present. services/svc-gateway/src/proxy.ts sets no identity headers at all in any case. Reword both to match the wording now used in claims.ts and oauth.ts (minted, no consumer yet).

🔧 Fix: plumb JWT signing keys into deploy; fix stale test comments
3 issues (2 warnings, 1 info) still open:

  • ⚠️ services/svc-multipass/src/routes/oauth.ts:80 - The newly documented way to supply the signing keys cannot actually work. The gateway's new importVerificationKey (packages/auth-tokens/src/token-validator.ts:74) normalises literal \n escapes before importSPKI, but svc-multipass imports the same env-sourced PEMs raw: routes/oauth.ts:80-81 importPKCS8(config.jwtPrivateKey, "ES256") / importSPKI(config.jwtPublicKey) and routes/auth.ts:88. Concrete failure: an operator follows the new deploy/.env.production:23-32 and .env.example:27-33 instructions, and since a docker-compose .env file cannot carry a multi-line value, puts the PEM in with \n escapes. The gateway imports its key fine; svc-multipass throws inside authRoutes/oauthRoutes at startup and the service never comes up, so every login and token request fails. If instead the operator sets only one of the two vars, the config.jwtPrivateKey && config.jwtPublicKey guard silently falls through to an ephemeral key pair and every token is then rejected by the gateway with 401 InvalidAuthToken - the exact failure the previous fix round set out to close. Smallest honest fix: run the multipass-side PEMs through the same normalisation (reuse the helper, or .replace(/\\n/g, "\n").trim()), and warn when exactly one of the pair is set instead of silently generating ephemeral keys.
  • ⚠️ scripts/lib/auth.ts:99 - scripts/lib/auth.ts:99 uses process.env.OPENFOUNDRY_AUTH_URL ?? baseUrl, which only falls back when the variable is absent, not when it is empty. .env.example:49-50 ships OPENFOUNDRY_AUTH_URL= empty, and scripts/sync-kelava.sh:24 / sync-kelava-incremental.sh:21 do set -a && source "$ROOT_DIR/.env", exporting it as an empty string to any tsx caller in that shell. Trace: OPENFOUNDRY_CLIENT_ID/_SECRET set, OPENFOUNDRY_AUTH_URL="" -> authUrl "" -> fetch("/multipass/api/oauth2/token") is not a valid absolute URL in Node, throws, is caught at line 113, logs "could not reach" and returns null -> the caller proceeds unauthenticated and gets 401s from an enforcing gateway, with the cause reported as unreachability rather than misconfiguration. The shell twin gets this right (${OPENFOUNDRY_AUTH_URL:-...} in scripts/lib/auth.sh:69), so the two are out of step. Use || baseUrl.
  • ℹ️ packages/auth-tokens/src/claims.ts:64 - parseRoles is added and exported from packages/auth-tokens (claims.ts:64, index.ts:6) but has no caller or test anywhere in the repo - the roles claim is deliberately left with no consumer until the permission-model work lands. Noting it so it is understood as a placed-ahead API rather than live code; no action needed unless you prefer it to arrive with its consumer.

🔧 Fix: share PEM normalisation, fix empty env fallback, drop parseRoles
2 warnings still open:

  • ⚠️ services/svc-gateway/src/server.ts:160 - Per-tenant rate limiting can never engage, and this change is what makes that visible. rateLimitPlugin is called on the root instance at server.ts:93, authPlugin at server.ts:160, and Fastify runs same-context onRequest hooks in registration order - so the rate-limit hook always runs before the auth hook. extractKey (services/svc-gateway/src/rate-limiter.ts:275) reads request.orgRid first, which is still undefined at that moment, and falls through to the Authorization-header hash. Concrete trace: a request with a valid token carrying org: "org-1" is bucketed as auth:<hash-of-that-token>, not org:org-1; an hour later the console refreshes its token and the same tenant gets a brand-new empty bucket, so the documented "per-tenant isolation" never holds and the limit is per-token instead. Before this change orgRid was never populated at all, so the dead branch was invisible; now middleware/auth.ts:130 comments that it extracts the org claim "for RLS and per-tenant rate limiting", which is not what happens. The remedy is a decision, not a mechanical edit: moving the authPlugin call above rateLimitPlugin makes the org bucket work but also means unauthenticated floods are rejected before they are counted, which is a deliberate trade-off about whether the 401 path stays rate-limited. Either reorder deliberately, or correct the claim in the auth comment and in extractKey's documented precedence.
  • ⚠️ services/svc-gateway/src/config.ts:94 - An empty AUTH_ISSUER or AUTH_AUDIENCE silently turns the corresponding check off instead of falling back to the default. env() (config.ts:80) uses process.env[key] ?? fallback, so an exported empty string survives as "", and middleware/auth.ts:103-104 then maps it through config.authIssuer || undefined - which jose reads as "do not verify this claim". Concrete path created by this change: the diff deletes the AUTH_ISSUER=openfoundry line from deploy/.env.production and adds a comment saying it is "deliberately not set". An operator following that comment by blanking the value they already have (AUTH_ISSUER=) rather than deleting the line gets a gateway that accepts any iss and any aud, with no warning and every test still green - the opposite of the intent of pinning the issuer to what multipass signs. This is the same empty-string-environment class as the OPENFOUNDRY_AUTH_URL fix in this branch, but here empty has a security meaning. Two honest readings exist - treat empty as absent (use the default, so the check always runs) or keep empty as an intentional escape hatch and document it - so which one is wanted is the author's call. Note that apps/app-console/src/config.ts was correctly left on ??: there an empty VITE_API_URL is a supported same-origin build that lib/authFetch.ts explicitly handles.

🔧 Fix: treat empty env vars as unset in gateway config
3 issues (1 warning, 2 infos) still open:

  • ⚠️ apps/app-console/src/context/AuthContext.tsx:129 - Once the gateway enforces, an expired console session leaves the UI authenticated but every request 401ing, with no path back to the login screen. Concrete trace: log in with the dev password flow (login() at AuthContext.tsx:167 calls setToken with a response that carries no refreshToken), leave the tab for over TOKEN_EXPIRY_SECONDS (3600), reload. The restore path at line 129 calls setToken(token, expiresAt) with a past deadline and unconditionally setCurrentUser(user) - it never checks the deadline it just restored - so isAuthenticated is true. authFetch's currentToken() then asks the manager, #isExpiringSoon() is true, #refresh() throws for want of a refresh token, and the catch falls back to readStoredToken(), which returns the dead token. Every page therefore renders useApi's "Authentication required (HTTP 401)" banner while the navbar still shows the user logged in, and nothing clears the session or redirects. This is reachable only because this change makes the gateway enforce, so the branch creates it. Flagged as ask-user because the remedy is new behavior rather than a correction: either treat a restored-and-expired session as logged out, or clear the session on a 401 from the interceptor - both are product decisions about what the user sees.
  • ℹ️ services/svc-multipass/src/config.ts:39 - The env() docstring added by the last fix round was copied verbatim from services/svc-gateway/src/config.ts and justifies itself with AUTH_ISSUER, which svc-multipass never reads - this env() serves JWT_PRIVATE_KEY, JWT_PUBLIC_KEY, DATABASE_URL and the expiry values. The rule it states is right; the example points at the wrong service. Rewrite the example in terms of a key this config actually reads (e.g. JWT_PRIVATE_KEY, where empty means "generate an ephemeral pair").
  • ℹ️ services/svc-gateway/src/middleware/auth.ts:29 - SKIP_AUTH_PREFIXES exempts /status/, /multipass/api/oauth2/ and the login endpoint, but not /metrics (registered via metricsPlugin at server.ts:150). With AUTH_PUBLIC_KEY set, a Prometheus scrape of the URL scripts/start-production.sh:220 advertises now gets 401. Nothing in deploy/ scrapes it today and the container healthcheck uses /status/health, so nothing breaks - noting it as a consequence to expect when a scraper is added, not a defect. Requiring a token on /metrics is a defensible posture.

🔧 Fix: clear unrecoverable expired console session on restore
1 info still open:

  • ℹ️ apps/app-console/src/context/AuthContext.tsx:124 - The last fix round was asked for a provider round-trip test (persist an expired, refresh-less session, mount the provider, assert not authenticated and storage cleared). What landed is apps/app-console/src/lib/session.test.ts, a pure unit test of the isRestorable predicate only. The predicate is correct, but the wiring that actually fixes the defect - the if (isRestorable(stored)) branch at AuthContext.tsx:124 calling clearStoredSession() instead of setCurrentUser(user) - has no test. Inverting that condition, or dropping the else arm, would leave the console authenticated-but-401ing again and every test still passes. Reported as ask-user because the remedy, not the defect, needs authorization: apps/app-console declares no jsdom, no @testing-library/react and no vitest.config.ts (it falls back to vite.config.ts, which has no test block), so a mount test means adding test dependencies and an environment - new machinery beyond this change's scope. The alternative is to accept the predicate-level coverage and say so.

🔧 Fix: extract and test console session restore decision
1 info still open:

  • ℹ️ apps/app-console/src/lib/session.ts:59 - When exactly one of the two persisted keys is present, restoreSession returns null at line 59 without clearing storage, so the unusable half-session survives the reload that found it - contrary to the function's own docstring. Concrete path: openfoundry_user is removed (evicted, cleared by a user, or written by a build that only set the token key) while openfoundry_token holds an hour-old access token. The provider shows the login screen, but readStoredToken() (apps/app-console/src/lib/authFetch.ts:65) still finds the token key and the interceptor attaches Authorization: Bearer <dead token> to every API request, which the gateway now rejects with 401 instead of the request simply going unauthenticated. Remedy is mechanical and stays inside this change: call clearStoredSession() on that !storedToken || !storedUser path when either key is present, and cover it with the same stubbed-storage test style already in session.test.ts.

  • ⚠️ apps/app-console/src/lib/session.ts:59 - restoreSession returns null at the if (!storedToken || !storedUser) return null; guard without clearing storage, contradicting its own docstring ("in which case the stored keys are cleared"). Concrete path: openfoundry_user is missing (evicted, cleared, or never written after a failed login) while openfoundry_token still holds an hour-old access token. The provider shows the login screen, but readStoredToken() (apps/app-console/src/lib/authFetch.ts:65) still finds the token key, so the interceptor attaches Authorization: Bearer <dead token> to every API request and the now-enforcing gateway answers 401 instead of the request simply going out unauthenticated. Remedy is mechanical and inside the change's scope: call clearStoredSession() on that path when either key is present, and cover it with the same stubbed-storage style already in session.test.ts.

  • ℹ️ services/svc-gateway/src/proxy.ts:91 - The header loop now has two consecutive skips for the same header: the pre-existing if (lowerKey === "content-length") continue; and the new if (BODY_DESCRIBING_HEADERS.has(lowerKey)) continue;. The second is unreachable for its only member, so the documented set is dead. Keep one: either drop the literal check and let the set own the rule (matching the new docstring), or drop the set.

🔧 Fix: clear orphaned half-sessions; drop duplicate content-length skip
2 issues (1 warning, 1 info) still open:

  • ⚠️ services/svc-gateway/src/server.ts:150 - await app.register(metricsPlugin) runs at line 150, before authPlugin(app, { config }) at line 160. Fastify snapshots the parent's hooks when it creates an encapsulated child context, and metricsPlugin is a plain async plugin (not wrapped in fastify-plugin), so the GET /metrics route it registers is created before the auth onRequest hook exists on the root instance and never inherits it. Concrete path: with AUTH_PUBLIC_KEY set, GET /metrics with no Authorization header returns 200 Prometheus text, while GET /metrics/:service (registered by healthRoutes after line 160) correctly returns 401. This directly contradicts the note this change added at AGENTS.md:122 ("SKIP_AUTH_PREFIXES does not exempt /metrics, so with AUTH_PUBLIC_KEY set a Prometheus scrape gets 401 - a deliberate posture"), so the one route escaping the new enforcement is the one documented as covered. The same encapsulation applies to the plugin's own request-timing hooks, which therefore record nothing for any real route - the same bug class this change exists to fix. Remedy is a posture decision, not a mechanical correction: either move the metricsPlugin registration below authPlugin (scrapes then need a token, as AGENTS claims) or add /metrics to SKIP_AUTH_PREFIXES and correct the AGENTS note; separately, making the metrics hooks actually cover routes needs fastify-plugin, which is beyond this change's scope.
  • ℹ️ deploy/.env.production:74 - ENFORCE_PERMISSIONS goes true -> false in deploy/.env.production and deploy/docker-compose.yml. The reasoning checks out against packages/permissions/src/middleware.ts: with the gateway deliberately asserting no X-User-* headers, true makes every guarded route 403 rather than enforce anything, so the previous value was deny-all, not protection. Noting the posture change explicitly: production guarded routes are now open to any authenticated caller, and the change documents that in three places (deploy files, GUIDE.md, AGENTS.md) as pending the permission-model unification. No action needed unless the captain wants deny-all retained until that work lands.

🔧 Fix: register gateway routes after auth guard; pin /metrics 401
1 info still open:

  • ℹ️ services/svc-gateway/tests/auth.test.ts:195 - The new "guards /metrics like every other route" test asserts only that an unauthenticated GET /metrics is 401. Fastify's 404 context inherits root onRequest hooks, so any unrouted path returns 401 too - the assertion holds whether or not /metrics exists, and would still pass if metricsPlugin were removed or its route renamed. It also cannot fail in the ordering it was written to pin, since _addHook recurses into kChildren and a root hook reaches earlier-registered plugins either way. Add the authenticated half: GET /metrics with a valid signed token must return 200 with Prometheus text.
✅ **Test** - passed

✅ No issues found.

  • npx turbo run test --filter=@openfoundry/svc-gateway --filter=@openfoundry/svc-multipass --filter=@openfoundry/auth-tokens --filter=@openfoundry/sdk-oauth --filter=@openfoundry/app-console (all green)
  • Manual E2E: booted svc-gateway/svc-multipass/svc-ontology/svc-objects/svc-actions with a shared ES256 keypair and AUTH_PUBLIC_KEY set; probed GET /api/v2/ontologies unauthenticated (401), with a garbage bearer (401), with client-asserted x-user-id/x-user-roles (401), with a real client_credentials token (200), and GET /status/health (200)
  • Before/after regression: temporarily restored the pre-fix await app.register(authPlugin, { config }) on the same running stack - unauthenticated request served HTTP 200; restored the fix - same request 401
  • bash scripts/demo/seed-pest-control.sh against the enforcing gateway with OPENFOUNDRY_CLIENT_ID/_SECRET set (full dataset seeded via of_curl)
  • npx tsx scripts/seed-dev-data.ts with and without credentials (401 without; 65 authenticated ontology/object writes with)
  • npx vitest run --config tests/vitest.config.ts tests/integration/kelava-sync.test.ts against the enforcing gateway, with credentials (9/9 pass) and without (fails with an explicit 401 credential message rather than skipping)
  • Browser (Playwright, 1440x900): dev login through the enforcing gateway, ontology data rendered, then backdated the persisted expiresAt on a refresh-token-less session and reloaded - landed on the login screen with both localStorage keys cleared

✅ No issues found.

  • npx turbo run test --filter=@openfoundry/svc-gateway --filter=@openfoundry/auth-tokens --filter=@openfoundry/svc-multipass --filter=@openfoundry/sdk-oauth --filter=@openfoundry/app-console (16/16 tasks, gateway 55 tests, console 51 tests)
  • Live stack: svc-ontology :18081, svc-multipass :18084, svc-gateway :18080 booted with a generated ES256 key pair in AUTH_PUBLIC_KEY / JWT_PRIVATE_KEY
  • curl $G/api/v2/ontologies unauthenticated -> 401 MissingAuthToken; -H 'Authorization: Bearer not.a.token' -> 401 InvalidAuthToken
  • curl $G/metrics -> 401 and curl $G/status/health -> 200 (guarded vs skipped prefixes)
  • OAuth2 client_credentials token exchange through the gateway, then curl -H "Authorization: Bearer $TOKEN" $G/api/v2/ontologies -> 200 with ontology data
  • source scripts/lib/auth.sh && of_curl $G/api/v2/ontologies with OPENFOUNDRY_CLIENT_ID/_SECRET -> 200 (shell caller credential path)
  • Before/after control: temporarily restored the pre-fix app.register(authPlugin, { config }) line, booted a second gateway on :18090 -> unauthenticated GET /api/v2/ontologies returned 200 (bug reproduced); file restored, worktree clean
  • Browser (chrome-devtools-axi) against console on :3000 with VITE_API_URL pointed at the enforcing gateway: unauthenticated /ontology -> /login; dev login admin/admin123 -> 200; /ontology renders live ontologies; inspected request headers showing the authFetch-attached Bearer token
  • Browser: removed the openfoundry_user key leaving openfoundry_token behind, reloaded /ontology -> redirected to /login and localStorage token cleared
🔧 **Document** - 1 issue found → auto-fixed ✅
  • ℹ️ AGENTS.md:42 - AGENTS.md records pnpm run test as green at 97/97, but this change adds four new test files (apps/app-console/src/lib/authFetch.test.ts, apps/app-console/src/lib/session.test.ts, services/svc-gateway/tests/auth.test.ts, services/svc-gateway/tests/trusted-identity.test.ts) and extends three existing suites, so that per-package task count is likely stale. I could not correct it because this phase is forbidden from running tests; the validation phase should re-run pnpm run test and update the number in AGENTS.md to the observed value.

🔧 Fix: verify AGENTS.md task and lint counts; no changes
✅ Re-checked - no issues remain.

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

Michael Santoso added 14 commits September 11, 2026 11:30
…edentials

The gateway never validated a JWT. `server.ts` called `app.register(authPlugin)`,
which encapsulates the plugin, so its `onRequest` hook applied only to routes
registered inside that child context - and there were none. healthRoutes,
v2Routes, v1Routes and multipassRoutes are siblings on the parent, so the hook
never ran: with AUTH_PUBLIC_KEY configured, GET /api/v2/ontologies still
answered 200 with real data, `request.claims` and `request.orgRid` stayed
empty, and nothing downstream could enforce anything.

Credentials land before enforcement does, so no commit leaves a caller broken.

Callers first:
- scripts/lib/auth.sh (`of_curl`) and scripts/lib/auth.ts (`authHeaders`)
  resolve a bearer token from OPENFOUNDRY_TOKEN, or from
  OPENFOUNDRY_CLIENT_ID/_CLIENT_SECRET via the client_credentials grant, read
  from the environment or the gitignored .env. Nothing is baked into the repo.
  With neither set they send no header, which is exactly today's behaviour.
- Wired into the Kelava sync scripts, every seed script, seed-dev-data.ts,
  seed-kelava-alerts.sh, test-integration.sh, the CI readiness probe and
  tests/integration/kelava-sync.test.ts.
- The console has ~80 bare `fetch` call sites, so one interceptor in
  lib/authFetch.ts attaches the stored token to gateway requests.
- bootstrap-admin.sh keeps no credential on purpose: it runs before one exists,
  against the signup route the middleware already exempts.
- .env.example registers AUTH_PUBLIC_KEY, JWT_PRIVATE_KEY/JWT_PUBLIC_KEY and
  the OPENFOUNDRY_* credential keys, all empty.

Then enforcement:
- Call authPlugin directly, like rateLimitPlugin, with a comment explaining
  why so the next person does not re-register it.
- Import AUTH_PUBLIC_KEY as a key object via the new
  `importVerificationKey`. The PEM's raw bytes were being handed to jose as an
  HMAC secret, which cannot verify an ES256 token, so enforcement alone would
  have rejected every caller. A malformed key now fails at startup.
- Default AUTH_ISSUER to "openfoundry-multipass", the only issuer this platform
  mints. The old "openfoundry" matched nothing.
- Plugin state moved into the closure; module-level state leaked between
  server instances in one process.

Trusted identity for downstream services:
- proxy.ts always drops client-supplied X-User-Id / X-User-Roles and re-asserts
  them from the verified claims, so role-based enforcement finally has an
  identity source a client cannot forge.
- Both headers are set together or not at all: the permission middleware denies
  a request with a user id but no roles, so a half-filled identity would 403
  every downstream route.
- Roles come from a new optional `roles` claim minted by svc-multipass. A token
  without it conveys no roles and downstream behaves exactly as before.
- The proxy also stops forwarding the incoming Content-Length, which no longer
  describes the re-serialised body; a pretty-printed payload was failing with a
  502 (this is why seed-kelava-alerts.sh never actually created its monitors).
- `local` used outside a function aborted sync-kelava-incremental.sh.

Verified against a gateway with AUTH_PUBLIC_KEY set: unauthenticated
/api/v2/ontologies is 401 and the same request with a token is 200; the seed
scripts, seed-dev-data.ts, the alert seeder and all 12 integration suites pass;
the console logs in and every page loads. The Kelava sync was exercised against
a local stand-in for Sanocare's Postgres - only its auth path is proven, not a
sync against the real source.
PR #9 landed the client-header strip and documented, accurately at the time,
that role enforcement still did not work: nothing set x-user-id / x-user-roles
from a verified token, `request.claims` was never populated, and issued tokens
carried no roles claim. This change is the trusted hop those notes were waiting
for, so AGENTS.md, deploy/.env.production and deploy/docker-compose.yml said
the opposite of what the code now does.

ENFORCE_PERMISSIONS=true no longer means "every gateway-proxied route answers
403"; it means the roles the caller's token carries are enforced, given
AUTH_PUBLIC_KEY and a token minted with a roles claim.
@Przyval
Przyval force-pushed the fm/openfoundry-auth-tidak-menegakkan branch from 011a44c to 16140e2 Compare September 11, 2026 04:56
@Przyval Przyval changed the title fix(gateway): enforce JWT auth after giving every caller credentials fix(gateway): enforce JWT auth and give every caller credentials first Sep 11, 2026
sdk-compat's Admin.Users.getCurrent() reads identity out of the bearer
token, so it 401s whenever the caller has none. The suite used to
hardcode admin/admin123; now that every caller resolves credentials from
the environment through scripts/lib/auth.*, CI has to supply them.

Set the seeded dev client_id/secret as job-level env so the readiness
probe, the seed script and the vitest suites all authenticate the same
way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Przyval
Przyval merged commit e376027 into master Sep 11, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant