Appearance
Agent authentication
How an AI client gets into a workspace, and why the pieces are shaped the way they are.
Two credentials, told apart by shape
mcp_ key | OAuth 2.0 access token | |
|---|---|---|
| Issued by | a workspace administrator | the authorization server, after a person consents |
| Names a person | no — it is a workspace credential | yes, in sub |
| Reaches | one workspace | one workspace |
| MCP ceiling | its own level, up to FULL | DATA_ONLY at most |
| View scoping | yes | no |
| Revocation | disable or delete the key | immediate, see below |
McpAccessGuard.authenticate() dispatches on the mcp_ prefix. Anything else is parsed as a JWT.
A session token is not an access token. Both are signed with the same key by the same issuer, so the signature proves nothing about intent. What separates them is the token_type: OAUTH2 claim, which only OAuth2AuthorizationService stamps. The absent workspace_id and the audience check are two further barriers. There is an integration test that presents a real session JWT to MCP and expects refusal.
Why workspace:write stops at DATA_ONLY. A user consenting to a scope called "write" is agreeing to let an application write rows, not to let it drop a column or rewrite a type on a database they own. Schema modification needs an mcp_ key, which an administrator issues deliberately. If that ever needs to be reachable by OAuth, it wants a scope that says so in as many words — workspace:schema — rather than a widening of this one.
The connector flow
A hosted client (Claude, ChatGPT) is handed a URL and nothing else. It cannot hold a key or a client_id in advance, which is what the rest of this is for.
POST /mcp → 401 + WWW-Authenticate: resource_metadata=…
GET /.well-known/oauth-protected-resource/mcp → names the authorization server
GET /.well-known/oauth-authorization-server → lists the endpoints
POST /api/oauth2/register → client_id, no secret (RFC 7591)
browser → /admin/oauth2/consent person approves
POST /api/oauth2/token → access + refresh token
POST /mcp?workspace=<uuid> → toolsThe 401 is produced by the Cloudflare origin proxy, not the backend. MCP is served by a Vert.x route where JAX-RS filters do not run, and the Quarkiverse extension binds its path exactly, so a Java-side filter would have meant a reroute() on the live MCP path. The edge is also where the challenge is cheapest. Consequence worth knowing: a direct request to the origin without a credential still gets a 200 with a JSON-RPC error — the origin is only reachable with X-Origin-Secret, so the edge is the public contract.
Which workspace
Three ways the workspace gets settled, and one check that runs in all three.
A client an administrator registered belongs to one workspace. A client that registered itself belongs to none, and either says which one it wants with the resource parameter (RFC 8707):
resource=https://schemastack.io/mcp?workspace=<uuid>or omits it, in which case the consent screen offers the workspaces the signed-in person can reach and sends the one they pick back as the same resource value.
A supplied resource arrives from whoever drove the user to the consent screen, so it is not trusted — and a value that came from our own picker is not trusted either, because by the time it arrives the two are indistinguishable. McpResource.parse requires this server's exact origin and path, rejects a fragment, requires a single UUID, and refuses ambiguity.
Then resolveWorkspace runs requireWorkspaceAccess for the signed-in user before consent is offered, on every path. This is the most important line in the flow, and it was wrong once: the administrator-registered branch treated "the client carries a workspace" as settling the question of whether this person may consent for it, and returned the workspace unchecked. Anyone with an account and a client ID — not a secret — could approve a connection to a workspace they had nothing to do with and exchange the resulting code for a working token. Both branches check now, and OAuth2ConnectorFlowIntegrationTest.refusesConsentForBoundClientFromOutsider fails if either stops.
The lesson generalises: which workspace and whether this person may act on it are two questions, and answering the first does not answer the second.
The picker must agree with the gate
WorkspaceRepository.findConsentableForUser restates requireWorkspaceAccess as HQL — elevated org roles reach everything including MAINTENANCE, everyone else needs a workspace membership and a workspace not in MAINTENANCE. Two statements of one rule is a liability, so everyOfferedWorkspaceSurvivesTheGate puts every offered workspace back through the real endpoint. Offering more than the gate allows makes a picker whose choices fail; offering less makes a workspace nobody can connect. Organisation membership status is deliberately unfiltered in both, because getOrganisationMembership does not filter it; if that gains a status check, the query needs the same one.
The workspace is fixed at consent, on the authorization code and on the tokens minted from it. Never on the client.
A token names both surfaces of its workspace: the Workspace API and that workspace's MCP resource. Same consent, same scopes, so it is not a widening — it saves an application that uses both from holding two tokens.
Immediate revocation
Access tokens are stateless JWTs and are not stored, so nothing can be done to one already issued. Revoking used to stop only the next token; the one in the client's hands kept working for up to an hour. For a client an administrator registered you could disable the client instead, but a self-registered client belongs to no workspace and cannot be disabled by the workspace whose data it is reading — so there was no immediate stop for exactly the case that needed one.
oauth2_grant_revocation is a watermark table: rows record a moment, and a token whose iat precedes the latest matching row is refused. A null user_email covers every user of that client. Append-only, so there is no upsert to get wrong and no row to contend for.
Deriving this from oauth2_refresh_token does not work, which is worth recording because it is the obvious first idea. Refresh rotation marks the old token revoked at the same instant the new access token is issued, so a rotation would be indistinguishable from a revocation and would invalidate the token it had just minted.
Comparison is in whole seconds because iat is. A token issued in the same second as a revocation is refused: the error always runs towards refusing a token that might be fine, never towards honouring one that is not.
Enforced in two places — GrantRevocation (metadata service) and OAuth2TokenValidationService.isRevoked (workspace-api). Those modules share no code by design, so the rule is restated, and both are unit-tested with the same cases so a change to one that is not mirrored fails a test.
Where the security decisions live
Everything the OAuth2 and MCP paths decide about a credential is a pure function in metadata-service/.../service/oauth2/, separate from the reactive lookups around it. That is so the properties can be asserted directly rather than inferred from an HTTP call that happened to fail: lookalike hosts, a host smuggled before an @, wildcard redirect URIs, scope escalation, a session JWT presented to MCP.
Add new decisions there, with the case that would have caught the mistake.
Routing
Three layers own different paths, which is easy to forget when something 404s:
| Path | Owner |
|---|---|
/mcp, /api/*, /sse/* | origin proxy Worker → Traefik → service |
/.well-known/oauth-* | origin proxy Worker → Traefik (wellknown-* router) → metadata-rest |
/.well-known/api-catalog, /.well-known/auth.md | Cloudflare Pages (static) |
/, /admin/*, /app/*, *.md, llms.txt | Cloudflare Pages _worker.js |
The discovery documents needed a Traefik router of their own — they sit at the root, and Traefik only routed /api, /sse and /mcp. That failure looked like Cloudflare's problem and was Traefik's, and no test could have found it, because tests do not go through Traefik.
Deploy order
Backend → Worker → Pages.
The Worker's routes send /.well-known/oauth-* to the origin, so they 404 until the backend is up; and auth.md on the site documents registration, so it should not describe an endpoint that is not answering yet.