Appearance
Status Page
Public status page at schemastack.io/status showing real-time health of all SchemaStack services.
Data Flow
metadata-rest GET /api/status (checks all services internally)
→ CF Worker cron every 5 min (calls origin, writes to KV)
→ CF Worker GET /status (reads KV, returns JSON)
→ Website Status.tsx (fetches /status, renders)
same cron, separate pipeline:
Tier 1 synthetics (5 live probes) → synthetic:state → Slack on transitionArchitecture
CF Worker (schemastack-deployment)
The existing origin-proxy worker (cloudflare/worker/) handles the status pipeline:
Cron trigger (*/5 * * * *):
- Calls
${ORIGIN_HOST}/api/statuswithX-Origin-Secretheader - On success: writes the response plus the appended history to KV as
status:state - On failure: writes a fallback snapshot (metadata-rest/consumer-worker = down, others = unknown)
- History is trimmed to
MAX_HISTORY = 288entries (24h at 5-minute intervals)
Snapshot and history are one key, written once — they were two keys until Aug 2026, which put the Worker over the free tier once synthetics joined the same cron. See Free Tier Budget.
GET /status endpoint:
- Reads
status:statefrom KV - Returns
{ current, history }withCache-Control: public, max-age=60and CORS headers — the shape the frontend has always consumed, so the key merge was invisible to it
Why the cron calls ORIGIN_HOST, not schemastack.io
A Worker's subrequest to a hostname on its own zone bypasses Worker routes and goes straight to that hostname's origin — for schemastack.io that is Pages, which serves the SPA shell and returns 200 for any path. A check written against the public hostname therefore passes while measuring nothing. Every cron-side request uses env.ORIGIN_HOST for this reason.
The trade-off: the Cloudflare edge and Traefik hops are not covered by these checks. Tier 2 (the hourly Playwright journey) is what exercises the public path.
KV Data Structure
status:state — { current, history } in a single value:
json
{
"current": {
"timestamp": "2026-03-12T14:30:00Z",
"services": {
"metadata-rest": { "status": "operational", "latencyMs": 42 },
"consumer-worker": { "status": "operational", "latencyMs": 28 },
"processor": { "status": "operational", "latencyMs": 156 },
"workspace-api": { "status": "operational", "latencyMs": 38 },
"postgres": { "status": "operational", "latencyMs": 12 },
"rabbitmq": { "status": "operational", "latencyMs": 8 }
}
},
"history": [
{ "t": "2026-03-12T14:30:00Z", "s": { "metadata-rest": "operational", "...": "..." } }
]
}History entries are deliberately compact (t/s, status only, no latency) — 288 full snapshots would not fit comfortably in one KV value.
Status values: "operational" | "degraded" | "down" | "unknown"
Free Tier Budget
KV writes are the binding limit at 1,000/day, and the cron fires 288 times a day. That leaves room for exactly three writes per invocation, so each pipeline gets one:
| Key | Writes/day |
|---|---|
status:state | 288 |
synthetic:state | 288 |
| Total | 576 (limit 1,000) |
The first synthetics build wrote four keys — 1,152/day — and would have started failing writes partway through each day. Any new state a cron-driven feature needs belongs inside an existing key, not a new one.
Reads are not a concern: /status is cached at the edge for 60s, well inside the 100,000/day limit.
Frontend (Status.tsx)
The component at projects/website/src/components/Status.tsx:
- Fetches
https://status.schemastack.ioon mount and every 60s - Maps
current.servicesto service cards with display names and icons (client-side mapping) - Maps
history[].s[serviceKey]to the uptime timeline bars (288 data points = 24h) - Calculates uptime percentage from history (operational count / total x 100)
- Shows loading spinner initially, "Unable to load status" on fetch errors
- When all services report
unknown, shows "Checking status..." state
Service Key Mapping
| API Key | Display Name | Group |
|---|---|---|
metadata-rest | API | Core Services |
consumer-worker | Real-time & Files | Core Services |
processor | Background Processing | Core Services |
workspace-api | Data API | Core Services |
postgres | Database (Primary) | Infrastructure |
rabbitmq | Message Queue | Infrastructure |
Backend: GET /api/status (metadata-rest)
Implemented in SystemStatusService (metadata-service, package …service.status). It runs server-side and checks all Docker-internal services:
| Service | Health Check |
|---|---|
| metadata-rest | Self (always operational) |
| consumer-worker- | http://...:8080/q/health/ready |
| processor- | http://...:8082/actuator/health |
| workspace-api- | http://...:8083/actuator/health |
| postgres | JDBC SELECT 1 on port 5432 |
| rabbitmq | http://rabbitmq:15672/api/health/checks/alarms |
Response format: the current object shown in the KV structure above.
Blue/green logic: for services with blue/green deployments, whichever instance is currently active is checked, and its status is reported.
It only sees the metadata database
checkPostgres() probes the metadata datasource. Workspace databases — including the managed demo_<org> ones — are not probed by any check, so connection-pool exhaustion there is invisible on this page. See Workspace Databases & Connection Pools.
Worker Secrets
Four values the Worker needs, none of which live in wrangler.toml. wrangler secret is write-only — there is no command that reads a secret back, so if you need to know a current value, look at wherever it originates (the server's .env.production.local, the SchemaStack UI, the Slack app page) rather than at Cloudflare.
| Name | Kind | Origin | Used for |
|---|---|---|---|
ORIGIN_HOST | plain var in wrangler.toml | — | Every cron-side request; see above |
ORIGIN_SECRET | secret | must match ORIGIN_SECRET in the server's .env.production.local | X-Origin-Secret header — Traefik rejects requests without it, so the origin can only be reached through the Worker |
SYNTHETIC_MCP_KEY | secret | a read-only MCP API key on the demo@schemastack.io account | The Tier 1 MCP probe. Read-only is deliberate: a writable key would let a failing probe mutate the public demo |
SYNTHETIC_ALERT_WEBHOOK | secret | Slack incoming-webhook URL | Alert delivery. If unset, checks still run and still write to KV — they just alert nowhere |
bash
cd schemastack-deployment/cloudflare/worker
wrangler secret put ORIGIN_SECRET
wrangler secret put SYNTHETIC_MCP_KEY
wrangler secret put SYNTHETIC_ALERT_WEBHOOK
wrangler secret list # names and modification dates only — never valuesBoth synthetic secrets are optional in the code (if (!env.SYNTHETIC_MCP_KEY) return;), which means a missing one degrades quietly rather than breaking the status pipeline. The flip side is that a typo'd secret name looks exactly like a healthy system: after rotating either, confirm the next cron actually produced a fresh synthetic:state.
bash
wrangler kv key get --binding STATUS_KV synthetic:state --remote | head -c 400Alerts fire on transition only — first failure and recovery, not every five minutes — so a silent channel means either nothing has changed state or the webhook is wrong. Tier 4 (a dead-man's switch that notices the monitoring itself dying) is not built, so silence is not yet proof of health.
Deployment
One-time setup (CF Worker)
bash
cd schemastack-deployment/cloudflare/worker
# Create the KV namespace
wrangler kv namespace create STATUS_KV
# Update wrangler.toml with the returned namespace ID
# (replace PLACEHOLDER_CREATE_WITH_WRANGLER)
# Deploy
wrangler deployVerification
bash
# Trigger cron locally
wrangler dev --test-scheduled
# Test endpoint
curl http://localhost:8787/status
# Website dev
cd schemastack-fe && npm run dev
# Visit http://localhost:3000/status