Skip to content

Status Page ​

Public status page at schemastack.io/status showing real-time health of all SchemaStack services.

Data Flow ​

metadata-rest GET /api/status (checks all services internally)
  → CF Worker cron every 5 min (calls origin, writes to KV)
    → CF Worker GET /status (reads KV, returns JSON)
      → Website Status.tsx (fetches /status, renders)

same cron, separate pipeline:
  Tier 1 synthetics (5 live probes) → synthetic:state → Slack on transition

Architecture ​

CF Worker (schemastack-deployment) ​

The existing origin-proxy worker (cloudflare/worker/) handles the status pipeline:

Cron trigger (*/5 * * * *):

  1. Calls ${ORIGIN_HOST}/api/status with X-Origin-Secret header
  2. On success: writes the response plus the appended history to KV as status:state
  3. On failure: writes a fallback snapshot (metadata-rest/consumer-worker = down, others = unknown)
  4. History is trimmed to MAX_HISTORY = 288 entries (24h at 5-minute intervals)

Snapshot and history are one key, written once — they were two keys until Aug 2026, which put the Worker over the free tier once synthetics joined the same cron. See Free Tier Budget.

GET /status endpoint:

  • Reads status:state from KV
  • Returns { current, history } with Cache-Control: public, max-age=60 and CORS headers — the shape the frontend has always consumed, so the key merge was invisible to it

Why the cron calls ORIGIN_HOST, not schemastack.io ​

A Worker's subrequest to a hostname on its own zone bypasses Worker routes and goes straight to that hostname's origin — for schemastack.io that is Pages, which serves the SPA shell and returns 200 for any path. A check written against the public hostname therefore passes while measuring nothing. Every cron-side request uses env.ORIGIN_HOST for this reason.

The trade-off: the Cloudflare edge and Traefik hops are not covered by these checks. Tier 2 (the hourly Playwright journey) is what exercises the public path.

KV Data Structure ​

status:state — { current, history } in a single value:

json
{
  "current": {
    "timestamp": "2026-03-12T14:30:00Z",
  "services": {
    "metadata-rest": { "status": "operational", "latencyMs": 42 },
    "consumer-worker": { "status": "operational", "latencyMs": 28 },
    "processor": { "status": "operational", "latencyMs": 156 },
    "workspace-api": { "status": "operational", "latencyMs": 38 },
    "postgres": { "status": "operational", "latencyMs": 12 },
      "rabbitmq": { "status": "operational", "latencyMs": 8 }
    }
  },
  "history": [
    { "t": "2026-03-12T14:30:00Z", "s": { "metadata-rest": "operational", "...": "..." } }
  ]
}

History entries are deliberately compact (t/s, status only, no latency) — 288 full snapshots would not fit comfortably in one KV value.

Status values: "operational" | "degraded" | "down" | "unknown"

Free Tier Budget ​

KV writes are the binding limit at 1,000/day, and the cron fires 288 times a day. That leaves room for exactly three writes per invocation, so each pipeline gets one:

KeyWrites/day
status:state288
synthetic:state288
Total576 (limit 1,000)

The first synthetics build wrote four keys — 1,152/day — and would have started failing writes partway through each day. Any new state a cron-driven feature needs belongs inside an existing key, not a new one.

Reads are not a concern: /status is cached at the edge for 60s, well inside the 100,000/day limit.

Frontend (Status.tsx) ​

The component at projects/website/src/components/Status.tsx:

  • Fetches https://status.schemastack.io on mount and every 60s
  • Maps current.services to service cards with display names and icons (client-side mapping)
  • Maps history[].s[serviceKey] to the uptime timeline bars (288 data points = 24h)
  • Calculates uptime percentage from history (operational count / total x 100)
  • Shows loading spinner initially, "Unable to load status" on fetch errors
  • When all services report unknown, shows "Checking status..." state

Service Key Mapping ​

API KeyDisplay NameGroup
metadata-restAPICore Services
consumer-workerReal-time & FilesCore Services
processorBackground ProcessingCore Services
workspace-apiData APICore Services
postgresDatabase (Primary)Infrastructure
rabbitmqMessage QueueInfrastructure

Backend: GET /api/status (metadata-rest) ​

Implemented in SystemStatusService (metadata-service, package …service.status). It runs server-side and checks all Docker-internal services:

ServiceHealth Check
metadata-restSelf (always operational)
consumer-worker-http://...:8080/q/health/ready
processor-http://...:8082/actuator/health
workspace-api-http://...:8083/actuator/health
postgresJDBC SELECT 1 on port 5432
rabbitmqhttp://rabbitmq:15672/api/health/checks/alarms

Response format: the current object shown in the KV structure above.

Blue/green logic: for services with blue/green deployments, whichever instance is currently active is checked, and its status is reported.

It only sees the metadata database

checkPostgres() probes the metadata datasource. Workspace databases — including the managed demo_<org> ones — are not probed by any check, so connection-pool exhaustion there is invisible on this page. See Workspace Databases & Connection Pools.

Worker Secrets ​

Four values the Worker needs, none of which live in wrangler.toml. wrangler secret is write-only — there is no command that reads a secret back, so if you need to know a current value, look at wherever it originates (the server's .env.production.local, the SchemaStack UI, the Slack app page) rather than at Cloudflare.

NameKindOriginUsed for
ORIGIN_HOSTplain var in wrangler.toml—Every cron-side request; see above
ORIGIN_SECRETsecretmust match ORIGIN_SECRET in the server's .env.production.localX-Origin-Secret header — Traefik rejects requests without it, so the origin can only be reached through the Worker
SYNTHETIC_MCP_KEYsecreta read-only MCP API key on the demo@schemastack.io accountThe Tier 1 MCP probe. Read-only is deliberate: a writable key would let a failing probe mutate the public demo
SYNTHETIC_ALERT_WEBHOOKsecretSlack incoming-webhook URLAlert delivery. If unset, checks still run and still write to KV — they just alert nowhere
bash
cd schemastack-deployment/cloudflare/worker
wrangler secret put ORIGIN_SECRET
wrangler secret put SYNTHETIC_MCP_KEY
wrangler secret put SYNTHETIC_ALERT_WEBHOOK

wrangler secret list          # names and modification dates only — never values

Both synthetic secrets are optional in the code (if (!env.SYNTHETIC_MCP_KEY) return;), which means a missing one degrades quietly rather than breaking the status pipeline. The flip side is that a typo'd secret name looks exactly like a healthy system: after rotating either, confirm the next cron actually produced a fresh synthetic:state.

bash
wrangler kv key get --binding STATUS_KV synthetic:state --remote | head -c 400

Alerts fire on transition only — first failure and recovery, not every five minutes — so a silent channel means either nothing has changed state or the webhook is wrong. Tier 4 (a dead-man's switch that notices the monitoring itself dying) is not built, so silence is not yet proof of health.

Deployment ​

One-time setup (CF Worker) ​

bash
cd schemastack-deployment/cloudflare/worker

# Create the KV namespace
wrangler kv namespace create STATUS_KV

# Update wrangler.toml with the returned namespace ID
# (replace PLACEHOLDER_CREATE_WITH_WRANGLER)

# Deploy
wrangler deploy

Verification ​

bash
# Trigger cron locally
wrangler dev --test-scheduled

# Test endpoint
curl http://localhost:8787/status

# Website dev
cd schemastack-fe && npm run dev
# Visit http://localhost:3000/status

SchemaStack Internal Developer Documentation