Appearance
Workspace Databases & Connection Pools
How a workspace reaches its data, why managed databases have a five-connection ceiling, and the failure mode that took the API down on 16 Aug 2026.
Two kinds of workspace database
| Kind | Provisioned by | Where it lives | Named |
|---|---|---|---|
| Bring your own | The customer, via the database form | Wherever they run Postgres/MySQL | n/a |
| Managed | DemoProvisioningService | The dedicated demo Postgres server | demo_<orgSlug> |
Managed databases are created by CREATE DATABASE … TEMPLATE acme_store — a template clone, so provisioning is close to instant. Workspace.managed records which kind a workspace is.
Managed databases are per organisation, not per workspace
java
String dbName = "demo_" + orgSlug.replace("-", "_");
if (exists) { LOG.infof("Demo database '%s' already exists, reusing", dbName); }Two consequences that surprise people:
- Two template workspaces in the same organisation share the same tables.
- They also share the database's connection budget.
This is currently by design rather than by decision, and is worth revisiting: including the workspace slug in the name would separate them, at the cost of a migration for existing managed workspaces.
The five-connection ceiling
DemoProvisioningService caps managed databases:
java
stmt.execute("ALTER DATABASE " + quoteIdent(dbName) + " CONNECTION LIMIT 5");That protects the shared demo server from one organisation monopolising it. It also means pool sizing has to respect it. WorkspaceDatabaseService sizes pools accordingly:
java
DEFAULT_WORKSPACE_POOL_SIZE = 10 // bring-your-own — the customer owns the limits
MANAGED_WORKSPACE_POOL_SIZE = 2 // managed — must stay under CONNECTION LIMIT 5Until 17 Aug every pool was 10, so a single managed workspace could never fill its own pool and a second workspace on the same database was guaranteed to be refused.
Pool lifecycle
Pools are cached per workspace UUID:
java
private final Map<UUID, Pool> workspacePools = new ConcurrentHashMap<>();getWorkspacePool populates the cache with computeIfAbsent; closeWorkspacePool removes and closes. Deleting a workspace must call closeWorkspacePool, which deleteByOrganisationSlugAndUuid now does after the soft delete is persisted. It deliberately never fails the caller: the delete has already succeeded by then, and a pool that will not close should not turn it into an error.
The 16 Aug outage — worth reading before changing any of this
closeWorkspacePool existed but nothing called it. Pools for deleted workspaces stayed cached and kept their connections open. The sequence:
- Synthetic runs created and deleted workspaces in the
syntheticsorg, each leaving a pool behind. - Five idle pools — with no workspace behind them — exhausted
demo_synthetics. - Every query against it failed with
53300 too_many_connections, which fell through the exception mapper as "Unmapped PgException returned as 500", and reached users as "An unexpected error occurred". - Vert.x then logged
IllegalStateException: Result is already completerepeatedly — the pool's connection-acquisition timer racing a promise that had already completed. Its stack is entirely insideSqlConnectionPool; it is a symptom of starvation, not a separate fault. - Hours later
metadata-restcould not obtain a connection for its own health check (FailingStreak: 715) and the API was down for about two hours until restarted.
Three fixes followed: release pools on delete, size managed pools below the database limit, and map 53300 to a 503 with DATABASE_CONNECTION_LIMIT and retryable: true.
Diagnosing it again
bash
# Limits vs actual usage
docker exec docker-postgres-demo-1 psql -U <admin> -d postgres -c \
"select datname, datconnlimit,
(select count(*) from pg_stat_activity a where a.datname=d.datname)
from pg_database d where datname like 'demo_%';"
# Who holds them, and for how long — idle pools with no workspace are the tell
docker exec docker-postgres-demo-1 psql -U <admin> -d postgres -c \
"select usename, application_name, state, count(*), max(now()-state_change)
from pg_stat_activity where datname='demo_<org>' group by 1,2,3;"application_name = vertx-pg-client in idle state for many minutes means leaked pools. To recover immediately:
sql
select pg_terminate_backend(pid) from pg_stat_activity
where datname='demo_<org>' and state='idle';What healthy looks like
A synthetic run should spike to at most MANAGED_WORKSPACE_POOL_SIZE and return to zero once the workspace is deleted. Connections that persist after a run, climbing hourly, mean the release path has regressed.
Caveat on the status page
SystemStatusService.checkPostgres() probes the metadata database, not workspace databases. No status check exercises a managed database, so this entire failure mode is invisible there — it was the synthetics that failed at 03:12 and 04:01, not the status page.