Skip to content

Workspace Databases & Connection Pools ​

How a workspace reaches its data, why managed databases have a five-connection ceiling, and the failure mode that took the API down on 16 Aug 2026.

Two kinds of workspace database ​

KindProvisioned byWhere it livesNamed
Bring your ownThe customer, via the database formWherever they run Postgres/MySQLn/a
ManagedDemoProvisioningServiceThe dedicated demo Postgres serverdemo_<orgSlug>

Managed databases are created by CREATE DATABASE … TEMPLATE acme_store — a template clone, so provisioning is close to instant. Workspace.managed records which kind a workspace is.

Managed databases are per organisation, not per workspace ​

java
String dbName = "demo_" + orgSlug.replace("-", "_");
if (exists) { LOG.infof("Demo database '%s' already exists, reusing", dbName); }

Two consequences that surprise people:

  • Two template workspaces in the same organisation share the same tables.
  • They also share the database's connection budget.

This is currently by design rather than by decision, and is worth revisiting: including the workspace slug in the name would separate them, at the cost of a migration for existing managed workspaces.

The five-connection ceiling ​

DemoProvisioningService caps managed databases:

java
stmt.execute("ALTER DATABASE " + quoteIdent(dbName) + " CONNECTION LIMIT 5");

That protects the shared demo server from one organisation monopolising it. It also means pool sizing has to respect it. WorkspaceDatabaseService sizes pools accordingly:

java
DEFAULT_WORKSPACE_POOL_SIZE = 10   // bring-your-own — the customer owns the limits
MANAGED_WORKSPACE_POOL_SIZE = 2    // managed — must stay under CONNECTION LIMIT 5

Until 17 Aug every pool was 10, so a single managed workspace could never fill its own pool and a second workspace on the same database was guaranteed to be refused.

Pool lifecycle ​

Pools are cached per workspace UUID:

java
private final Map<UUID, Pool> workspacePools = new ConcurrentHashMap<>();

getWorkspacePool populates the cache with computeIfAbsent; closeWorkspacePool removes and closes. Deleting a workspace must call closeWorkspacePool, which deleteByOrganisationSlugAndUuid now does after the soft delete is persisted. It deliberately never fails the caller: the delete has already succeeded by then, and a pool that will not close should not turn it into an error.

The 16 Aug outage — worth reading before changing any of this ​

closeWorkspacePool existed but nothing called it. Pools for deleted workspaces stayed cached and kept their connections open. The sequence:

  1. Synthetic runs created and deleted workspaces in the synthetics org, each leaving a pool behind.
  2. Five idle pools — with no workspace behind them — exhausted demo_synthetics.
  3. Every query against it failed with 53300 too_many_connections, which fell through the exception mapper as "Unmapped PgException returned as 500", and reached users as "An unexpected error occurred".
  4. Vert.x then logged IllegalStateException: Result is already complete repeatedly — the pool's connection-acquisition timer racing a promise that had already completed. Its stack is entirely inside SqlConnectionPool; it is a symptom of starvation, not a separate fault.
  5. Hours later metadata-rest could not obtain a connection for its own health check (FailingStreak: 715) and the API was down for about two hours until restarted.

Three fixes followed: release pools on delete, size managed pools below the database limit, and map 53300 to a 503 with DATABASE_CONNECTION_LIMIT and retryable: true.

Diagnosing it again ​

bash
# Limits vs actual usage
docker exec docker-postgres-demo-1 psql -U <admin> -d postgres -c \
  "select datname, datconnlimit,
          (select count(*) from pg_stat_activity a where a.datname=d.datname)
   from pg_database d where datname like 'demo_%';"

# Who holds them, and for how long — idle pools with no workspace are the tell
docker exec docker-postgres-demo-1 psql -U <admin> -d postgres -c \
  "select usename, application_name, state, count(*), max(now()-state_change)
   from pg_stat_activity where datname='demo_<org>' group by 1,2,3;"

application_name = vertx-pg-client in idle state for many minutes means leaked pools. To recover immediately:

sql
select pg_terminate_backend(pid) from pg_stat_activity
where datname='demo_<org>' and state='idle';

What healthy looks like ​

A synthetic run should spike to at most MANAGED_WORKSPACE_POOL_SIZE and return to zero once the workspace is deleted. Connections that persist after a run, climbing hourly, mean the release path has regressed.

Caveat on the status page ​

SystemStatusService.checkPostgres() probes the metadata database, not workspace databases. No status check exercises a managed database, so this entire failure mode is invisible there — it was the synthetics that failed at 03:12 and 04:01, not the status page.

SchemaStack Internal Developer Documentation