This is a deep dive from 6 October 2026, published by Rory Malcolm at incident.io. He credits the work to himself and Jonathan Donaldson on the reliability team. The capacity figures are incident.io’s account of its own load tests.
incident.io says Postgres is its main transactional database, with about 900 tables. The On-call “ticker” checks each active escalation and advances it when someone acknowledges or declines. incident.io says that happens about 15 million times a day, through a Postgres queue. In load tests, with rate limits off, acquiring work began to dominate, and more ticker workers made it worse. The post’s opening result is a 3.5 times improvement in peak escalation capacity at the current size. The close describes the same tests as 3.5 times escalation throughput under peak load.
How the savepoint became the wait
A SAVEPOINT marks a point inside the current transaction so later commands can be rolled back alone. PostgreSQL says a subtransaction starts inside a transaction, can commit or abort without affecting the parent, and can be started with SAVEPOINT. incident.io says each tick ran in its own subtransaction after the batch was claimed, so one failure did not dirty the rest of the batch.
The claim was SELECT ... FROM escalations ... FOR UPDATE SKIP LOCKED. SKIP LOCKED skips any selected row that cannot be locked immediately, so workers pass one another. incident.io says the savepoint meant the row’s xmax held a MultiXact id, Postgres’s label for a row locked by more than one transaction, rather than a single transaction id. Looking that id up reads pg_multixact through an LWLock. Every skipped row was another lookup. Every new worker locked more rows and added more readers to that lock. Sampling pg_stat_activity showed LWLock:MultiXact. incident.io says Cloud SQL Query Insights showed only the wait class. This is its account of this queue. It does not say every savepoint in Postgres behaves this way.
What incident.io says it changed
It moved the queue from the wide escalations table to a slim escalation_jobs row. incident.io says that tuple is about 15 to 60 times smaller. Logical reads of the old table from shared_buffers had peaked at 136 GiB/s, which it says counts pages touched, not bytes read from disk.
It dropped the per-tick subtransaction. One statement sets a lease, claimed_until = now() + interval '10 seconds', on a batch chosen with FOR UPDATE SKIP LOCKED, and returns the ids. Each escalation is then processed in its own transaction. MultiXact waits still appeared, because other subtransactions remain elsewhere. In the heaviest test they no longer dominated time in the database. A crash after the claim and before the tick leaves the row claimed for up to 10 seconds. A graceful shutdown finishes claimed work.
It left claimed_until unindexed and set fillfactor = 85, so inserts pack 85 percent of each page. A heap-only tuple update applies when no indexed column changes and the new version fits on the same page, which avoids new index entries. Fillfactor is that packing percentage, from 10 to 100. next_due_at stays indexed, so incident.io says autovacuum still has to be strict.
It also stopped every worker from claiming for itself. One dispatcher goroutine per pod fills a channel, and workers block on that channel. Claim queries drop from workers times pods to one per pod. If a job waits until more than half the lease is gone, the worker renews it. The post says that almost never happens.
Practical takeaway
On incident.io’s account, a savepoint around each queued job turns the lock into a MultiXact, and SKIP LOCKED then looks that MultiXact up on every row it passes. Their replacement is a short lease in an unindexed column, one transaction per job, free space for a HOT update, and one claimer per pod. The 3.5 times figure is their peak load test after these changes and others they defer to a later post. It is their result on their queue.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.