A conversation-indexing worker goes offline while the application keeps writing new events to PostgreSQL. The worker’s replication slot still marks the oldest write-ahead log (WAL) it may need. Disk use on the database server can keep growing even though the consumer has stopped doing useful work.
A slot preserves a restart position for a replication consumer. PostgreSQL’s replication settings say a slot can retain WAL in pg_wal, and the default max_slot_wal_keep_size = -1 permits unlimited slot retention. This matters for an AI product whose indexing, analytics, or agent-memory pipeline consumes database changes: an idle downstream service can become a storage risk upstream. The slot’s existence alone does not prove a problem; its activity and required WAL position do.
Start with a read-only inspection:
SELECT slot_name, slot_type, active, restart_lsn,
wal_status, safe_wal_size
FROM pg_replication_slots
ORDER BY slot_name;
The slot view defines restart_lsn as the oldest WAL that might still be needed. active tells you whether a slot is currently streamed. wal_status distinguishes retained WAL from a slot whose required WAL is due for removal or already lost. safe_wal_size estimates how many more WAL bytes can be written before a finite limit puts the slot in danger; it is null when the limit is unlimited or the slot is lost. Check the consumer’s health and ownership before changing any slot.
Setting a finite max_slot_wal_keep_size bounds what slots may retain at checkpoint time. It is a recovery policy, not an exact instantaneous disk quota. If a consumer falls farther behind than the limit, required WAL can be removed and that consumer may be unable to resume from its slot. wal_keep_size has a different role: it sets a minimum amount of past WAL kept for standby streaming, and other needs such as archiving or checkpoints can retain more. The standby documentation describes archive recovery for a physical standby when the needed segment is available there; do not assume a logical change consumer has the same recovery path.
Choose the retention limit from a measured outage window and WAL generation rate, with room for bursts and checkpoint timing. Alert on inactive slots, their WAL status, and disk headroom. If a consumer has exceeded the chosen window, plan its reinitialization or other supported recovery before expecting it to reconnect. Keeping every byte protects resumption until the disk fills; bounding retention protects the server while making that resumption conditional.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.