--- title: Monitoring and Troubleshooting description: Monitor replication health, resolve apply failures, and recover from WAL buildup canonical: https://www.paradedb.com/docs/operate/deploy/logical-replication/monitoring-and-troubleshooting --- Monitor both the publisher and ParadeDB to catch replication failures before they affect query freshness or fill the publisher’s disk. ## Monitoring the Publisher Permanent logical replication is operationally safe only if you watch the publisher, not just the subscriber. The most important signal is how much WAL a logical slot is retaining. ```sql SELECT slot_name, active, restart_lsn, confirmed_flush_lsn, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal, wal_status, safe_wal_size, inactive_since FROM pg_replication_slots WHERE slot_type = 'logical'; ``` Watch for: - `retained_wal` growing steadily because the subscriber is not acknowledging WAL quickly enough - `inactive_since` becoming non-`NULL` for longer than expected - `wal_status` showing that the slot is under pressure - Filesystem usage on the volume that contains `pg_wal` To reduce blast radius, configure `max_slot_wal_keep_size` on the publisher. This caps how much WAL a slot may retain, but it can also invalidate a lagging subscriber, so it should be paired with alerting and a reseed plan. ## Monitoring the Subscriber Use the subscriber to confirm that apply workers are healthy and that errors are not accumulating: ```sql SELECT subname, worker_type, received_lsn, latest_end_lsn, latest_end_time FROM pg_stat_subscription; SELECT subname, apply_error_count, sync_error_count FROM pg_stat_subscription_stats; ``` If `latest_end_time` stops advancing or `apply_error_count` increases, inspect the subscriber logs immediately. ## Troubleshooting Apply Failures One common cause of apply-worker failures is schema drift between the publisher and subscriber. Two common log patterns for schema drift are: ```text logical replication target relation "public.doctor" is missing replicated columns: "personnel_id", "role_function_id" ``` ```text logical replication apply worker for subscription "paradedb_subscription" has started background worker "logical replication apply worker" (PID 2570238) exited with exit code 1 ``` The first message is the root cause. The second means the apply worker crashed after hitting that error and Postgres will try to restart it. When you see these messages: 1. Inspect the subscriber logs for the first schema-mismatch error, not just the worker restart message 2. Compare the affected table definition on the publisher and ParadeDB 3. Apply the missing DDL on ParadeDB 4. Re-enable or refresh the subscription if needed 5. Rebuild any ParadeDB indexes affected by the schema change Another common cause of apply-worker failures is a logical replication conflict. For example, a duplicate key, a permissions failure on the target table, or row-level security on the subscriber can stop replication even when the schemas match. ```text ERROR: duplicate key value violates unique constraint ... CONTEXT: processing remote data during INSERT for replication target relation ... ``` When you suspect a replication conflict: 1. Inspect the subscriber logs for the first conflict error and note the finish LSN and replication origin if Postgres logged them 2. Resolve the underlying issue on the subscriber, such as conflicting local data, missing privileges, or row-level security policy interference 3. Resume replication normally once the conflict is removed 4. Only if you intentionally want to discard that remote transaction, use `ALTER SUBSCRIPTION ... SKIP` with care Skipping a conflicting transaction can leave the subscriber inconsistent, so it should be treated as a last resort rather than the default fix. For conflict types and the Postgres recovery workflow, see the [Postgres logical replication conflicts documentation](https://www.postgresql.org/docs/current/logical-replication-conflicts.html). ## Emergency: WAL Keeps Accumulating on the Publisher If the logical slot on the publisher is filling disk and ParadeDB cannot catch up quickly enough, the priority is protecting the publisher. 1. First, fix the subscriber if the issue is simple and recent, such as a schema mismatch or networking issue 2. If the publisher is running out of disk and the subscriber can be rebuilt, remove the subscription or drop the logical slot so the publisher can recycle WAL again 3. Recreate the subscription and reseed ParadeDB once the publisher is safe Disabling the subscription is not an emergency fix for WAL buildup. A disabled subscription still leaves the logical slot behind on the publisher, and that slot can continue retaining WAL. If the subscriber is reachable and healthy enough to cleanly tear down, dropping the subscription is the cleanest path: ```sql DROP SUBSCRIPTION paradedb_subscription; ``` To protect the publisher from continued `pg_wal` growth when you are intentionally giving up the current replica state, drop the slot on the publisher: ```sql SELECT pg_drop_replication_slot('paradedb_subscription'); ``` After either step, ParadeDB must be reinitialized from a fresh schema and data copy before it can resume as a logical subscriber. ## Common Pitfalls - Starting with pre-populated subscriber tables while using `copy_data = true` - Applying DDL on only one side of the replication link - Forgetting that new tables must be added to the publication and refreshed on the subscription - Writing directly to subscribed tables on ParadeDB, which can create conflicts with incoming replicated changes - Leaving a broken logical slot unattended on the publisher until `pg_wal` fills disk - Assuming `ALTER SUBSCRIPTION ... DISABLE` relieves publisher-side WAL pressure