A backup tells you where your data was at one moment. Point-in-time recovery lets you choose the moment. That distinction stops being academic at 10:33 on a Tuesday, when someone runs a DELETE without a WHERE clause and your last full backup is from midnight.
PostgreSQL supports point-in-time recovery, usually shortened to PITR, by combining a physical base backup with a continuous archive of write-ahead log files. This article walks through three approaches — a manual base backup plus WAL archiving, pgBackRest, and how both compare with pg_dump — and ends each one with a restore, because a backup you have never restored is not a backup.
Everything below was run on a disposable lab VM. Do not practise archiving or recovery settings on a production server.
How point-in-time recovery works
PITR replays a recorded history of changes on top of a starting copy, and stops at the moment you choose. Four pieces make that possible.
- Base backup. A physical copy of the whole data directory, taken while the server is running —
pg_basebackupdoes this. - WAL archive. Every write-ahead log segment, 16 MB by default, copied somewhere safe as it fills, using
archive_command. - Recovery target. Where replay stops: a timestamp, a transaction ID, an LSN, a named restore point, or simply the end of the archive.
- recovery.signal. An empty file in the data directory telling PostgreSQL to start in archive recovery, fetch WAL through
restore_command, and replay up to the target.
After a successful recovery PostgreSQL starts a new timeline, so WAL from the old history is never overwritten and a second attempt at a different target is still possible.
One rule decides whether any of this works: the recovery target must fall after the end of the base backup and inside the WAL you archived. Both edges of that window produce errors, and both are covered below.
Your recovery point objective — how much data you can lose — is whatever has not yet been archived. It is bounded by how quickly WAL segments reach the archive, which is why archive_timeout exists: it forces a segment switch after a set period even when the server is quiet.
The scenario that proves it
Every method below is tested the same way, against a database holding 1,000 customers and 50,000 orders:
- Take a backup.
- Insert a marker row, then record the time with
SELECT now();. - Cause the accident:
DROP TABLE orders;. - Run
SELECT pg_switch_wal();so the segment containing the accident is archived. - Restore to the recorded time, and confirm the table and the marker row are back.
The marker row matters. Without it you can prove the table returned, but not that it returned with the last transaction before the accident still in it.
Method A: base backup plus WAL archiving, by hand
This uses only what ships with PostgreSQL. It is slower and more error-prone than a dedicated tool, and it is the best way to understand what the tool is doing for you.
Turn on archiving
SHOW wal_level; -- must be replica or higher
ALTER SYSTEM SET archive_mode = 'on';
ALTER SYSTEM SET archive_command =
'test ! -f /var/lib/postgresql/wal_archive/%f && cp %p /var/lib/postgresql/wal_archive/%f';
In archive_command, %p is the path of the WAL file and %f is its name. The test ! -f refuses to overwrite a segment already in the archive. Changing archive_mode needs a restart; changing archive_command only needs a reload.
Then confirm archiving is actually running, because a silently failing archive is the most common way this whole mechanism turns out to be useless:
SELECT pg_switch_wal();
SELECT archived_count, last_archived_wal, failed_count FROM pg_stat_archiver;
You want files appearing in the archive directory and failed_count at zero. When archiving fails, WAL accumulates in pg_wal and will eventually fill the disk and stop the database. Fix it before going further.
A plain cp is fine in a lab. It does not flush to disk or verify what it wrote, which is one of several reasons production estates use a purpose-built tool.
Take the base backup
pg_basebackup -D /var/lib/postgresql/backups/base1 -Fp -X stream -c fast -P -v
-Fp writes plain files, -X stream includes the WAL needed to make the backup internally consistent, and -c fast starts the checkpoint immediately rather than waiting. The target directory must not already exist.
Restore to just before the accident
Stop the server, move the damaged data directory aside rather than deleting it, and put the base backup in its place with the right ownership and a mode of 700. Then tell PostgreSQL where to fetch WAL and where to stop:
restore_command = 'cp /var/lib/postgresql/wal_archive/%f %p'
recovery_target_time = '2026-10-09 10:33:44+05:30'
recovery_target_action = 'promote'
Use the now() value recorded before the accident, including the time zone offset. promote makes the server read-write once the target is reached; the default, pause, leaves it read-only so you can inspect the result before committing to it — which is the safer choice when you are not certain the target is right.
Create the trigger file, start the server, and read the log:
touch /var/lib/postgresql/16/main/recovery.signal
pg_ctlcluster 16 main start
tail -n 30 /var/log/postgresql/postgresql-16-main.log
A successful recovery logs starting point-in-time recovery, then restored log file ... from archive, then recovery stopping before commit of transaction ..., then selected new timeline ID: 2. Verify with a row count and a query for the marker row, then remove the temporary recovery settings so a later restart does not try to recover again.
The two errors that stop a restore
These are worth more than the successful run, because they are what you will actually meet. Both come from the same place: the recovery target falling outside the window the archive can serve.
must specify restore_command when standby mode is not enabled
PostgreSQL started in recovery because recovery.signal existed, but nothing told it how to fetch archived WAL. On Debian and Ubuntu the usual cause is a recovery settings file that is not where it is expected, or not named with a .conf extension so include_dir never picks it up.
Confirm PostgreSQL can actually read the setting before starting the server:
postgres -C restore_command -c config_file=/etc/postgresql/16/main/postgresql.conf
recovery ended before configured recovery target was reached
The target is later than the last transaction in the archive. PostgreSQL replayed everything available, ran out, and refused to promote because it never reached the point you asked for. The log line last completed transaction was at log time tells you exactly where the archive ends.
Pick a target inside the archived range. If you did not record now() before the change, pg_waldump will show you the commit times.
Reading the real error
pg_ctlcluster and systemctl will only tell you the service failed. The cause is always in the PostgreSQL log.
| Log message | Cause | Fix |
|---|---|---|
| must specify restore_command | Recovery settings file missing or not read | Create the .conf file; check include_dir |
| recovery ended before configured recovery target was reached | Target is later than the archived WAL | Pick an earlier target |
| requested recovery stop point is before consistent recovery point | Target is earlier than the end of the base backup | Pick a time after the backup finished |
| data directory has invalid permissions | Mode is not 700 | chmod 700 the data directory |
| Permission denied | Wrong owner | chown -R postgres:postgres |
Two log lines look alarming and are not: cp: cannot stat ... .history and a failed fetch of the next WAL segment. PostgreSQL asks for the file after the last one, does not find it, and ends recovery. That is how it knows it has finished.
One thing that is not obvious: a failed recovery can leave the data directory half-recovered. Restore the base backup again before retrying rather than reusing what is there.
Method B: pgBackRest
pgBackRest is a dedicated backup and restore tool for PostgreSQL. It uses the same underlying mechanism, so the difference is entirely operational: full, differential and incremental backups, parallel and compressed transfers, retention policies, checksum verification, asynchronous archiving, and repositories on S3, Azure Blob Storage or Google Cloud Storage.
A stanza is the name pgBackRest gives a PostgreSQL cluster. After installing and writing a short configuration file, point PostgreSQL at it:
ALTER SYSTEM SET archive_command = 'pgbackrest --stanza=main archive-push %p';
Note what that replaces. Method A set archive_command with ALTER SYSTEM, which writes to postgresql.auto.conf — and that file overrides anything in conf.d. Replacing it the same way is what stops the old command quietly winning.
pgbackrest --stanza=main stanza-create
pgbackrest --stanza=main check
pgbackrest --stanza=main --type=full backup
pgbackrest info
check forces a WAL switch and confirms archiving and the repository work end to end — a verification step Method A simply does not have. The restore is one command:
pgbackrest --stanza=main --delta --type=time \
"--target=2026-10-09 10:33:44+05:30" --target-action=promote restore
pgBackRest writes the recovery settings and creates recovery.signal itself, which replaces every manual step in section A. --delta rewrites only the files that differ, which on a large database is the difference between minutes and hours.
pg_dump and Barman
pg_dump is not point-in-time recovery. It exports the logical contents of one database, so the only moment you can return to is the moment the dump started. That is a real limitation and it is also sometimes exactly what you want: pg_dump restores across major versions and platforms, can restore a single table, and needs no WAL archive at all. For migrations, small databases and object-level restores it is the right tool.
Barman manages physical backups and WAL for many PostgreSQL servers from one central host, with a catalogue and retention policies. It suits teams that want a dedicated backup server. pgBackRest is more often installed alongside each database, though it can also use a separate repository host.
Comparing the four
| Aspect | pg_dump | Manual base backup + WAL | pgBackRest | Barman |
|---|---|---|---|---|
| Backup type | Logical | Physical | Physical | Physical |
| Point-in-time recovery | No | Yes | Yes | Yes |
| Data you can lose | Since the last dump | Since the last archived WAL | Since the last archived WAL | Since the last archived WAL |
| Incremental backups | No | No | Yes | Yes |
| Compression and parallelism | Yes | Not built in | Built in | Built in |
| Retention management | Manual | Manual | Policy based | Policy based |
| Integrity checks | Restore errors only | Manual | Checksums and verify | Built in |
| Restore effort | Low | High, many manual steps | One command | One command |
| Restore to a different major version | Yes | No | No | No |
| Single table restore | Yes | Restore all, then extract | Restore all, then extract | Restore all, then extract |
| Setup effort | Very low | Medium | Medium | Higher |
What the table means in practice: pg_dump answers what did this database look like at one moment. The physical methods answer what did it look like at this second. Manual archiving, pgBackRest and Barman all use the same PostgreSQL mechanism, so the choice between them is operational rather than technical — and under pressure, the tool that automates retention and verification is the one that is far less error-prone.
What to carry into production
- Test the restore, and time it. A backup counts once you have restored it and checked the data. The timing is what turns a recovery time objective from an aspiration into a number.
- Record the time before risky changes.
SELECT now();, or better,SELECT pg_create_restore_point('before_change');and recover to the named point. - Keep the target inside the window. After the end of the base backup, before the end of the archive.
- Monitor archiving. Watch
failed_countinpg_stat_archiver. A failing archive fillspg_waland eventually stops the database. - Know which setting wins.
ALTER SYSTEMwrites topostgresql.auto.conf, which overridesconf.d. - Keep the archive off the database disk. Separate storage or a remote repository, with restricted permissions and encryption.
- Keep enough history. Retention has to hold WAL back to the oldest backup you might need, or the backup is unrecoverable.
- Write the runbook. Under pressure, a tested checklist beats memory every time.
PITR is not a feature you switch on. It is a chain: base backup, continuous archive, a target inside the window, and a restore somebody has actually performed. Every link has to hold, and the only way to know they do is to break something on purpose and put it back.
If you are running PostgreSQL in production and have never tested a restore, that is the gap worth closing first. We do that work as part of PostgreSQL consulting, and as ongoing cover under remote database support — including building the archive, setting the retention, and rehearsing the recovery so the first real one is not the first one.

