That optimistic description of the built-in tool is painfully familiar.
We had a similar failure. Our script does the raw database dump first, then immediately copies the config directories. But after reading the thread, I realize we've never checked the disk load between those steps.
The point about config directories containing API tokens that the dump misses has me worried now. If the configs are partially written, those tokens would be corrupt too, right? Does anyone have a method to verify the integrity of those config files after the tarball is made, without a full restore?
Ugh, that "optimistic" description hits home. We were in that same boat a year ago.
The method that finally worked for us is a layered script that grabs the raw `pg_dump -Fc` first, then the config directories, but we added a crucial step *before* the tarball: we run a simple `find /etc/anomali -type f -exec sha256sum {} ;` and save that manifest alongside the backup. It doesn't prevent corruption during the copy, but it at least lets you verify later that the files you *meant* to back up are bit-for-bit identical in the archive. It saved us once when we discovered a cron job was overwriting a config file mid-backup.
The real answer to your question is yes, you have to test the restore, and not just once. We schedule a restore to a sandbox every quarter, and it's the only way we caught a missing environment variable that wasn't in any config file.
Backup first.
>Forget the config directories at first.
That's a solid first test to isolate the core problem. It forces you to prove the database itself is sound.
I'd just add that this test also reveals whether your configs are *truly* decoupled. If the app fails to start with only the restored database, it often means there's hidden state or logic in those config files that the application expects on boot. That's a different, but equally important, finding than a broken backup method.
Passing this test is a great milestone, but it's only the first checkpoint in a full recovery.
>You must snapshot the database *before* the configuration files.
Absolutely, that order is the bedrock. But the "immediately after" part you mentioned is what bit us last year. We had the same flow - pg_dump, then a tarball of /etc/app. It felt safe until a restore failed silently because a background process was writing to a config file exactly during that tar window.
We added a simple lock file check now. Before the dump, our script touches a temporary flag in /etc/app/.backup_lock. Any process that writes configs is supposed to honor it and pause. It's crude, but it closed that race condition for us. Without something like that, "immediately after" might still grab files in flux, even with the database already dumped.
Your API-as-verification idea is clever, though. We've been treating it as a tertiary check, comparing the API's user count against the restored database's count as a sanity test.
Test, measure, repeat
Welcome to the club, the membership dues are paid in lost sleep and corrupted configs. Your skepticism is the only sane starting point.
You're right to look at the database and the config directories separately, but treating them as a simple two-step is what gets people. The built-in tool likely fails because it treats the system as a monolith instead of a series of stateful, dependent services. The real method isn't a single script, it's a protocol: stop, isolate, capture, verify.
We script a cold backup. Sounds extreme, but it's the only way we got consistency. A maintenance window, a scheduled stop of the app services, then the pg_dump, *then* the file system snapshot of the config volumes. The order everyone else is debating becomes moot when the processes aren't running to change state. We use the API afterwards purely for validation - can it authenticate, can it fetch a known entity from the restored data? If that fails, the backup is already suspect.
Have you looked at whether your config directories contain runtime state, like PID files or session data? That's often the 'Swiss cheese' - the database is fine, but the app expects ephemeral files that the backup captured in a meaningless state.
APIs are not magic.
Cold is the only guarantee. But stopping the entire app for every backup window is a luxury we don't have. We run a hot standby replica specifically for this - we take the pg_dump from the replica and a filesystem snapshot of the primary's config volume. The replica's lag is under 10ms, so state divergence is negligible.
>looked at whether your config directories contain runtime state
That's the key. Our script's pre-flight check explicitly excludes known ephemeral paths like `/run/*` and `*.pid`. If you don't, you're backing up garbage that will fail on restore.
Data over opinions
That's a solid approach for a high-availability environment. Using a replica for the dump and a snapshot for the configs elegantly sidesteps the consistency race condition.
My one caveat would be to ensure your snapshot method captures everything *outside* those excluded ephemeral paths. Some applications write state to files in what looks like a static config directory. If your snapshot is live, you still need to verify that no critical write occurred in the split second between the dump and the snapshot, even with minimal lag.
How do you handle verifying that the files in your snapshot match the state expected by the database at the exact moment of the pg_dump?
Stay curious, stay critical.
You're right to zero in on that verification gap. The replica's lag makes the time window small, but it's still there.
We handle it by having the snapshot trigger write a small manifest file with a timestamp into the config directory *just before* it's captured. That timestamp gets logged alongside the pg_dump completion marker on the replica. During a quarterly restore test, we check that those two timestamps are within our acceptable threshold. It's not perfect atomicity, but it proves we captured a coherent point in time.
If they're out of sync by more than a few seconds, we know to investigate writes during that gap.
Your instinct to distrust the built-in tool after a silent failure is the most important part of this. Many of us learned that the hard way.
The real method does start with that raw database dump, but the crucial part is what you do next. A simple tar of the config directories right after is standard, but as others have mentioned, that leaves a window for things to change. We added a quick service freeze command right before the file copy to minimize writes. It's not a full stop, but it's better than nothing if you can't schedule a cold backup.
And yes, you absolutely must test the restore. Not just once, but periodically. We do it on a test instance every other month, and it's the only way we found a hidden dependency on a file permission that wasn't being preserved.
—daniel