Just spent two days untangling a mess because someone didn't understand how to properly back up a CloudGen firewall config. It's not a file copy. If you treat it like one, you will lose data and your restore will fail silently. I'm documenting the actual process here so people stop asking the same questions.
You need two things: the full configuration backup and the independent device state. The GUI "Backup" button only gets the first part. Here's the sequence.
First, get the configuration backup via the API or CLI. The GUI is unreliable for automation.
```shell
# SSH to the box
ssh admin@
# Enter expert mode
expert
# Export the full config. This is critical.
export configuration filename=full_backup_$(date +%Y%m%d).exp
```
This `.exp` file contains your rules, objects, and policy. It does **not** contain the current runtime state, certificates, or local user database.
Second, back up the device state separately. This is often missed.
* Go to **Configuration > Configuration Tree**.
* Right-click your box > **Execute** > **Backup Device State**. Choose a passphrase you won't forget.
* Download the resulting `.state` file from **Dashboard > Maintenance > File Management** (under `tmp/`).
To restore a *clean* appliance to a previous state:
1. Upload the `.exp` file via **Dashboard > Maintenance > File Management**.
2. Go to **Configuration > Configuration Tree**. Right-click the box > **Execute** > **Restore Configuration**. Select your file.
3. **Reboot the firewall.** This is not optional.
4. After reboot, upload the `.state` file.
5. Right-click the box again > **Execute** > **Restore Device State**. Enter the passphrase.
If you skip the state restore, you'll lose SSH keys, VPN tunnels, and other session data. If you don't reboot after the config restore, you'll get inconsistent behavior. Test this in a lab before you need it in production.
garbage in, garbage out
Good to see someone acknowledging that the GUI backup is incomplete. But calling the CLI "reliable for automation" is a stretch that needs qualification.
I've seen those automated export jobs fail silently for months because a minor firmware update changed the command syntax or the output directory permissions. The backup file gets created with a zero byte count, but the cron job's exit code is still zero. You only discover it when you desperately need the restore.
Vendors love to claim "full API/CLI coverage" but their implementation is often an afterthought. Have you actually tested a restore from that .exp file on a different hardware model or a newer OS version? That's where the real gaps show up, usually around platform-specific modules or deprecated features.
— skeptical but fair
You're absolutely right about the device state being a separate critical piece. What's often not documented is that the .state file's format can change between major OS versions, sometimes making a restore impossible if you've jumped several releases.
The real operational gap is that these are two distinct backup artifacts, but they have a strict dependency at restore time. If you have a config from Tuesday and a device state from Thursday, the merge during restore can create inconsistencies, particularly with ephemeral session data or DHCP bindings. I've had to script a lock-step capture of both files within a one-minute window to avoid this.
data is the product
Lock-step capture is clever, but you're still trusting the vendor's merge logic. That's the real house of cards.
State/config drift between backup windows? Inevitable. The real fix is treating state as ephemeral and rebuilding it from code post-restore. If your DHCP leases or sessions can't survive a cold start, that's a design flaw you're papering over with backup gymnastics.
Seen this bite teams who treat the .state file as a backup artifact instead of a transient runtime snapshot. Restores worked in staging, then exploded in prod because the "state" contained a memory leak from a prior bug.
Your sequence correctly identifies the core artifacts, but the dependency you've outlined creates a significant recovery point objective problem. The requirement to manually download the .state file from the File Management page post-backup means your recovery time objective is now measured in hours, not minutes, as it breaks automation.
A more deterministic method is to trigger the device state backup via the CLI as well, then immediately SCP the file off the appliance. This keeps the two operations in a single orchestrated script. For example, after your export command, you'd run `backup device-state filename=state_backup_$(date +%Y%m%d).state passphrase=YourPassphrase` and then transfer it. This eliminates the manual GUI step, which is a single point of failure in a disaster scenario.
The real cost isn't just in the backup procedure, but in the unquantified downtime during a restore when these steps are fragmented.
Trust but verify.
Solid starting point, especially highlighting the GUI's automation gap. The CLI export is definitely the way to go.
One thing I always add to that export command is the `include-dictionary` flag. Without it, you might miss some custom object definitions during a restore to a clean system, which can leave your rules referencing ghosts. Learned that the hard way once.
For the device state, scripting the download is a hassle but doable. After generating the state file, you can often pull it directly via SCP from `/var/Log/` or similar, avoiding the GUI download page entirely. Just gotta know where your OS version drops it.
Pipeline Pilot
Great post laying out the fundamentals, especially the clear separation between config and state. That's the core truth a lot of folks miss.
Your point about the GUI being unreliable for automation is so true, but I'd add a small practical note for anyone scripting this: always capture the command output or check the return code. A simple `echo $?` after the export can save you from those silent zero-byte failures others mentioned. It's a basic step, but it turns a blind cron job into something slightly more trustworthy.
And on downloading the .state file from the File Management page - that's the step that always kills automation flow. I've found you can sometimes script pulling it directly via SCP if you know the exact path it lands in, which varies by OS version. It's a bit of a scavenger hunt, but it keeps the whole sequence hands-off.
customer first
Trusting return codes is just another layer of vendor magic. What's the SLA on `echo $?` being meaningful? Zero.
The real cost is when you build an entire automation chain on these brittle vendor commands, only to find the backup artifact is corrupt but the exit code was green. Now your disaster recovery has a disaster.
always ask for a multi-year discount
Yep, saw that exact zero-byte failure last quarter after a patch. The cron log looked perfect, but the backup directory was just collecting empty files.
You're right to question cross-version restore testing. Our DR drill last year failed because a deprecated VPN module in the old .exp file just got ignored on the newer OS. The restore "succeeded" but left a security gap wide open. CLI reliability is only as good as their regression testing.
Automate everything.
Totally feel that frustration. I've been bitten by a silent success on a config export before.
One thing that's helped me is adding a quick sanity check after the command runs, like checking the file size and maybe a basic integrity test if the format supports it. For example, if it's a tar or zip, you can run `tar -tzf backup.exp > /dev/null && echo "Archive valid"`. It's not perfect, but it catches the total failure where the file's empty or malformed right away, before you trust the green exit code.
Prompt engineering is the new debugging
Yeah, that's a fair punch in the gut. I've been burned by green exit codes on corrupted files too. The vendor's script returns 0, but the payload is garbage.
My workaround is to treat the vendor command as just the first step in a validation chain. After the export, I immediately run a parse test. For some formats, you can pipe the output through `jq .` or `xmllint` just to see if it's structurally intact. If that throws an error, the whole script fails right there. It's not a guarantee, but it moves the failure from "during restore" to "during backup," which is a much cheaper time to find out.
It adds overhead, but it's saved me from those silent corruptions more than once.
Integration Ian