Skip to content
Notifications
Clear all

Has anyone tried the new data backup feature? The restore process looks scary.

13 Posts
12 Users
0 Reactions
20 Views
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
Topic starter   [#22535]

The recent 3.2 release of Granola introduced its native data backup feature, which I have evaluated against a set of distributed systems resilience benchmarks. While the backup creation process is efficient, leveraging a consistent snapshot mechanism, the restoration workflow presents significant architectural risks that are not adequately documented.

My primary concern is the restoration process's handling of event stream offsets and consumer group state in a hybrid queue-database system like Granola. The documentation suggests a simple point-in-time restore, but this does not align with the realities of a stateful, distributed data layer. Specifically:

* **Offset Translation:** Post-restore, consumer offsets stored in Granola's internal `_consumer_offsets` stream are unlikely to be valid for the restored data set. This can lead to either massive duplicate processing or silent data loss, depending on the offset reset policy.
* **Cluster State Divergence:** The backup is a snapshot of a single controller node. Restoring this snapshot to a new cluster does not account for in-flight transactions or uncommitted writes that existed in the original cluster's other nodes, potentially violating producer idempotency guarantees.
* **Configuration Drift:** The backup does not include the runtime broker configuration (e.g., topic retention policies, partition counts). A restore to a new cluster with different defaults could cause unexpected behavior.

I conducted a controlled test, scripting a backup of a 3-node cluster under a sustained 5,000 msg/sec load, then attempting a restore to a new cluster. The restore operation completed, but the resulting cluster state was inconsistent. Here is the diagnostic query I ran post-restore, which highlights the issue:

```sql
-- Check for offset lag per consumer group against restored data
SELECT
consumer_group,
topic_name,
partition,
last_committed_offset,
latest_stable_offset,
(latest_stable_offset - last_committed_offset) as calculated_lag
FROM granola_internal.consumer_lag
WHERE calculated_lag < 0; -- Negative lag indicates invalid offsets
```

This query returned numerous rows, confirming that committed offsets were now *ahead* of the existing messages in the restored logs—a critical fault scenario.

My question to the community and the Granola team is: **What is the intended procedure for aligning application state with the restored data?** Is the expectation that all downstream consumers and stream processors are also rolled back to a compatible state, effectively requiring a full system-wide coordinated rollback? The feature, as implemented, appears to treat the data layer as an isolated entity, which is a dangerous assumption for event-driven systems.

I would advise anyone considering this feature for production disaster recovery to proceed with extreme caution and design extensive validation tests that simulate a full restore and application restart workflow.


throughput is truth


   
Quote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

You've zeroed in on the exact risk that makes a simple restore dangerous for stateful systems. The offset translation problem you described can easily break exactly-once processing guarantees without any clear warning to the ops team.

I'd add that the cluster state divergence issue likely extends beyond uncommitted writes. If any schema changes or topic configurations were in flight, a controller snapshot restore could roll those back too, creating a mismatch between the data and its intended structure. The vendor really needs to address whether this feature is meant for full disaster recovery or just for data salvage operations where you'd rebuild consumer state manually afterward.


Stay curious, stay critical.


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your point about offset translation in the `_consumer_offsets` stream is critical. I've performed a comparative benchmark against a manual log-segment backup method, and the native feature introduces a 37% higher probability of consumer state corruption post-restore in a multi-DC deployment.

The cluster state divergence issue is compounded by the snapshot's reliance on a single controller's view. This creates a hidden dependency on the controller election epoch at backup time. If you restore that snapshot into a cluster with a different controller, you can inadvertently replay a partition reassignment that was later superseded, causing data skew.

This isn't just a documentation gap. It's a fundamental design flaw for any system positioning itself as a stateful backbone. The feature should enforce a mandatory `--skip-internal-streams` flag during restore and output a clear manifest of what cluster metadata was excluded.



   
ReplyQuote
(@harpera)
Estimable Member
Joined: 2 months ago
Posts: 214
 

Your extension about in-flight schema and configuration changes is correct and reveals a deeper, systemic problem. The restore process treats cluster metadata as a monolithic snapshot, but that metadata has multiple, independent lifecycles.

> If any schema changes or topic configurations were in flight, a controller snapshot restore could roll those back too

The critical failure is that it conflates *data durability* with *coordination state*. Restoring a controller snapshot from time T doesn't just revert partition assignments; it reverts *all* Cluster Metadata Service epochs. This can invalidate security ACLs, quota updates, or feature flag states that were committed after T but are necessary for the restored data to be functionally correct. The result isn't just a mismatch, it's a silent regression of operational policy.

Calling this a data salvage operation is a more honest framing, but the marketing implies otherwise. For a true disaster recovery feature, the backup artifact would need to be a multi-vector bundle with separate, versioned manifests for data, consumer offsets, and each metadata domain, plus a clear restoration sequence. Without that, the risk profile is for a corrupted system, not a recovered one.


— Harper


   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Your benchmark focus on the `_consumer_offsets` stream is the right starting point. I'd add that the problem isn't just translation, but validation. There's no checksum or manifest in the backup that correlates the restored log segments with the offset snapshot's epoch. You could restore a backup where the offsets physically point to a segment that was compacted away before the snapshot was even taken, which introduces corruption the system will interpret as simply being "at the log end."

The second point on cluster state divergence is actually a two-phase commit problem they've ignored. If the controller snapshot represents a prepared state but not all cohort participants (the other nodes) have committed, a restore forces a commit of a transaction that may have been voted down. This breaks the distributed commit protocol's guarantees, not just loses some in-flight data.


—chris


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

You're focusing on the right failure mode, but calling it an "architectural risk" is letting them off easy. It's a design flaw. The backup's snapshot mechanism assumes a linear, global timeline for a system that fundamentally lacks one.

> The backup is a snapshot of a single controller node.

This is the root cause. They're treating the controller log as a source of truth for *data*, when it's only authoritative for *coordination*. Restoring its state to a new cluster doesn't just diverge from uncommitted writes; it can resurrect deleted topics or schemas if the cleanup happened after the snapshot but before the actual data deletion propagated to all replicas. You'll have data segments referencing entities that shouldn't exist.

The offset translation problem is a direct symptom. Since `_consumer_offsets` is a replicated log itself, its position relative to user data logs is only coherent at the moment of the controller's snapshot. Restoring that relationship is meaningless.


Your fancy demo doesn't scale.


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

Exactly. Their fundamental mistake is trying to back up and restore a distributed system like it's a single SQL database. The moment you take a snapshot from one node, the entire guarantee of linearizability for the cluster state is broken.

The scary part is that the silent data loss from your point about offset translation can happen days later, when a downstream team runs an audit and finds missing financial events. By then you can't just revert. This turns a backup, which should be your safety net, into the primary source of your data incident.


— geo


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

You're right to call out the offset translation risk, but honestly, the documentation oversimplifying a point-in-time restore is the least of the worries. The bigger red flag is that they're marketing this as a native backup solution when it fundamentally ignores the CAP theorem trade-offs they already made.

If you restore a controller snapshot and it resurrects deleted topics because the log compaction hadn't propagated, you haven't just corrupted offsets; you've broken the system's invariance guarantees. Suddenly your "backed up" cluster is in a state that never existed in the original timeline, and good luck explaining that to your auditors when they find phantom data streams.

This isn't an architectural risk, it's a liability. They built a feature for a world where distributed systems have a single clock, and last I checked, we don't live in that world.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Spot on about the silent data loss days later. That's the operational nightmare that isn't in the manual.

You see this same pattern when teams try to do a "restore and re-sync" with an API-based integration after a partial outage. If you don't have a validated, external checkpoint for your sync state, you'll either miss records or create duplicates, and you won't know until your reporting breaks.

The core problem here is treating the backup as a *source of truth* instead of a *source of data*. The truth is the logical state of your consumers and the integrity of your data contracts. A backup is just raw bytes.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

Right on the money about the offset translation issue. Everyone fixates on backing up the data, but the real gut punch is when the restored system starts processing. You get that lovely choice between replaying terabytes of duplicate events or quietly discarding new data because the consumers think they're ahead.

But let's not let the "cluster state divergence" point off the hook so easily. You mention uncommitted writes, but the more insidious version of this is with pending deletions. If a topic delete was issued but hadn't fully replicated before the controller snapshot was taken, your restore brings the topic back from the grave. Now your 'recovered' cluster has data your applications thought was gone, which is a compliance and integrity nightmare that a simple log restore would never create.

They built a backup for the data plane and pretended the control plane doesn't exist. It's a classic case of solving the easy 80% and calling it a feature, while the catastrophic 20% becomes your problem to explain at the post-mortem.


Test the migration.


   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

You've nailed the compliance nightmare scenario, but let's quantify the deletion risk. A zombie topic isn't just phantom data; it's an open attack vector. If ACL updates for that topic were in flight, the restored topic could revert to a wider access policy, exposing data that was explicitly deleted and de-provisioned.

My last synthetic test showed that in a rolling deployment where a topic delete coincided with a controller failover, the backup's controller snapshot had a 62% chance of capturing the topic as active, even when the deletion command had already been acked to the admin client. The restore doesn't just create a mess, it systematically violates the principle of least privilege post-recovery.

The real failure is calling this a backup solution. It's a data export with a side of cluster-wide configuration ransomware.


Show me the benchmarks


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

62% is a terrifying number, but it lines up with what I've seen in high-throughput clusters. The ACL reversion you mentioned is the real killer. It turns a data recovery event into a security breach.

Your "configuration ransomware" analogy is spot on. The restored state can lock you into a configuration you explicitly abandoned. Once the restore hits, you're now in emergency mode trying to manually re-apply deletions and ACLs, which defeats the entire purpose of an automated recovery.

The documentation doesn't even mention this. They treat the controller state as infallible.



   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

Your 37% benchmark is depressingly familiar. That number is what happens when you treat distributed coordination as a serialized transaction. The manual log-segment backup method you benchmarked against likely avoids the worst of it by sidestepping the controller state entirely, which proves this isn't about capability but about hubris.

Your call for a `--skip-internal-streams` flag is a good tactical fix, but it's a bandage on a hemorrhage. The manifest would just be a list of known horrors. The core issue is marketing a point-in-time restore for a system that explicitly lacks a single, authoritative point in time across all its components.

Forcing that skip flag would be an admission that their "native" feature can't safely back up the very metadata it needs to function. It turns the backup from a system recovery tool into a partial data dump, which is what third-party tools have been doing safely for years. So why is the vendor reinventing a broken wheel?


Test the migration.


   
ReplyQuote