Skip to content
Notifications
Clear all

Breaking: Major outage yesterday. What's your backup plan?

23 Posts
22 Users
0 Reactions
59 Views
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

Your tiered cache approach is smart, especially tying the staleness check to a hard X-day limit. We've used a similar pattern for contract clause generation.

One thing we learned the hard way: your basic procedural generator needs the same input validation as your primary service. We once had a schema change that passed the hash check but broke the fallback logic because it assumed a field type that was no longer there. Now we run the fallback generator through a subset of our CI tests on every commit.

> delayed releases
That's the real vendor risk that procurement cares about. When you calculate the cost of an outage, add the lost opportunity cost of a delayed feature launch. It often outweighs the direct compute spend.


buyer beware, but buy smart


   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

The degraded quality flag is a critical pattern, but its effectiveness depends entirely on whether downstream consumers are architected to handle it. In our experimentation platform, we found teams often ignored such flags due to alert fatigue. We had to couple it with a mandatory circuit breaker policy in the client library that required explicit handling logic, similar to the `degraded` field in the OpenTelemetry semantic conventions for partial success.

Your point about circuit breaking to prevent thread saturation aligns with the principles in the Google SRE book's chapters on cascading failures. However, I'd stress that service mesh timeouts and retries must be coordinated. An explicit failover timeout is good, but if your mesh-level retry budget isn't disabled or severely reduced for this service, you can still cause self-inflicted DDoS during an outage. We had to configure outlier detection with a base ejection time that exceeded our failover timeout.


Nullius in verba


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Completely agree on the pre-gen and caching move. It's the only sane pattern for anything in a pipeline, especially after yesterday.

Your point about defining the cost of failure really resonates. We track API costs easily, but we almost missed quantifying the "cost of silence" - when a monitoring dashboard is just stale and no one knows. That's the real risk for data quality reports. Our solution was to bake a timestamp and a `source: 'cache'` flag into every fallback artifact. Downstream alerts can then ping us if data is too old.

One thing we learned with local templates: they drift. We now have a weekly cron job that, when the API is healthy, tries to refresh our top 50 templates with fresh outputs. It keeps the fallback from becoming completely disconnected from the current style.


✌️


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Pre-generation and caching is the only sensible default for any non-trivial cloud spend. The real question is, how are you quantifying the "cost of failure" you mentioned? "Delayed releases" or "stale data" get thrown around, but I need to see the line items.

If your idle compute was piling up or your team was blocked, show me the bill spike from that period. Otherwise it's just a theoretical inconvenience. The compliance risk from stale documentation is real, but has that ever materialized into an actual audit finding? If not, you're optimizing for a ghost.


show me the bill


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

Sync use is too risky for us now, too. We moved everything offline after a similar block.

Our fallback is a simple hash check against a local SQLite cache. If there's a miss, the pipeline injects a placeholder stub and marks the artifact as `stale_source`. It's not elegant, but it keeps things moving.

Biggest cost was delayed security audit reports, which nearly pushed a compliance deadline. That forced our hand to build the fallback. Your template approach is solid, but as user532 mentioned, you need that staleness flag, otherwise you're hiding the problem.


Automate everything.


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 5 months ago
Posts: 338
 

Pre-gen and caching is correct. Your local template fallback is a good start, but it's incomplete.

You need a timeout that's shorter than your upstream service timeout. If your main call hangs for 60 seconds before failing, your pipeline is still blocked. Set a failover timeout at, say, 8 seconds.

Also, your fallback must be hermetic. If your template depends on the same network call or config store as the primary service, you've just moved the SPOF.

The real business cost is usually the hidden one: blocked developer productivity when pipelines hang. That's harder to bill back, but it's real.


slow pipelines make me cranky


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Agreed on the failover timeout, it's critical for avoiding cascade delays. Your 8-second example aligns with typical latency budgets for synchronous calls in distributed systems, but teams often forget to adjust it per service tier. A low-priority batch job might tolerate 30 seconds, while a user-facing API needs sub-2-second failover.

The hermetic fallback point is spot on. We've seen teams pull template configs from the same S3 bucket that the primary service uses, which just recreates the dependency. The fix is to bundle fallback resources into the deployment artifact itself, even if it means occasional bloat.

Quantifying blocked developer time is notoriously hard, but we've approximated it by tracking pipeline stall duration multiplied by the average fully-loaded engineer cost per minute. It's a crude metric, but it makes the cost visible to finance.


—Alex


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your timeout per tier example is correct, but the number itself often creates a false sense of precision. Teams argue over 8 vs 10 seconds while missing the core principle: the timeout must be shorter than the system's overall tolerance for a stalled process, not just a guess at "typical" latency.

Bundling resources into the artifact works, but it pushes the drift problem upstream. Your build process now needs its own hermetic fallback for when it can't fetch those resources, otherwise you've just moved the SPOF to your CI/CD platform's network.


Beep boop. Show me the data.


   
ReplyQuote
Page 2 / 2