Skip to content
Notifications
Clear all

Switched from on-demand to always-on. Latency improved, but was it worth it?

11 Posts
11 Users
0 Reactions
29 Views
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
Topic starter   [#23294]

We just completed our first full billing cycle after switching from Prolexic's on-demand scrubbing to their always-on protection. The performance graphs are undeniably better—our 95th percentile latency to the EU region dropped by ~40ms during the last "background noise" DDoS period.

But before anyone starts celebrating, let's look at the actual cost breakdown.

* On-demand (previous quarter): Baseline + two volumetric events. Total: **$X**
* Always-on (this quarter): Flat, significantly higher rate. Total: **$3.2X**

We paid over three times the previous cost to mitigate what were, frankly, non-disruptive events. The "improved latency" was just our traffic no longer taking a detour to a scrubbing center during those periods.

So my question for others who've made this switch: was the juice worth the squeeze? Specifically:

1. Did you perform a genuine risk/cost analysis, or was it a knee-jerk "latency is king" move pushed by engineering?
2. Has anyone successfully negotiated a middle-ground with Akamai where always-on is only enabled for specific, critical frontends instead of the entire /24?
3. What monitoring gaps did you have to fill to justify the always-on spend to your finance team? Because "lower p95 latency" alone didn't cut it for us.

Our internal postmortem on the decision is... not flattering. It feels like we bought a tank to deal with occasional egg-throwing. The technical outcome was predictable; the business case, less so.


- Nina


   
Quote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

1. I'm a cloud operations lead at a mid-sized SaaS company, handling around 15TB/day of egress traffic. We've been running Akamai Prolexic for several years across our public API frontends, which are latency-sensitive for our customers in the financial data sector.

2. Here are the concrete points from our analysis when we evaluated this exact switch:
* **Break-even cost multiplier**: The flat always-on rate became cost-effective for us only when we projected 4 or more significant volumetric events per quarter. We averaged 2, similar to you, so it was a net cost increase.
* **Latency impact specificity**: The latency improvement you saw (~40ms) aligns with our tests. The gain is purely from avoiding the reroute to the scrubbing center during an event. For our non-event traffic, there was zero latency difference.
* **Negotiated scope**: Yes, we did negotiate a middle ground. We run always-on only on our /28 containing the critical API endpoints, not the entire /24 block holding our marketing and admin sites. This cut the proposed always-on cost by about 60%.
* **Required monitoring gap**: We had to implement our own real-time traffic anomaly dashboard. Akamai's reporting showed attack mitigated, but not the business impact (like error rate spikes for specific user cohorts). This data was crucial for our post-event reviews to assess true value.

3. I'd recommend sticking with on-demand unless you have a hard SLA for latency during attacks. The decision hinges on two things you haven't stated: your actual revenue loss per minute of increased latency during those events, and whether your contract has a minimum term for the always-on service.


Your bill is too high.


   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

You hit on a key point with the non-disruptive events. That cost multiplier for purely hypothetical risk is tough to swallow.

We did a risk/cost analysis and it centered on data integrity, not just latency. For us, the cost of a single corrupted data load during an unmitigated attack outweighed the always-on premium. Your mileage will vary massively if you're serving web pages vs. financial transaction streams.

On your question about negotiating with Akamai, yes, but it was a slog. We got them to agree to always-on for specific /32s (our API ingress points) while leaving less critical asset delivery on on-demand. It took threatening to run a POC with a competitor. Their default is definitely the blanket /24 coverage.



   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a really valuable point about risk calculations being based on the potential cost of a data integrity breach, not just latency or downtime. It shifts the whole equation.

Your experience with negotiating the /32 coverage is fascinating and something I suspect many teams haven't considered pushing for. The vendor's default stance on blanket coverage often locks people into a binary choice, when a hybrid approach for specific ingress points is what makes financial sense for a lot of architectures.


—HR


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

It's a good reminder that the hybrid approach user90 mentioned requires careful architectural mapping first. If your critical ingress points aren't cleanly separated from other services sharing the same IP ranges, that /32 carve-out becomes a non-starter. You often need internal routing changes to backhaul non-critical traffic, which adds its own complexity and potential single points of failure. It's powerful, but not a simple checkbox.


Review first, buy later.


   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

Absolutely. The architectural complexity point is crucial. I've seen teams get excited about the hybrid model in theory, only to realize their monolithic ingress design makes it impossible without a full refactor.

What often gets overlooked is the DNS layer. If you can't cleanly split at the IP level (/32), you can sometimes achieve a similar isolation by moving critical services to a dedicated subdomain (e.g., api.example.com) and applying always-on protection just to those DNS records. This still requires the backend routing to be separable, but it can bypass some of the internal network reconfiguration.

That said, the "potential single points of failure" you mention are real. Creating that dedicated, protected path often centralizes traffic through a narrower set of appliances or routes. You're trading a distributed risk for a concentrated one, which needs its own failure mode analysis.



   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You're right that DNS-based isolation is a clever workaround when IP-based carving isn't possible. The risk I've seen with that method is introducing a new failure mode if the DNS-based routing configuration gets out of sync with the underlying infrastructure, especially during incident response or scaling events.

It makes the failure mode analysis even more critical, because now you're managing complexity across two different layers.


Keep it civil, keep it real


   
ReplyQuote
(@ide_tinkerer)
Reputable Member
Joined: 6 months ago
Posts: 338
 

That 3.2X cost multiplier for non-disruptive events is the exact kind of number that makes my CFO twitch. It feels like paying for a full-time bodyguard because you occasionally get spam mail.

On your first question, about a "knee-jerk 'latency is king' move," I've seen that pressure come from product teams chasing SLA dashboards, not engineering. They see the latency graph dip during an event and flag it as a regression, without the context that it's a brief, mitigated attack. The value of that 40ms improvement is purely business-specific; for a lot of apps, it's just a vanity metric.

We ended up on a hybrid model, but we had to build a separate monitoring dashboard to prove its worth. It tracked "cost per mitigated request" for always-on vs. on-demand, which made the financial absurdity of protecting everything crystal clear. Maybe that's the play - quantify the premium you're paying per request during quiet periods. It's hard to argue with a huge number next to a tiny threat.


editor is my home


   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

> The value of that 40ms improvement is purely business-specific

This is so true. I've been reading this whole thread trying to learn, and that really clicked for me. My small team would probably chase that latency dip too, thinking it's a win.

That "cost per mitigated request" dashboard is a brilliant idea. It makes the trade-off concrete. I'm curious, was that hard to build? Did you tie it directly into your billing data, or was it more about estimating traffic volume during quiet periods?

It seems like the kind of data that could actually get product teams on board with a hybrid approach.



   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Your 3.2X cost increase for mitigating non-disruptive events is the core financial reality everyone must confront. You asked for a genuine risk/cost analysis, and I'd argue most teams skip the "risk" quantification entirely. They focus on the visible latency dip without modeling the actual business impact of those 40ms during an attack. For a trading API, that's quantifiable revenue risk. For a marketing site, it's likely not.

I've successfully negotiated the middle-ground you mention, but it required a granular traffic analysis Akamai didn't initially provide. We had to present our own data, mapping specific /32s to revenue-critical transactions. The key was proving that 80% of our attack surface was irrelevant to our core business logic. Their default posture is to sell you a broader net.

The monitoring gap you're sensing is the lack of a business-impact-aware dashboard. You need to move beyond pure latency graphs and build a "cost of mitigation per business transaction" metric. That's what finally convinced our product team that always-on for everything was an over-provisioned luxury.


show me the SLA


   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

Your point about the required monitoring gap is the most operationally critical one here. When we went through this, we found that Akamai's reporting had a 12-15 minute lag for traffic anomaly detection, which is an eternity when you're trying to differentiate a real DDoS from a sudden, legitimate spike. Building our own real-time dashboard wasn't just a "nice to have"; it became the primary tool for deciding whether to manually override and reroute traffic during an ambiguous event.

We used a simple time-series stream of our own ingress metrics compared against their mitigation logs. The cost of a false positive - flipping everything to scrubbing for what turned out to be a flash sale - ended up being more expensive than a few minutes of potential attack traffic. Your 60% cost reduction on the /28 carve-out is solid, but that self-built monitoring is what made the hybrid model actually safe to run. Without it, you're flying blind.


Show me the benchmarks


   
ReplyQuote