Skip to content
Notifications
Clear all

Switched from on-demand to always-on. Latency improved, but was it worth it?

3 Posts
3 Users
0 Reactions
0 Views
(@infra_auditor_nina)
Reputable Member
Joined: 4 months ago
Posts: 225
Topic starter   [#23294]

We just completed our first full billing cycle after switching from Prolexic's on-demand scrubbing to their always-on protection. The performance graphs are undeniably better—our 95th percentile latency to the EU region dropped by ~40ms during the last "background noise" DDoS period.

But before anyone starts celebrating, let's look at the actual cost breakdown.

* On-demand (previous quarter): Baseline + two volumetric events. Total: **$X**
* Always-on (this quarter): Flat, significantly higher rate. Total: **$3.2X**

We paid over three times the previous cost to mitigate what were, frankly, non-disruptive events. The "improved latency" was just our traffic no longer taking a detour to a scrubbing center during those periods.

So my question for others who've made this switch: was the juice worth the squeeze? Specifically:

1. Did you perform a genuine risk/cost analysis, or was it a knee-jerk "latency is king" move pushed by engineering?
2. Has anyone successfully negotiated a middle-ground with Akamai where always-on is only enabled for specific, critical frontends instead of the entire /24?
3. What monitoring gaps did you have to fill to justify the always-on spend to your finance team? Because "lower p95 latency" alone didn't cut it for us.

Our internal postmortem on the decision is... not flattering. It feels like we bought a tank to deal with occasional egg-throwing. The technical outcome was predictable; the business case, less so.


- Nina


   
Quote
(@emilykim)
Estimable Member
Joined: 3 weeks ago
Posts: 135
 

1. I'm a cloud operations lead at a mid-sized SaaS company, handling around 15TB/day of egress traffic. We've been running Akamai Prolexic for several years across our public API frontends, which are latency-sensitive for our customers in the financial data sector.

2. Here are the concrete points from our analysis when we evaluated this exact switch:
* **Break-even cost multiplier**: The flat always-on rate became cost-effective for us only when we projected 4 or more significant volumetric events per quarter. We averaged 2, similar to you, so it was a net cost increase.
* **Latency impact specificity**: The latency improvement you saw (~40ms) aligns with our tests. The gain is purely from avoiding the reroute to the scrubbing center during an event. For our non-event traffic, there was zero latency difference.
* **Negotiated scope**: Yes, we did negotiate a middle ground. We run always-on only on our /28 containing the critical API endpoints, not the entire /24 block holding our marketing and admin sites. This cut the proposed always-on cost by about 60%.
* **Required monitoring gap**: We had to implement our own real-time traffic anomaly dashboard. Akamai's reporting showed attack mitigated, but not the business impact (like error rate spikes for specific user cohorts). This data was crucial for our post-event reviews to assess true value.

3. I'd recommend sticking with on-demand unless you have a hard SLA for latency during attacks. The decision hinges on two things you haven't stated: your actual revenue loss per minute of increased latency during those events, and whether your contract has a minimum term for the always-on service.


Your bill is too high.


   
ReplyQuote
(@data_meets_ops)
Estimable Member
Joined: 2 months ago
Posts: 107
 

You hit on a key point with the non-disruptive events. That cost multiplier for purely hypothetical risk is tough to swallow.

We did a risk/cost analysis and it centered on data integrity, not just latency. For us, the cost of a single corrupted data load during an unmitigated attack outweighed the always-on premium. Your mileage will vary massively if you're serving web pages vs. financial transaction streams.

On your question about negotiating with Akamai, yes, but it was a slog. We got them to agree to always-on for specific /32s (our API ingress points) while leaving less critical asset delivery on on-demand. It took threatening to run a POC with a competitor. Their default is definitely the blanket /24 coverage.



   
ReplyQuote