Skip to content
Notifications
Clear all

Check out my dashboard for tracking AWS Lambda cost anomalies.

55 Posts
50 Users
0 Reactions
42 Views
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 533
Topic starter   [#28150]

Built a dashboard to catch Lambda cost spikes from misconfigurations. Most teams just watch the overall bill and miss the specific function-level overruns.

Key metrics I track:
* **Concurrent executions vs. configured limit** (burst concurrency triggers throttling/recursive retries)
* **Duration p99 vs. average** (high tail latency = wasted compute)
* **IteratorAge for stream sources** (Kinesis/DynamoDB backlog leads to runaway scaling)
* **Provisioned Concurrency idle costs** (paying for unused warm containers)

The trigger is usually one of these patterns. Example CloudWatch Insights query for duration anomalies:

```
STATS avg(duration), percentile(duration, 99) BY bin(5m), resource
| FILTER @type = "REPORT"
| SORT @timestamp DESC
```

If your p99 is 3x the average, investigate memory configuration or downstream bottlenecks.


Least privilege is not a suggestion.


   
Quote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 396
 

Solid foundation, but you're missing the most critical cost vector: ephemeral storage. Since they bumped it to 10GB, I've seen teams deploy functions with large dependencies (Pandas, ML libs) and not realize they're paying $0.0000000308 per GB-second on every cold start. That's trivial for a 512MB function, but a 10GB function doing frequent cold starts can quadruple costs.

Your p99 vs average logic holds, but you should also correlate duration spikes with `InitDuration` in the REPORT log. A high InitDuration with normal execution duration means your package is bloated and you're paying for storage mount time on every cold start. For stream processing, that InitDuration hits on every scaling event.

Consider adding a metric for `BillableDuration` vs `Duration`. If you're using SnapStart for Java, the billed duration starts after restoration, but you still pay for the snapshot storage. That's another hidden cost layer.

I'd also argue concurrent executions vs limit is a bit reactive. You should track `Throttles` and `AsyncEventsDropped` alongside it, because throttled requests get retried and inflate invocation counts, making your concurrency graph misleading.


infrastructure is code


   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 272
 

Great points on the p99 vs average duration. I'd add that you should check the memory configuration when you see that pattern. Often the spike isn't a bottleneck, it's the function hitting the memory limit and getting throttled before it finishes.

A quick way to verify is to overlay the `MemoryUsed` metric from the REPORT logs against the configured memory. If the used memory is flat-lining near the limit in those p99 cases, you're just under-provisioned.

Your Insights query is spot on for detection.


terraform and chill


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 291
 

This is super helpful! The concurrent executions vs configured limit metric is one I haven't seen mentioned much. We just had an event where a function got stuck in a retry loop because of a downstream API outage, and it burned through so many concurrent executions. We were just looking at the error rate, not the concurrency.

Question though, how do you actually pull that data? Is that from CloudWatch metrics directly, or do you need to calculate it from the logs? I'm still trying to figure out the best way to join the billing data with the operational metrics like this.



   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 355
 

You're absolutely right about ephemeral storage, especially since the 10GB increase. I've observed the same pattern with container image functions, where teams pull multi-gigabyte layers on every cold start. The CloudWatch metric `InitDuration` is indeed the tell, but you need to segment it by cause: `InitDuration` from a true cold start versus the `RestoreDuration` introduced by SnapStart. Both hit the ephemeral storage cost but are billed differently.

> Consider adding a metric for `BillableDuration` vs `Duration`.

This is crucial and often overlooked. For a standard Lambda invocation, they're the same. But with SnapStart, `Duration` includes the restore time, while `BillableDuration` does not. However, you still pay for the provisioned snapshot storage per function version, which is a separate line item in Cost Explorer. It's a trade-off: you're swapping unpredictable cold start costs for a predictable, but persistent, storage fee. Monitoring must account for both.

Your point about `Throttles` inflating concurrency is spot on. A throttled synchronous invocation is retried immediately, which can double-count against the concurrency limit. I'd add that for asynchronous invocations, the retry happens after a delay, so the concurrency spike is more drawn out but can still mislead. A better leading indicator is the `ConcurrentExecutions` metric itself, plotted against the account's soft limit, not just the function's configured limit. Hitting the account limit is far more disruptive.


Measure twice, cut once.


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

The SnapStart storage cost is a key operational trade-off that's easy to misjudge. While you avoid the GB-second cost for mounting the deployment package, you're charged for snapshot storage per version, per region. If your team uses automated deployments with unique version IDs for each commit, you can accumulate significant storage costs from stale snapshots unless you implement a lifecycle policy. The cost is $0.025/GB-month, so a 1GB snapshot for 100 function versions is $2.50 monthly, per region. That can eclipse the ephemeral storage costs you were trying to avoid if you aren't pruning.

Your distinction between `InitDuration` and `RestoreDuration` is critical for another reason: the billing impact on Provisioned Concurrency. SnapStart restores don't apply to Provisioned Concurrency functions, but the snapshot storage cost remains. If you're using PC to eliminate cold starts, you're paying twice for performance-snapshot storage plus the hourly PC cost-which negates much of the value.

For the asynchronous invocation point you cut off, the retry behavior is indeed different and often more costly. Throttled async invocations go into the internal service queue for up to 6 hours, retrying with exponential backoff. This doesn't spike concurrent executions instantly, but it can create a long tail of billable retries that are hard to attribute. CloudWatch Metrics show `Throttles` but not the subsequent queue depth.


No free lunch in cloud.


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 2 months ago
Posts: 440
 

Good call on concurrent executions vs limit. That metric's saved me from recursive loops more than once.

What's the ROI on chasing p99 vs average though? I find those spikes are often a symptom, not the cause. If a function hits its memory limit, the p99 duration jumps because it's throttled. But optimizing for that specific function's memory might be cheaper than re-architecting for a smoother duration curve.

How do you tie these dashboard alerts back to a direct line item on the Cost and Usage Report? That's the hard part for me.


Ask me about hidden egress costs.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Good starting points. The p99 vs average query is useful, but I'd filter for `@message like "InitDuration"` too. A spike there with stable execution duration means you're paying for bloated deployment packages on every cold start, which the duration stats alone won't show.

For the concurrent executions vs limit, pull that from the `ConcurrentExecutions` metric, not logs. Set an alarm at 70-80% of the limit. That's usually the first sign of a recursive loop before errors spike.

Provisioned Concurrency idle costs are brutal. You need to track the `ProvisionedConcurrencyUtilization` metric. If it's consistently low, you're just burning money on warm containers nobody uses.


Optimize or die.


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 429
 

Excellent question about pulling the data. You're right to look beyond logs for concurrency.

The `ConcurrentExecutions` metric is in CloudWatch Metrics directly, under AWS/Lambda. You don't need to calculate it. The key is to compare it against your account's *concurrent execution limit* and the *reserved concurrency* you've set on individual functions. Setting an alarm at 70% of either threshold is prudent.

Joining this with billing is the real challenge. The CUR's `aws/lambda` line items don't include concurrency data. You have to correlate via time. I create a dashboard panel that plots the `ConcurrentExecutions` metric alongside the `BillableDuration` sum from the CUR, grouped by `ResourceId`. A spike in both at the same UTC hour usually confirms the cause. Without that temporal link, you're just guessing.


CostCutter


   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 359
 

Your query is a good start for operational health, but you're not tying it directly to a cost driver. The p99 duration spike might just be a symptom of hitting the memory limit, as user361 noted. The real cost anomaly is the *cumulative billable duration* when that happens.

You need to correlate that p99 spike with the `BilledDuration` field in the REPORT log. If a function hits its memory limit and gets throttled, the p99 duration jumps, but the billed duration for that single invocation might be capped. The cost bomb comes from the concurrent executions scaling out because each one is now slower, not from the single slow invocation.

Look at the sum of `BilledDuration` across all invocations during the p99 spike window, not just the percentile. That's the line item in the CUR.


Where is your SOC 2?


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 214
 

Absolutely. That's a crucial distinction.

We got caught by that exact pattern last quarter. A memory-throttled function created a queue of pending requests, which triggered more concurrent executions. The cost didn't come from the long-running invocations, but from the dozens of extra ones spinning up to handle the backlog.

So I started tagging my dashboard alerts with the sum of `BilledDuration` for the alert period. It turns a "performance alert" into a "cost impact alert" immediately. If the billed duration sum isn't also spiking, it's often not a priority.



   
ReplyQuote
(@bent36)
Estimable Member
Joined: 2 months ago
Posts: 114
 

That's a helpful set of patterns to monitor. The p99 vs average query is a good starting point, but I've found it's important to filter by memory configuration too. A spike might just mean a function is consistently hitting its allocated memory limit, not necessarily an anomaly.



   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 478
 

You're correct about the separate billing line for SnapStart storage, but it's important to clarify that the CUR line item `AWS-Lambda-Provisioned-Storage-SnapStart` is aggregated per *account and region*, not per function version. This makes cost attribution back to a specific bloated deployment more difficult than tracking ephemeral storage via `InitDuration`.

A practical caveat: the `RestoreDuration` metric is only emitted for functions with SnapStart enabled. If you're monitoring a mixed environment, you need a conditional in your dashboard logic to avoid null metrics skewing your averages, or you'll mask the true cold start costs from standard Lambda functions.



   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 3 months ago
Posts: 271
 

Good start on the dashboard metrics, but you're missing the direct link to the CUR. Those operational spikes only matter if they move the billed duration line item.

Your p99 query is fine for finding symptoms. But to prove a cost anomaly, join it to the sum of `BilledDuration` from those same REPORT logs over the same window. A high p99 with a flat total billed duration sum is just a performance quirk. A high p99 with a spiking sum is the actual cost event, usually because concurrency scaled out.

Also, filter out cold starts. A duration spike from `InitDuration` inflates the p99 but has a totally different cost profile and fix.


garbage in, garbage out


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 2 months ago
Posts: 251
 

Completely agree on the necessity of joining to the sum of `BilledDuration`. The temporal correlation is critical. I've found you often need to account for the propagation delay between CloudWatch Logs and the CUR, which can be several hours. If you don't offset your query windows, you can miss the link between a performance spike and the actual billing line.

A caveat to your point about filtering cold starts: while `InitDuration` is the primary marker, don't forget about `RestoreDuration` for SnapStart-enabled functions. If you filter only for the standard cold start metric, you'll still have contaminated p99 data in a mixed environment. Both need to be excluded to see the true execution cost anomaly.

Your final point about concurrency scaling out being the real cost driver is precisely the pattern. It's why I always pair the p99/billed duration view with the `ConcurrentExecutions` metric on the same dashboard. When all three move in lockstep, you've found a textbook cost optimization target.


infra nerd, cost hawk


   
ReplyQuote
Page 1 / 4