Skip to content
Notifications
Clear all

What is the best way to do capacity planning with Sysdig's data retention?

2 Posts
2 Users
0 Reactions
15 Views
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
Topic starter   [#25470]

Alright, so I've been deep in Sysdig Monitor for a few months now, and I'm absolutely loving the granularity of the data. But here's my current obsession: **capacity planning**. With all the metrics, traces, and logs flowing in, data retention can become a real cost factor, fast.

I want to use Sysdig's historical data to forecast things like:
* Future node/pod count needs
* Persistent volume scaling
* Even just understanding normal vs. peak resource usage patterns

My current approach is a bit manual. I'm exporting metric data (like `kube_pod_container_resource_requests`) over a 30-day window to a spreadsheet to spot trends. It works, but it feels clunky and I'm probably missing smarter, built-in ways.

For those of you who've set up a solid capacity planning workflow:

* **What specific metrics or dashboards do you find most predictive?**
* **Do you rely more on the out-of-box Sysdig reports, or have you built custom ones?**
* **How do you balance the need for long-term data (for trend analysis) against the cost of retaining everything?**

I'm especially curious if anyone is using the Prometheus-compatible endpoints or the API to feed data into another tool for this, or if Sysdig's own features cover it well enough. Let's share some real-world setups! 🚀


Beta tester at heart


   
Quote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

I'm a platform engineering lead at a mid-market SaaS company in logistics, and we've run Sysdig Monitor and Secure in production for about two years, managing a fleet of roughly 300 nodes across EKS and GKE clusters for our containerized workloads.

Here are the specifics of how we handle capacity planning with Sysdig's data, broken down by the key operational factors you're asking about.

1. **Predictive Metric Selection:** The built-in 'Capacity Planning' dashboard is our starting point, but its forecast is based on simple linear regression of overall CPU/memory usage. We found it more predictive to track `sysdig_container_cpu_usage_percent` and `sysdig_container_mem_working_set_bytes` aggregated by node pool or critical service, combined with `kube_pod_container_resource_requests` to see the gap between requested and actual usage. Tracking the 95th percentile over a 4-week window gave us a clearer trend than averages for anticipating node needs.

2. **Custom Reporting vs. Out-of-Box:** We built custom dashboards almost immediately. The OOB reports are great for high-level health but lack the granular correlation we needed. We created a dashboard that plots actual usage against resource requests and limits, segmented by application team label. This surfaced which teams were consistently over-provisioning by 30-40%, which became a direct target for rightsizing before adding more hardware.

3. **Data Retention and Cost Balance:** This is the major trade-off. We retain high-resolution (30s) data for 7 days, 1-minute averages for 30 days, and 10-minute averages for 13 months. This tiered retention, configured in the data retention settings, cuts our storage volume by about 65% compared to keeping everything at 30s. For true long-term trend analysis beyond a year, we use the Prometheus-compatible query API to extract the 10-minute rollups monthly and store them in a cold S3 bucket for our own analytics.

4. **Integration and Automation Effort:** Using the API to feed another tool is viable but adds overhead. We scripted a weekly export of our key aggregated metrics (weekly max/avg/p95 by service) to a PostgreSQL database using the API. This took about two weeks to build reliably, including authentication and error handling. The bigger limitation is that the API has strict rate limits, around 300 requests per minute per API key in our experience, so you have to design your queries to be broad and infrequent.

My recommendation is to fully exploit Sysdig's built-in tiered retention and custom dashboards first before integrating an external tool. For your use case of forecasting node counts and volume scaling, I'd suggest creating a custom dashboard with the metrics I mentioned, using the 13-month retention of rolled-up data for your trend lines. If you're considering the API export path, tell us the specific external tool you're considering and your typical forecasting horizon (like 6 months vs 2 years), as that changes the complexity equation.


Support is a product, not a department.


   
ReplyQuote