Skip to content
Notifications
Clear all

What is the best way to do capacity planning with Sysdig's data retention?

2 Posts
2 Users
0 Reactions
16 Views
(@marketing_ops_priya)
Trusted Member
Joined: 5 months ago
Posts: 41
Topic starter   [#2642]

Capacity planning with Sysdig often feels like a secondary consideration, overshadowed by its core security and monitoring use cases. However, its data retention policies directly impact your ability to forecast infrastructure needs and cost projections accurately. Relying solely on default retention windows can lead to reactive, rather than proactive, planning.

The most effective method leverages Sysdig's data export capabilities combined with external analytical tools. Sysdig's native dashboards are excellent for real-time and recent trend analysis, but for true capacity planning, you need longitudinal data that often exceeds the standard retention period for high-resolution metrics.

My recommended approach involves a three-step process:

1. **Identify Critical Time Series:** Not all metrics are created equal for capacity planning. Focus on exporting time-series data for:
* `cpu.used.percent` (per host/container)
* `memory.bytes.used` (per host/container)
* `disk.bytes.used` (per mounted volume)
* Key application throughput and latency metrics from your own services

2. **Establish an Export Pipeline:** Use the Sysdig Secure/Monitor API or a scheduled script to regularly export these specific metrics. The goal is to pull data at a daily or weekly aggregate level (e.g., 95th percentile, average, max) to reduce volume. Store this aggregated data in a cost-effective, queryable data store like a cloud data warehouse (BigQuery, Snowflake, Redshift) or even a dedicated time-series database you control.

3. **Model and Forecast:** With your historical dataset, apply standard forecasting models (e.g., ARIMA, simple linear regression) against the aggregated metrics. This external analysis allows you to:
* Predict when resources will be exhausted based on growth trends.
* Model the cost impact of adding nodes or increasing cloud instance sizes.
* Correlate infrastructure consumption with business metrics (e.g., user growth, campaign launches) for more accurate attribution.

A common pitfall is attempting to do this analysis solely within Sysdig's UI, which is constrained by its retention policy. The "best way" is to treat Sysdig as the authoritative real-time data *source*, but not the analysis *platform* for long-term capacity planning. This decoupled strategy mirrors best practices in marketing analytics, where raw engagement data from HubSpot is exported and modeled elsewhere for attribution and lifetime value forecasting.


Show me the data


   
Quote
(@startup_ops_lead_jen_2)
Eminent Member
Joined: 7 months ago
Posts: 16
 

Hey, I actually went through this exact process last quarter when we were trying to forecast our cloud spend. I'm the ops lead at a 25-person SaaS startup, and we run our core app on EKS with monitoring handled by Sysdig Monitor. We needed to plan our node scaling for the next six months.

Based on that project, here's how the "export and analyze" method plays out in reality:

**Data Granularity Loss**: The biggest catch is that exported data, in my experience, is aggregated. You often lose the container-level detail after 30 days, getting rolled up to pod or node-level metrics. For long-term planning that's usually fine, but you can't go back and do a deep dive on a specific container's growth trend from six months ago.
**Pipeline Effort Isn't Trivial**: Setting up the export isn't just a quick API call. To get it reliable, we had to write a script that handles pagination and Sysdig API rate limits, then ship it to S3, which took me and a part-time engineer about a week to build and test. The hidden cost is the ongoing maintenance of that pipeline.
**Tooling Choice Drives Cost**: Where you send the data changes everything. We used BigQuery because we're on GCP, and querying a year's worth of metrics costs us maybe $5-$10 a month. I looked at doing this in Datadog, but ingesting that volume of historical data there would have been prohibitively expensive for our budget, like adding hundreds to our monthly bill.
**Forecasting is Still Manual**: Even with the data in a database, Sysdig doesn't provide forecasting tools for this use case. We had to build simple linear regression models in Looker Studio dashboards ourselves. It works, but it's a custom solution that needs tending.

For a startup like mine where every dollar counts, I'd recommend the export-to-your-own-data-warehouse path. It's the only way to get the longitudinal data without a massive ongoing observability tax. My pick is to use the Sysdig API, land the critical metrics in a cheap columnar store (BigQuery, Redshift, Snowflake), and build your charts there.

That said, if you're already paying for a high-end analytics platform, tell us which one. And how much historical data do you really need - are we talking 6 months or 2+ years? The effort shifts a lot based on that.


learning daily


   
ReplyQuote