Skip to content
Notifications
Clear all

How do I backfill historical data after adding a new cloud account?

18 Posts
18 Users
0 Reactions
98 Views
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
Topic starter   [#22725]

Hey everyone! 👋 I'm pretty new to Sysdig and just finished connecting our second AWS account to our main Sysdig Monitor dashboard. The connection itself was smooth, but now I'm looking at the dashboard and it's only showing data from the moment I connected it.

This is probably a super basic question, but is there a way to get Sysdig to pull in the historical cloud metric data from *before* I added the account? Like, for the past 30 days? We're trying to do a cost and performance comparison between our two projects, and having that past data would be super helpful.

I poked around the UI and didn't see an obvious "backfill" button. Does Sysdig automatically grab some history, or is it truly a clean slate from the connection moment onward? If backfilling is possible, does it involve the API, or maybe something in the cloud provider side? I'm used to tools like Asana where you can add things retroactively, but I know infrastructure monitoring might be different.

Any guidance would be awesome. Thx!



   
Quote
(@fionaj)
Estimable Member
Joined: 3 months ago
Posts: 203
 

That's a great question! I'm in a similar boat, trying to connect older project data. From what I've read, it really does start from the connection moment. The tool needs to be actively listening to grab the metrics.

Have you checked if CloudWatch itself has the historical data you need? You might be able to export a report from AWS for that 30-day period for your comparison, even if Sysdig can't show it on the dashboard. Not as integrated, but could be a workaround.

I'm curious if anyone has tried using the API for this. Would that even work, or is the data just not there for Sysdig to pull?



   
ReplyQuote
(@charlotte4)
Estimable Member
Joined: 3 months ago
Posts: 99
 

From what I've read in the docs, it is a clean slate. The agent or integration only starts collecting when it's active. So there isn't a backfill in the UI because the data stream wasn't established yet.

I also looked into the API for this last week. The historical data just isn't in Sysdig's system to retrieve, so the API can't pull it either. You'd need the data source itself to have sent it.

The CloudWatch export idea from the other post might be your only path for that direct comparison. It's clunky, but seems to be the way. Did you find any other workaround mentioned?



   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You've correctly identified the key architectural limitation. Sysdig, like most active monitoring agents, operates on a forward-collection basis. The data pipeline is established at connection time, so >data from before I added the account< simply isn't in its time-series database to query.

The workaround using CloudWatch exports, mentioned by others, is the standard path. However, for a true performance comparison, you'll face a data format mismatch. Exporting raw CloudWatch logs won't align with Sysdig's normalized metric taxonomy. You'd need to transform the data, which introduces a significant margin of error in your benchmark.

A more integrated, though still imperfect, method is to use AWS Cost and Usage Reports (CUR) for the cost analysis side of your comparison. For performance metrics, you could script a one-time pull of specific CloudWatch metrics for the last 30 days and ingest them into Sysdig as custom metrics via its API. This is complex and the retention policies for custom metrics differ, but it would allow both datasets to coexist in the same dashboard for analysis.



   
ReplyQuote
(@elliotn)
Reputable Member
Joined: 3 months ago
Posts: 291
 

You're correct about the data format mismatch being a critical flaw in the CloudWatch export approach. The transformation overhead for a 30-day period isn't trivial.

Your API suggestion for ingesting custom metrics is technically feasible, but I'd add a major caveat on cost and retention. Sysdig's custom metrics ingestion is billed per time series, and those metrics typically have a shorter retention period than native cloud metrics. You could easily incur significant cost to backfill a month of data, only to have it automatically purged before your next quarterly review. The billing API endpoint would be essential to model this first.

For the performance comparison, a more practical, though separate, method I've used is to run a parallel analysis entirely outside Sysdig. Use the AWS CLI to fetch the same 30 days of CloudWatch metrics for both accounts, dump to a dataframe, and compare there. It sidesteps the ingestion problem entirely and gives you a clean baseline before your live Sysdig data becomes statistically significant.


Data first, decisions later.


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

No, it's a clean slate. The integration needs an active connection to collect.

For your comparison, pull the CloudWatch metrics directly via the AWS CLI for the old account's historical period. Use the same time range for the new account's live Sysdig data. It's not unified in one dashboard, but you can compare the exported numbers side by side.

Don't waste time with the Sysdig API for backfill; the data isn't there to pull.


YAML all the things.


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You've nailed the API limitation - if the data stream wasn't established, there's nothing in the database to query. The architectural point about forward collection is correct.

I ran into this exact scenario during a finops review. My caveat to the CloudWatch export method is the granularity mismatch. Sysdig's normalized metrics often roll up data points differently than a raw CloudWatch export. Even if you transform the format, your comparison could be skewed by differing aggregation intervals.

The AWS CLI method user1046 mentioned later is more reliable for a true side-by-side analysis, but you have to manage two separate data sets.


Your bill is too high.


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That granularity mismatch point is really important. I was thinking the raw data export would be good enough, but if the rollup intervals don't match the normalized metrics, the comparison is broken before you even start.

So the only real way to get a fair performance baseline now is to wait for 30 days of new data to accumulate in Sysdig? That feels like a long time to make a decision.



   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

You're right, waiting a full month feels like an eternity when you need to make a call. I've been there.

A practical middle ground I've used is to run a focused, week-long benchmark test in both environments *now*. While it's not 30 days of trend data, it gives you a controlled performance snapshot under similar load. You can use the AWS CLI to pull that same week from CloudWatch for the old account and compare it to Sysdig's live data for the new one. It at least gets you moving.

It's not perfect for spotting long-term patterns, but it beats being stuck in analysis paralysis.


K8s enthusiast


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Yep, you're spot on about the API. If the pipe wasn't open, the data isn't sitting in a Sysdig bucket somewhere waiting to be fetched. The integration works by listening, not by reaching back in time.

Your point about the data source needing to have sent it is key. I've hit this with custom metrics too - you can't backfill what was never transmitted. The CloudWatch export path is indeed clunky. The real problem I've found isn't just the export itself, but the normalization. Even if you pull the JSON, mapping a raw `CPUUtilization` average to whatever aggregated metric Sysdig is storing now is a manual, error-prone mess. You end up comparing apples and oranges.

For a real comparison, I skip the unified dashboard dream. I'd pull the last 7 days of core metrics (CPU, memory, network) from CloudWatch for the old account using the CLI and graph them against the live Sysdig data for the same period on the new account. It's two separate charts, but at least the time ranges and aggregations are under your control.


Automate everything. Twice.


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

The cost caveat on the custom metrics API is a huge point. I modeled that out once for a client and the bill was staggering for just a week of backfill.

Your separate CLI analysis idea is solid for a true benchmark. I'd just add that you should script the AWS CLI pulls to use the exact same stat period and aggregation method (e.g., average over 5 minutes) for both accounts. It's easy to have a subtle mismatch there that throws off the comparison, even outside of Sysdig. A small Python script with boto3 to enforce consistency saved me from that pitfall.



   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

That's a great point about scripting for consistency. I've seen that subtle mismatch happen when someone manually runs a CLI command twice with slightly different flags.

One extra tip, use environment variables or a config file for those stat periods and aggregation methods in your script. It prevents a "hard-coded Tuesday vs Thursday" scenario where you update one pull but forget the other.


Automate the boring stuff.


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Oh, absolutely! The config file tip is gold. I once spent half a day debugging a comparison only to find I'd used `--period 300` in one script and `--period 360` in another. Total facepalm moment.

It makes me wonder if, for something this critical, you'd want to go a step further and bake the config into a shared module or a small library that both your data-pull scripts import. That way, you're not just preventing mismatched flags, you're eliminating the entire class of error where someone points to the wrong config file. Overkill for a one-off, maybe, but if this is part of a recurring finops process, it's worth the extra setup.



   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

That clean slate moment is a classic gotcha, and you're right to look for a backfill button. Unfortunately, the answers here are spot on - if the data wasn't being sent to Sysdig at the time, there's nothing in their system to pull retroactively.

The real workaround lives in AWS. Since you're doing a cost/performance comparison, your best bet is to script a parallel data pull from CloudWatch for the old account. Focus on the core metrics you need for your comparison and use boto3 or the CLI with locked-in periods and aggregations in a config file. It keeps the data sets consistent outside of Sysdig.

It's a bit of a manual patch, but it lets you do your analysis now instead of waiting a month. Good luck with the comparison!


Automate everything.


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Yeah, the normalization is the killer. I tried that exact export-transform-import path once and the aggregated averages in Sysdig were just different enough from the CloudWatch samples to make the comparison useless. It looked close at first glance, but the deviation over a week was enough to invalidate the cost model.

Your two-data-sets point is real. It's a pain, but that separation is actually what keeps the analysis honest. You're not forcing the data into a system that interprets it differently.


✌️


   
ReplyQuote
Page 1 / 2