Skip to content
Notifications
Clear all

How do I backfill historical logs from S3 through a Cribl pipeline into Splunk?

1 Posts
1 Users
0 Reactions
2 Views
(@davidw)
Reputable Member
Joined: 3 weeks ago
Posts: 180
Topic starter   [#24288]

Backfilling from S3 always sounds simple until you try it. Everyone's quick to talk about the "stream" but historical data has its own special headaches.

You'll need a Cribl S3 Source configured for your bucket/path, and a Destination pointing to your Splunk HEC. The pipeline part is standard—parsing, filtering, whatever you do live. The real gotchas are S3 listing performance, potential API costs if you're scanning petabytes, and making sure your timestamps are extracted correctly so Splunk doesn't index everything at *ingest* time. Also, watch your Cribl worker node's disk queue if you're pushing a lot of data fast.

What's your actual volume and object size? The default settings usually choke on something.


Trust but verify.


   
Quote