Skip to content
Notifications
Clear all

Complete newbie - can I run a meaningful test with less than 100GB of logs?

12 Posts
12 Users
0 Reactions
12 Views
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
Topic starter   [#25146]

I'm evaluating Anomali ThreatStream for my org. We're a mid-sized company, not a Fortune 500. The sales rep is pushing for a full-scale POC with months of data ingestion, claiming we need "at least 100GB of logs per day" to see any real value.

That smells like vendor nonsense to me. I've been through this with SIEMs and other security platforms before.

My question is straightforward: can I run a meaningful, technically valid test of the platform's core detection and investigation capabilities with significantly less data? Let's say 10-20GB per day for a two-week period.

I need to verify:
* The correlation engine actually works on our type of data (cloud infra, proxy, EDR).
* The UI and investigation workflows are efficient for my analysts.
* The data parsing and enrichment is accurate.

I'm not trying to model our year-long peak seasonal traffic. I'm trying to see if the product fundamentals are sound before we commit to a massive deployment project.

What specific tests or use cases would you recommend for a limited-data POC? What are the key pitfalls in setting up a small-scale test that would make the results misleading? I'm particularly concerned about them later saying "you didn't ingest enough, so of course feature X didn't work well."



   
Quote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Yes, you can absolutely run a meaningful test with 10-20GB daily. The sales requirement is excessive for validation. Focus your limited POC on data quality and analyst workflow, not volume.

Force them to use your proposed data volume. Have your team run specific investigations: trace a simulated phishing email through your proxy and EDR logs, or hunt for a specific IOC across all data sources. If the correlation breaks on small, targeted data sets, it'll fail at scale.

The main pitfall is letting them configure default, noise-heavy rules. You'll get flooded with meaningless alerts and they'll blame your data volume. Instead, define 3-5 precise use cases with known outcomes. If the platform can't accurately parse, enrich, and connect events for those, you have your answer.


—AF


   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

I completely agree with your instinct that 100GB/day is excessive for a proof-of-concept. You're right to focus on validating fundamentals over scale at this stage.

The key pitfall you mentioned is the biggest one: them blaming the data volume later. To prevent that, get them to sign off on your POC success criteria *before* you start, explicitly tied to your 10-20GB data set. Frame three concrete scenarios that mirror your real analyst work:
- A cloud IAM role assumption chain from your cloud trail logs.
- Tracing a specific user's web and EDR activity for a potential insider threat.
- Enriching and correlating a known-bad IP across proxy and firewall logs.

If the engine can't connect those dots in your small, clean data set, it won't magically work with more noise. Your plan is sound; just formalize the test cases and get vendor agreement upfront.



   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Absolutely. The formal agreement is the only way to protect your time. I'd take it a step further and make the cost of the POC resources part of that agreement.

If they insist you need 100GB/day to see value, ask them to provide the extra 80GB of synthetic or lab data at their expense to meet their own requirement. Their reaction to that request is often more telling than the test itself. If the core logic is sound, it should work on your 20GB. If it isn't, padding with their data won't fix it.

Also, benchmark the ingestion and processing costs for your 20GB/day volume during the POC. That's your real unit economics, not the hypothetical "at scale" number they'll quote.



   
ReplyQuote
(@cloud_cost_hawk)
Reputable Member
Joined: 3 months ago
Posts: 250
 

You're right, it's nonsense. They're trying to anchor you to a huge infrastructure bill from day one.

For a 20GB/day POC, test the data pipeline first. Feed it a known-bad IP and a known-bad hash. Your proxy, firewall, and EDR logs for that day should all connect those dots in a single timeline. If they can't, the correlation is broken.

The real pitfall is letting them use their own, pre-canned detection rules. You'll get a flood of false positives and they'll say you need more data to "tune" them. Write your own three specific correlation rules based on your defined scenarios. If the engine works, it'll fire on your small, curated data. If it needs 100GB of noise to function, the product is just a very expensive log dump.


cost optimization, not cost cutting


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Spot on about the pre-canned rules. That's where they get you.

I'd even take it a step further and ask them to disable ALL default detections for the POC. Force the test to be about the raw engine and your specific rules. If they push back, you know the value is in their prepackaged alerts, not the platform's core strength.

The cost anchoring is real, too. Great call.


Trial first, ask later.


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

Totally agree on disabling the defaults. That's the only way to test the platform and not their content team.

But here's a practical snag - their professional services team might fight you on it because their deployment playbook is probably built around turning on those default packs. They'll say you need the "baseline". Push back hard.

It also reveals their pricing model. If the core engine is weak and the value is in their managed detection rules, you're just renting a very expensive threat intel feed, which you could get elsewhere.


Spreadsheets > marketing slides.


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 3 months ago
Posts: 294
 

>If the core engine is weak and the value is in their managed detection rules, you're just renting a very expensive threat intel feed

Exactly. So run the POC to find out which one it is.

Force the issue on the defaults. If they insist their 'baseline' detections are mandatory, ask them to itemize the licensing cost for the platform engine separately from the threat intel feed and rule packs. They won't.

That's your answer.


Doubt everything


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

Agreed, that data requirement does sound excessive for validation. One thing I haven't seen mentioned yet is the data *mix*. For your 10-20GB test, make sure it proportionally represents all your critical sources - cloud, proxy, EDR. If you feed it 90% firewall logs because they're bulky, you won't test the correlation across sources properly.

Could you test the enrichment separately? Like, feed it a handful of known-bad IPs and domains from a recent threat report and see if the context it adds is actually useful for your team. That doesn't need volume at all.

What's their response when you push back on the 100GB?



   
ReplyQuote
(@calebs)
Reputable Member
Joined: 3 months ago
Posts: 318
 

The data mix point is critical. We had a vendor claim correlation was broken because our POC data was "unbalanced". It wasn't. Their enrichment engine couldn't handle sparse fields across sources unless the volume was huge.

>test the enrichment separately
Do this, but also check the latency. Feeding it a list of 10 known-bad IPs is a valid test. If the context pops up in 5 seconds, good. If it takes 5 minutes for the lookup to complete, it's useless during an investigation.

They haven't responded to the pushback yet. Their silence usually means the requirement is arbitrary.



   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Your instinct is correct, the 100GB requirement is often a scaling test masquerading as a validity test. You can absolutely assess the core functionality with 10-20GB/day.

The primary risk in a small-scale test isn't statistical validity like in an A/B test, it's representational validity. You must ensure your sample isn't *systematically* biased. A few critical checks:

* **Log Diversity, Not Just Volume:** Manually verify your 20GB/day includes proportional representative samples from each critical log source. Don't let it become 90% verbose firewall logs. Pull a day's worth of actual CloudTrail, proxy, and EDR logs, then calculate the mix. Replicate that mix.
* **Known-Bad Injection:** Seed your test data. Insert a controlled, low-volume IAM escalation sequence into your CloudTrail logs and a corresponding outbound call to a known-bad IP in your proxy logs from the same source IP. If the correlation engine can't stitch that together across 20GB, it's fundamentally broken. Volume won't fix logic.
* **Parsing Fidelity:** This is your silent killer. Export 100 random raw log lines from each source. After ingestion, query for those exact lines in the platform. Spot-check 10. Are all fields parsed correctly? A single mis-mapped field (e.g., `src_ip` going into a `destination` field) will cripple correlation at any scale.

The pitfall you mentioned - them later blaming the data - is mitigated by defining success as the platform correctly connecting your *injected* attack sequences and accurately parsing your sampled logs. If they say "you need more data to tune," your response is that tuning is a separate phase; detection of a known-bad pattern in clean data is a binary capability check.

If the engine works, it will work on your small, clean dataset. If it requires 100GB of noise to generate a signal, you're buying a statistics platform, not a detection platform.


p-value < 0.05 or bust


   
ReplyQuote
(@hannahm)
Reputable Member
Joined: 3 months ago
Posts: 217
 

Yeah, that requirement seems off to me too. For testing the fundamentals, 10-20GB a day sounds totally reasonable.

I'm curious about one thing people haven't mentioned yet. What happens if you test with *less than a day* of that 20GB sample? Like, a single hour's worth of logs where you've injected a known attack pattern? If the correlation engine can't connect the dots across your cloud, proxy, and EDR logs in that tiny sample, then throwing 100GB of data at it definitely won't fix the logic.

Wouldn't that be an even faster way to pressure-test the core engine before building out the full two-week pipeline?


Just my two cents.


   
ReplyQuote