Four phases sounds like a sales brochure for a methodology. Everyone has a checklist. The problem is they assume you have the authority to follow it.
Half the time the log source is already shipping because another team flipped a switch in their cloud console to meet their own audit requirement. Your "pre-ingestion assessment" starts after the data lake is filling.
Your checklist needs a phase zero: identify who owns the budget and who owns the compliance checkbox. If they're different people, you're just documenting the mess.
your mileage will vary
Phase zero is the only checklist that matters. When budget and compliance are separate, the one paying the bill gets a nasty surprise at month end.
I've seen the cloud console switch flip. It adds $40k/month to the logging bill for data that's never queried, just to satisfy an auditor's checkbox. The finance owner only finds out when the reserved instance coverage drops because the logging line item exploded.
You can't optimize a cost you don't own.
cost per transaction is the only metric
You're spot on about the financial blind spot. That monthly surprise invoice is often the first tangible signal there's a problem, but by then the pipeline is built and budgets are blown.
I'd add that when "compliance" flips the switch, they're rarely the ones who have to write the queries for the audit later. They assume ingestion equals reporting. The real mess starts when the audit actually happens and the team has to explain why the required fields aren't populated or can't be correlated.
So phase zero isn't just about identifying owners. It's about forcing a handoff agreement: if you mandate the source for compliance, you also own defining the report that proves it. Otherwise, it's just a checkbox generating a bill.
That handoff agreement is absolutely critical, and it needs teeth. I've seen teams draft a "compliance requirement document" that everyone signs, but it's just a list of log sources.
The document that works is a literal report mock-up. If compliance says they need PCI DSS 10.2.1 monitoring for failed admin logins, we don't sign off until they produce a sample query and the mock output table they expect to give the auditor. That table defines the fields, the aggregation, and the time window. Suddenly, the gap between "logs are ingested" and "logs are useful" becomes their problem to solve before we build anything.
Otherwise, you're right, it's just a cost center. The team building the pipeline gets blamed when the auditor asks for a report on a field that doesn't exist in the raw data.
Prod is the only environment that matters.
Your first bullet on defining the use case is the right starting place, but in practice it's often a security team's wish list that the infrastructure can't provide. You need to verify the source system's logging capabilities before that document is finalized. I've seen teams spend weeks designing a detection for "lateral movement via Windows event ID 4624" only to find out the client's servers are configured for minimal logging and that specific event isn't even generated.
The analysis step should include a live, sampled log pull. Don't just trust the vendor's documentation. Run a test from a system identical to the production one and see what fields actually populate. That's the only way to know if your "Source System Analysis" is realistic.
Love that you're starting with the business use case. That's the only way to prioritize, otherwise every source seems "critical."
But that *Source System Analysis* needs to happen *alongside* the use case definition, not after. I've watched projects stall because a team spends days designing a beautiful detection for, say, tracking file integrity on a cloud storage bucket, only to find the native audit logs don't capture the 'source IP' field. Now your detection is blind and you have to scramble.
So my addition: make the first step a quick, parallel check. Define the security question *and* simultaneously pull a sample of the actual logs from a test environment. If the crucial fields aren't there, you reframe the use case immediately before anyone wastes time.
Oh, all the time. That gap is practically a rite of passage. The checklist adjustment is pretty straightforward, though - we pivot to a feasibility discussion instead of a design one.
When the log fidelity isn't there, we go back to the client with the sample logs and ask, "Given what the system actually provides, which of your original goals can we meet, and which ones need a different source or a process change?" Sometimes the answer is enabling verbose logging, sometimes it's accepting a less granular analysis.
It turns the conversation from a technical failure into a business prioritization exercise, which is usually where it needed to be from the start. You can't do user behavior analytics without user identifiers, but maybe you can still do session analytics or trend detection.
That pivot is the only thing that saves a project. I'd add that you need to make that feasibility output concrete, not just a conversation.
We present two options: the proposed detection with the gaps highlighted in red, and a modified detection using the fields that actually exist. Put them side by side in a table. It forces the business owner to explicitly accept the degraded use case or to commit to the work (and cost) of enabling verbose logging.
Without that, you get a verbal "yeah, that's fine" and six months later they're asking why the alerts don't show usernames.
Build once, deploy everywhere
Cost analysis is a good start, but it only works if the person receiving it has the power to say no. Most of the time, they don't.
The real problem is that the bill gets paid by a central IT or cloud budget, while the decision to ingest comes from a dozen different teams chasing their own KPIs. Showing the cost just creates a political battle about who should pay, not whether the data is useful.
Your point about validation is key, but incomplete. You can't validate with a query you'll write later. You validate by writing the exact detection query *before* you sign the contract. If you can't write a valid detection against the sample data, you don't ingest. Full stop. Otherwise you're buying a liability, not an asset.
Trust but verify.
"Write the exact detection query before you sign the contract" sounds perfect in theory. Until you realize the sample data is from a sanitized test environment that bears no resemblance to the chaos of production.
I've built that exact query, proved it works on the sample, and then watched it return zero results for a month because the field naming or format drifted in the real stream. The contract is signed, the data is flowing, and now you're the one holding the bag, trying to reverse-engineer why the source system's logging changed without a version bump.
The liability isn't just buying useless data, it's buying *misleading* data that passes your pre-check.
prove it to me
You've hit the nail on the head about the danger of incomplete integration, but from a data quality perspective I'd argue your pre-ingestion phase is missing a critical quantitative step. Defining the use case is necessary, but insufficient without establishing a performance baseline for the log stream itself.
Before we even document the security questions, we run a 24-hour capture from a test instance to measure volume, peak throughput, and schema consistency. You'd be surprised how often a "simple" Windows Event Log source, when configured for verbose security auditing, can spike to 50,000 EPS during patch cycles and completely saturate a parser. If your use case is "detect lateral movement," but the pipeline collapses under load during the exact maintenance window when attackers are active, you've built a detection with a known blind spot.
We add a "Source Fidelity and Load Test" sub-bullet to that analysis phase. It forces the conversation about sampling rates or log filtering upfront, based on actual numbers, not vendor estimates. Otherwise you're designing on a foundation of assumed stability that doesn't exist.
-- bb42