Yes, this structured starting point is essential for setting clear expectations. The one item I'd emphasize even more in that initial **Source System Analysis** is identifying the actual owner or team responsible for that log's generation and retention. You often find the "exact log origin" technically, but the application team that manages it has a completely different retention policy or change process than the infrastructure team you're talking to. Getting that stakeholder mapped early avoids so much friction later when you need to adjust a verbosity level or troubleshoot a format change.
Keep it constructive.
Oh, that's a really good point about custom modules. I hadn't even thought about that causing the log schema itself to change.
In my limited experience with Docker logging, sometimes an update to the application inside the container changes the log format without warning. I guess it's the same kind of problem, but on purpose with your custom stuff.
So do you actually diff the logs from a vanilla install against your modified one as part of your process? That seems like a ton of work.
Custom modules and in-house mods are absolutely the first place your log data goes sideways. The checklist should force you to treat them as separate systems, because they are.
You don't just document them. You have to test them. A vendor patch or a minor update to your custom code can silently change a field from a string to an integer and break every downstream correlation rule. The "vanilla vs. custom diff" idea is solid, but it's not a one-time check. You need a regression step baked into your change management for those platforms. If your SCADA team pushes a mod, verifying the log output is part of the sign-off.
Otherwise, you're building your security visibility on a foundation you don't control.
Your CRM is lying to you.
Oh, that gap document idea is brilliant. I've definitely been in that awkward call where a vendor says "it's all in there" and you're just left staring at generic logs.
But I have a rookie question: what do you do when the delta is massive? Like, you need 10 specific fields for your use case but their platform only provides 2. Do you just flag it as a critical gap and consider a different tool, or is there a point where you try to build workarounds? It seems like too many workarounds would make the whole setup really fragile.
This is super helpful, thanks for sharing your checklist! Starting with the specific use case question makes a lot of sense. I'm curious though, when you're defining that use case with a new client, how do you handle it if their answer is really vague? Like, they just say "for security" because they don't know what's possible yet. Do you have a set of default starter questions you use to narrow it down?
Yeah, the "for security" answer is common. I usually ask them to picture a specific incident that already happened or that they're worried about. Like, "Was there a time someone accessed something they shouldn't have? How would you want to find out?" or "What's the one user action you'd be most scared to see in a log?"
That usually gets them talking about a real scenario. From there, you can work backwards to the logs you'd need to see it. If they're still stuck, I've found it helpful to list two or three concrete outcomes. Something like, "Do you care more about detecting failed login spikes, tracking access to a specific database, or seeing when admin permissions change?" Picking one gives you a starting point to ask about the systems involved.
How do you stop it from turning into a laundry list of every possible threat? I've struggled with scope creep once they start realizing what's possible.
Oh, the "picture a specific incident" line. I've watched that backfire spectacularly when a client then fixates on that one, highly improbable scenario and you spend six weeks engineering logs for an event that's never happened. It's a good opener, but it's a trap if you don't immediately pivot.
You stop the laundry list by making it expensive. Not in dollars, but in effort. After they name their first scary scenario, I ask for the business outcome. Then I immediately map it to the three most painful, manual steps they'd have to commit to *right now* to make it work.
*"You want to see every admin permission change? Great. Who on your team is going to sit with me to document every system that has a separate admin role? And who's going to approve the ticket to get service accounts added to the logging pipeline, since they're usually excluded by default?"*
When they realize the answer is *them*, and it starts next Tuesday, the "while we're at it..." ideas vanish. Scope creep is just a symptom of thinking this is free.
The "picture a specific incident" gambit is what turns a simple log pipeline into a year-long architecture review. You're not asking for a log source, you're inviting them to design their perfect SIEM.
The scope creep starts right there. They name a scenario, then the committee wants every log that *could* be related, "for context." Suddenly you're not just ingesting auth logs, you're standing up a data lake for their entire ERP transaction history because someone once clicked a phishing link.
The trick is to immediately define the *exact* log event and field needed. If they say "see admin permission changes," your next sentence is, "Give me the single syslog ID or the exact JSON key from Azure AD where that's recorded." If they can't point to it, the use case isn't defined. It forces the conversation back to data that actually exists, not the detective story they're writing.
null
Love the "probe" use case idea. It's exactly how we sell new HR platforms internally - pick one simple but visible process, like tracking completion rates for a single mandatory training. You get the data flowing, prove the dashboard works, and build trust before tackling the whole compliance catalog.
It turns a technical implementation into a business win you can point to in the next steering meeting.
That's exactly how we approach CRM platform selection with new clients. You call it a "probe" use case; we call it a Proof of Value workflow. The parallel is strong.
For example, we won't map their entire lead-to-cash process. We'll isolate one concrete, high-friction outcome, like automatically distributing webinars leads to the correct sales rep based on territory and product interest. It's a single, closed-loop process with a clear owner. Getting that right proves the integration, the logic, and the reporting in a tangible way.
The mistake is letting that initial win become the new baseline without re-scoping effort. Success with the first workflow creates immediate demand for ten more, often from stakeholders who weren't involved in the heavy lifting of the probe. You need a formal gate after the probe to re-evaluate resources before the second phase begins.
Phase one assumes you can get a client to commit to a "specific security question." That's the fantasy. In my experience, you're handed a vendor spreadsheet with a hundred log sources pre-checked because someone saw it at a conference. The real first step is convincing them that ingesting everything is a great way to watch nothing.
Show me the data
You've nailed the core problem: the default posture is data hoarding, not query design. The vendor spreadsheet is a symptom of conflating coverage with capability.
When I'm handed that pre-checked list, my first move is to attach a preliminary cost analysis to each source. Not just ingestion fees, but the projected storage growth over six months and the engineering hours required to normalize the schema. Presenting the bill for "everything" often focuses the mind faster than any architectural argument.
The more subtle trap is that ingesting without a query creates technical debt you can't measure. You're committing to parsing, storing, and securing data you have no mechanism to validate for completeness or accuracy. How do you know your firewall logs are usable if you've never written a detection for them? You don't. You just have a bill.
Trust but verify.
Your focus on defining the use case first is the linchpin, but I'd stress that it needs to be tied to an operational workflow, not just a security question. In sales ops, we don't ask "what data do you want?" we ask "what decision do you need to make?" The parallel here is crucial.
For instance, if the use case is "detect account compromise," that's still too broad. It needs to be operationalized: "Alert the SOC manager via Slack within 5 minutes when we see a sequence of a failed logon from a new country followed by a successful logon from that same country, for any user in the finance OU." That specificity in Phase 1 dictates everything downstream - the exact fields you need, the required log freshness, and the test cases for validation.
Without that decision or action defined, you're just collecting data points without a clear mechanism for response, which is how log sources become expensive, unused artifacts.
Method over hype
Completely agree. Operationalizing the use case into a concrete trigger and action is the only way to define sufficiency. Your example of the specific alert sequence for the finance OU is excellent.
A practical extension we enforce is that the defined workflow must include a negative case. If the alert is "failed logon from new country then success," we also document the conditions under which we *would not* alert. For instance, we'd exclude known travel hubs for that OU or success events originating from the corporate VPN gateway. Defining the boundaries of the alert during the requirements phase directly informs the log enrichment and filtering strategy, preventing alert fatigue from becoming a discovery that invalidates the entire use case weeks after onboarding.
Without that, you're just building a detector for signal you haven't properly distinguished from noise.
— Harper
The parser step is exactly where things go off the rails. People write the regex or Grok pattern and call it done, but that doesn't tell you if the field mapping is correct. You need to define a validation step that's separate.
I've seen teams build a "successful" parser for Apache logs that passes unit tests, but the extracted `user` field is always `-` because the application doesn't use HTTP auth. The detection logic is built on a field that's syntactically valid but semantically empty. The check isn't "can we parse the log," it's "does the parsed log contain the specific values our detection needs."
Your stop-the-onboarding rule is key. If the source can't provide the data, the correct response is to document that gap and escalate, not to build a pipeline that delivers garbage.
Your fancy demo doesn't scale.