Successfully onboarding a new log source into a SIEM is often the most critical yet under-documented phase of detection engineering. A rushed or incomplete integration directly undermines data quality, leading to unreliable alerts, missed detections, and increased operational overhead during investigations. Over numerous client engagements, my team has refined a methodical, four-phase checklist that prioritizes context and actionability over simple volume ingestion.
This guide outlines our standardized process, moving from initial assessment to full production monitoring. The goal is to ensure each log source provides validated, structured, and useful data for security operations.
### Phase 1: Pre-Ingestion Assessment & Requirements Gathering
Before any connector is configured, we conduct a business and technical assessment.
* **Define the Use Case:** Document the specific security questions this log source must answer. Is it for user behavior analytics, threat hunting, compliance evidence, or application-specific attack detection?
* **Source System Analysis:**
* Identify the exact log origin (e.g., specific cloud service, on-prem server version, appliance model).
* Determine the log generation mechanism (syslog, agent, API pull, file export).
* Review the vendor's logging documentation for event codes, sample formats, and volume estimates.
* **Log Sampling:** Obtain a statistically significant sample of raw log data (minimum 24 hours, ideally covering a business cycle). This is non-negotiable.
### Phase 2: Log Analysis & Parsing Specification
Using the log sample, we perform a hands-on analysis to define the parsing logic.
* **Field Mapping:** Identify and map all relevant fields to a normalized schema. We create a mapping document that aligns source fields to common information model (CIM) or internal schema fields (e.g., `user`, `src_ip`, `event_id`, `result`).
* **Identify Critical vs. Redundant Data:** Not all fields are necessary. Flag fields that are essential for detection (timestamps, IDs, actions) versus those that are merely informational or redundant.
* **Develop Parsing Logic:** Draft the exact parsing rules (regex, key-value, JSON, CSV) for the chosen SIEM or parsing pipeline. We prototype this in a controlled environment.
```python
# Example: Prototyping a regex for a custom application log
import re
sample_log = '2023-10-01T12:00:00 host=web01 user="john.doe" action="file_upload" file="report.pdf" status=OK size=2048'
pattern = r'^(?Pd{4}-d{2}-d{2}Td{2}:d{2}:d{2})s+host=(?PS+)s+user="(?P[^"]+)"s+action="(?P[^"]+)"'
match = re.search(pattern, sample_log)
if match:
print(match.groupdict()) # Validate extracted fields
```
* **Define Event Breaks & Correlation IDs:** Determine how to identify discrete events and any inherent identifiers for session or transaction correlation.
### Phase 3: Staged Ingestion & Validation
We never go directly to production ingestion. A staged approach is used.
1. **Non-Production Ingestion:** Ingest logs into a development or staging SIEM instance.
2. **Parsing Validation:** Run the parsed logs against the mapping document. Verify field extraction, data types, and normalization. Check for parsing errors on edge cases.
3. **Volume & Performance Baseline:** Monitor the impact on the SIEM ingestion pipeline. Confirm the volume aligns with estimates and does not exceed licensed or performance thresholds.
4. **Tagging & Classification:** Apply necessary metadata tags (sensitivity, department, environment) at ingestion time.
### Phase 4: Use Case Activation & Operational Handoff
The final phase ensures the log source delivers security value.
* **Develop Initial Detection Rules:** Create at least one high-fidelity detection rule or correlation search based on the defined use case. Test it against historical data in the staging environment.
```sql
-- Example: Skeleton of a detection search for failed administrative logins
index=main sourcetype="new_app_auth" action="login" result="failure"
| stats count by user, src_ip
| where count > 5
```
* **Documentation:** Update the internal runbook with:
* Source owner and contact.
* Onboarding details (parser location, ingress endpoint).
* Sample queries and detection examples.
* Troubleshooting steps for common log collection failures.
* **Alert Integration:** Define how alerts from this source will be triaged and integrated into existing SOAR playbooks or analyst workflows.
* **Schedule a Review:** Set a 30- and 90-day review to assess log source stability, detection efficacy, and to refine or expand use cases.
This structured approach significantly reduces rework, ensures data quality from day one, and aligns log ingestion directly with detection engineering objectives. The key is treating logs as structured data products rather than simple event streams.
This is a strong start. I've seen too many projects skip that use case definition and just start piping logs.
One question from an outsider's perspective: how often do you run into a situation where the client's stated use case, like "user behavior analytics," doesn't match the actual log fidelity available from their specific source system? You mention identifying the exact origin, but I'm curious about the gap between the ideal and the possible, and how you adjust the checklist when you find it.
Your question about the use case versus log fidelity gap hits on a real operational tension. We see it frequently, particularly with legacy systems or niche SaaS applications where the vendor's logging is an afterthought.
The adjustment happens in the requirements phase. When we identify a fidelity gap, we don't downgrade the use case. Instead, we document the exact delta - for instance, a UBA requirement needs user IDs and action types, but the source only provides IP addresses and generic success/failure. This becomes a formal gap document for the client, outlining the detection blind spot and proposing compensating controls, like a parallel log source or a change request to the application team. It shifts the conversation from "we can't" to "here's the risk and the path to mitigation."
This often uncovers a deeper issue: the business purchased a tool for a capability its logs can't actually support. Having that data-driven gap analysis prevents endless cycles of trying to make poor logs fit a premium use case.
Latency is a liability
That gap documentation is crucial for another reason: it creates an artifact for capacity planning. If a compensating control requires a secondary log source or a custom parser, you've now quantified the incremental ingestion volume and processing overhead before you build anything. It prevents the "why is our log storage bill up 40% this month?" surprise.
We also log the gap document's ID as a tag on the eventual SIEM source configuration. During alert tuning or incident review, seeing that tag immediately signals to an analyst that the data has known limitations, which stops them from chasing ghosts.
Sounds good in theory. But your checklist assumes you're being brought in *before* the procurement signature. In my experience, the log source is often already bought, deployed, and on a premium license tier before anyone asks if the logs are usable.
The real first phase is checking the vendor contract for data export rights and any egress fees for pulling logs to your SIEM. You'd be surprised how many "cloud-native" platforms charge extra for API access or charge per GB for log forwarding. That gap document doesn't mean much if the cost to fill it breaks the budget.
Read the contract
Documenting that delta is exactly how you avoid the "magic black box" problem some vendors sell. I've had to pull gap documents out during vendor escalation calls when they claim their platform "fully supports" a use case, but the logs they emit are just timestamps and generic event codes. Having the specific fields you need versus what they provide in writing shuts that down fast.
The compensating controls part is where it gets real. Tagging the gap doc in the source config is good, but I also add it to the runbook for any related alert. That way the on-call engineer isn't just told the data is limited, they're told *what specifically to check elsewhere* when the alert fires.
Sleep is for the weak
This is incredibly detailed, and I appreciate you sharing the structured approach. I work in a manufacturing environment where we're finally getting serious about SIEM for our industrial control systems and ERP. The emphasis on "context and actionability over simple volume ingestion" really hits home.
In our case, a major hurdle I see is that initial source analysis. We run older, heavily customized versions of systems like NetSuite and our SCADA platforms. You mention identifying the exact origin, like the server version. Does your checklist include a step for documenting custom modules or in-house modifications that might completely change the log schema from the vendor's standard? I've found that's where the first major data fidelity gap usually appears for us, long before we even think about parsing.
Oh, custom modules. That's the best part. The vendor's own documentation becomes useless and you're left reverse-engineering some internal dev's logging from 2012.
Your gap document better have a whole section titled "Undocumented Customizations." I've seen a "successful login" log from a bespoke module that just spits out "OK" with no user or source IP. Good luck building a detection on that.
The fun starts when you ask the app team for the schema and they say "We don't know, the dev who built it left five years ago."
—aB
Exactly. We always hit that first in the assessment. "What server version?" leads to "What custom plugins are installed?" and finally "Who modified the logging config last year?"
We add a field in our source analysis template for the "last dev to touch the logs." If that person left, we flag it as a high risk for log parsing failure. It's saved us from building parsers against stale sample data more than once.
Automate everything.
Good flag. We also tag that "stale sample data" risk to the data pipeline's monitoring. A spike in parsing errors from a source with that flag triggers an immediate pause and review instead of letting it silently fail and consume budget.
cost per transaction is the only metric
Starting with use cases is the right call. We push it a step further by linking each one directly to a git branch and a corresponding test in our CI pipeline. For example, a "compliance evidence" use case gets a branch with the required log samples, and the pipeline validates they parse and contain the mandated fields. It forces the requirement to be machine-readable from the start.
git push and pray
I absolutely love the idea of linking use cases to git branches and CI tests. It turns a static document into a living part of the deployment pipeline.
My team tried something similar for a PCI DSS requirement around admin access logging. We built the test case first, with the exact field list, and then gave that spec to the application team. It forced a conversation about data availability much earlier than usual.
The only hiccup we ran into was managing the test data when the source system's log format drifts over time, like after a patch. How do you handle versioning or updating those log samples in the pipeline when the source changes?
Pipeline is king.
We treat those log samples in CI like any other test dependency. When we get a patch notification for a source system, that triggers a scheduled task to pull fresh log samples and run a diff. If the format changed, it fails the existing tests and we get a ticket.
The key is building a relationship with the app team so they know to alert you *before* a logging change hits prod. We've had success adding them as reviewers on the pull request to update the test branch. That way, the conversation about any missing fields or parsing issues happens in the PR comments, not after a detection breaks.
Data is sacred.
Starting with use cases is fine, but teams skip the next step. Define what a "successful" detection looks like before you write a parser. If the use case is "detect account takeover," you need the user field populated and a reliable IP address. If the source can't provide that, stop the onboarding. You've just saved weeks of work on a detection that will never fire correctly.
Too many projects ingest first and ask questions later. That's how you get 200TB of logs that answer no security questions.
Trust, but audit.
You're overcomplicating it. If a patch changes your log format, your detection is already broken. The CI test is just telling you the sad news.
That "relationship with the app team" is a fantasy for most of us. By the time they review a PR, the update's already in prod. I treat log format as a hard contract. If they change it without telling me, their logs get dropped until they provide the new schema. Costs less than maintaining a versioned museum of samples.
Stops the drift real quick.
CRM is a means, not an end.