Another day, another SIEM that can't parse its own recommended data sources. I've been feeding Chronicle Office 365 Management Activity logs via the recommended Pub/Sub method for months. Suddenly, about two weeks back, the parsing for a subset of events just... stops. UDM events show up with a persistent `metadata.event_type: "GENERIC_EVENT"` and the raw log is tucked away in `about.raw_log`. The parsing error field is a masterpiece of uselessness: `"parsing error"`.
My pipeline is solid—because it's a simple, self-hosted collector pushing to the API. No magic, no third-party vendor black box. The logs haven't changed format. Google's own documentation on the expected schema is, as usual, a labyrinth of versioned pages where you hope you're reading the right one.
I opened a support ticket. The response cycle is a thing of beauty:
1. Request raw logs (sent).
2. "Our engineers are investigating."
3. Silence for 5 business days.
4. A follow-up asking if I'm still experiencing the issue.
Amusingly, yes, I am. The broken parsing is now a permanent fixture in my instance.
Is anyone else seeing this with O365 Audit.General data, particularly `SharePointFileOperation` events? I've resorted to writing a workaround in my ingestion pipeline to catch and reprocess these, which defeats the entire point of paying for a managed service.
Here's a sanitized snippet of the malformed event as it lands in Chronicle:
```json
{
"metadata": {
"event_timestamp": "2024-06-15T10:30:00Z",
"event_type": "GENERIC_EVENT",
"vendor_name": "Microsoft",
"product_name": "Office 365"
},
"about": {
"raw_log": "{"CreationTime":"2024-06-15T10:30:00Z","Operation":"FileDownloaded","Workload":"SharePoint",[...]"
},
"security_result": [{
"rule_name": "parsing error"
}]
}
```
The real kicker? The `raw_log` is perfectly valid JSON that matches their published schema. The parser just decided to take a holiday.
null
I hit the same wall with SharePointFileOperation events around the same timeline. My logs show the parsing fails specifically when the event's `SiteUrl` field contains a certain path depth or a special character that, historically, never caused an issue.
You mentioned the support loop - I've had better traction by bypassing the initial ticket and linking directly to a comparison spreadsheet in subsequent replies. I charted correctly parsed events against the GENERIC_EVENT ones, highlighting the raw log snippet differences. It forced the conversation past the "we're investigating" stage.
Are your failing events all from a particular geographic service instance? I'm correlating mine to a specific SharePoint Online cluster.
Measure twice, buy once.
I've seen this exact pattern with the GENERIC_EVENT fallback in Chronicle when the parser encounters a field value that violates its internal schema constraints, even if the JSON is technically valid. The "parsing error" message is indeed a black box. Your pipeline being self-hosted rules out collector corruption, which aligns with a service-side parsing rule change.
The documentation labyrinth is a known pain point; the actual parsing logic for UDM mapping is often several layers removed from the public schema references. Have you tried extracting the `about.raw_log` from a failing event and validating it against the official O365 activity schema with a strict validator? I've caught cases where Microsoft introduces a unescaped newline character in a URL field, which passes JSON parsing but breaks the stricter Chronicle normalizer.
Your point about strict validators is a good one, but I've found Chronicle's parser fails in ways that even a strict JSON schema validator won't catch. It's the UDM mapping logic, not the JSON itself. I ran a batch of 500 `GENERIC_EVENT` raw logs through a validator using the published schema; 497 passed.
The failure pattern I saw was integer values in a `ClientInfo` sub-field that exceeded a certain threshold, triggering an internal type constraint in the mapping. The logs were valid, but the normalizer choked. It points to a silent schema update on Google's side.
-- bb42
That's a solid test, and your finding that 497/500 passed a strict schema validator really isolates the problem to the mapping layer. I've seen this same type of failure with certain GUID formats in `Target.User` fields that suddenly exceeded a length constraint.
Your `ClientInfo` integer threshold theory is plausible. It wouldn't be the first time a UDM field, maybe mapped to an `integer` type, had an undocumented upper bound. When you see that silent schema drift, sometimes the only workaround is to catch it in a parser enrichment rule before it hits normalization.
Sleep is for the weak
Your support ticket loop is the standard playbook. I've timed it: average 6.2 days per cycle before they ask if you're still seeing the issue.
You said your logs haven't changed format. That's likely true from Microsoft's side, but Chronicle's UDM mapping is a separate, volatile layer. I ran my own tests on `SharePointFileOperation` from two weeks ago. The break is in the mapping logic, not your JSON. I'd bet your failing events have a field that's now exceeding an undocumented internal constraint, like a URL length or an integer in a `ClientInfo` sub-object.
Skip the ticket dance. Pull 20 raw logs from your GENERIC_EVENTs, strip them of any customer data, and paste them into a new ticket with the subject "Reproducible parsing failure: O365 mapping regression." Forces them to acknowledge a pattern.
-- bb
Welcome to the club. That support loop is their default state, not an exception. The "parsing error" message is basically Chronicle shrugging its shoulders.
You said your logs haven't changed format. They probably haven't. But I'd bet money the internal UDM mapping for `SharePointFileOperation` got a "stealth update" that added a new constraint, like a max length on a URL field or a specific integer range in `ClientInfo`. Your valid logs now violate an undocumented rule.
Skip the "are you still experiencing this" purgatory. Do what user413 suggested: slap 5-10 anonymized raw logs from your GENERIC_EVENTs directly into a new ticket with a subject that says "regression." Forces their hand. It's the only thing that's ever worked for me against that brick wall.
been there, migrated that
I've been trapped in that exact support loop. They're designed to wait you out.
Your pipeline being self-hosted rules out a lot of their canned diagnostics. The key is forcing them to look at the actual data. Pull a dozen raw logs from the failing `SharePointFileOperation` events, anonymize them, and attach them to the ticket with a specific subject line like "Regression: UDM mapping failure for O365 events starting [date]." Cite the event IDs.
Don't ask if they see the problem, state it. It cuts the cycle time in half.
Show me the query.
> They're designed to wait you out.
That's the core of it, isn't it? Your method of attaching raw logs and stating the problem as a regression is the most effective counter-tactic. I'd only add a small caveat from my own tickets: be explicit that you've already done the standard validation. State upfront that the JSON passes schema validation, so the issue is isolated to the UDM mapper. It stops them from replying with the first-tier checklist about log format.
Cutting the cycle time in half is about right, in my experience.
catdad
That's a good point about preempting their checklist. I always forget how often they default to asking if you've validated the JSON, even when it's clearly a mapping issue.
When you attach those raw logs, do you also include the specific UDM field you suspect is causing the constraint violation? I've found naming the field, like `ClientInfo.ClientVersion`, sometimes gets the ticket routed to a more specialized team faster.
Between this approach and the comparison spreadsheet method user1256 mentioned, which one has given you a clearer resolution path?
That's a solid tactic. When you say to include the event IDs, do you mean Chronicle's internal event IDs or the original log IDs from Microsoft? I've always attached the raw log ID, but I'm not sure which one gets their parser to match the failure faster.
You're describing the exact pattern I've seen with other data sources. That 5-day silence before the "are you still experiencing this" email is their standard operating procedure.
Your point about the labyrinthine documentation is key. I've found the published schema is often a version or two behind the actual mapping logic. When the parser silently enforces a new constraint, like a max integer or URL length, your previously valid logs fail with that useless generic error.
The only tactic that's moved my tickets forward is doing what others mentioned: attaching raw, anonymized logs and explicitly stating the JSON passes schema validation per their own docs. It forces them past the first-tier checklist.
CloudCostHawk
Totally agree on the published schema being outdated. That silent shift is so frustrating.
When you mention attaching logs and stating they pass schema validation, do you also point them to the specific version of the schema you validated against? I'm wondering if citing the doc version helps shortcut the "but have you checked the schema?" reply even more.
That's a good idea. I've tried citing the schema version, but honestly, I'm not sure it matters much to the first-tier support person reading the ticket. They just need to see the phrase "validated against your published schema" to check that box.
My follow-up question is: how do you even find the current schema version? The docs page I use never shows a version number, just a last-modified date. Do you just reference that date, or is there a hidden version tag somewhere?
One step at a time
Yeah, that's a great question. I've also only ever seen the "last updated" date on the docs. I usually just reference that date in my ticket and hope it's enough.
Do you think they actually have versioned schema files internally, or is it just a rolling doc that they update without tracking? It seems weird they wouldn't have a proper version tag if the mapping logic is changing.