Filtering at the source is the only way to win. We did the same thing with a particularly chatty cloud service - realized we were paying a premium to ingest JSON bloat for fields we never queried.
The user story exercise is smart, but teams often lie to themselves. "Oh, the SOC *might* need this for an investigation" becomes a license to keep every data faucet wide open. You have to be brutal and cut the models that don't have a signed-off playbook.
Hourly batch for non-critical anomalies is the real pro-tip. The vendors never suggest it because they're selling real-time. Most "behavior" doesn't change minute-to-minute anyway.
been there, migrated that
You're right about mapping sources to pricing tiers. It's the step that shows you where the real costs are hiding, not in the headline "per-GB" rate.
But the second point is the real trap. Vendors sell you on the idea that you need all their behavioral models. We audited ours and found 60% of them were analyzing behavior for which we had no defined response playbook. Turning those off cut our processing load in half. The dashboard got simpler, but it actually became useful.
What's your process for deciding which custom log source is worth the parsing fight versus just aggregating it externally first?
Show me the query.