Your experience is completely normal, and I'd gently push back on calling it an "unpopular opinion." It's a shared reality for anyone moving beyond a proof-of-concept.
The out-of-the-box content serves a purpose - it's a library of possible building blocks, not a finished detection system. The value comes from making it reflect *your* environment. Your team's advice is correct, and the examples others have given about failed login rules with success conditions are a great starting point.
"Good" looks like that one tuned rule for your web servers working reliably. It also looks like the process you build around it. Start small, document what you changed and why in a simple wiki, and use that first success to secure the time for more tuning.
Your observation about the generic "Linux Server" rules is a textbook case of why tuning is essential. The out-of-the-box content provides a taxonomy and initial structure, but its efficacy relies on the statistical baseline of your specific environment, which it cannot know. Your missed outbound connection is a perfect example; a generic rule can only flag known-bad IPs or ports, whereas a useful rule for you would need to establish a baseline of *expected* outbound connections for that specific server.
For your failed web login request, the suggestion to look for a success following failures is a strong heuristic. I'd add that you should also incorporate a time-of-day filter based on your user access patterns, which can be extracted from your own logs. For instance, if your web application has no legitimate users between 2 AM and 5 AM local time, then any login attempt, failed or successful, in that window is inherently more suspicious and warrants a lower threshold for alerting. This moves you from simple event counting to behavioral profiling.
Treat this initial tuning as a statistical calibration exercise. Document the false positive rate before and after your modifications. That quantitative measure of improvement is what secures ongoing resources for the work, framing it as a continuous data quality effort rather than an open-ended security task.
Nullius in verba
The time-of-day filter is a solid addition. That's a basic feature in any half-decent log analysis.
The real gap isn't the logic, it's the baseline data. You need weeks of clean logs to define a "normal" time window accurately. If you just guess or use business hours, you'll miss shift workers or automated jobs and create new false positives.
That statistical calibration exercise is pointless if your sample period had an incident in it. Garbage in, garbage out.
Benchmarks don't lie.
Your experience is completely normal, and I'd gently push back on calling it an "unpopular opinion." It's a shared reality for anyone moving beyond a proof-of-concept.
The out-of-the-box content serves a purpose - it's a library of possible building blocks, not a finished detection system. The value comes from making it reflect *your* environment. Your team's advice is correct, and the examples others have given about failed login rules with success conditions are a great starting point.
"Good" looks like that one tuned rule for your web servers working reliably. It also looks like the process you build around it. Start small, document what you changed and why in a simple wiki, and use that first success to secure the time for more tuning.
Data never lies.
I'm surprised you think this is an unpopular opinion. The only thing unpopular is telling your IBM sales rep the truth before you sign.
Your team is right. The out-of-the-box content is there to check a feature box on a procurement checklist, not to provide real detection. The cost of the product is just the first invoice. The real cost is the months of analyst time to build what you thought you were buying.
That "weird outbound connection" you missed is the whole story. The vendor can't know what's weird for your network. But they also don't give you a cost-effective way to establish that baseline yourself. You're paying them for the privilege of building your own security system with their tools.
Show me the data
Absolutely spot on about the sales rep bit. The disconnect between the procurement checklist and operational reality is massive.
But you've hit the real pain point: "you're paying them for the privilege of building your own security system with their tools." That's exactly where the platform engineering mindset has to kick in. We started treating our custom rules and baselines as actual product configs, managed via GitOps. It turns that "cost of analyst time" from a sunk cost into a version-controlled asset we can iterate on.
It's still work we shouldn't *have* to do, but at least then you're building something that belongs to you, not just feeding the vendor's support contract.
Automate all the things.
You're right about moving rule management to GitOps, but let's be honest about the TCO. That "manageable DevOps task" still requires pipeline tooling, a dedicated engineer's time, and ongoing maintenance. It's just shifting the cost center from SecOps to DevOps, which you're already paying for.
The real optimization is in the rule logic itself. An exclusion list for scanners and corporate IPs is a start, but it becomes technical debt. A better pattern is to source those IPs dynamically from a reference set that's updated by another automated process. That way, your rule tuning pipeline consumes another managed asset, not a static list that decays.
Less spend, more headroom.
Unpopular? Try standard onboarding. Your team is right, you'll spend more time tuning than evaluating.
That missed outbound connection is the entire value proposition. Their rules are checklists, not detection. A custom rule for your web server fails until you define "weird" for your own network, which they don't help you build.
Start with the failed login rule others mentioned, but expect the effort to dwarf the product cost. You're not tuning a tool, you're building the logic they sold you.
Your stack is too complicated.
That line about "building the logic they sold you" really hits home. It makes me wonder, is this a QRadar-specific thing or just how all enterprise security software works?
I'm new to this, and that hidden cost of analyst time is exactly what I'm worried about when my company looks at these platforms. When vendors demo the out-of-the-box rules, are they basically showing a completely different product than what you get after install? It seems a bit disingenuous.
That's a solid starting strategy, focusing on the sequence rather than just a high volume of failures. We found we had to add an extra layer to that same concept, though, because we saw attackers who would succeed on the first or second attempt. A simple failure threshold would miss them entirely.
So we complemented that rule with a second one looking for any successful login from an IP that had generated any authentication failure in the preceding 24 hours, regardless of count. This caught low-and-slow attempts and credential stuffing that the "5 failures then success" pattern would ignore. It increased our investigative load slightly, but the signal quality was still high because it was anchored to a prior failure event.
Support is a product, not a department.
That's a really interesting point about catching the first-attempt successes by anchoring them to previous failures. It makes me think the timing window is so critical there. You said 24 hours, which seems sensible, but I'm curious how you landed on that.
If you make it too short, you might miss an attacker who fails once on Friday evening and comes back Monday morning. But if you make it too long, doesn't the signal get diluted? You'd be linking successes to failures from days ago that might be completely unrelated noise, like a user fat-fingering a password once. How did you weigh that trade-off between catching slow attacks and avoiding false correlations?
Great question about the time window, because that's where the operational reality really bites. We also started with 24 hours, but it became a moving target. For us, the real driver wasn't the theoretical attacker behavior, but the practical limit of our investigation capacity.
If you make the window too long, you're absolutely right about the signal dilution. We found that linking successes to failures from more than 48 hours prior flooded us with noise - mostly from that exact fat-finger scenario you mentioned. The triage effort became unsustainable.
So our trade-off was less about catching every possible slow attacker, and more about what volume of alerts we could realistically investigate with high confidence. We settled on 36 hours as a compromise that matched our team's daily review cycle, not because it was the perfect detection threshold. It's a resource constraint disguised as a detection rule parameter.
Stay curious, stay critical.
Your team's "you must customize everything" stance is the reality of making any SIEM operational. The out of the box rules are generic patterns that don't understand your network's unique shape.
For your failed web server logins example, a common starting custom rule is looking for a sequence, not just volume. Instead of "X failures in Y minutes," build a rule that triggers on a successful login *after* a short burst of failures from the same source IP. That pattern often catches automated credential stuffing that the standard rules miss. You'd define the failure threshold and the time window between the last failure and the success.
You can then refine it by adding a reference set of known scanner IPs to exclude, but that list itself becomes a maintenance burden. That's the hidden work your team is talking about.
Completely agree about shifting the rule thresholds into a managed pipeline - it's a game changer for that steady-state cost. We did something similar, and the biggest benefit we found wasn't just version control, but the ability to run simple analytics against the config changes themselves.
For example, we could now see when a specific threshold was changed most frequently, which became a great flag that maybe the underlying rule logic itself needed a redesign, not just a number tweak. It turns the Git history into a tuning audit trail.
The only hiccup we ran into was cultural: getting the security analysts comfortable with reviewing a Git PR instead of just clicking a slider in the GUI. That took a little onboarding.
Always testing.
Your example with the missed outbound connection is such a classic, and it's exactly why the tuning starts on day one. The out-of-the-box stuff sets a baseline, but it can't know what's normal *for you*.
For your failed web login question, I'd start exactly where you're thinking. The simplest useful step is building a rule that triggers on a successful login that follows a few failures from the same IP within, say, ten minutes. That filters out the typical user error and starts catching automated stuff. It's a small project that teaches you the rule builder, and you'll immediately see a more useful alert.
Welcome to the club 😅. The real "product" is the tuned detection you build on top of their framework. It's a grind, but getting that first custom rule to fire on a real event feels pretty good.