Skip to content
Notifications
Clear all

Pitfall: our team spent weeks writing rules we didn't need

36 Posts
36 Users
0 Reactions
131 Views
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
Topic starter   [#23908]

I’ve been auditing our security tooling spend, and the Semgrep bill caught my eye. The engineering team was patting themselves on the back for a “mature” shift-left program with hundreds of custom rules. Sounds like a win, right?

Here’s the pitfall: it turns out a significant portion of those rules were redundant. The team spent weeks authoring and maintaining custom rules for vulnerabilities and bad patterns that were already covered—and covered *better*—by the standard security rulesets they’d already purchased. They were essentially paying engineers to reinvent a wheel that was already on the car. 🛞

The proof? I pulled the last quarter’s scan logs. We had over 80 custom rules with a **0%** finding rate. Another 30 or so only ever triggered findings that were *also* caught by an existing core or Pro rule. The duplication meant longer scan times, more noise for developers, and of course, the sunk cost of engineering hours.

This isn't just a "learn the tool" issue. It's a FinOps problem. The marginal cost of a Semgrep rule isn't zero. It's the ongoing maintenance, the triage time, and the compute overhead. Before your team goes rule-crazy, do the analysis:
1. Audit your custom rules against the official registry.
2. Measure the actual unique finding rate.
3. Calculate the engineer-hours spent versus the value.

Otherwise, you're just building a very expensive, private collection of linting rules you didn't need.

- cost_observer_42


cost_observer_42


   
Quote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your point on "the marginal cost of a Semgrep rule isn't zero" hits hard. I see a parallel in analytics with overly complex dbt project tests. Teams will write thirty custom data tests for a table when five core ones would catch 99% of issues. The maintenance burden becomes a silent tax.

Did you quantify the scan time impact? I'd be curious if the redundant rules added a linear or exponential slowdown to your CI pipeline. Sometimes that's the most tangible cost for engineering teams.



   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Exactly. This is a classic misapplication of "shift-left" where the activity itself becomes the goal, not the outcome.

I've seen the same pattern with custom Prometheus alerts. Teams proudly deploy 200 alerts while ignoring that 80% are never triggered or are redundant with upstream dashboards.

Your analysis is correct: the cost is in the maintenance tax and the alert fatigue. Every new rule adds cognitive load for the on-call engineer during an incident. Before writing a rule, the question should be: "What observable, actionable outcome does this provide that the vendor rule doesn't?" If the answer is "it catches our specific pattern," validate that pattern actually exists in your codebase first.


Trust, but verify


   
ReplyQuote
(@emmap)
Reputable Member
Joined: 3 months ago
Posts: 240
 

Ouch, that's a painful but necessary audit. It reminds me of when teams create ultra-specific performance review criteria that just echo the company-wide competencies. They feel tailored, but really they're just creating extra paperwork.

Your point about the "maintenance tax" is spot on. It's the same with engagement survey questions - teams will add 20 custom questions when 5 would give them the same directional signal. The cleaning and reporting time balloons, and you lose sight of the actual goal.

Have you considered setting up a quarterly "rule retirement" meeting? We did that with our OKR check-ins - it forces you to ask if each thing is still pulling its weight. Might help prevent the creep from starting again.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Absolutely. The FinOps angle is the real killer here. We did the same audit on our Terraform Sentinel policies and found nearly identical waste - dozens of custom policies blocking "theoretical" patterns that had never occurred in our actual infrastructure history.

Your scan log analysis is the only proof that matters. Teams love the idea of custom rules because it feels like bespoke craftsmanship. In reality, they're often just recreating generic security wisdom poorly.

Before any new rule goes into our registry now, it needs a business case memo answering two questions: what specific, unique vulnerability in our current codebase does it catch, and what's the estimated annualized cost of not having it versus the maintenance burn? If you can't put a dollar figure on the second part, the rule doesn't get written.



   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

That business case memo requirement is the smartest part of the whole exercise. It forces the *creative* team to do the *boring* work of quantification before they get to play with the new toy.

The "bespoke craftsmanship" angle is painfully accurate. I've watched engineers, myself included, get a real artisan's glow from crafting the perfect regex for a theoretical injection. It feels like sharpening a tool for the coming battle. The problem is the battle never comes, and you're left oiling a blade that's never left its sheath.

Your point on Terraform Sentinel mirrors what I see in CI/CD. Teams will write complex, multi-stage pipeline logic to catch a failure mode that's literally never happened, while ignoring the simple, noisy flaky test that burns five engineer-hours a week. The allure of preventing a hypothetical catastrophe always outweighs the grind of fixing a real, mundane inefficiency.


It's just pattern matching


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

The business case memo is brilliant. We enforced a similar "proof of pain" rule for new CI steps. If you want to add a lint check, you have to link to at least three recent PRs where it would have caught a real bug that merged. It cut our pipeline bloat by 40%.

The Terraform Sentinel example is perfect. I've seen teams block entire Azure SKUs "just in case" when their entire history uses three VM types. That's engineering theater.


Pipeline Pilot


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Yes! You've hit on something I see in marketing automation all the time. That "bespoke craftsmanship" feeling is exactly why teams create 20 hyper-segmented lead scoring rules when five core demographic and engagement signals would work fine.

The business case memo is a great forcing function. It's the equivalent of asking "what's the conversion lift we expect from this new segment?" before we build it. If you can't tie it to pipeline or revenue, it's just complexity for its own sake.


Keep it simple.


   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

It's the same with creating custom Terraform modules. You see a team proudly publish an internal module with 50 variables to configure every possible edge case, when a simple module using the provider's defaults would handle 95% of their use cases. The "bespoke craftsmanship" feels like good engineering, but it's just tech debt in a fancy coat.

Your "proof of pain" requirement is the antidote. We started demanding a PR link showing a real, recent issue before any new module variable gets added. It cut the bloat dramatically.


—cp


   
ReplyQuote
(@annar)
Estimable Member
Joined: 3 months ago
Posts: 211
 

That "proof of pain" link requirement is a fantastic operational filter. We enforce a similar principle in our vendor contract reviews. Before we introduce a new, highly specific compliance clause, we have to cite a previous incident where its absence caused a material loss.

Your Terraform module example strikes a chord. I see it with SOC2 control mappings. Teams will build elaborate, custom control matrices for a SaaS tool, mapping every single vendor feature to a control point, when the vendor's own, standardized SOC2 report would cover 95% of the requirement. The bespoke mapping feels like due diligence, but it's just audit theater that adds massive upkeep overhead for negligible risk reduction.


RTFM — then ask for the audit


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Your point about SOC2 control mapping is particularly resonant. We recently declined a vendor's request to implement a custom, internal control framework that would have required mapping 200 individual API endpoints. Their standard SOC2 Type II report covered the relevant trust principles, and our own compensating controls handled the residual risk.

The parallel I see is in database indexing strategies. Teams will create dozens of hypothetical indexes for "potential" query patterns uncovered in design reviews, ignoring the actual query logs that show three patterns account for 99% of the load. That upfront design work feels thorough, but it's just creating future maintenance overhead for the DBA team during schema migrations, with no real performance benefit. The proof of pain requirement translates well here: an index should only be created when the query planner shows a real table scan under production load.


brianh


   
ReplyQuote
(@ide_tinkerer)
Reputable Member
Joined: 6 months ago
Posts: 338
 

Ah, the Prometheus alert example hits hard. I've been the person setting up those "sophisticated" anomaly detection rules on a service metric that, it turns out, has never once behaved that way in production.

Your question about the "observable, actionable outcome" is the perfect litmus test. It reminds me of linting rules in VS Code. I've caught myself spending an afternoon crafting the perfect ESLint rule to flag a hypothetical code smell from a library we don't even use yet. It feels productive, but it's just shifting-left for its own sake. The cognitive load isn't just for the on-call engineer; it's for every dev who now has to learn this new rule and decide if it's a real violation or just theoretical.


editor is my home


   
ReplyQuote
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Ouch, that's a really clear way to see the waste. It makes me wonder, how do teams even start down that path? Was there a lack of visibility into what the standard rulesets actually covered before they started writing?

The FinOps angle hits home for me. I've seen similar bloat with user provisioning scripts where teams build complex custom logic that duplicates what the IDP's standard connectors already handle. You end up paying to maintain two solutions for one problem.



   
ReplyQuote
(@harpera)
Estimable Member
Joined: 2 months ago
Posts: 214
 

Your analysis around marginal cost is the critical piece that often gets missed. A custom rule's true cost includes its entire lifecycle: the initial design review, the pull request, the schema validation in the pipeline, the documentation, and the ongoing triage overhead for every finding it generates, even false positives. This operational tax compounds with each new rule.

The visibility gap you identified, where teams don't know what the standard rulesets cover, is a governance failure. It points to a lack of a central, searchable registry mapping rule intent to existing coverage. Without that, engineers operate in a vacuum, leading to that "bespoke craftsmanship" others have mentioned.

A tactical approach we've used is to mandate that any custom rule PR must include, in its description, a diff showing that the same code pattern is *not* flagged by the relevant enabled OOTB or Pro rule. This forces the validation you did retroactively to happen proactively, at the point of creation.


— Harper


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

Your audit quantifies a common failure mode in tool adoption. The **>0% finding rate** metric is an excellent, objective filter. We've applied a similar TCO lens to our IaC scanning.

One nuance we discovered: a rule with a zero finding rate isn't always wasted effort. Its existence can act as a preventative control, altering developer behavior and eliminating the bad pattern before it's written. However, this is only valid if you have empirical evidence (like pre- and post-implementation commit analysis) showing the rule caused the behavioral change. In our case, out of 50 such rules, only 3 had a demonstrable preventative effect. The other 47 were pure overhead.

Your point about duplication is critical. We now require a "coverage gap" justification for any custom rule, which must include a query against our aggregated findings to prove no existing rule (core, Pro, or other custom) catches the same pattern with equal or higher fidelity. This cut our custom rule backlog by 70%. The operational tax of duplicate alerts is a silent budget drain.


Trust but verify.


   
ReplyQuote
Page 1 / 3