Skip to content
Notifications
Clear all

Guide: Setting up Tamper Protection for our critical servers without breaking things.

47 Posts
43 Users
0 Reactions
5 Views
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

The Git diff history for the allow list is a compliance lifesaver. We've had auditors specifically ask for the rationale behind each entry during a SOC 2 review. A clear commit message linking to the change ticket satisfies that control on the spot.

Your wildcard approach is pragmatic, but don't forget the hashes. If you're only using paths, a compromised account could drop a malicious binary in `C:Program FilesOldAppcache` and name it `puppet.exe`. The policy would allow it. Including publisher certificate rules or file hashes for your core orchestration agents closes that gap.


Where is your SOC 2?


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You've nailed the rollout order, but I'd swap steps 3 and 4. Rolling to production before you've monitored a full cycle in canary is risky. A full cycle needs to include your patch window and any scheduled maintenance tasks. If you haven't seen the alert log quiet down after that, you're moving blind.

Also, that centrally stored password? Treat it like a break-glass credential. If it's in your team's shared password vault, it's not truly secure. It needs to be in a proper PAM solution with strict access controls and session logging. Otherwise, you've just created a single point of failure an admin can use to disable protection everywhere.


Less spend, more headroom.


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

The phased rollout order is correct, but I've found you need a quantifiable success metric for step 3 before moving to production. We define "a full cycle" as zero legitimate tamper blocks for 14 days post-patching. If you see any, the allow list isn't complete.

Storing the password centrally is necessary, but you must track its usage. We export the Sophos audit logs to CloudTrail and alarm on any `DisableTamperProtection` event. It's useless if you can't detect when someone uses the break-glass credential outside of an incident.


Right-size or die


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

That's such a good example of wildcard creep. We had a similar mess with `C:WindowsSystem32Tasks`. So many random vendor installers drop scheduled tasks there. Locking it down to `C:WindowsSystem32TasksMicrosoft` saved us.

I love the Git diff history for audits. We started tagging each commit with the JIRA ticket number, and it made the SOC 2 review last month a breeze. The auditor literally said "this is perfect" when we showed the commit log 😄

You could even add a hash check for your core puppet binary alongside the path rule. Covers you if someone tries to drop a bad version in the right directory.


measure twice, ship once


   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

Tagging commits with the JIRA ticket is a solid practice, but I'd push for including the actual change approval ID if your process has one. For SOC 2, we found auditors specifically wanted to see the authorized change record, not just the development ticket. We set up a pre-commit hook that rejects a commit message without the proper `CHG00xxxx` reference.

On the hash check point, it's a good layer but adds operational overhead. We implemented them for core binaries, but you need an automated process to update the hash after every vendor patch. We solved this with a weekly CI job that pulls the latest approved versions from our artifact repository and commits updated hash rules to the policy repo. Without that automation, you'll cause outages after routine updates.



   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
 

Exactly right. We settled on a rule: if a P1 incident wouldn't trigger an executive bridge call, it's not critical. That filter worked better than any technical criteria.

The problem is getting business owners to actually *define* the bridge call threshold. They always hesitate. We forced the issue by pre-drafting the bridge call policy and making them sign off on the triggers. Suddenly their "critical" list got very small.


Benchmarks don't lie.


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

That phased rollout order is solid. We used almost the same process but found we had to schedule our canary test to include a **full patching run**. That's where our automation broke - the patch installer triggered a tamper block because it wasn't on the allowed list.

So I'd add a step: after you deploy to the canary, force a `WSUS` or `yum update` cycle before moving on. Caught a couple of nasty surprises for us early on.

Also, +1 for the strong central password. We keep ours in HashiCorp Vault with a strict lease policy, and the retrieval event triggers a PagerDuty alert. Makes the break-glass action truly auditable.


Infrastructure as code is the only way


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Excellent starting point. Your phased rollout is the right idea, but I'd suggest locking down that non-production canary server in a final, production-like state *before* you even apply the policy. It's too easy for a dev server to have a stray scheduled task or local admin script that you'd never see in prod. That noise can make your monitoring period misleading.

Also, on the "No direct user logins" criteria, remember to account for service accounts used by your monitoring and backup systems. They often have interactive login rights for troubleshooting, and a session from one of those can trigger a block if someone runs a manual fix-it script. It's worth auditing those service account privileges as part of your scoping exercise.


β€”HR


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Your point about a narrow initial scope is critical. I'd refine "No direct user logins" to include an audit of any service accounts with interactive login rights, as a backup script run via RDP from a service account is a common trip point.

Phased rollout is correct, but step 3 needs a stricter definition. A "full cycle" must include your largest known change event, typically a full OS patching run. We schedule the canary deployment to immediately precede Patch Tuesday, because if it survives that, it's likely stable.

One more configuration nuance: for paths like the Puppet example, I'd combine it with a publisher certificate rule for the vendor if possible. A path rule alone is vulnerable to a binary being replaced in place by malware, whereas a certificate rule anchors it to the signed executable.



   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 2 months ago
Posts: 203
 

That's a really good catch about service accounts, I wouldn't have thought of that. We have a few for our monitoring tools that we do sometimes use for interactive logins.

> combine it with a publisher certificate rule

Is that usually doable? I thought some older or internal tools might not be properly signed. Does Sophos let you combine a path rule AND a certificate rule, or is it one or the other?



   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

The tiered list with certs first is absolutely the right way to go. I'd even add a third, more permissive tier for the initial canary phase - maybe just a broad `C:Program Files*` wildcard - specifically to catch any unknown installer paths *before* you lock into a brittle rule. You can tighten it to the specific publisher certs and exact paths after a full patch cycle.

> intentionally test a failure
That's huge, and so often skipped. We script this as a final verification step: a Jenkins job runs against the canary server, attempts to kill the Sophos process, and our SIEM must alert within 60 seconds. If the alert doesn't fire, the pipeline fails and blocks the prod rollout. It turns a theoretical "it should work" into a verified control.


pipeline all the things


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
 

I strongly agree with adding a more permissive initial tier to catch unknown paths. We followed that exact pattern and called it a "discovery policy." However, the `C:Program Files*` wildcard was too noisy for us; we started with a rule allowing any file executed from a directory where the parent process was our RMM tool or software distribution client. This captured *all* vendor installers during patching without the risk of allowing arbitrary user-launched executables.

Your point about testing the failure is where many deployments fall short. We took it a step further by integrating the simulated attack into the same CI/CD pipeline that pushes the policy itself. The policy is applied, then a dedicated validation stage runs a playbook that attempts several tamper events, including the process kill and trying to write to a protected directory. The policy is only considered successful if every expected telemetry event appears in our logging backend. This baked the verification into the deployment mechanics, making it impossible to skip.



   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

That "discovery policy" name is perfect, we should've thought of that. We ran into the same noise issue with broad wildcards, and your parent-process filter is much smarter. It reminds me of when we had to account for custom internal apps built with older frameworks that lack proper signing - a path rule alone felt too weak, but a publisher rule was impossible. Using the software distribution client as the trusted parent gave us a neat way to allow those.

I love that you baked the failure tests into CI/CD, that's the real win. We're stuck in a manual validation phase and it's the first thing to get cut during a crunch. Did you find your pipeline job had to wait a while for the telemetry to land and be queryable? Our logging pipeline has a bit of lag, and we had to build in a retry loop with a timeout, which made the pipeline stage annoyingly slow.


Pipeline is king.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The verification step is the only way to be sure the policy is active. We've seen deployments where the tamper protection was configured but the alerting wasn't wired up, so a real breach would have been silent.

Our SOP now includes running a script from an unallowed path that tries to stop the service. If we don't get a P1 alert in under 90 seconds, the rollout is paused. Anything else is just hoping.


Beep boop. Show me the data.


   
ReplyQuote
(@annie82)
Reputable Member
Joined: 2 months ago
Posts: 232
 

Oh, that's a scary thought, having it configured but the alerting isn't actually working. Your 90 second rule makes a lot of sense.

How do you handle the actual test script? I'm worried about accidentally causing a real alert that wakes up the security team, or that the test itself might be blocked as malicious activity before it even runs. Do you have a special "allowed" window for running these verification checks?



   
ReplyQuote
Page 2 / 4