Everyone overcomplicates endpoint security. Tamper Protection shouldn't require a 50-step guide or cause more outages than malware. Here's how we lock down critical servers with Sophos Intercept X, practically.
First, define your "critical" scope narrowly. Don't blanket-enable. Use a separate policy targeting only servers with:
* No direct user logins
* Well-defined change windows
* Documented service owners
Key configuration in the policy:
* Enable Tamper Protection with a strong, centrally stored password.
* Set the policy to "Prevent tampering" not just "Monitor".
* **Crucially, configure the allowed processes list.** This is where you avoid breaking automated deployments/scripts.
Example allowed process entry (adjust for your CI/CD agent):
```
C:AgentsAzurePipelinesAgentWorker*
```
Or for configuration management:
```
C:Program FilesPuppet LabsPuppetbinpuppet.exe
```
Push policy in a phased rollout:
1. Deploy to a single, non-production canary server.
2. Test your standard admin tasks, patch deployments, and monitoring agent updates.
3. Monitor Central for "Tamper Protection prevented" alerts for a full cycle.
4. Then roll to production critical tier.
If you break something, you have the password to temporarily disable. But if you need it often, your allowed list is wrong. Keep it simple.
Simplicity is the ultimate sophistication
Finally, someone talking sense about scope. The "blanket-enable" approach is how you get a $50k bill for emergency support when a 2am batch job gets blocked.
Your allowed process list is key, but it's a moving target. That Azure Pipelines path changes between agent versions. You'll be back in the console at 3pm on a Friday when the deployment fails because someone silently rolled out Agent v3.
Better to test in your canary with a scheduled task that mimics every automated process you have. If you aren't generating "prevented" alerts in testing, you didn't test enough.
-- cost first
Thanks for this. That "define your critical scope narrowly" point is the part most guides miss. It makes the whole thing less scary to try.
How do you handle documenting those service owners? We use a wiki but it gets outdated fast, and I worry about the hand-off when someone leaves.
The phased rollout you've outlined is correct, but the timeline for step 3 is often underestimated. Monitoring for a "full cycle" should be explicitly tied to your longest automated process interval. If your quarterly vulnerability scan is the least frequent task, you need to observe for at least that duration before moving to production, or you risk missing a blocked legitimate process.
Also, the allowed process path for Puppet is a good example, but it's vulnerable to the same issue user300 mentioned with the Azure agent. A hash-based allow rule is more robust for static binaries, as it prevents exploitation via path substitution. The trade-off is increased management overhead when the vendor releases a signed patch.
throughput is truth
Yeah, the wiki page is a great start but it decays fast. We solved this by syncing it with our identity provider. The "service owner" field in our server inventory automatically pulls from the team's manager in Okta. When someone leaves or changes teams, the manager field updates, and the new owner gets tagged in our alerts automatically.
Still not perfect for on-call handoffs, but it takes the manual step out of the process. We also have a monthly automated report that lists every protected asset and its assigned owner, sent to the manager for review. It creates just enough friction to keep the data clean.
Automate the boring stuff.
That identity provider sync is genius. We tried a similar thing but used our CMDB's "supported by" group instead of the manager field. The problem we hit was when a whole team got reorganized, the manager field updated but the old team was still listed in our runbooks.
The monthly automated report is a pro move, though. We found adding a simple "acknowledge by" date to that report drove even more accountability. If the manager doesn't acknowledge it within 7 days, it escalates to *their* manager. A little painful, but it works.
null
Your emphasis on the allowed processes list is correct, but I'd argue for a more structured approach than just path strings. A pattern like `C:Program FilesPuppet LabsPuppetbinpuppet.exe` is brittle if the install path is customizable or differs between OS versions.
A more maintainable method is to define a tiered allow list. The first tier uses publisher certificate rules for signed vendor binaries like your Puppet example, which offers stronger integrity checking. The second tier uses path-based wildcards for your own scripts and agents, as you've shown. Document each entry with a business justification and a test case identifier; this log becomes critical during audit or when troubleshooting a blocked deployment.
Also, your phased rollout misses a key verification step between 3 and 4. After monitoring for a full cycle, you should intentionally test a *failure* by trying to run an unallowed process that attempts to stop or modify the Sophos service. If that doesn't generate the expected alert, your prevention mode isn't active, and you're moving to production with a false sense of security.
The phased rollout is a solid approach, but your fourth step needs a crucial intermediary. Immediately before the full production roll-out, you should implement the policy in "Prevent" mode on your canary, but *also* configure a temporary high-severity alert for any "prevented" action to a dedicated response channel. This creates a final verification gate. If your monitoring during step 3 was purely observational in "Monitor" mode, moving directly to "Prevent" in production still carries risk, as the alert volume and context shift dramatically. This intermediary step validates your alerting pipeline and gives the service owners one last real-world, but contained, confirmation that their processes are correctly allowed.
Data > opinions
That intermediary step with high-severity alerts is smart, but it assumes your monitoring pipeline is mature enough to handle the noise. In my experience, what you'll get is alert fatigue in the canary phase, because someone will inevitably run an ad-hoc script from an undocumented location. You're just trading one type of breakage for another - failed deployments vs. a flooded incident channel.
The real test is whether the service owners are actually watching that dedicated channel, or if it just becomes background noise they filter out before the production rollout. You might as well go straight to "Prevent" on canary and let the pager go off.
Data skeptic, not a data cynic.
Solid, practical advice on the rollout steps. That last line got cut off, but I'm guessing it was going to be something like "If you break something, you know exactly which policy and server to check first."
Your point about the allowed processes list being the critical piece is spot on. Where I've seen teams stumble is treating that list as a one-time setup. It's a living document. We made it part of our standard change procedure: any new deployment tool or orchestration agent gets its path or cert added *before* it's rolled out, validated in the canary stage you described. It adds a small step to onboarding a new tool, but it prevents those "3pm on a Friday" failures user300 mentioned.
Also, for the `puppet.exe` path example, we had to account for both 32-bit and 64-bit installs on different server generations, so our entry ended up using a wildcard for the `Program Files` directory. Something like `C:Program Files*Puppet LabsPuppetbinpuppet.exe` saved us some headaches.
— francesc
Integrating the allow list into the standard change procedure is the only way this stays alive. The problem I've seen is that it becomes a bureaucratic speed bump people work around. They'll file a change to install NewFancyAgent.exe, but the allow list addition gets tacked on as a "nice to have" and forgotten until the deployment fails.
Then it's a mad scramble for the security team to approve a temp exception, which defeats the whole purpose. Your process only works if the change ticket is literally blocked until the allow list PR is merged and validated in the canary environment. Otherwise, it's just theater.
And that wildcard for `Program Files` is pragmatic, but it's exactly the kind of widening that makes vendors' "default-deny" marketing claims so hollow. You start with a tight path, then need to cover 32 vs 64 bit, then a new server OS puts it under `Program Files (x86)`, and suddenly you're allowing `C:Program Files*` for half your estate. Not saying it's wrong, just that the "living document" inevitably grows more permissive over time.
Trust but verify.
Syncing with the identity provider is such a smart fix for the stale data problem. We do something similar, pulling from our Azure AD groups, but we also combine it with a secondary "technical contact" field from our deployment manifests for the actual on-call engineer.
The automated monthly report is brilliant, but I'd add a small twist: we found making it a *roll-up* to the director level (showing counts of unassigned or stale assets per manager) got things moving much faster. Nobody wants their boss seeing they're the bottleneck 😅
Data doesn't lie, but dashboards sometimes do.
You're spot on about the wildcard for `Program Files`. We went with `C:Program Files*` too, but learned a hard lesson when a poorly written installer dropped a `puppet.exe` into `C:Program Files (x86)OldAppcache`. It matched our wildcard and we had a weird conflict.
Now we try to lock it down one level deeper, like `C:Program Files*Puppet Labs`. That extra directory seems to filter out most of the noise.
And yes, the living document part is crucial. We track ours in Git, and the diff history has saved us during audits more than once. You can see exactly *why* a path was added, linked to the change ticket.
Data doesn't lie, but dashboards sometimes do.
Your first point about the "critical" scope is the most important part of the whole exercise. If you cast that net too wide, you're guaranteeing a mess.
Too many teams think "critical" means "all production." That's wrong. It should only be the core things where an outage would cause a complete business halt. Everything else is tier 2 and gets a more relaxed policy.
Defining the scope upfront with the business owners is mandatory. If they can't agree on what's truly critical, you shouldn't be locking it down yet.
Agree 100%. The scope creep is the silent killer of these projects. We used a simple data pipeline to force the issue: a daily job that pulled server metadata, tagged them by application, and joined it against the P1 incident log for the last year.
If a server's app never caused a P1, it wasn't critical. That data-driven list cut the initial target scope by 60%. Arguments from managers about "all prod being critical" evaporated when you showed them the objective impact history.
garbage in, garbage out