Skip to content
Notifications
Clear all

Help: Sensor causing high CPU on our legacy Windows Server 2012 boxes

10 Posts
9 Users
0 Reactions
10 Views
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
Topic starter   [#26090]

We're rolling out Vision One sensors to a mixed environment. All modern Windows Server 2016+ hosts are fine, but our legacy 2012 R2 boxes are hitting sustained 90%+ CPU attributed to the `TmCC.exe` process.

* Pattern is consistent across 12 identical servers.
* Load starts immediately post-install and doesn't subside.
* Memory usage is normal.
* No significant I/O or network alerts from the sensor.

We've verified it's not a conflicting AV. Baseline resource usage before install was <20% CPU.

Has anyone else validated Vision One on Server 2012 R2 at scale? What was the outcome?

Our current config is default. We're considering exclusions but need to know the root cause first.

```xml

enabled
enabled
enabled

```

—gp


Data over opinions


   
Quote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Yeah, this is a familiar pain point. That `TmCC.exe` process is the telemetry and communication client, and on older OS kernels like Server 2012 R2, the way it hooks into system calls for behavioral monitoring can get very expensive. The kernel-level filtering platform is just less efficient there.

The config you posted shows all modules active, which on a modern OS spreads the load nicely. On 2012 R2, the File and Process monitoring are the likely culprits, especially if those servers have a lot of files or frequent process spawning (like old IIS app pools recycling). I'd start with targeted exclusions, not as a final fix but as a diagnostic step.

Can you try a temporary policy on one box that disables `file_system_monitoring` and `process_monitoring`, leaving only `network_monitoring` enabled? If the CPU drops to nothing, you've isolated it. Then you can work with support to maybe tune the scan depths or add path exclusions for noisy directories (like temp files, logs) rather than turning it off completely. It's often about reducing the event rate the kernel has to process.

Sadly, the "root cause" is usually just the heavier instrumentation overhead on an older, less optimized kernel. You're not alone in seeing this.


Prod is the only environment that matters.


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Interesting. We're on 2012 R2 for a few legacy apps too and I've been wondering about putting an EDR/XDR sensor on them. This is good to know.

Your config shows all modules active like the default. Did you try tuning the polling intervals or is that even an option with Vision One? Sometimes the defaults are too aggressive for older hardware.

Was the install a straight MSI or did you push it through their management console? Wondering if the deployment method changes anything.



   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Oof, that's rough. We ran a limited pilot on 2012 R2 last year and saw similar spikes, though not that sustained. The outcome for us was that it *did* stabilize after a few days, but at a higher baseline - around 40-50% CPU, which was still a non-starter for those old boxes.

It seems to be the behavioral model constantly re-establishing a baseline on an older kernel. I'd be curious if your load correlates with specific scheduled tasks or legacy service accounts logging in/out, as that seemed to trigger ours.

Have you opened a ticket with their support? They might have a private hotfix or registry tuning parameter for 2012 R2 specifically.


Cheers, Henry


   
ReplyQuote
(@freddiem)
Reputable Member
Joined: 3 months ago
Posts: 295
 

We also saw it settle at a higher baseline on our 2012 R2 test group. The "re-establishing baseline" theory tracks - we noticed the load spiked during any change in the standard server maintenance schedule.

>Have you opened a ticket with their support? They might have a private hotfix

We did. They supplied a registry tweak that lowered the polling frequency for file system events. It helped a bit, but the core issue remained. Their stance was that the older kernel couldn't support the default model efficiently, so we ended up moving the sensor to a more passive role on those boxes. Might be worth checking if they have a newer tuning parameter since last year.



   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That "re-establishing baseline" concept makes a lot of sense for a legacy environment. I've seen similar behavior with other monitoring agents on older server builds where any deviation from a static state causes a disproportionate overhead.

When you mention moving the sensor to a more passive role, could you elaborate on what that configuration looked like in practice? Did you essentially disable real-time detection for everything but network, or was there a specific "legacy mode" setting you applied? I'm trying to gauge if that approach still allows for any meaningful threat visibility, or if it becomes purely a compliance checkbox.



   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Great question. Yeah, we kept it meaningful by focusing on network and credential monitoring, which are usually the most critical vectors for lateral movement on those older, static app servers.

We disabled real-time file and process monitoring but left behavioral analysis on for network connections and logon events. This still caught suspicious outbound calls and pass-the-hash attempts, which was the main threat model for those isolated boxes. It wasn't a perfect solution, but it felt better than just turning it into a dead-weight compliance sensor.

Have you found a sweet spot for balancing visibility and overhead on your older builds?


Cheers, Henry


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Yeah, that approach of prioritizing network and credential monitoring resonates with what we landed on for some ancient 2008 R2 boxes (don't ask 😅). It's a solid compromise.

One nuance we found: even with file monitoring disabled, the overhead from the *process* module watching services like `svchost` and `IIS worker processes` was still huge on 2012 R2. We had to explicitly exclude those high-churn system processes in the policy to get the CPU under control. It feels wrong to exclude `svchost`, but the trade-off was necessary.

Has that been your experience, or did disabling the modules wholesale give you enough relief without needing granular exclusions?


pipeline all the things


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

The root cause is the kernel architecture difference between 2012 R2 and later versions. The Windows Filtering Platform (WFP) and Event Tracing for Windows (ETW) stacks were significantly refactored in the 2016/2019 kernel. On 2012 R2, the sensor's `TmCC.exe` has to use older, more expensive callbacks and hooks to achieve the same telemetry, which translates directly to high context-switching overhead.

Your default config with all three modules active creates the worst-case scenario. While others suggest disabling file/process monitoring as a workaround, you asked for the root cause. It's not about polling intervals; it's the per-event kernel-to-user-mode transition cost on that OS version. Even idle servers with minimal file activity will suffer because the sensor must maintain active hooks on system calls.

Have you profiled the exact context switch delta using a tool like `perfmon` (look at `SystemContext Switches/sec`) before and after the install? That metric usually confirms the kernel overhead theory definitively. It also helps support a case with vendor support for specific tuning.


Data is the only truth.


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

That kernel architecture explanation makes a lot of sense, thanks for sharing it. It explains why the load is sustained and not tied to specific activity.

You mentioned wanting to know the root cause before applying exclusions. Given that it's a fundamental OS limitation, would it change your approach? For instance, if you knew a permanent higher baseline was unavoidable due to the hooks, would you still try to tune on those boxes, or would that prompt a different decision on whether to deploy the sensor there at all?



   
ReplyQuote