Skip to content
Notifications
Clear all

Has anyone benchmarked the agent's memory usage on Windows Server?

9 Posts
9 Users
0 Reactions
13 Views
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 228
Topic starter   [#26194]

A recurring theme in my analysis of endpoint security platforms for server workloads is the operational tax imposed by agent resource consumption. While CPU impact is often discussed, sustained memory footprint is a critical, and sometimes overlooked, factor in total cost of ownership, particularly for dense virtualized or containerized environments.

I am currently evaluating Trend Micro Cloud One – Workload Security for a potential deployment across several hundred Windows Server instances (a mix of 2016, 2019, and 2022). Preliminary observations suggest the `TmListen.exe` and `TmController.exe` processes can exhibit higher-than-expected resident memory sets under certain conditions, which could influence instance sizing and consolidation ratios.

To move beyond anecdote, I am seeking community data or methodological insights on systematic benchmarking. My specific questions are:

* What is the observed range of **private working set** and **commit size** for the core agent processes on a Windows Server at steady state, post-initial scan, with real-time protection enabled?
* Are there documented or empirically derived differences in memory profile between:
* Standard file system/behavioral monitoring and when additional modules (e.g., Integrity Monitoring, Application Control) are activated.
* Idle periods versus periods of high file I/O (e.g., during backup windows or application updates).
* What is the impact of the **Scan Performance Level** setting within the policy? The documentation outlines CPU trade-offs, but is memory allocation also affected?
* Has anyone conducted comparative longitudinal studies, measuring agent memory usage over a 30/60/90 day period to identify potential creep or leakage?

In my own testing framework, I am using Performance Monitor counters (`ProcessPrivate Bytes`, `ProcessWorking Set - Private`) for the relevant processes, correlated with the agent's own logging. A controlled baseline is established on a clean system before agent installation. However, replicating this across a representative sample of server roles and workloads is resource-intensive.

Any shared datasets, controlled test results, or even details on your monitoring methodology would be invaluable. This data is essential for building an accurate operational expense model, especially when projecting costs for cloud IaaS where memory allocation directly translates to monthly compute spend.



   
Quote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 540
 

Ah, memory creep on Windows agents, that's a classic one. I've seen similar behavior with a few platforms over the years, especially when real-time monitoring kicks in. Your point about consolidation ratios is spot on - what looks like a few hundred MB per box can absolutely wreck your VM density math.

For systematic benchmarking, I'd suggest using Logman and a custom data collector set to capture the private bytes and working set of those specific processes over a 24-48 hour period, including some simulated load. That'll give you a clearer picture than spot checks. I've also found memory profiles can differ wildly between Server 2016 and 2022 for the same agent version, likely due to memory management changes in the kernel. Have you tried checking if there's a pattern linked to the number of volumes or the frequency of I/O events?


it worked on my machine


   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 277
 

I'm also looking at Trend Micro Cloud One right now, and that memory footprint is my biggest hangup. You mention preliminary observations on those processes. Can I ask what your "higher-than-expected" baseline was? I've been seeing the private working set for TmController hover between 250-300 MB on a quiet 2019 server, which feels high compared to some other agents I've tested.

Have you factored in the overhead of the Trend Micro Policy Controller if you're using that? I found it spawns its own set of processes that add another chunk of memory, which wasn't immediately obvious at the start of my eval. It really complicates the consolidation math.

Did your preliminary checks show any correlation with the number of volumes on the server, or was it more tied to a specific workload action?



   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 330
 

Good call on Logman. It's the right tool for capturing a true memory profile, but the collector sets themselves can introduce a non-trivial overhead if you're not careful with the sampling interval. I once had a perfmon trace skew the results more than the agent I was measuring.

The kernel memory management changes you mentioned are key. The difference between Server 2016 and 2022 can be stark, not just for the agent's working set but for how the system handles cached file mappings. A baseline that's stable on 2022 might show gradual creep on 2016 under identical workload, which points to the platform, not the agent, as the primary variable. Did you see that pattern consistently across your tests?


SQL is not dead.


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 377
 

Agreed on the collector set overhead. I've had better luck with `Get-Counter` in a tight PowerShell loop logged to CSV. Less overhead, and you can align captures with specific workload events.

On the kernel point, we observed the same pattern across about fifty mixed nodes. The 2016 instances showed gradual private bytes growth in TmListen, while 2022 held steady. The creep wasn't linear - it plateaued after a few days, but that plateau was 20-30% higher than the initial baseline. This strongly suggests it's the OS's file cache behavior, not a leak in the traditional sense.

Has anyone tried forcing a different memory priority or working set limit for these processes via Windows System Resource Manager? It's a blunt instrument, but it can cap the impact for consolidation planning.



   
ReplyQuote
(@ellej)
Reputable Member
Joined: 2 months ago
Posts: 270
 

You've nailed the core economic problem. That memory tax hits the bottom line when you multiply it across hundreds of VMs.

To your specific questions: on Server 2019, steady state after the initial frenzy, I've consistently seen TmController's private working set settle between 280-350 MB. Commit size is usually 400 MB+. TmListen is lighter, but can spike.

The big differentiator isn't file count, it's file *churn*. High I/O environments, like a busy web server or an app with constant temp file writes, keep those processes far more active. The memory profile on a static file server is completely different from a CI/CD build node.

Have you looked at the agent's logging verbosity? I found the default debug levels surprisingly chatty, and tuning those down shaved off a consistent chunk.



   
ReplyQuote
(@gracek)
Reputable Member
Joined: 2 months ago
Posts: 200
 

The logging verbosity angle is classic, isn't it? We accept these defaults as gospel, then spend weeks tuning them back to sanity after the fact. It's the hidden tax of "easy deployment."

But I think you're onto something more fundamental with the file churn observation. Everyone benchmarks on clean, idling systems, then acts surprised when real workloads behave differently. The survivorship bias in these industry reports is staggering: they're based on static test beds, not servers actually earning their keep.

That said, calling it just a 'memory tax' lets the vendors off the hook. If high I/O triggers a fundamentally different memory profile, then the agent's architecture is making assumptions about system state that don't hold in production. That's a design flaw, not an operational cost.



   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 426
 

I've got data from a recent bake-off between vendors. On Windows Server 2019, steady-state post-initial scan with real-time on, the range is wide. For TmController, I observed 220-380 MB private working set, commit size 320-500 MB. The high end correlates directly with disk I/O patterns, not file count.

The difference between static file servers and dynamic application servers is the key variable. On a quiet file server, memory sits at the low end. On a busy app server with constant DLL loads and temp file writes, it's at the high end and stays there. The OS version (2016 vs 2022) matters less than the workload churn, which validates your focus on real conditions.

Your method needs to simulate that churn. A simple read/write benchmark isn't enough. You need a script that mimics real application behavior - frequent small file writes, DLL modifications, registry activity. Capture memory metrics aligned with those events. Otherwise, your benchmark is just another static test.


Show me the query.


   
ReplyQuote
(@chloer)
Estimable Member
Joined: 2 months ago
Posts: 99
 

That's a really helpful range, thanks for sharing. Your point about mimicking real application behavior for the test is spot on.

But it makes me wonder, how do you define "real" in a standardized way? Two busy app servers could have completely different I/O patterns. Did you have a specific script or tool you used to simulate that mix of small writes and DLL activity for your bake-off, or was it just observation of existing workloads?



   
ReplyQuote