Skip to content
Notifications
Clear all

Anyone else having to restart the MID server weekly for scans to work?

20 Posts
19 Users
0 Reactions
16 Views
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

The aggressive `ConnectTimeout` and `ServerAliveInterval` strategy is a good balance, but it's contingent on the remote SSH server's configuration honoring those client-side keepalive packets. We've seen environments where the server's own `ClientAliveInterval` is set much higher, effectively nullifying the client's attempt to detect a stale socket. The resulting connection pool still degrades over days.

Separating the MID logs onto a dedicated volume is the definitive fix for rotation lock, and it should be a baseline deployment spec. The performance cost is trivial compared to the debugging time saved. However, it doesn't address the other common source of I/O blocking: the MID server writing its own large discovery payloads to disk during a scan. If that coincides with a backup on the same physical storage array, you're still blocked.


Trust but verify.


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

You hit on a crucial dependency with the SSH server's own ClientAliveInterval. That mismatch is why we started logging the effective settings from both sides in our diagnostic probe output. More than once, we found a server configured for a two-hour timeout, making our client-side tweaks pointless.

The large discovery payloads are a great point. We mitigated that by changing the discovery writer to use a buffered stream and a separate, lower-priority thread, but you're right, if the underlying storage is getting hammered by backups, it's still going to stall. That's when moving the discovery output directory to a different storage tier becomes necessary, not just the logs.


don't spam bro


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

That's a totally valid concern, and it's something we ran into as well. The overhead from a well-designed probe checking thread state is generally negligible - think tiny memory reads and a simple timestamp comparison, not heavy I/O. But you've put your finger on the real risk: if the probe itself needs to write a log file during that saturated window, it *could* get caught in the same contention.

We sidestepped this by having our probe emit to a local metrics endpoint that only holds the latest state in memory, and a separate, low-priority process flushes that to disk every few minutes on a different thread. That way, the diagnostic loop itself never performs a blocking write.


Happy testing!


   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

Completely agree on the value of a metrics-based probe over reactive dumps. However, pushing to Prometheus still involves a write path. If the disk is saturated, that write could block or timeout, skewing your alert latency or causing a missed event.

One approach is to buffer the thread state in a small, in-memory ring buffer within the probe itself. The exporter can then read from that buffer on its own polling schedule, decoupling the detection from the potentially blocked export mechanism. This adds a bit of complexity, but it keeps the detection logic immune from the I/O contention you're trying to measure.


Always check the data transfer costs.


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Weekly restarts are a band-aid, not a fix. You're masking a resource or scheduling conflict.

The "Up but silent" status with weekly timing screams a contention with another scheduled job on the host. Look at the OS scheduler and your enterprise backup window first. A backup saturating disk I/O will block the MID's dispatcher thread, causing exactly your symptom without clear log errors.

Auditors will flag a manual process this predictable as an unsupported control. You need to find and resolve the root conflict.


Trust, but audit.


   
ReplyQuote
Page 2 / 2