We’re implementing Trend Micro Cloud One Workload Security across our AWS auto-scaling groups (Linux, Amazon Linux 2). The goal is to have the agent install and register cleanly as part of the instance launch via User Data. However, we’re seeing a roughly 30% failure rate on scale-out events, where the agent either fails to install or fails to register with the Deep Security Manager, leaving instances unprotected.
Our current User Data script follows the standard API-driven install pattern:
```bash
#!/bin/bash
yum install -y https://dsm.trendmicro.com/software/package/amzn2/x86_64/agent.rpm
sleep 10
/opt/ds_agent/dsa_control -r dsm://:4120/ "tenantID:" "token:"
```
The failure manifests in two ways, captured in `/var/log/ds_agent.log`:
1. Package install times out (network issue from the instance to the Trend Micro package repo).
2. The registration command executes before the agent service is fully ready, returning a non-zero exit code and halting the User Data script prematurely.
We’ve attempted to add retry logic and longer sleeps, but this feels brittle and increases our instance boot time unacceptably. Has anyone successfully built a resilient, production-grade install for auto-scaling environments? Specifically, I’m looking for:
- A robust way to wait for the agent service to be ready for registration.
- Handling of temporary network failures during package fetch.
- Whether a pre-baked AMI with the agent installed (but not registered) is a better pattern, though we’d still need to handle registration at boot.
Any insights into the proper readiness check would be appreciated.
—J
Classic vendor install pattern. They document the happy path, not the reality of flaky networks and race conditions.
Your retry logic is just hiding the core issue: their agent is brittle at scale. I've seen this exact failure mode. The registration command often fails because the service manager hasn't actually started the agent yet, even after the RPM install returns.
Throw away the sleep. You need a proper health check that polls for the agent socket or service status before attempting registration. Even then, their package repo is notoriously slow. You'll probably need to cache the RPM in an S3 bucket you control.
Your stack is too complicated.