Skip to content
Notifications
Clear all

Switching from AppDynamics to Dynatrace. What should I watch out for?

5 Posts
5 Users
0 Reactions
28 Views
(@kubernetes_wrangler_42)
Estimable Member
Joined: 4 months ago
Posts: 64
Topic starter   [#10724]

Hello everyone. I've been working with AppDynamics in Kubernetes environments for several years now, but my current organization is strongly considering a switch to Dynatrace. I'm tasked with leading the evaluation and potential migration. While I've done some lab testing, I know there are always pitfalls that only reveal themselves at scale in production.

I'd like to share some of my initial observations and ask this knowledgeable community for the hidden challenges I should be watching for. My primary concerns aren't the basic feature lists, but the operational and architectural nuances that impact day-to-day SRE life.

From my hands-on testing so far, the fundamental shift is moving from a traditional sidecar/Java agent model (AppDynamics) to Dynatrace's OneAgent operator and dedicated ActiveGate pods. The deployment paradigm is completely different.

**Key differences I've noted:**
* **Installation & Management:** AppDynamics controllers/agents vs. Dynatrace's Kubernetes Operator managing the OneAgent lifecycle. The operator pattern is more cloud-native, but it's another custom resource definition (CRD) to manage.
* **Data Ingestion:** AppDynamics is quite selective in what it sends by default. Dynatrace's OneAgent seems to take a "capture everything" approach, which is powerful but immediately raises questions about cost control and data noise.
* **Kubernetes Context:** Dynatrace's built-in Kubernetes awareness seems deeper, with automatic topology mapping that feels more integrated than our current setup of piecing together AppDynamics metrics with kube-state-metrics.

**My specific questions for those who have made this journey:**

1. **Cost Surprises:** Everyone talks about licensing, but what about the ancillary costs? Did Dynatrace's richer data ingestion lead to unexpected increases in your log storage or network egress costs within your cloud provider?
2. **Alerting Transition:** We have hundreds of custom health rules and policies in AppDynamics. How did you manage the migration of alerting logic? Was it a manual rebuild, or did you find any tools or methods to translate policies?
3. **Operator Resilience:** In practice, how does the Dynatrace Operator behave during cluster upgrades or network partitions? Does it ever become a single point of failure for monitoring itself?
4. **Custom Metrics:** We instrument a fair number of custom business metrics via the AppDynamics APIs. What's the development experience like for pushing custom metrics to Dynatrace, and do you see any performance impact at high cardinality?

I'm particularly interested in any configuration snippets or Helm chart customizations you found essential. For example, I already learned the hard way in the lab to carefully define resource requests/limits for the ActiveGate pods to prevent them from being evicted.

```yaml
# Example: A snippet from my Dynatrace Operator values.yaml for resource constraints
activeGate:
resources:
requests:
memory: "512Mi"
cpu: "500m"
limits:
memory: "1024Mi"
cpu: "1000m"
```

Any war stories, recommended practices, or even reasons why you might have decided *against* such a switch would be incredibly valuable. Thank you in advance for sharing your hard-earned experience.

kubectl apply -f


yaml is my native language


   
Quote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Good point about the operator being more cloud-native. That CRD does become a single point of failure for your monitoring config, though. I've seen teams get caught when a `kubectl apply` on the DynatraceOperator configuration accidentally wipes out custom OneAgent arguments for specific namespaces.

You mentioned AppDynamics being selective with data. Dynatrace's default "full stack" mode is the opposite - it's incredibly verbose. Without fine-tuning the `OneAgent` spec in your Kubernetes yaml (especially the `processModuleArgs`), you can get swamped with low-value process and log data, hitting ingest limits faster than you'd expect. Start with a focused profile, maybe just app and service level, and expand deliberately.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

Absolutely. That shift from a sidecar model to the operator is fundamental, and it changes your failure domain. With AppDynamics, a bad agent config might crash a single pod. With the Dynatrace operator, a flawed `OneAgent` CRD can theoretically disrupt instrumentation across entire node pools if it's configured to auto-inject.

You'll also need to rethink your pipeline for agent updates and rollbacks. The operator automates OneAgent updates, which is great, but you lose the granular, pod-by-pod control you might be used to. I'd recommend implementing a strict canary process for the operator itself, deploying it to a non-critical cluster or namespace first, before letting it manage your production nodes. The operator's health becomes as critical as your monitoring data.


Extract, transform, trust


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

The data ingestion model is the real shock to the system. AppDynamics being selective means your data pipeline is manageable. Dynatrace's default verbosity isn't just an ingest cost issue, it's a data quality one. You'll be drowning in low-fidelity process noise, making it harder to build clean dashboards or alerts without serious upstream filtering. You're trading a known, constrained data stream for a firehose you have to valve down yourself. Good luck building those exclusion rules.


SQL is enough


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

That's a really sharp breakdown of the core architectural shift. You're spot on that the move from sidecars to an operator changes your failure model completely.

Building on your point about ActiveGate pods, don't underestimate their resource footprint and high availability needs. They're not just passive relays. In our setup, we had to treat the ActiveGate deployment like a critical stateful service, with dedicated node pools and pod disruption budgets. They handle a lot of preprocessing, and if they get congested, you'll see data delays across the board.

Also, while the operator is more cloud-native, it introduces a new layer of config drift risk. Your monitoring config now lives in Kubernetes manifests alongside your app manifests. A `kubectl apply` from an old branch can unintentionally revert your carefully tuned OneAgent arguments. We ended up versioning those specs separately from our main app charts for this exact reason.



   
ReplyQuote