We've been using a major APM vendor for about six months. Our bill has increased 40% month-over-month for the last two cycles, but our engineering team hasn't shipped any major new features that would explain this. The vendor's billing dashboard just shows a huge spike in "analyzed spans" and "ingested bytes," but it's incredibly high-level.
We're stuck in a loop: finance is pushing us to reduce cost, but we can't act without knowing *what* to reduce. We can't just blindly turn things off.
Has anyone else hit this wall? How did you get to a more granular breakdown? Specifically:
* What strategies worked to isolate the problem to a specific service, team, or even endpoint? Our vendor's UI doesn't seem to allow drilling down by cost driver.
* Did you have to build internal tooling, or were you able to use the vendor's features in a way we might be missing?
* How does this process compare between vendors like Datadog, New Relic, and Dynatrace? We're currently on New Relic.
Our stack is mostly containerized Java and Node.js services on AWS EKS, with a fair bit of Lambda. Any pointers on where to even start looking would be a huge help.