Skip to content
Check out what I ma...
 
Notifications
Clear all

Check out what I made: A checklist for first-time tool evaluators.

2 Posts
2 Users
0 Reactions
14 Views
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
Topic starter   [#15966]

Everyone asks "how do I pick a monitoring tool?" They get lost in feature lists and marketing. You need a decision framework, not another sales pitch.

I built a checklist. It forces you to define concrete requirements *before* you look at a single product demo. Use it to avoid picking a shiny toy that can't handle your actual scale.

**Core Requirements (Answer First)**
* **Primary Use-Case:** APM? Infrastructure? Logs? Real user monitoring? Be specific.
* **Data Volume:** Estimate per second/minute/hour. Not "a lot." Give me numbers.
* Example: `~50k metrics/min, 10 GB logs/day, 100 traces/sec`
* **Retention & Cost:** How long do you need data? What's the actual budget? Include infra costs if self-hosted.
* **Non-Negotiables:** Must have Prometheus compatibility? Must be SaaS? Must support OpenTelemetry?

**Technical Evaluation Checklist**
* **Ingestion:** OTLP? Specific exporter? Pull vs. push?
* **Querying:** PromQL? LogQL? Latency under load?
* **Storage:** Read/write performance benchmarks for your volume.
* **Reliability:** How are HA and data durability handled?
* **Integration:** How does it fit with your existing stack (k8s, cloud, CI/CD)?

Skip the "easy setup" demo. Test with your worst-case load pattern. If they can't provide benchmarks for your scenario, move on.

—DD


Metrics don't lie.


   
Quote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Agreed on quantifying data volume first, but I'd stress you also need a benchmark for query concurrency. Tools that handle your base ingestion rate can still fall over when five engineers run complex PromQL queries simultaneously.

You mentioned "Read/write performance benchmarks for your volume." That's critical. Too many teams just trust vendor claims. You should run a synthetic load test that mirrors your actual cardinality and query patterns - I've seen systems degrade by 40% when you introduce high-cardinality labels.

What's your plan for testing the retention period's impact? Query latency often degrades as hot storage fills, which a short PoC can miss.


-- bb42


   
ReplyQuote