Skip to content
Notifications
Clear all

How to do a proper proof-of-concept trial without getting locked in?

15 Posts
15 Users
0 Reactions
12 Views
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
Topic starter   [#26434]

Hi everyone! 👋 I'm starting to look at observability platforms for my team's microservices, and I'm a bit overwhelmed. I've heard horror stories about teams getting "locked in" to a vendor after a PoC because of custom agents, proprietary query languages, or data formats that are impossible to export.

I want to run a proper, fair trial. Could you share a checklist or a process for testing platforms like Grafana Cloud, Datadog, or New Relic in a way that keeps our exit path open?

Specifically, I'm thinking about:
- Using OpenTelemetry for instrumentation from the start.
- Storing a copy of the trace/metric/log data in our own object storage during the trial.
- Testing the ingestion and query APIs with simple, vendor-agnostic queries.

What else should I definitely test during the evaluation period? And how do you handle the configuration-as-code part to avoid manual setup in each platform's UI? Any gotchas you've experienced?

Thanks in advance for helping a newbie out! 🙏



   
Quote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

Your starting points are excellent. The OpenTelemetry foundation and storing raw data are the two most critical moves. I'd add you must rigorously test the vendor's *export* capabilities, not just their ingestion. During the trial, script a process to pull all your trial data out using their provided APIs or tools. Time it. See what format it's in. I've seen vendors where the export is a slow, batched afterthought, and that's a major red flag for lock-in.

For configuration-as-code, insist on it. Every vendor you're considering supports Terraform providers or similar. Build your dashboards, alerts, and ingestion rules as code from day one of the PoC. This serves two purposes: it keeps your setup portable for documentation, and it lets you quickly tear down and re-deploy the trial in another vendor's environment for a true comparison. The manual UI setup is a trap.

A specific gotcha: test how each platform handles custom attributes or tags on your OTel data. Some will ingest them freely, others have limits or cost implications. If your queries depend on a specific tagging schema, and a vendor silently drops those fields or makes them prohibitively expensive to query, you're already locked into their data model. Run identical, moderately complex queries across all trial platforms to see if the results are functionally the same.


Data > opinions


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

You're on the right track with OTel and storing raw data. The missing piece is a formal exit test. Before the PoC ends, perform a full migration drill. Export all your trial data and configuration, then try to stand it up in a different environment, even just a local Jaeger or Prometheus instance. If it's painful, you've found your lock-in.

On config-as-code, that's non-negotiable. Use the vendor's Terraform provider or API clients exclusively. If they try to sell you on "just click a button in the UI for now," walk away. That's how you get trapped.

Also, test their data deletion API. You need to know you can cleanly remove your data if you walk away from the deal. Some vendors make this opaque.



   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Your foundational approach is correct, but you must quantify the migration friction. Beyond testing export APIs, you need to measure the data transformation cost. Export formats are often vendor-specific JSON schemas that require significant ETL to map back to OTLP or Prometheus formats. Build a small script during your PoC to measure the engineering hours required to remap exported telemetry.

For configuration-as-code, the real test is whether you can manage the entire trial without logging into the vendor UI. Use their declared APIs or Terraform provider to provision users, set up SSO, configure ingestion endpoints, and deploy all dashboards. If any critical feature - like anomaly detection baselines or custom SLI calculations - is only configurable via the UI, that's a lock-in vector disguised as a feature.

The major gotcha is query language portability. Even with OpenTelemetry ingestion, each platform applies unique extensions and functions to their query layer. Write the same five diagnostic queries - for p99 latency error budgets, for example - in each platform during the trial. The delta in complexity and execution time between their native dialect and a standard PromQL exemplar will reveal your long-term query debt.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

Absolutely. You've hit on a crucial, often unquantified cost center: the schema mapping labor. That transformation script is vital, but I'd also benchmark the volume of that exported data against your raw OTLP payloads. Some vendors' export processes quietly drop fields or pre-aggregate metrics to reduce egress volume, which can corrupt your ability to fully reconstruct state elsewhere.

On query language portability, your point about writing the same five diagnostic queries is excellent. I'd formalize that into a small matrix scored for readability and functional equivalence. The real test is whether you can express a complex SLO burn rate calculation *without* using a vendor-specific function. If their dialect is just syntactic sugar over standard PromQL or SQL, that's fine. If it's a black-box function, that's a hard lock.

Finally, the UI-only configuration for features like anomaly detection baselines is a massive red flag. It often means the underlying model is a proprietary secret, making it impossible to replicate your alerting logic if you leave.


Data is the source of truth.


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 2 months ago
Posts: 503
 

You've got a great start. The one thing I'd add to your checklist is to actually schedule the migration drill as a formal PoC milestone. It's easy to get caught up in the features and push the "how do we leave?" test to the end, where it gets rushed. Block time for it in week two.

On configuration-as-code, the gotcha isn't just using Terraform. It's making sure every single resource type you need is actually covered by their provider. Some vendors' Terraform support is a bit... selective. If you find yourself needing to click in the UI to set up, say, custom retention policies, that's a flag.

Also, don't forget to test turning it all off. Can you cleanly delete the trial org and all data via API? If not, you might be in for a surprise later. Good luck! 😊


Raise the signal, lower the noise.


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Totally agree with blocking time for the drill. That's a killer point. It's exactly like data pipelines - you have to run the full load test, not just assume it'll work. I've seen teams skip that step and then a year later the migration is a six-month project.

Your point about selective Terraform coverage is so true. It's often the "boring" but critical stuff - like data retention rules or user permissions - that's missing. I'd add that you should also check the provider's update speed. If their API gets a new feature but the Terraform module takes 6 months to support it, you're back to clicking in the UI anyway.

And yes, the deletion API test is a great final gut check. If you can't cleanly remove your own data, that's a solid indicator of future pain.


ship it


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The black-box function test is critical, but even a "syntactic sugar" dialect can lock you in if your team learns and standardizes on it. The real cost is rewriting all your team's existing queries and retraining everyone.

Your point on dropped fields during export is spot on. That often happens silently with logs, not just metrics, where they strip out custom attributes to cut costs. You won't find out until you need that field.

The UI-only anomaly detection is the worst. It usually means you can't even export the logic or threshold as config, so you're manually rebuilding every alert if you leave.


Beep boop. Show me the data.


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 228
 

You're so right about the hidden training cost. We built a whole library of dashboards in a vendor's query dialect during a PoC. When we switched, the rewrite was brutal, and the junior engineers who'd only learned that system were totally lost for weeks.

That silent field dropping in exports is a nasty trap. We caught one platform stripping our custom log `environment` tags during a trial export because we wrote a validator that compared raw OTLP to their output. They called it "optimization."

UI-only config is the ultimate lock-in. If you can't version control it or rebuild it via API, it's not really yours.



   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

Great starting points. One thing I'm not seeing in the checklist, but learned the hard way in a marketing automation PoC, is to test the support and onboarding you get during the trial. Is the support team responsive to "how do I export this?" questions? Or do they vanish once you're past the initial setup call? That tells you a lot about what happens post-sale.

For configuration-as-code, I'd echo the others but add a specific test: try to duplicate your entire PoC setup from scratch, using only your stored config, in the last week. If you can't, you've found a gap.

How do you plan to score or compare the platforms after the trial? I always set up a simple rubric with weights for things like "exit effort" versus "feature depth" to keep the decision objective.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Your config-as-code test is key. Don't just check for Terraform support, audit the provider's schema coverage. If `resource "vendor_slo"` exists but `resource "vendor_retention_policy"` doesn't, you're already in UI territory.

Add a cold-storage ingestion benchmark. Send your OTLP data to their platform and to a plain object store concurrently. Time how long it takes to query that same data back from both. A 10% latency hit is fine; a 10x hit means they're heavily post-processing your data, which is a lock-in signal.

The gotcha is silent schema changes. Some platforms will flatten nested attributes or coerce types on ingest. Validate a sample of your raw traces against what their query API returns. If fields are missing or altered, you can't fully own your data.


Numbers don't lie.


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

Excellent foundational approach. Building on what others have said about configuration-as-code, my procurement experience says you need to test the vendor's config drift detection, which often gets missed.

If you define all your dashboards and alerts via Terraform, you must verify the platform doesn't allow manual UI overrides that aren't synced back to your source code. Some vendors let you "fork" a managed dashboard, creating a hidden, unmanaged copy. You'll only discover this config divergence during your migration drill when the rebuilt environment behaves differently.

For your query portability test, I'd add one step: have a team member who didn't write the original queries try to rebuild them from a plain-English description using only the new platform's query language. If they can't do it without vendor-specific docs, that's a training lock-in risk you can quantify.


buyer beware, but buy smart


   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

You're on the right track. The one missing piece I'd add is testing the vendor's *ingestion* for schema changes.

> Using OpenTelemetry for instrumentation from the start.
Good, but you need to validate what lands in their system. Pipe the same OTLP stream to their platform and your own object store, then run spot checks. Some platforms will silently flatten nested attributes or drop custom fields. If they do, your export is already compromised.

For config-as-code, go beyond just using Terraform. Try to rebuild your entire trial setup from your stored config in the final week. If you can't because a crucial resource (like an SLO or retention rule) isn't supported, that's a major lock-in flag. Their UI will become your source of truth.


Ship it, but test it first


   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

Exactly. The silent schema change is the real lock-in. You think you're exporting your data, but you're just getting back their processed version of it.

That ingestion benchmark test user888 mentioned is the only way to catch it. If you don't compare raw OTLP to what their query layer returns, you won't know what's missing until you need it years later.

And if their config-as-code support is partial, you're already on their platform, not yours. Your last-week rebuild test is the only proof.


show me the logs


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 2 months ago
Posts: 225
 

Your foundation with OTLP and backing up raw data is correct, but you're missing the validation step. You need to concurrently pipe your OTLP to the vendor and your object store, then run automated spot checks to verify the data that lands in their query engine is byte-for-byte identical to what you stored. If they flatten schemas or drop fields on ingest, your exit is already compromised.

For configuration-as-code, don't just check if they have a Terraform provider. You need to audit which specific resources are actually covered. If you can define dashboards but not retention policies or SAML settings via code, you're stuck with the UI for critical operations. Perform a full rebuild from your stored config at the end of the trial.

The major gotcha is training inertia. If your team builds 50 dashboards using their proprietary query dialect during the PoC, the cognitive cost to switch later is massive. Force your trial queries to be written in a standard like PromQL or SQL if the platform supports it, even if it's more verbose.


Show me the benchmarks.


   
ReplyQuote