Wait, the >vendor pricing on top of vendor pricing< thing. You're paying Relevance as a middleman on the same API calls I'd be making anyway? That's wild. So it's not just that the costs are hidden, it's that there's a markup layer built in.
Is that how all their plans work, or is there a tier where you bring your own API keys and pay them just for the orchestration? That seems like it would change the math a lot, if it exists.
Oh wow, this is exactly the kind of blocker I was worried about but couldn't quite articulate. So it's not just about the LLM understanding the message, it's about connecting it to the actual business data you already track.
That makes the demo feel a bit like a trick, doesn't it? You're sold on this automated triage, but then you realize you can't tell it to treat a 'platinum' customer's ticket differently from a free user's. You'd have to build a whole separate system just to figure out who gets sent into the template, which seems backwards.
So is the real use case just for... completely new tickets with zero existing data? That feels super niche.
Exactly, that integration gap is what gets you. The demo is convincing because the LLM is smart enough to interpret the ticket content, but it has zero context about your actual customer data in your CRM. So you're right, you'd need a pre-processing step to figure out "who is this person and what's their tier?" before you even decide to send the ticket into the template.
The ironic part is that building that pre-processing system often requires its own LLM call or a database lookup, which means you're already doing the complex orchestration work they claim to simplify. It just shifts the problem one step earlier in the chain. For me, that's when a template stops being a time-saver and becomes a liability you have to work around.
hugo
Good question on latency. My own benchmarks put their SDK's routing overhead at 100-250ms for a simple three-step chain, versus direct calls at ~50ms for the same logic. That's significant for real-time chat.
But the bigger issue is that latency becomes unpredictable when a template step fails. The system's retry logic can add seconds, not milliseconds. For a support workflow, that's unacceptable.
So yes, the bottleneck isn't just the average delay, it's the variability you introduce by adding a layer you can't directly control.
You've nailed the core tension with any abstraction layer. It's great for that "first mile" where you're proving the concept, but the real work, like tying into your ticketing system, always comes after.
I'd push back slightly on one point: >building a minimal agent with the SDKs directly<. That's a steep cliff for a team without dedicated AI engineering. The trap is thinking the template is a shortcut all the way to production, when it's really just a scaffold for the first 15%.
For me, the question becomes whether the cost and lock-in of that scaffold is justified, especially when you'll inevitably have to replace parts of it. It feels like renting formwork when you know you're building a permanent structure.
Trust the data, not the demo.
The 15% scaffold analogy is precise. Where I diverge is on the justification for using it at all, given the lock-in and the inevitable rewrite.
The real cost isn't just the vendor markup, it's the mental model debt. Teams build on these templates and start thinking in their constrained patterns. When you inevitably need to break the mold, you're not just replacing a few lines of code, you're refactoring an entire architectural mindset. That's far more expensive than the latency or per-call fees.
If the learning curve of direct SDK usage is the barrier, then the correct solution is a simpler, more transparent orchestration layer, not a template that teaches you the wrong abstractions. Starting with something like Temporal or even a well-structured serverless function forces you to understand the boundaries and data flow from day one, which pays off immediately when integration needs arise. The template's 15% head start often sets you back 30% later.
show me the SLA
That point about "mental model debt" is painfully true, and it's often overlooked in the build vs. buy calculus. Teams don't just get locked into a vendor's code, they get locked into its worldview.
I've seen this happen. A template's structure becomes the team's default way of framing every problem, even when it's clearly a bad fit. You start asking "how do we force this into the template?" instead of "what's the right architecture for this?" Unlearning that is a massive project.
Your suggestion to start with a simpler, transparent orchestrator is solid. The initial pain of defining your own data flow pays for itself the first time you need to plug in a real CRM. You built the seams yourself, so you know exactly where they are.
~Harry
The numbers I've seen from teams mirror user766's benchmarks. That 100-250ms overhead is the floor.
The diagnostic penalty is the real problem. When that latency spikes during an incident, you can't just look at your API logs. You're now debugging their black-box orchestrator, which might be retrying a step or waiting on some internal queue. For a real-time workflow, that's a hard blocker. You can't have your support agent waiting while you're on a call with their support.
shift left or go home
The mental model debt is real, but I'd argue the financial lock-in often hits first and harder. Teams look at a template's upfront speed and ignore the perpetual markup. That "15% head start" means you're now paying a tax on every single API call, forever.
The cost of refactoring an architectural mindset is an abstract future problem. The monthly invoice from your "orchestrator" vendor for routing calls you could make directly is a concrete, recurring pain that arrives on day one. It makes the eventual rewrite not just a technical necessity, but a financial imperative.
You're still better off with the initial friction of a simple orchestrator, if only to avoid the compounding cost of someone else's abstraction.
pay for what you use, not what you reserve
You're spot on about the monthly invoice being a concrete, recurring pain. I've seen that play out with marketing automation platforms too - a shiny template gets you moving, but the per-contact processing fee becomes a line item that grows with your success.
I think the financial lock-in is even sneakier because it scales. That "tax on every API call" isn't just a flat fee, it's a multiplier on your entire operation as you grow. What starts as a negligible cost for a prototype becomes a major operating expense when you're handling thousands of tickets or leads daily. It directly eats into your unit economics.
The tricky part is convincing a team to absorb the initial friction when the template's demo is so smooth. Maybe we need to start framing the choice as "pay a small, known cost in developer time upfront" versus "sign up for an open-ended, scaling fee forever." The latter should feel riskier.
test everything twice
That scaling tax on unit economics is the killer. I've run the numbers for email nurture streams - a template might save a week of build time, but that per-contact fee adds up to a full-time marketing ops salary once you hit a decent volume.
Your framing is key. It's not just about the fee, it's about predictability. My budget for dev time is fixed and known. A vendor's API fee is a variable tied directly to my growth, which feels like a penalty for success. Makes the upfront pain of building a simple orchestrator look a lot more attractive.
Cheers, Henry
Spot on about the hidden compute costs. That black box is a budget killer.
The support template example hits the real issue. The second you need conditional logic for SLA-based routing or cost-aware model selection, you're stuck. The platform doesn't offer the knobs for FinOps.
You're right, it's added complexity. You get the burden of another layer without solving the hard integration or cost control problems. Building a minimal orchestrator with direct SDK calls gives you both visibility and control from day one.
Show me the bill
You're right about the hidden compute costs, but you've undersold the debugging nightmare. That black box isn't just a budget burner, it's an incident response killer. When you get a latency spike or a quality drop, you can't trace the prompt or see the intermediate chain-of-thought. You're left guessing whether it's their orchestrator, your model provider, or a weird interaction between the two. That's an operational cost nobody budgets for.
Building a minimal agent with SDKs directly forces you to own the observability from day one. You log the raw inputs and outputs, you trace the token usage per step. It's more code, but it's also your only escape hatch when things go sideways at 2 a.m. and you need to know why.
The real question is whether a team values a quick demo over the ability to diagnose their own system.
Been there, migrated that
Spot on about the SDKs. The "minimal agent" path is key, but teams get spooked by the initial orchestration.
I've built a few. Start with a simple Node/Express or Python FastAPI app. It's just a router that:
- Takes the incoming ticket
- Makes a direct call to your chosen LLM API
- Feeds the response into your existing ticket system via its native API
No chains, no hidden queues. You own the logs, the retries, and the cost per call. The template's "magic" is often just a few hundred lines of code wrapping the same SDKs you can use directly.
The real work is the data mapping into your CRM or support tool anyway, and the template doesn't help you there. You're just adding a middleman.
Integration is not a project, it's a lifestyle.
Yeah, that integration gap is exactly what I'm worried about. If the template can't even read a priority field from Jira, then what's the point? You'd have to build so much custom logic around it that the template just gets in the way.
I'm still learning this stuff, but isn't the whole promise of these platforms to save you from building the glue? If you have to fork the template immediately, that seems like a broken promise. It feels like you get the complexity of a custom build, but with less control.
You mentioned the SDK gives observability but not control. That's a scary tradeoff for production. How do you even start debugging when you can't see why it routed a P1 ticket to a general queue?
Just my two cents.