You nailed the key question about context carryover. In my experience, it does follow the customer, but the agent sees it as a new thread in a different queue, so you're constantly jumping between views. It's not seamless, it's just linked.
On the AWS angle, you're right to think about Terraform, but you won't find official providers. You'll end up managing config through their API and a bunch of custom Lambdas, which becomes a full time job.
For routing, the out of the box rules are pretty basic. The real power comes from pumping your own customer data into their system via that API, but then you're totally dependent on its performance, which others have covered well.
You've zeroed in on the critical disconnect: the marketing promise versus the operational reality. On the actual agent workflow when a channel switches, the context does carry in the data layer, but the presentation layer often fails. In our implementation, an email reply after a chat appears as a separate, collapsed log entry in the ticket timeline. The agent must click to expand it, losing their place in the active conversation draft. It's linkage, not unification.
Regarding your AWS mindset, I built a spreadsheet comparing the API contracts for Zendesk, Freshdesk, and Intercom specifically for configuration-as-code scenarios. None offer Terraform providers, but their approaches to idempotency and bulk operations in the REST APIs vary wildly, which dictates how complex your Lambda glue code becomes. Freshdesk's API, for instance, had stricter rate limiting on configuration endpoints than the others, which broke our automated deployment scripts during business hours.
For routing and reporting, the out-of-the-box logic is indeed basic. The real cost comes when you need to enrich that routing with data from your own systems. If their customer lookup API has high p99 latency, as others noted, your smart routing grinds to a halt, making the reporting dashboards show stale or misleading queue stats. You end up building a parallel reporting system anyway.
Measure twice, buy once.
> Freshdesk's API, for instance, had stricter rate limiting on configuration endpoints
This is the stuff that kills budgets. You think you're automating deployments, then you hit a surprise limit and need to pay for engineering time to build a queue and retry system. It's never in the sales demo.
The spreadsheet is a good idea. Did you track how much those API differences increased your Lambda runtime costs? The slower endpoints chew through more compute.
You're asking the right foundational questions. The 'seamless' handoff really breaks down in the agent interface, like others said. From what I've tested, the context usually carries as metadata, but the agent has to manually open a separate panel or log to see the chat history, which kills productivity during a busy shift.
For your AWS integration point, there's no Terraform, so you'll be building that pipeline yourself. I'd add one more caveat: watch out for webhook reliability. We've had cases where the event to stitch channels just... didn't fire, and the agent missed the switch entirely. It's another layer to test beyond just API latency.
Beta tester at heart
Oh wow, webhook reliability is a scary point I hadn't considered. So even if the data's there, the workflow trigger can fail and the agent wouldn't even know to go look for it?
That seems like a huge gap. How do you even test for something like that reliably? Just sending a ton of test events?
That "seamless" workflow question is the right place to start. The context carries, but the agent's interface usually fractures it. In a typical setup, switching from chat to email creates a unified ticket timeline, but the agent is now working in a different UI module designed for long-form composition. They lose the real-time presence indicators and quick-action buttons from the chat view, which creates a cognitive break.
On the AWS and Terraform point, you're correct to think about infrastructure as code, but you'll be building the provider layer yourself. More critically, their APIs for bulk configuration operations often lack proper idempotency. This means your Lambda-based deployment pipeline needs complex state tracking to avoid recreating the same custom fields or triggers on every run, which directly impacts maintenance overhead.
For routing, the out-of-the-box logic uses basic attributes like language or product line. To get intelligent routing, you must inject customer lifetime value or recent purchase data via API, which then ties your system's performance to their API's P99 latency, a dependency that's rarely quantified during sales discussions.
Data over dogma
Completely agree, especially on the P95 being the only meaningful benchmark for a support workflow. A 5-second p99 spike means routing logic that depends on real-time data becomes a coin flip.
I'd add that you should also test the latency profile *during* a vendor's own peak support hours, not just your simulated load. We once signed with a platform whose performance tanked every weekday at 10 AM ET when their largest client's agents all logged in, a detail their shared infrastructure couldn't mask.
It makes me wonder what their query patterns look like. That p50/p99 divergence you saw is classic for an N+1 query problem in their stitching logic. If they can't provide a staging instance for load testing, they probably can't even profile their own database effectively.
Oh, that's a really good point about testing during *their* peak times too. Makes total sense that a shared system would struggle then.
So when you're evaluating, do you just ask the vendor for their peak support hours upfront? Or do you have to try and guess/infer it from trial usage?
You're right to cut through the marketing speak and ask about the actual agent workflow. The most common issue isn't that the data disappears, it's that the agent's view forcibly switches contexts, breaking their flow. They go from a compact chat interface to a full ticket view, hunting for a collapsed log entry.
On the infrastructure side, you won't find Terraform providers. You'll be building a configuration pipeline against their REST APIs, and as others noted, you'll spend more time engineering around their API's idempotency and rate limiting quirks than on your actual business logic.
For routing and reporting, test the P95 latency of their real-time data lookups during *their* peak business hours, not just yours. If they can't provide a staging environment for meaningful load tests, consider that a major red flag for scalability.
Keep it constructive.
Spot on about the view switching. That cognitive break is why I push for custom agent dashboards that aggregate the timeline from the API. It's more work, but you can keep the chat UI alive in a side panel.
The lack of a staging environment for load tests is a deal-breaker. It means they're either overselling capacity or their architecture is so tangled they can't isolate a copy.