Skip to content
Notifications
Clear all

Rolled out Kustomer to 150 agents - what broke in year one

6 Posts
6 Users
0 Reactions
20 Views
(@claireb)
Reputable Member
Joined: 3 months ago
Posts: 250
Topic starter   [#9650]

Our revenue operations team led the procurement and implementation of Kustomer to serve our 150-agent support organization, with the primary goals of consolidating three legacy ticketing systems and enabling a true omnichannel customer view. After a year in production, the platform has delivered on several core promises, but we have encountered significant, unforeseen operational friction that I believe is essential to document for any organization considering a similar-scale deployment.

The most acute issues manifested in three primary areas:

**1. Performance Degradation with Complex Routing Logic**
While the initial setup of omnichannel queues and basic business rules was straightforward, layering in the nuanced, multi-condition routing required for our specialized product lines caused severe latency. A routing rule exceeding 15 conditions (e.g., customer tier + product SKU + last agent assignment + issue category + language preference) could take 8-12 seconds to evaluate during peak volume, leading to ticket assignment delays and agent idle time. We ultimately had to decompose these into sequential, simpler rules, which introduced a maintenance burden and potential gaps in logic.

**2. Reporting Limitations at Scale**
The out-of-the-box reporting dashboard struggled with our data volume. Generating an agent performance report for a 30-day period, filtered by channel and team, would frequently time out. More critically, we found the data model for custom reporting to be inflexible. Attempting to join conversation data with backend product-version data via the API for a churn-risk analysis required a separate ETL process we had not anticipated, negating the hoped-for "single view" benefit for analytics.

**3. Hidden Costs in "Unlimited" History and Integrations**
The promise of unlimited conversation history became a double-edged sword. While technically unlimited, performing full-text searches across historical data beyond 6 months became prohibitively slow, effectively forcing us into a paid archive solution to maintain performance. Furthermore, while the platform boasts numerous native integrations, the operational version required for our scale (e.g., bi-directional sync with Salesforce for account updates) often fell into a "premium connector" category, adding 20% to our initial projected costs.

A summary of our key pain points versus initial expectations:

| Expectation | Reality at 150-Agent Scale |
| :--- | :--- |
| Intelligent, real-time routing | Latency spikes with multi-condition logic; required rule simplification |
| Consolidated analytics & reporting | Built-in reports failed on large datasets; advanced analytics needed external ETL |
| Omnichannel customer timeline | Performance degraded with >2 years of dense conversation history |
| Predictable total cost of ownership | Significant add-ons for performance (archive) and critical integrations |

In retrospect, our evaluation period focused heavily on feature parity and agent interface usability, but did not sufficiently stress-test the platform's backend performance under our specific load and data complexity. For organizations of a similar size, I would strongly recommend commissioning a performance benchmark using a replica of your most complex routing scenarios and historical data volume *before* finalizing any contract. The question I'm left with is whether these are inherent limitations of the platform's architecture at scale, or simply a failure of our specific configuration.


Method over hype


   
Quote
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
 

That latency with multi-condition rules is really surprising. I'm looking at similar omnichannel routing for a smaller team, and your point about breaking complex rules into sequential ones is helpful, but the maintenance gap you mention worries me.

Did you find any logging or reporting issues when you split the rules? I'd be concerned about losing visibility into why a specific ticket took a certain path.



   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 5 months ago
Posts: 403
 

Lost visibility exactly once before we forced a naming convention. Now each sequential rule appends a custom field called `routing_trail` with its own ID. Ugly but you can at least rebuild the path from the audit log.

Splitting them turned into a version control nightmare though. We keep the rule configs in a git repo now, with a pre-commit hook that validates dependencies. Otherwise someone inevitably reorders a rule and breaks the whole chain.



   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Yeah, that latency hits home. We saw something similar when trying to route based on a customer's predicted LTV combined with active campaigns. The system choked.

A weird workaround we found was to pre-compute some logic. We added a custom field populated by a webhook on ticket creation that combined, say, the product SKU and customer tier into a single value. Then the routing rule just checked that one field. It's not elegant, but it cut our evaluation time down from those 8+ second ranges to under 3. The trade-off is now you've got an extra data pipeline to maintain.


✌️


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

That pre-compute trick is clever, kind of like denormalizing for performance. I wonder if you could push that logic out to a sidecar service and feed it into Grafana for a dashboard? It's an extra pipeline, but at least you'd get visibility into the performance impact vs your old 8-second evaluations.



   
ReplyQuote
(@kubernetes_knight)
Estimable Member
Joined: 7 months ago
Posts: 68
 

Ouch, that latency is painful. We saw something similar, not in Kustomer but in a custom routing engine we built. The sequential rule decomposition is a classic workaround, but it introduces a distributed transaction problem across your rule chain - if a rule in the middle fails its condition, the whole sequence halts.

Have you considered a rules-as-code approach? We ended up moving our critical routing logic into a dedicated, scalable microservice. We could then:
- Version and test the logic in isolation
- Cache expensive lookups (customer tier, product SKU mapping)
- Add detailed tracing, so you *know* which condition took 4 seconds

It's more infrastructure, but it turns a black-box performance problem into something you can observe and scale horizontally. Did your team evaluate pushing any logic outside the platform?


YAML is not a programming language, but I treat it like one.


   
ReplyQuote