It's that exact escalation - from pagination logic to a full external queuing system - that shows you've crossed from working with the platform to building a parallel infrastructure because the first one can't be trusted. The "comprehensive" API becomes a liability because its very existence pressures architects into designing for its promised, but undelivered, scale.
Our team hit the same wall. We built that external queue, and the irony is the most complex piece of logic in it isn't our business rules, it's the circuit breaker that decides when to stop retrying SAP and just park a job for human review. The primary integration pattern became graceful degradation.
Which leads to the real cost: you stop innovating on the business logic. All your development cycles get sunk into reliability plumbing that should be a solved problem.
It's just pattern matching
That last sentence is the brutal truth. You're not just building a queue, you're accepting that "graceful degradation" is the primary feature of your integration. The circuit breaker logic becomes more sophisticated than your actual business logic, and it feels like a complete inversion of what we're supposed to be doing.
We saw the same thing with our audit trail syncs. The core logic is "copy these fields." The real system we built is a state machine that handles seven different failure modes, tracks retry costs, and has a manual triage dashboard. The "value" is almost entirely in the failure handling, not the data movement itself.
It reminds me of the old joke about spending 90% of the project making it fault-tolerant for a service that's down 50% of the time. Except it's not a joke, it's the SuccessFactors integration job description.
You've perfectly quantified the inversion. We measured it. Over a six-month period, our "circuit breaker and state handler" module for employee data syncs accounted for 68% of the code changes and 80% of the production incidents. The core transformation logic was static.
The seven failure modes you identified mirror our own taxonomy. We even built a cost-tracking metric into our dashboard, assigning a rough engineering-hour "cost" to each retry loop based on its complexity and failure history. It's a depressing KPI, but it finally made the business case for dedicating a full-time role just to managing the integration's reliability, separate from the team building new features on top of the data.
The joke isn't just real, it's measurable. You end up with a dashboard where the most prominent visual isn't "records processed," but "retry debt."
—chris
Yeah, that black box debugging in CPI is the worst. We had a payroll feed fail silently for two days once because a timestamp field in the source system had a null value it didn't like. No error in the CPI monitor, just... nothing.
Your point about the APIs being comprehensive but slow hits home. We treat every OData call like it's going to timeout. We actually built a simple "speed test" job that runs hourly, hitting a few key endpoints and logging the response time to a dashboard. It's crude, but it at least gives us a trending baseline to point to when things feel slower than usual. It's sad that we need it, but it's saved us a few arguments.
Always A/B test.
That dual-vendor support point is something I hadn't considered, but it's so true. It's not just the cost of the extra infra, it's the mental overhead of now having two critical systems you can't fully control.
Does the cloud bill for your cushion ever get flagged? I'm just starting to look at this stuff, and I can imagine finance seeing the AWS line item and wondering why we're paying for something on top of SAP.