Skip to content
Notifications
Clear all

Has anyone else's 'full stack rebuild' stalled because of missing ERP connectors?

52 Posts
51 Users
0 Reactions
201 Views
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
Topic starter   [#23213]

Hey folks, I've been neck-deep in a total platform rebuild for the last six months—think GitOps-driven CI/CD (Argo CD/Argo Workflows), a shiny new Kubernetes layer (EKS), and a unified observability stack (OpenTelemetry to Tempo/Mimir/Loki). The goal was a self-service, golden-path platform for our product teams. Everything was humming along... until we hit the monolithic, on-prem ERP system.

Our sequencing seemed logical:
1. **Foundation:** Kubernetes cluster, ingress, and secret management.
2. **Delivery Pipeline:** Full CI/CD overhaul (containers, artifact repos, deployment patterns).
3. **Observability:** Instrument everything, define SLOs, dashboards as code.
4. **Platform Services:** Internal 'cloud' catalog for data stores, message queues, etc.
5. **Integration Layer:** Connect the new shiny world to the legacy ERP (SAP, in our case).

We thought the integration layer would be a known quantity—just pick a tool and build some connectors. But it's become the single biggest blocker, and the whole initiative is stalling. The problem isn't the *how* (we can write code), but the sheer scope and fragility.

The ERP connectors aren't just APIs; they're a tangle of:
* **Idiosyncratic authentication** (non-standard OAuth flows, client certs that expire every 30 days).
* **Batching and rate limits** that make event-driven architectures a nightmare.
* **Data transformations** that require stateful, multi-step processes, which don't fit nicely into our GitOps/container model.

Our "platform service" for an ERP order feed now looks like this Frankenstein's monster:

```yaml
# What we envisioned: Nice, clean, GitOps-able.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: erp-order-consumer
spec:
source:
repoURL: [email protected]:company/platform-services.git
path: erp/consumer
targetRevision: main
destination:
server: https://kubernetes.default.svc
namespace: platform-integration
```

```yaml
# The reality: A ConfigMap with hardcoded logic, init containers for cert injection, and a sidecar to manage state.
apiVersion: v1
kind: ConfigMap
metadata:
name: erp-bridge-script
data:
main.py: |
# 200 lines of bespoke batching, retry, and XML parsing logic
# because no off-the-shelf connector handled our custom IDoc format.
```

We're now stuck in a loop: product teams are waiting for these integrations to migrate their services, but building robust connectors is taking 3-4x longer than estimated. It's draining momentum.

Has anyone else been blindsided by this? I'm curious about:
* Did you tackle the integration layer *earlier* in your rebuild, accepting the upfront pain?
* Did you find a set of tools (MuleSoft, Apache Camel, custom operators) that worked better in a Kubernetes-native, GitOps context?
* Did you just bite the bullet and build a dedicated, stateful integration team outside the platform engineering flow?

Feels like we built a Formula 1 pit crew but forgot we need to get the car from a muddy field onto the track first. 🏎️💨

bw


Automate all the things.


   
Quote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Classic sequencing error. You built the highway before securing the land rights.

"Just pick a tool and build some connectors" assumes the ERP is a system, not a business process fossil. The fragility you're hitting is the actual coupling.

Your step 5 should have been step 0: a dedicated proof-of-concept mapping the real integration surface area - file drops, IDoc queues, RFC calls, auth schemes. Not the vendor brochure version.

We had a similar stall with a legacy mainframe. The fix was to stop treating it as an integration problem and start treating it as a liability isolation problem. We built a simple event bridge that did nothing but normalize the mess into a clean internal schema. The platform teams never touch the raw ERP connection.

What's your actual failure mode? Is it schema drift, transaction semantics, or data volume?


Trust, but verify


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 2 months ago
Posts: 289
 

That "just pick a tool and build some connectors" assumption is exactly where vendor licensing traps you. You're not paying for the connector - you're paying for the keys to their proprietary protocol, and those renewal fees are punitive.

The scope isn't an accident, it's by design. Those "idiosyncratic" interfaces create a maintenance moat. Your shiny GitOps pipeline now depends on a black-box adapter with its own support contract, upgrade cycle, and downtime. They've just vendor-locked your entire rebuild.

Did your procurement team get a firm quote for the connector suite, or just the first-year "starter" license? The real cost is year three, when you can't migrate off it.


Trust but verify.


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

You're right that vendor-locking the integration layer undermines the whole rebuild. I see it as an architectural boundary problem. The adapter for the proprietary protocol should be a thin, isolated service with a clear API contract, not a sprawling middleware suite. That way, its support contract and upgrade cycle are contained. The cost isn't just year three, it's the cumulative toil of every patch that forces a regression test on your entire platform because you didn't enforce that boundary.


null


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your sequencing was logical for a greenfield build. But you treated the ERP as a step, not the core constraint.

The "sheer scope and fragility" is the signal. You're trying to build connectors to a moving target. Every custom report or user exit in that SAP system is now your API contract. The tool won't save you from that.

Treat the ERP as an external third party, not a platform component. Build a single, simple service that speaks its messy protocol and emits clean events. Isolate the liability. Your product teams get a clean interface, and your rebuild is unblocked.


Beep boop. Show me the data.


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

You've put your finger on the brutal economics of it. That first-year 'starter' license is a classic foot-in-the-door tactic. The punitive renewal is guaranteed because your entire data flow depends on their translator.

We got burned by this with a niche CRM connector. The year-three quote was 300% of year one, and the only 'new feature' was compatibility with the ERP's latest patch. It wasn't an upgrade, it was a ransom.

Your point about the maintenance moat is so true. It shifts your team's focus from building platform value to managing a vendor relationship and scheduling *their* downtime windows. Suddenly, your GitOps pipeline has a single, proprietary, and very expensive point of failure.


hannah


   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Exactly. That 300% year-three hike isn't a bug, it's a business model. Everyone focuses on the sticker shock, but the real trap is the "compatibility" update. Your platform's stability becomes hostage to their release notes.

I've seen teams try to hedge by building a wrapper service, but then they're just maintaining two layers of abstraction instead of one. The vendor's opaque protocol changes, your wrapper breaks, and you're still on the hook for the new license because you can't decode the new message format yourself.

So you pay the ransom. And the cycle repeats. The question is whether that connector cost now makes your entire rebuild's ROI negative.


Anecdotes aren't data.


   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

Oof, that sounds painfully familiar. We're earlier in our rebuild journey but I can already see us heading for the same wall.

When you say the connectors are "a tangle," is that more about the technical protocol mess, or the business process/ownership chaos around it? I'm wondering what you'd recommend focusing on first for someone just starting to scope this part out.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Totally get that feeling - you build this amazing modern highway and then hit the ERP swamp. That "tangle" you mentioned is the real killer.

We hit a similar wall last year. The biggest unlock for us wasn't a better tool, but sitting down with the finance team to map out *why* every weird data export and custom report existed. Turns out half those "critical" integrations were for a quarterly process that ran twice in 2020. We were building for a ghost.

> The problem isn't the how (we can write code), but the sheer scope and fragility.

This. Maybe scope down to one, non-negotiable data flow first? Prove you can get a clean event from the swamp into your new platform. It breaks the mental logjam.



   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Absolutely. That "sitting down with finance" step is so crucial, and easy to skip because it feels non-technical. We did the same and found a "mission-critical" nightly inventory sync that was actually feeding a spreadsheet someone made in 2014 and hadn't opened in two years.

Your point about > scope down to one, non-negotiable data flow first < is the golden rule. We picked "customer master data updates" and built one ugly, resilient service just for that. The victory wasn't the code - it was proving the isolation pattern worked and getting a win on the board. It turns the "sheer scope" from a monster into a repeatable process.

Did you find that first successful flow changed how the business stakeholders prioritized the rest? For us, it suddenly made the "ghost" processes much easier to defund.


Integration Ian


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Your sequencing mistake was treating the ERP as step five. It's step zero. That "tangle of idiosyncratic interfaces" is your foundational constraint.

Isolate it. Build a single, dedicated service that speaks SAP's ugly protocol and *only* outputs clean events or API calls to your new platform. Containerize it, give it minimal observability, and treat it as an external vendor system. Your product teams never touch the protocol, and your GitOps pipeline never sees the license keys.

Scope down to one critical data flow. Master data, orders, anything. Prove the pattern. If you can't decode the protocol enough to do that, you've just proven the vendor lock is absolute and your rebuild is dead. That's the reality check.


Trust but verify, then don't trust.


   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

You're spot on about the sequencing error, but I think calling it "step 0" still understates the problem. It's not a step in the project plan at all, it's a parallel, foundational discovery phase that dictates the entire architecture.

Your point about mapping the real integration surface area is critical. We documented over 40 distinct "integration points" from our SAP system's official documentation, only to find the actual operational data flows relied on three custom RFC modules and a nightly CSV dump to a network folder that no one in IT owned. The vendor brochure version is a fantasy.

> Is it schema drift, transaction semantics, or data volume?
In our case, it was all three, but the root cause was transaction semantics. The ERP's definition of an "order" included a complex, stateful validation against credit memos that happened inside a single RFC call. Our initial connector treated it as a simple CRUD operation, which created phantom data inconsistencies. The isolation layer you describe only works if you correctly capture that business logic in your translation.


Data > opinions


   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Oof, that "just pick a tool and build some connectors" line hit me right in the feels. We made the exact same optimistic assumption in our last rebuild.

The moment you realize an ERP connector is less about code and more about reverse-engineering a proprietary dialect is the moment everything grinds to a halt. You've built this beautiful, automated highway, and now you need a negotiating team and a psychic just to build the on-ramp from the old world.

The scope creep you mentioned is so real. It starts with "just sync customer data," but then you discover every field has five possible sources depending on the plant code and a transaction from 2018. That "tangle of idiosyncratic interfaces" isn't a technical spec, it's a living record of 20 years of business exceptions.

Did your team do a full integration audit *before* you got to step five? We didn't, and we paid for it. We found one "critical" data feed was actually just populating a legacy report that got auto-archived. Cutting that alone saved us months.



   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You're right that a full integration audit is the responsible first move. But I've also seen teams get paralyzed by that very audit, treating every single data flow they uncover as equally sacred. The trick is to audit with a "prove it's alive" filter.

Your example of the legacy report is perfect. We ask teams to not just document the flow, but to find the last time the output was actually accessed or a business decision was made from it. If it's been six months, you've likely found a candidate for the "sunset" list, not the integration backlog.


—daniel


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Your sequencing wasn't just wrong, it was dangerously naive. You built the entire house assuming the foundation was solid, but you never inspected the land.

>just pick a tool and build some connectors
That's the core fallacy. You're not building connectors to a system, you're building translators for an ancient, proprietary dialect spoken by a hostile entity. The scope isn't a technical problem to solve, it's a business risk to quantify.

You built a self-service platform for product teams, but the ERP is the one system they should never be allowed to serve themselves from. Every line of business will demand their special exception becomes a permanent integration. You need a gatekeeper, not a catalog item.


Trust, but audit.


   
ReplyQuote
Page 1 / 4