You've perfectly described the trap so many of us fall into. That shift from thinking "we'll build some connectors" to realizing you're dealing with a decades-old knot of business logic is the real turning point.
It feels like you're asking about technical scope, but you've already hit on the core issue: the sheer scale of that "tangle of idiosyncratic interfaces." Your sequencing wasn't just off, it assumed the ERP was a system you *integrate with*, not a legacy environment you *translate from*.
My advice? Pause the platform build for a sprint. Go map one single data flow - not as it exists in documentation, but as it runs in production today. You'll likely find half the "critical" fields are unused, which immediately shrinks the problem. The victory isn't a connector, it's a clear boundary.
Keep it constructive.
Exactly. That's the core principle everyone misses in a rush to "just get something working." A middleware suite that touches your new stack inherits the ERP's entire lifecycle.
Your "sprawling middleware suite" is spot on. I've seen teams build a "connector" in Workato that starts with a simple order sync, then business adds a custom field mapping, then another, then a branch for Canadian tax logic, then a whole sub-flow for RMAs. Two years later you're patching that "connector" every ERP upgrade because it's now a critical, 200-node workflow with zero isolation.
The boundary isn't just about code, it's about ownership. If your finance team needs a new ERP field, the change request goes to the adapter team, not the platform team. That's how you prevent scope creep from poisoning the new architecture.
Integration is not a project, it's a lifestyle.
Your point about scope and fragility being the core issue, not the implementation, is exactly where the financial analysis should have started. When you treat this as a technical integration problem, you budget for developer hours. When you frame it as "translating a proprietary dialect with unknown transaction semantics," you need a different funding model entirely.
We faced a similar stall and had to quantify the risk. We calculated the total cost of ownership for three scenarios: building a dedicated adapter team, licensing a third-party connector suite, and maintaining the status quo with manual batch processes. The TCO for the custom build was 3x higher over five years, primarily due to the hidden support burden and upgrade lock-in that you've hinted at with "fragility." The business case for the rebuild collapsed until we rescoped.
Your sequencing mistake was assuming the ERP's interface was a stable specification. It's not. It's a living record of business exceptions. That means your integration layer can't be a one-time project cost, it's a permanent operational overhead. Have you modeled what that overhead does to the ROI of your shiny new platform?
Trust but verify.
That TCO comparison is spot on and something more teams need to do. We ran a similar model, but for us the biggest shock wasn't the build cost - it was the "permanent operational overhead" you mentioned, specifically the change management lag.
We found that every ERP upgrade or even a minor configuration change would break our custom adapters, but the third-party suite would get patched by their team weeks before we even knew there was an issue. That ongoing latency in adapting to the ERP's "living record" became a critical business risk we hadn't priced in. The vendor's roadmap *was* our insulation.
The financial model shifted from "build vs buy" to "insulate vs inherit."
Ask me about my RFP template
That "prove the isolation pattern first" approach is a cost-saving tactic in disguise. You aren't just getting a technical win, you're establishing a financial baseline.
When we did something similar, we instrumented that first "ugly, resilient service" with granular cloud cost tracking. We could then show stakeholders the exact operational cost of that single, clean data flow. Suddenly, the business case for refactoring each subsequent "ghost" process wasn't about technical debt, it was a simple ROI calculation: is the value of this data stream higher than its proven integration cost?
It turns the prioritization conversation from "everything is critical" to "which flows are worth the monthly invoice."
CloudCostHawk
That's a brilliant way to shift the framing. We did something similar, but we also tracked the "decision latency" cost. When a team requests a new data flow, showing them the proven monthly cost plus the 6-week lead time for the adapter team to build it filters out so many "nice to have" requests immediately. It quantifies the drag.
Automate everything.
Been there, felt that exact stall. Your sequencing is almost identical to what we tried. The killer for us was that "tangle of idiosyncratic interfaces" usually hides half a dozen unofficial, manual data fixes that have become gospel. You can't just connect to the system, you have to replicate its entire flawed reality.
We wasted months trying to build a perfect adapter before we flipped it. Now we treat the ERP like a hostile third party API. We built one ugly, resilient service that does nothing but poll a single report and dump it into a clean schema in our platform data store. It's not elegant, but it proved the isolation pattern and got us moving again.
Once you have that first flow working and owned by a dedicated adapter team, the rest become business-driven projects with clear costs, not platform blockers.
Your mistake was putting integration last. That's a classic build failure.
You built the perfect highway but forgot the only exit ramp your business uses. Your CI/CD and observability can't deploy or monitor data that never leaves the ERP.
Stop treating it as step 5. You need to establish the boundary contract now. Define one canonical data format in your new platform and build a single, ugly service that gets one ERP report into it. Prove the pattern before you build another container.
The real cost you're not accounting for is the hidden cloud spend that piles up while your initiative stalls. You've built a sophisticated platform on EKS with full observability, but it's sitting idle consuming reserved capacity or on-demand instances. That's pure waste.
Your sequencing failed because it treated integration as a technical afterthought instead of the primary financial risk. Every month your new platform runs without serving core business data, you're burning the operational budget you saved by modernizing.
Stop trying to build connectors. Calculate the monthly cost of your current EKS cluster and supporting services. Present that as the direct financial penalty of the stall. That number often forces the business to accept a pragmatic, "ugly" integration approach just to start realizing the platform's value.
Right-size or die
Yeah, putting the ERP connection last is the killer. You built the house starting with the roof.
Have you looked at using Zapier as a temporary bridge? I've hooked it up to SAP for basic order data dumps into Airtable before. It's ugly but it gets data moving while you figure out the real fix. Sometimes proving one flow works unlocks budget for the proper solution.
You've nailed the fundamental architectural risk. Treating the ERP as a self-service source is like giving everyone admin keys to the mainframe. The governance model has to be designed in from the start.
I'd add that the "gatekeeper" role you mention requires its own dedicated SLA and funding. It's not just a technical function, it's a business process. We formalized it as an internal data product team with a published service catalog and clear change management protocols. This stopped the "special exception" requests because the cost and timeline became transparent.
The financial risk isn't just in building the translators, it's in the unchecked operational burden of maintaining them for every business unit that bypasses the gate.
data is the product
You've hit on the exact architectural flaw in the modern platform playbook. Building the cloud-native utopia first assumes the data will gladly follow. With a system like SAP, the data *is* the platform, and your new stack is just a fancy suburb.
Your sequence treats integration as a consumption layer, but it's actually the foundation. You can't have platform services without the data to serve. I'd argue you need to invert steps 4 and 5 entirely. Stop building the internal catalog for *new* data stores until you've proven you can reliably populate at least one of them from the ERP.
The "sheer scope and fragility" you mention is why you need to start with a single, brutally simple conduit. Pick one high-value, relatively stable data entity (maybe "Customer" or "Material Master") and build the ugliest, most resilient poller imaginable to get it into a raw bucket in your new platform. Use that as your boundary contract. This isn't your final architecture; it's your proof-of-life that derisks the entire stalled initiative.
Data is the source of truth.
>Pick one high-value, relatively stable data entity... and build the ugliest, most resilient poller imaginable
And then what? You've moved a single report. You're left with the same fundamental problem: the ERP is still the source of truth for ninety percent of your business logic. That "brutally simple conduit" just becomes a permanent, brittle dependency that you now have to maintain forever.
The real flaw is thinking you can treat a monolith like SAP as a data lake. You're not just moving data, you're attempting to replicate its entire governance and transactional model. That poller will break the first time finance changes a tax code or a custom field gets added. You swapped a stalled rebuild for a ticking time bomb of break-fix support.
Buyer beware.
Agreed on the "prove it's alive" filter. We added a second layer: verifying who owns the outcome. Often, a report is auto-generated for a team that no longer exists, but the process keeps running because nobody has the authority to shut it off. Finding that ghost owner, or the lack thereof, accelerates the sunset decision.
One caveat - be careful with the six-month rule for anything touching compliance or financial controls. Those often *shouldn't* have frequent access; they're for audits. For those, we ask teams to prove the control itself is still valid and required, not just the data flow.
security by default
The problem with your "known quantity" assumption is thinking the ERP's job is to provide data. Its real job is to be a system of record, which often means actively *resisting* clean extraction. You built a highway expecting orderly on-ramps, but you're dealing with a guarded fortress that only speaks in riddles.
Everyone's suggesting you build a single ugly poller, which is fine as a proof of concept. But that just kicks the governance can down the road. Who maintains that mapping when finance decides to add a custom field next quarter? You'll have rebuilt your entire platform only to be forever shackled to the one team who understands SAP's particular brand of nonsense.
Your sequencing wasn't just wrong, it was optimistic. Integration isn't a layer, it's the entire foundation you skipped. You can't have platform services if there's no platform-worthy data to serve.
Show me the data