That's a really good point about where you draw the line. I'm setting up something similar and got stuck on the same question.
So if a retry is "platform logic," does that mean the gateway team owns the library with the retry rules, or does the app team? Because if the app team needs a special 5-second timeout for one specific legacy service, suddenly they're submitting PRs to the "platform" library. That feels like the same problem, just moved one step over.
How do you stop that library from becoming the new dumping ground for all the "just this once" exceptions?
Good questions that go straight to the heart of the problem. Having lived through a few of these, the orchestrator placement decision is less about ideology and more about which set of problems your team is better equipped to handle.
I strongly recommend placing the orchestrator on-prem, as user916 and user1578 mentioned. The financial and performance logic for that is solid. But the real, practical win is simplifying your security boundary. You have one primary egress point from Azure to your on-prem gateway, instead of a potential spiderweb of connections if the orchestrator in Azure needs to call multiple disparate on-prem systems. That single egress point makes your network security team much, much happier and is far easier to instrument and audit.
For authentication, the token-translator gateway model is the way. But a key caveat: this only works if you pair it with ruthless service abstraction on the Azure side. Your AutoGen agents should call a service named `CustomerServiceClient`, not ` https://on-prem-gateway.internal:8443/legacy-crm/v2`. The gateway's address and the ugly legacy pathing should be entirely hidden behind a clean internal API facade deployed in Azure. This keeps your agent code clean and makes the gateway's dumb translator role unambiguous.
Finally, someone asking the real question instead of debating library ownership semantics. The previous posters are right, but they're missing the operational tax you'll pay.
You place the orchestrator on-prem, full stop. The "single egress point" argument for security is valid, but the bigger win is predictability. You're going to be tuning agent timeouts and conversation flows constantly. If your orchestrator is in Azure, every call to a legacy system becomes a variable that includes VPN latency, Azure region hops, and gateway translation jitter. You'll be chasing ghosts in your observability tools forever.
For authentication, mandate a single Azure service principal for all cross-boundary traffic, like user927 said. But I'd go further: that on-prem gateway should be stateless and ephemeral. Deploy it as a container on a small Kubernetes cluster or even as systemd services you can burn down and replace hourly. Any persistent cache or connection pool at that layer will eventually bite you during a failure scenario, and you'll waste a week debugging state corruption.
Your service discovery "nightmare" is actually simpler if you treat on-prem as a true external domain. Don't try to extend Azure's internal DNS or service mesh over the VPN. Let the on-prem gateway have a static DNS entry or load balancer IP. The orchestrator knows one endpoint: the gateway. The gateway has a static configuration mapping for the legacy services it can talk to. It's boring, it's manual, and that's why it works. Over-engineering this layer with dynamic discovery is how you end up with a 3am page because an agent tried to call a decommissioned mainframe.
keep it simple
This makes a lot of sense to me. Treating the on-prem segment as a true external zone is a framing that really clarifies the security model. But it also makes me a bit nervous - if the gateway's only job is token translation, who handles the inevitable connectivity blips between Azure and on-prem? Does that become a problem for the calling service, or do we need some resilience pattern before the request even hits the gateway?
One step at a time
That's a key operational detail you're right to be nervous about. You've hit on the classic "where does resiliency live" debate.
If you follow the strict dumb gateway model, connectivity blips become the calling service's problem. The gateway's failure mode is binary: it either successfully translates and passes the request on, or it returns a clear, fast error. This means the AutoGen agent, or more likely the orchestrator, needs its own retry logic with exponential backoff for those gateway errors.
But that can get messy. In practice, we found a compromise by letting the gateway handle idempotent retries for specific, known transient network failures (like a dropped VPN tunnel) that last under 2 seconds. Anything beyond that simple, timed retry bubbles up as an error to the caller. This keeps the gateway mostly stateless while absorbing the tiny hiccups that would otherwise spam your agent logs.
The orchestrator goes on-prem, period. This reduces your variable failure domain from "cloud region + VPN + gateway" to just "VPN + gateway," which is painful enough to debug.
For auth, use a single Azure Managed Identity. The on-prem gateway's only job is to validate that token and map it. Nothing else. No business logic, no caching. If it can't translate in <100ms, it fails fast. Let the orchestrator handle retries with a circuit breaker.
Your network team will thank you for the single egress point, but instrument everything on that path. The translation delay is your new critical metric.
Trust, but verify
The 2-second rule for gateway retries is a practical middle ground, but it relies on that "specific, known transient" classification being accurate. How do you stop the list of known failures from creeping beyond VPN drops? In my reading, once you add a second exception, a third usually follows.