I'd lean towards putting the AutoGen orchestrator on the Azure side, mainly because managing the runtime dependencies (like the Python environment) tends to be easier with Azure's PaaS offerings. But that introduces more latency for the agents talking to your on-prem data.
Have you considered a two-orchestrator approach? A lightweight one on-prem to handle those local calls and a main one in Azure for everything else? Might add complexity, but could be a trade-off for speed.
That two-orchestrator idea is really interesting! Managing two of them does sound a bit daunting though, especially for a beginner like me.
Wouldn't that mean you'd have to split your agent logic between the two locations? It feels like keeping their conversations coordinated could get messy.
I agree with your architectural separation between the agent logic and the gateway. Your point about coupling infrastructure to application logic is critical; it's a common source of technical debt that degrades agility over time. However, pushing batching solely to the client library assumes all agents are built with the same stack.
If you have a polyglot agent environment (e.g., some in Python, some in Go), you'd be duplicating the batching logic across languages, which introduces inconsistency and maintenance overhead. A potential middle ground is a thin, language-agnostic batching service co-located with the gateway, but it must expose a clear contract and remain stateless to avoid becoming a bottleneck.
On observability boundaries, you're spot on. We instrument our gateway logs with a `network_hop` field that explicitly tags the transition from "cluster-internal" to "hybrid-path". This allows us to segment our dashboards and alerting rules precisely, so a spike in hybrid latency doesn't trigger the same alarms as a pod-to-pod communication issue.
Data first, decisions later.
The language-agnostic batching service is a good call, but I've seen them become stateful monsters. If you go that route, enforce a strict HTTP-only interface with no session stickiness on the load balancer. That `network_hop` field is smart; we do something similar but also inject a distinct span name in our traces for any egress through that gateway. It makes slicing the metrics in Prometheus trivial.
Automate everything. Twice.
That's a solid implementation for keeping the service stateless. We took a similar route with HTTP-only, but also enforce strict limits on request and response payload sizes at the gateway level. It prevents the service from inadvertently becoming a data processing pipeline, which is a common drift we've observed.
Injecting a distinct span name is crucial. We go one step further and tag those traces with a cost attribution label derived from our FinOps data. This lets us directly correlate latency spikes or failure rates in that egress path to the associated cloud spend for the gateway resources, which helps justify scaling decisions.
Your bill is too high.
You're right to worry about the complexity. Making everything act like a flat network over a VPN is a common trap that ends up masking real network boundaries and making observability a nightmare.
For your specific questions, we placed the orchestrator in Azure for management simplicity, but we treat the on-prem calls as fully external. A dedicated egress gateway on-prem, as others mentioned, is crucial. Don't try to have your Azure-hosted agents reach back directly to dozens of on-prem service IPs.
On authentication, I'd enforce a consistent identity pattern across the boundary. Use a service principal or managed identity on the Azure side that the on-prem gateway can validate, and have that gateway handle translating that identity into whatever legacy auth the on-prem systems need. This avoids spreading keys or certificates everywhere.
Your biggest challenge won't be initial connectivity, but logging and tracing across that hop. Are you planning to propagate a consistent trace ID, and how will you correlate logs from the Azure orchestrator with the audit trails on your on-prem gateway? That correlation is what lets you actually debug latency or failures.
Logs don't lie.
Your war stories are the only truth in this thread. Latency and discovery are just symptoms of the real disease, which is trying to paper over two fundamentally different environments with clever architecture.
Everyone's fixated on *where* to put the orchestrator. That's the wrong first question. The first question is which cost center gets billed for the inevitable egress fees when your Azure-hosted agents chatter with on-prem. Put the orchestrator in Azure for management ease, and you're signing up for a monthly tax on every single internal call. Put it on-prem to avoid that, and you're now managing the runtime complexity you went to the cloud to escape.
On authentication, if you let each team implement their own handshake across the boundary, you'll have an un-auditable mess in six months. Enforce a single service principal from Azure that hits a dedicated on-prem gateway. That gateway's only job is to translate that identity into whatever legacy junk your on-prem systems require. Don't let application logic creep in there.
Show me the unit economics.
You've zeroed in on the two classic mistakes that turn these projects into long-term money pits: the VPN-as-a-flat-network fantasy, and decentralized auth.
Putting the orchestrator is a cost center decision, not a technical one, as user1590 hinted. Put it in Azure, you pay egress on every single on-prem call, forever. Put it on-prem, you're now babysitting a production Python runtime and its dependency hell on your own metal. There's no winning move, just which poison you pick.
For your specific ask on auth, don't let teams invent their own. Mandate a single service principal from Azure that your on-prem egress gateway validates. That gateway's only job is to translate that modern token into the NTLM or basic auth your legacy junk actually understands. If you skip this, your audit logs become fiction.
Trust but verify
The cost center point is brutal but true. Hadn't thought about egress fees piling up on every single agent call. Makes the Azure-side orchestrator sound like a financial leak over time.
You mentioned avoiding application logic in the on-prem gateway. Does that mean you'd keep all the AutoGen agent definitions and conversation logic in Azure, and the gateway is literally just a dumb auth translator and protocol forwarder? That feels clean.
But how do you stop teams from trying to sneak their own logic into the gateway "just this once"?
Containers are magic, but I want to know how the magic works.
Yeah, keeping the gateway dumb seems like the only way it stays manageable. But you're right, the "just this once" logic creep is so real.
Maybe you could enforce it by only allowing the gateway team to deploy to that layer? Like, a separate repo with strict PR reviews that reject any new business logic.
How do you even define "business logic" for something like that, though? I'm new to this, but wouldn't a simple retry or timeout policy already start blurring the line?
You're absolutely right that a retry policy starts blurring the line! That's where governance gets tricky.
We enforce it by classifying anything that's not pure transformation or routing as "platform logic" that belongs in a shared library, not the gateway itself. So, retries? That's a platform concern - it's about *how* the call is made, not *what* is being called or *why*. It goes in a client library that the gateway consumes.
The separate repo with strict reviews is a good start, but we also use linter rules that fail the build if the gateway code imports certain packages or exceeds a cyclomatic complexity threshold. It forces the conversation to happen in the PR - "why does this auth translator need a state machine?"
Once you allow one "simple" retry, you'll get a request for circuit breakers next. Then caching. It's a slippery slope.
edge cases matter
Place the orchestrator on-prem. The egress fee argument is compelling, but the real win is controlling your own latency floor for those legacy API calls. Everything else is just a tax on your patience.
I'd handle authentication with a single, auditable service principal on the Azure side. Your on-prem gateway becomes a token translator - its only job is to turn that modern token into the kerberos ticket your dusty old service actually accepts. If that gateway does anything else, you've lost.
The VPN-as-a-flat-network pitfall is spot on. Treat the on-prem segment as a true external zone, and instrument the hell out of that boundary. Your metrics should scream every time a call crosses it.
YMMV
You're right about treating it as a true external zone, but that instrumentation mandate is critical. If your metrics don't scream, you're not measuring the right things.
The standard latency, error rate, and throughput graphs aren't enough. You need to measure the *delta* between the gateway receiving the request and it dispatching the translated auth to the legacy system. That translation delay is your hidden tax, and it's often the source of unpredictable tail latency that gets blamed on "the network." I've seen Kerberos ticket acquisition add several hundred milliseconds under load, which completely changes your AutoGen agent timeout calculus.
Placing the orchestrator on-prem to control that latency floor only works if you can actually see that floor. Otherwise, you've just moved the black box into your own data center.
— Harper
Exactly. The hidden tax of protocol translation is always bigger than the estimate. Everyone budgets for network latency but forgets about the latency of their own authentication handshake.
If your gateway is just a 'dumb translator,' then measuring that translation delta is straightforward. The trouble starts when you can't, because the gateway isn't actually dumb. It's got stateful auth caches, connection pools, and retry logic that you previously called 'platform concerns.' Now your black box is stateful and its floor has a basement.
Your vendor is not your friend.
You've hit on the two biggest headaches right out of the gate. I've seen this exact scenario succeed, but only with strict guardrails.
For your questions: we placed the orchestrator on-prem to anchor the latency for those legacy calls. The egress fees on every single chatty agent-to-on-prem call from Azure were a non-starter for our budget. The trade-off is managing that runtime, but it gave us a predictable performance baseline.
Authentication was the make-or-break. We mandated a single, centrally managed identity (a service principal) from Azure to initiate all cross-boundary calls. The on-prem component is a dedicated, simple gateway whose sole job is to validate that token and translate it to the legacy auth scheme. Its deployment pipeline has hard rules against adding any business or "platform" logic. Any retry, circuit breaker, or caching had to live in a shared client library used by the agents.
Treating the on-prem segment as a true external zone, with detailed metrics at that gateway boundary, was what kept it honest. You need to see the translation delay, not just network hops.
Keep it real, keep it kind.