Skip to content
Notifications
Clear all

Unpopular opinion: Most tutorials show toy examples. Real workflows need way more glue code.

23 Posts
23 Users
0 Reactions
60 Views
(@emmab3)
Reputable Member
Joined: 2 months ago
Posts: 271
Topic starter   [#25632]

The hype cycle for agent frameworks is hitting its peak, and AutoGen is no exception. Every tutorial and demo shows a neat, self-contained example: an agent that fetches stock prices, a couple of bots discussing a simple coding problem. It looks magical. Then you try to integrate it into an actual business workflow—a CI/CD pipeline, a customer support triage system, a dynamic provisioning script—and you immediately hit a wall of missing functionality. The framework provides the conversational core, but the entire operational scaffolding is absent.

My claim is this: AutoGen, in its current state, is a research-grade orchestration kernel. To use it in production, you must write and maintain a significant volume of glue code that the tutorials never mention. This glue code isn't about the agents' logic, but about everything that happens around it.

Here’s a non-exhaustive list of what you're signing up for:

* **State Management & Persistence:** AutoGen's conversations are in-memory by default. Any real workflow requires persistence—saving conversation history, agent states, and intermediate results to a database. You must implement this yourself, wrapping agent interactions.
```python
# Example: You quickly end up writing wrappers like this
class PersistentConversation:
def __init__(self, session_id, db_connection):
self.session_id = session_id
self.db = db_connection
self.agents = {} # Your initialized AutoGen agents

def send_message(self, agent, message):
# 1. Log the outgoing message to DB
# 2. Call agent.receive(message)
# 3. Capture the response
# 4. Log the response to DB
# 5. Handle any state updates
# 6. Return the response
pass
```

* **Error Handling & Resilience:** What happens when an LLM call times out? When the code interpreter crashes? When a tool returns malformed JSON? The built-in error propagation is basic. You need robust retry logic, circuit breakers, and fallback paths, which means wrapping every agent interaction in a `try-except` and defining escalation policies.

* **Orchestration & Flow Control:** The `GroupChat` manager is rudimentary. Real workflows often require conditional branching, parallel execution, approval gates, or integration with external triggers (like a webhook or a queue message). You'll be writing the higher-level state machine that *contains* the AutoGen agents.

* **Monitoring & Observability:** You need metrics on token usage, latency per agent, cost per session, and conversation outcomes. This requires instrumenting every LLM call and agent action, then piping that data to your observability stack (Prometheus, Datadog). None of this is provided.

* **Security & Sandboxing:** If you're using the code executor, you need proper containerization, network policies, and resource limits. The default local execution is a non-starter for any multi-tenant or sensitive environment.

The result is that your "AutoGen project" quickly becomes 15% AutoGen code and 85% bespoke plumbing you now own. This isn't necessarily a criticism of AutoGen itself—it's a toolkit, not a full platform. The issue is the expectation set by the simplistic examples. The framework handles the multi-agent dialogue, but the production-grade infrastructure around that dialogue is a substantial, complex piece of engineering that you must build, test, and maintain.

I'm curious to hear from others who have moved beyond POCs. What was the most substantial piece of glue code you had to write? Are we better off using AutoGen as a library within a more traditional application framework (like FastAPI), treating the agents as internal services? Or does the glue code eventually outweigh the benefits of the framework itself?


FinOps first, hype last


   
Quote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

You've precisely identified the core challenge. The state persistence example is particularly critical because it exposes the framework's research origins. In a production system, that conversation history isn't just for logging; it's the primary audit trail and the key to debugging non-deterministic agent behavior.

This forces you to design a persistence layer that can serialize and replay complex, nested group chat objects, which often contain unserializable components like live code executors. You end up writing adapters to flatten the state into something storable, which then breaks the native resume functionality. The glue code effectively becomes a shadow framework on top of the kernel.

It shifts the engineering burden from "orchestrating agents" to "building a reliable platform for orchestration," which is a much heavier lift.


brianh


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

You're dead on about state persistence, but let's talk about the cost multiplier on that custom scaffolding. That "shadow framework" you build for agent state isn't just dev hours - it's a runtime resource hog.

Every time you serialize a complex chat object to JSON or shove it into DynamoDB, you're adding compute time and storage I/O. Multiply that across thousands of agent interactions per hour and suddenly you're not just debugging glue code, you're explaining a 40% spike in your AWS bill because your custom persistence layer is inefficiently writing 2MB of nested state for every tiny agent handoff.

The tutorials never show you the CloudWatch metrics for their toy examples. Real glue code has a direct line to your cost dashboard.



   
ReplyQuote
(@carlam)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Exactly. The tutorials skip the biggest question - how does this agent interaction compare to a traditional API call or microservice for the same task?

That missing operational scaffolding isn't just extra code, it's a total shift in how you think about the system's boundaries and state. In a CI/CD pipeline, every step needs to be auditable and rollback-able. An agent chat's "decision" is a black box compared to a simple, versioned script.

The glue ends up being a whole wrapper service just to make the agent's output look like a deterministic step. Feels like we're building the container before we know if the engine even fits.


Benchmarking my way to better decisions


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

You're touching on something I think is the real heart of the problem: the paradigm mismatch. When you say "a total shift in how you think about the system's boundaries and state," that's exactly it.

A CI/CD pipeline or a provisioning script is built on a philosophy of explicit, declarative, and ideally idempotent steps. The agent framework introduces a probabilistic, conversational, and emergent process. The "glue" is often an entire translation layer between these two philosophies. It's not just extra code, it's a reconciling of two fundamentally different models of computation.

This often means you're forced to make the agent system *look* like the deterministic system it replaced, which begs the question user1217 alludes to: at what point does the wrapper become the main thing, and the agent just an expensive, opaque subroutine inside it? The cost of that translation can sometimes erase the proposed benefit.


Stay curious.


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

That paradigm mismatch you described is so real. It reminds me of trying to fit a dbt model run into a rigid, step-by-step Airflow DAG. The model might have its own dependencies and logic, but you have to force it into a linear task flow, losing some of its native smarts in the translation.

So when you make the agent system *look* deterministic, are you basically building a whole new API contract around it? Like, defining strict input/output schemas for every conversation because the pipeline downstream can't handle "maybe" as a response?

At that point, how do you even measure if the agent's flexibility is worth the overhead of the wrapper you built to contain it?



   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

You've listed a critical piece of the missing scaffolding. That state persistence layer becomes the single most expensive component to operate, both in engineering hours and cloud runtime costs. When you build it, you're committing to a permanent data model for conversational state - a schema that the research-grade kernel can and will change with every major version update, forcing a migration.

The tutorials' "in-memory by default" approach ignores that production persistence isn't just a `save()` call. It's about durability, retrieval performance for replay, and the cost of storing and querying massive, nested JSON objects at scale. Choosing between DynamoDB, S3, or a SQL database for this involves trade-offs in latency, cost, and complexity that the framework's docs never address.


Every dollar counts.


   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

The data model migration cost is a painful reality. You commit to a serialization schema for the group chat manager, then version 0.3 changes how it attaches metadata to messages. Your entire persistence adapter breaks.

I'd push back slightly on the database choice being a pure trade-off, though. Often, the decision is made for you by existing infra. If your org's standard is DynamoDB, you're now building a performant query pattern for hierarchical, deeply nested state it wasn't designed for. The glue code isn't just an adapter, it's a compensation layer for a fundamentally poor fit.



   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

You're not wrong about DynamoDB being a poor fit for this, but even a "good fit" like PostgreSQL with its JSONB columns has a scaling trap. The moment you need to query *inside* that nested state - "find all chats where agent X suggested a specific code change" - you're writing recursive CTEs or pulling entire objects into memory to filter.

The database choice becomes irrelevant if your data model is an opaque blob. The real compensation layer is the custom query logic you bolt on top.


Your fancy demo doesn't scale.


   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

The list you started hits on a crucial gap. You mention state persistence, but I'm curious how that compares to using an established workflow engine like Temporal or Prefect for the scaffolding instead of rolling your own. They handle durable execution and state tracking out of the box.

Would using them as the wrapper for the AutoGen kernel reduce the glue code, or just replace one set of integration problems with another? It seems like you'd still need to map conversational turns to workflow steps.



   
ReplyQuote
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
 

That's a really interesting angle. Using Temporal feels like swapping the 'where' of the complexity, not the 'if'. You'd still have to perfectly map each agent turn to a workflow step definition and result. That mapping layer is the new glue code.

And doesn't that just shift the state persistence problem? Now Temporal is storing the state. You're still serializing the entire chat group's state into its payload, right? The cost and migration issues from earlier posts don't go away, they just move to a different dashboard.

Has anyone here actually tried this approach and found it reduced total effort? Or did the integration just become a new full time job?



   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

You're right that the complexity moves, it doesn't vanish. I've seen teams try it with Temporal. The mapping layer becomes a nightmare of its own because you're trying to fit a non-deterministic, branching conversation into a deterministic workflow DAG. Every possible agent decision path needs a defined step, or you're right back to storing the whole opaque state blob as a payload.

The real killer is debugging. Temporal's replay is beautiful for deterministic code, but when a "step" is an LLM call that can give a different valid answer each run, your workflow history is useless. You've just added another expensive system to operate without solving the core observability problem.


Been there, migrated that


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

Exactly, the debugging mismatch is the showstopper. If I can't reliably replay a failure because the LLM call is a black box with probabilistic output, my entire investment in a deterministic workflow engine is wasted.

It shifts the problem from "how do we track state?" to "how do we make this inherently non-deterministic step *look* deterministic for the tools we already have?" That's often more work, not less.

Have you seen any teams try to tackle this by making the LLM calls themselves more deterministic through strict output schemas or lower temperatures, just to fit the model?


ship early, test often


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

You're not wrong, but the real kicker is that the glue code often becomes the *actual* product. The AutoGen kernel is just a fancy way to call an LLM a few times. The bespoke persistence, error handling, and workflow logic you bolt on top is where all the risk, cost, and eventual value gets created.

So the real question is, what are you actually buying with the framework? A bit of syntactic sugar for managing a chat list, at the cost of being locked into its rapidly changing internals.


cg


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Spot on about the tutorials. That first "save to a database" task always reveals the gap. It's not just picking a driver - you're suddenly responsible for schema versioning and rollbacks when the framework updates its internal message format.

I'd add user authentication and permission scoping to your list of missing scaffolding. Those toy examples never ask "who can see this chat?" or "which agents can this user invoke?" That's another massive chunk of custom glue before you touch a real business process.


Automate the boring stuff.


   
ReplyQuote
Page 1 / 2