Skip to content
Notifications
Clear all

Anyone actually using AutoGen in production with real customers?

21 Posts
19 Users
0 Reactions
75 Views
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

Both. The JSON schema is your first and strongest guardrail to catch malformed output. But for a customer-facing summary, you also need content validation. We combine a schema with a keyword check.

The schema ensures structure. A Pydantic model that enforces fields like `period_start`, `period_end`, `top_incidents` as a list. It fails the run if the agent invents a field.

Then we run the generated narrative text through a second check against a list of required terms (e.g., "uptime", "response time") and banned terms (e.g., "I apologize", "I cannot", "as an AI"). Missing a required term triggers a rewrite. A banned term fails the job entirely.

It adds maybe 50ms, but it stops nonsense from going out the door.


Build once, deploy everywhere


   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
 

I agree with the core premise about log grouping, but I think your distinction between deterministic tasks and narrative generation is the crucial architectural boundary. The mistake is letting the agent make *any* logical decision.

> treat the LLM like a templating engine with extra steps

This is exactly right, but I'd phrase it as a constrained text transformation. The "narrative structure" you build outside is a JSON schema. The agent's prompt becomes "Fill this schema with natural language based on this data." Its entire world is the input data and the output schema. There's no room for it to decide to group logs or calculate trends; that's all done before it's called.

The validation then isn't just a whitelist; it's a schema validation plus semantic checks against the source data. If the summary mentions a "spike in errors" but the source data shows no error delta, the run fails.


Data over dogma


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Moving the cost check to the session manager is the right call. We handle cost similarly, but we also factor in historical token counts for specific prompt templates as part of the estimate, not just the static data sizes. This catches edge cases where a particular query pattern tends to generate longer reasoning chains.

The per-session budget is definitely more effective for runaway conversations, but have you found it necessary to include a mid-session checkpoint for unusually long, but valid, sessions? We have a read-only agent that can intervene with a final summary if the token count approaches the per-session limit, to provide a graceful stop instead of a hard cut-off.



   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 3 months ago
Posts: 292
 

>testing with human_input_mode="NEVER". That's a full stop.

It absolutely is, but I think the underlying problem is conceptual. You're framing your agent as an autonomous analyst. In production, you must frame it as a single-purpose function call with a language interface. The conversation pattern is a development and debugging convenience that you later lock down.

For your snippet, the production version wouldn't be an agent loop at all. It would be a Python script that first runs `query_prometheus()` directly, structures the data, and only then makes one call to a configured `AssistantAgent` with a system prompt like "Write a two-sentence summary for this data." You set `max_consecutive_auto_reply=0` so it cannot start a chain. The cost and hallucination control comes from that extreme constraint, not from hoping a `human_input_mode` setting will save you.

Your concerns about safeguards are spot on; you're essentially building a deterministic pipeline where the LLM inhabits one, heavily fortified cell.



   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

>you must frame it as a single-purpose function call with a language interface.

You're both circling the actual point but still building complexity on a shaky premise. The conversation you're trying to lock down shouldn't exist in the first place. Why is there a 'conversation pattern' at all, even for development?

You're describing building a deterministic data pipeline, piping it to a text generator, and then adding layers of validation to protect yourself from that text generator. That's not an agent. That's a glorified, expensive, and unreliable printf statement. If you need a predictable summary from structured data, just write a damn template.

The overhead you're adding - schema design, validation runs, pre-processing scripts, token estimation - is already more work than solving the actual problem with simple code. You're not controlling the agent; you're working around its fundamental unsuitability for the task.


monoliths are not evil


   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

You're right about the safety checks, but I think the "complex enough to need a narrative summary" gate is where people trip. That's another decision point. If you're already deciding whether to call the agent, you've already structured the data enough to use a simple template 90% of the time. The remaining 10% where you genuinely need nuance is the only place the agent's unpredictability is worth the validation headache. Otherwise you're just paying for the option to be surprised.


Data over dogma.


   
ReplyQuote
Page 2 / 2