I'm trying to build a review workflow where a manager agent checks the quality of responses from my support agents. Right now, my "manager" just approves everything. It doesn't seem to catch errors or provide useful feedback.
What are the best practices for setting up a manager that can actually evaluate work? I'm using AutoGen for a SaaS helpdesk scenario. Should I be using specific evaluation functions, or is it more about the prompt design? Any concrete examples would be really helpful.
That's a classic AutoGen setup problem. I've seen it happen when the manager agent's instructions are too vague. The key is giving it a concrete rubric to score against, not just a "check quality" directive.
For a SaaS helpdesk, I'd prompt the manager to evaluate three things: accuracy against your knowledge base, clarity for a non-technical user, and whether the solution includes the next step or escalation path. Make it assign a score, like 1-5, for each category. Only approve if all scores are above a threshold you set. This forces a granular check instead of a binary yes/no.
You can also feed it a list of common past mistakes as examples of what to reject. That usually sharpens its focus.
Connecting the dots.
I've been down this exact road for automated alert quality reviews. The prompt design is 90% of it, but you need a structured output from the manager to make evaluation actionable.
Take user1243's scoring idea and force a specific format. For example, require the manager's verdict to be a JSON block:
```json
{
"approved": false,
"scores": {"accuracy": 2, "clarity": 4, "next_step": 1},
"feedback": "The solution suggests restarting the service but doesn't mention checking error logs first, which is step one in our KB."
}
```
Then you can programmatically parse that and only pass forward the approval if `approved` is true. This moves it from a vague opinion to a checkable result. Give the manager a list of your top five most common critical errors and tell it to reject any response containing those patterns.
Sleep is for the weak
The real problem isn't prompt design. It's that you're trying to automate a judgement call without a real definition of quality.
You're asking a language model to "evaluate" another language model's output, based on instructions written by... you. If your own criteria are fuzzy, the manager has nothing to work with. The JSON and scoring systems people are suggesting are just elaborate ways to document the manager's *guesses*.
Before you write a single line of prompt, define a failure condition you can write as a rule. "If the agent suggests a password reset for a billing inquiry, that's a rejection." Do that for three things that actually matter in your SaaS. Otherwise you're just building a theater of oversight.
— skeptical but fair
You mentioned your manager agent just approves everything. I ran into something similar. Can you share what your manager's prompt looks like now?
I wonder if it's about having the manager compare against a specific source of truth. For a SaaS helpdesk, do you have a public knowledge base or internal runbook the manager can reference when checking the agent's answer? That might help it spot deviations.
That's a good idea about the source of truth. My current manager prompt is pretty basic. It just says "Review the support agent's reply for correctness and clarity. Reply with APPROVED or REJECTED with a reason."
I don't have a specific knowledge base URL in the prompt for it to reference. It's probably just making a general assessment. How do you actually give it that runbook? Do you paste the whole thing into the prompt, or is there a better way to attach a reference document?
"Elaborate ways to document the manager's guesses" is a solid way to put it, honestly. I've built these loops and watched them rubber-stamp nonsense because the underlying rule set was just a vibe.
That said, I think the "theater of oversight" line is where you lose me a bit. Defining three hard failure rules is a great start, but for a real workflow, you still need the prompt to enforce those rules and a way to handle the grey area that isn't a hard fail. The JSON isn't *just* documenting a guess - it's structuring the output so the next system action (approve, kick back for rewrite, escalate to a human) has something to parse. The failure conditions are the guardrails, but you still need the road.
been there, migrated that
Your basic prompt puts all the work on the manager to define "correctness" on the fly, which leads to that rubber-stamp behavior. You need to build that definition for it.
Start by embedding a concrete checklist in the prompt, drawn from your actual support guidelines. For example: "Check 1: Does the agent's answer cite the correct help article? Check 2: Is any technical jargon explained? Check 3: Does it provide a clear action for the user?" This turns evaluation from a vibe into a verifiable process.
I'd also suggest making the manager agent "think aloud" in its evaluation before giving a final verdict. Ask it to first list which checklist items pass or fail, then decide. That often surfaces hesitations a simple approve/reject hides.
—daniel
Totally been there! The rubber-stamping manager is a classic early-stage problem. The key is moving from a vibe check to a defined process.
What finally worked for me was combining the checklist idea with a "show your work" step. I forced my manager agent to output an evaluation table before its verdict, like:
- **KB Accuracy:** [Quotes the relevant KB article number]
- **Actionable Step:** [Yes/No - Was a concrete user action provided?]
- **Tone Check:** [Pass/Fail - No blaming the user]
Only after that table does it give an APPROVED/REJECTED. This makes its reasoning traceable and often catches oversights. Also, bake a few example rejections into the prompt. Show it what a bad answer looks like.
For your SaaS scenario, you could embed a snippet of your most critical support policy right in the prompt, so it has a real source of truth to reference.
Keep deploying!
That point about starting with clear failure conditions is probably the most important thing I've read in this thread, and it's easy to skip over. It's tempting to just start building.
Your "theater of oversight" comment made me pause, though. I think you need both. Those three hard rules are non-negotiable. But in a real marketing automation context I work with, there are also quality goals that aren't binary failures. Like, "does this email draft maintain our brand voice?" or "is the proposed subject line likely to avoid spam filters?"
So I'm wondering if the right sequence is exactly what you said: define the hard failure rules first (the "rejection" criteria). Then, use a structured output, like the JSON people mentioned, to also assess the softer, graded aspects of quality. That way you're not just documenting a guess - you're separating the absolute must-fail conditions from the performance scoring. Does that dilute the principle, or is it a practical next step after the rules are set?
Checklists are a solid step. But they just move the problem. The manager now guesses if "technical jargon is explained" instead of if the answer is "correct".
Where's the validation that the checklist itself works? You're still relying on the manager's interpretation of your rules, not an objective measure. I'd add a second layer: track how often a human reviewer overrides the manager's checklist-based verdict. If it's more than 10%, your checklist is the new vibe.
If it's not a retention curve, I don't care.