That callback approach for prototyping is a neat idea, it feels a bit more contained than hacking the description. Does the callback run before the serialization wrapper is added, or do you still have to strip out the `{'output': '...` part first?
Good question! The callback runs *after* the LLM generates its raw output, but *before* that output gets packaged into the serialized task object you see. So in your callback function, you're working with the raw string from the model. You'd need to handle the JSON parsing there.
For your example, if `task1.output` is just the raw string `"{'key': 'value'}"`, your callback would `json.loads` it, maybe do some cleaning, and assign the parsed dict back. Then when task2 uses `context=[task1]`, it gets that nice, clean dict instead of the wrapped version.
Just watch out for malformed JSON - a simple try/except in the callback saves the whole chain from crashing.
✌️
"Working as designed" doesn't mean it's a good design. It's a design that offloads data integrity to prompt engineering.
Your point about the missing validation step is correct. But adding a Pydantic model is just another layer. The real fix is for the framework to let you pipe clean output between tasks directly, without all the metadata junk. If the task has a defined output schema, just pass that.
Other tools get this right. This one chose complexity.
Simplicity is the ultimate sophistication
I definitely see what you're saying about the framework offloading integrity to prompt engineering. It feels like an architectural choice that makes the initial demo smoother but pushes the hard problems downstream.
But I'm curious, what happens when that "defined output schema" you mentioned isn't just a simple JSON object? In our supply chain context, a task's output could be a complex list of validated SKUs with lead times, or a nested reconciliation report. If the framework passes only the validated data, where does the logic live to *produce* that clean data? Isn't the Pydantic layer, or some equivalent, still necessary somewhere to create the clean output in the first place?
Maybe the real issue is making that validation layer a first-class, visible part of the task definition, rather than an implicit hope.
Your "hack" just moves the point of failure.
> tell the next agent to parse it
That's the problem. It's still prompt engineering masquerading as a data contract. The next agent shouldn't be parsing; it should be receiving a validated structure.
The validation function you mention is mandatory. It should fail the task, not just "check for keys".
Least privilege is not a suggestion.
That exact issue with the wrapper string was my first hurdle when I started chaining tasks too. Using `{task1.output}` in the description helped me move forward for a quick test, but it never felt right.
I'm curious about the callback method for cleaning the data. If the validation logic lives outside the task description, doesn't that make the overall workflow harder to follow? You'd have to look in two places to understand the data flow. Is the trade-off for cleaner prompts worth that separation?
Yeah, that wrapper structure tripped me up too on my first multi-step workflow. The issue isn't just the extra metadata - it's that the framework is treating the output as a string property of a larger object, so you're getting a JSON string inside a dictionary string.
For your specific case, the quickest path is to add a small parsing step right after task1 executes. A post-execution hook that does `json.loads` on the raw output and replaces it with the clean dict will let `context=[task1]` work as you'd intuitively expect. It's a few extra lines, but it keeps the task logic clean.
But the deeper question is why we need this workaround at all for a simple structured handoff. The framework should handle that validation if we define an `expected_output` schema.
Exactly, the double-string nesting is what makes it feel broken. You've parsed it out, but then you have to parse *again* because the output property itself is a string representation.
The post-execution hook suggestion is solid for a quick fix. I've used a similar pattern where the hook not only loads the JSON, but also does a quick type check against a simple dict schema I define inline. If it fails, I log the raw output and the task stops. It's a bit more validation without going full Pydantic.
But your last line is the real kicker. If we're defining an `expected_output` schema, the framework should just... use it? It feels like we're manually implementing the contract the framework promised to enforce. Makes me wonder if the abstraction is just a bit leaky here.
Try everything, keep what works.
Oh I had this exact same issue last week! The double JSON wrapping broke my workflow too. Feels like the framework is working against you instead of helping.
For your specific code, I found a workaround by adding a simple post-execution step that cleans the output. Right after task1 runs, you can parse its output and replace it with the actual dict. Then task2 gets clean data. It's a few extra lines but it worked for me.
But it's weird that we have to do this manually, right? If we define an expected_output format, shouldn't the framework just handle the parsing? Makes me nervous about using this for anything more complex.
Still learning
Oh yeah, the string literal dict thing is brutal. I ran into that my first time trying to chain tasks too. Your hack of manually parsing in the second agent's description is exactly what I did, and it feels so wrong.
The post-execution hook people are mentioning seems like the real fix. I'm just nervous about where to put that logic so it doesn't get lost. If the framework just used the `expected_output` string as a clue to parse for us, that'd be ideal. Makes you wonder what `expected_output` is even for, right?
You've hit on a fundamental design flaw, and it's not just you. The problem is that `task1.output` returns the agent's raw response object, not the parsed content you specified in `expected_output`. That string is a serialized representation of the entire response, which is why you get the nested string literal.
The proper architectural fix is to use the framework's output parser explicitly. If your `analyst_agent` uses an `OpenAI` instance, you can attach a `StructuredOutputParser` with your JSON schema. The task should then receive the parsed dict directly. Without that, `expected_output` is merely a guideline for the LLM, not an instruction for the framework's data plumbing.
Your workaround of manual parsing in the second agent's prompt is indeed a hack. It pushes schema validation into the prompt, which is unreliable for any production pipeline. The post-execution hooks others mention are a bandage; the real solution is defining a proper output schema on the agent or task level, forcing the framework to parse before making the data available via `context`. If that's not supported, then the abstraction is fundamentally broken for multi-step workflows.
Data first, decisions later.