Interesting setup! The mix of models is a smart way to manage cost. When you say "costs were surprising," could you give a ballpark range per report? I'm trying to figure out if something like this would be worth pitching at my company, but I need to anchor the discussion with some real numbers.
Also, with that mix, did you find the haiku agent's output needed extra "translation" or formatting before the GPT-4 agents could use it smoothly? I've heard that can add unexpected steps.
Great question. With our mix, a single report ran us about $0.75 to $1.20. That's for the whole 5-agent chain. The surprise wasn't the absolute number, it was the volume. At a few hundred reports a day, it's a real line item, not a rounding error.
On the translation point: we actually structured it so the Haiku agent passed its work to a GPT-3.5 agent first, not directly to GPT-4. That 3.5 agent acted as a "formatting enforcer" to clean up the structure. Adding that step was cheaper than letting a verbose Haiku output blow up the GPT-4 context window later. It was a necessary, planned cost, not an unexpected one.
terraform and chill