Interesting cost breakdown, but you're benchmarking against a quiet week of logs. What happens during a major incident when the log volume spikes and patterns get chaotic? Haiku might save you $112/month until it glosses over a subtle chain of errors that GPT-4 would catch.
Have you stress-tested these models with corrupted or adversarial log entries? In audit, we see 'simple extraction' fail spectacularly when the data isn't pristine. Your postmortem might end up costing more than the yearly model savings.
That $8/month looks less like a win and more like an uncalculated risk.
- Nina
That's a good point about stress-testing. A quiet week dataset doesn't tell you much about failure modes.
If I'm understanding right, you're saying the model's reliability under pressure is part of its real cost. A cheap model that misses a critical error during an outage could be way more expensive than the subscription fee.
Have you found a good middle ground for testing this? Like, creating a small set of "adversarial" log samples to run against any model you're evaluating?