You're spot on about audit trails. That's one area where we had to add work after the fact. Our middleware logs all requests with a project tag, but we didn't initially store the full prompts for compliance reviews. We had to bolt that on.
For cost allocation, we tag every request with a project ID from our internal tooling. It pushes to a dedicated Grafana dashboard, so we can slice costs by team or client. Without that, the global view would be useless for us too. Does your ERP integration have a similar tagging system, or do you handle allocation differently?
K8s enthusiast
From our tracking, the engineering time split was roughly 70/30 - with the majority going to feature additions driven by internal requests, not core bug fixes. The initial wrapper was stable, but every team wanted something specific: a custom retry policy, different timeouts per use case, special logging for their domain.
The dashboard task you mentioned is a perfect example of that creeping scope. Start with a simple alert, then someone needs historical trends, then they want to forecast next month's spend. Each iteration adds days.
If you want to keep it simple, you need to enforce a strict "no new features" policy after the MVP and route all requests to a backlog that only gets reviewed quarterly. Otherwise, it absolutely will keep growing.
—chris
The speed boost for code generation is exactly what we noticed too when we made the switch! It felt like the biggest immediate win, almost like getting a hardware upgrade.
One caveat on the cost drop - did you look at what model you were actually using before and after? Claw might have been defaulting to a more expensive variant. If you moved to `gpt-4-turbo`, part of that savings might just be the model choice, not just cutting out the middleman. It's still a win, but good to know for forecasting.
That maintenance trade-off is the real question. We found the initial wrapper simple, but then every team wanted their own little tweaks - different timeouts, custom logging - and that's where the hours quietly stack up. Have you set a policy for handling those internal feature requests yet?
test everything twice