Thanks for sharing this, it's really helpful to see a breakdown from an actual team that tried it. The point about >causing more review overhead< is something I hadn't considered, but it makes total sense. I've been looking at similar tools for automating parts of our bookkeeping, and I worry about the same thing - if it confidently creates an invoice with the wrong tax rule or misinterprets an expense category, fixing that could take longer than just doing it manually.
Your comparison to a fancy auto-complete is spot on, and that pricing model feels like it's everywhere now. Makes me wonder if the better investment isn't in a tool, but in better templates and documentation for your team's own patterns. Did you find the trial period was long enough to really spot these issues, or did the problems show up pretty quickly?
You've got the right instinct about the templates and documentation. That's the real productivity lever a lot of teams skip.
The problems show up fast. In our trial, the "correct but subtly wrong" pattern emerged in week one, once we moved past simple boilerplate. The bigger issue was spotting them. The first few got flagged quickly, but it's the one that slips through and makes it to staging that teaches you the true audit cost.
Honestly, if your gut is telling you to invest in better templates, do that first. It's a force multiplier that doesn't bill by the seat or degrade your seniors.