I've seen this idea floated around a few marketing blogs, usually with suspiciously perfect results. So I tried it myself. The premise is simple: dump your receipt line items into ChatGPT and ask it to generate compliant, descriptive business purposes.
My quick and dirty test was a mess. The model is predictably bad with ambiguity. For a client lunch receipt, it defaulted to "Business development meeting" which might fly, but for internal compliance, we need the actual project name or code. When I prompted for that, it hallucinated one.
The bigger issue is the lack of audit trail. If your finance team ever questions a description, "the AI made it up" isn't a valid defense. You're still on the hook for verifying every line. So what's the actual time savings? You're now doing prompt engineering and fact-checking instead of just typing "Team sync, Project Phoenix" yourself.
I'm skeptical it works for anything beyond the most generic, low-risk expenses. Has anyone here run a proper test? I'd be interested in a controlled comparison: time per report with and without the LLM, plus the error rate finance kicks back. Without those numbers, this feels like a solution in search of a problem.
Data skeptic, not a data cynic.