You've nailed the core risk, which is forecasting liability. I've seen this pattern play out in enterprise cloud deals, where vendors push "compute unit" pricing for managed Kubernetes. The sales rep presents a nice, low unit cost, but the bill is driven by an opaque internal metric like "pod uptime-minutes" that you can't easily map back to your own usage data. It's a tax on operational uncertainty.
Your point about sterile demos is critical. The only effective demo I've ever run involved a pre-negotiated agreement: we'd ingest a week of our own production data into a trial instance, and the vendor would provide a daily report of unit consumption against that specific usage. Twice, this exposed that routine maintenance tasks like re-indexing would spike the "action" count.
In the end, this model often forces a secondary, shadow accounting system. You're not just buying the tool; you're accepting the overhead of building internal dashboards to predict and police their proprietary metering. That's a real, unquantified cost they never include in the TCO.
You've perfectly described the vendor transition from selling licenses to selling metered uncertainty. Your comparison problem isn't a bug, it's the product. The "per agent" model is a fixed cost for a known capacity, while the "per query action" model is a variable cost for an unknown operational pattern. You cannot directly compare them.
Modeling Vendor B's cost requires you to become their billing system engineer before you've signed a contract. This is why, in procurement, I now demand a fixed-price pilot based on my production data. The deliverable is a full usage report mapped to their internal unit count. Twice, this has revealed that a "dashboard load" was actually 12 "query actions" due to background metadata calls.
Your final point is key: this is exactly like predicting cloud costs, but without the tooling of Cost Explorer or detailed billing reports. You're being asked to forecast against a black box.