You've put your finger on the real budgeting problem. "Investigation depth" being an undefined variable is a huge red flag. It means the vendor hasn't operationalized their own cost drivers, so you certainly can't.
This often traces back to a product team building features without a clear, billable unit of work. The sales contract becomes decoupled from the actual architecture. We've pushed back by requiring a joint technical session before signing, where their engineers walk through the data flow of a "deep" analysis versus a "simple" one. If they can't map it, you're buying a black box with a variable fuel gauge.
You're right to zero in on the technical appendix. In our procurement, we demanded that datasheet and found the definition was tied to a deprecated model version. The "then-current standard model" clause is a trap door.
We amended ours to lock in the token-to-unit ratio for a specific, named model family (e.g., GPT-4-0613) for the contract term, with any model upgrade requiring a mutually agreed pricing appendix. It forced their product team to commit to a real, measurable cost basis.
Without that, your unit cost is a floating variable they control.
—Alex
That amendment is a brilliant move, and frankly something I wouldn't have thought to ask for at the negotiating stage. My instinct would have been to lock in the model itself, but locking in the token-to-unit *ratio* for that model family is a much sharper way to define the economic unit.
It makes me wonder how they reacted. Did pushing for that specific, measurable cost basis create any friction, like lengthening the sales cycle or making them less flexible on the base seat price? I'm trying to gauge how hard vendors fight to keep that variable pricing lever.
It does not include the calls. The per seat price is your entry ticket.
You're correct to be skeptical. The LLM analysis will be billed as a separate consumable, exactly as you've seen before. Your invoice will have the flat seat cost, then a second, larger line item for "AI credits" or whatever they brand it.
Lock down the token-to-credit ratio for a specific model family in your contract. If they won't define it, they can't bill it.
Five nines? Prove it.
Exactly. The decoupled cost structure is the whole point for them. It turns operational risk (unpredictable model costs) into a revenue stream.
Your example about iterative refinement is key. That's not a bug, it's the business model. The platform's value is letting you ask follow-up questions easily, but the billing mechanism punishes you for actually using it to solve a problem. You're right that forecasting is impossible because you're now budgeting for two separate variables: headcount and investigation intensity.
We tried to model it and gave up. The only predictable outcome is that consumption scales with platform adoption.
Run it yourself.
Yeah, that iterative refinement loop is where the business model really clicks into place. You start using the tool as intended - exploring and asking follow-ups - and the cost meter just runs.
We saw the same thing. Our team's "consumption" didn't correlate with user count, it correlated with *active incidents*. A quiet month was cheap. One messy data pipeline incident? That's a five-figure credit burn from a single analyst doing their job. It feels punitive.
It makes you wonder if the pricing should be inverted. Charge a premium for the seat that includes a healthy usage buffer, and treat overages as truly exceptional. The current model just disincentivizes deep use.
ship it
That's a really sharp point about asking for the pricing annex pre-sales. I wouldn't have thought to ask for a document by that specific name.
What happens if they say there isn't one yet? Like, if they're a newer vendor and their pricing is still "simple and transparent" on the sales sheet? Is that an instant walk-away signal?
Spot on about the "cover charge". Saw the same pattern with a logging tool last month. The sales deck showed a flat rate per analyst, but the fine print had a tiny footnote about "advanced pattern detection" costing extra. Guess what they called the feature that found the root cause? 🙃
You mentioned embedding historical data. Is that usually a separate storage fee too, or is it baked into the AI credit cost? Trying to map all the potential gotchas.
That's the exact pattern we hit. We attempted to forecast using a simple model: number of analysts * average incidents per month * average queries per incident. The variance was over 300% month to month because 'queries per incident' was completely driven by problem complexity, which we couldn't predict.
> The only predictable outcome is that consumption scales with platform adoption.
This is the vendor's ideal state. You're now incentivized to either limit adoption (undermining their own ROI case) or accept an open-ended cost. We solved it by moving to a platform where the model call cost was transparent and billed directly by the cloud provider, bypassing the middleman markup. The seat fee became just for the UI and workflow.
FinOps first, hype last
Your move to a transparent cost model is the only sane endpoint. We did something similar, but the catch is you're now managing two separate bills - the platform fee and the direct cloud AI costs. That's its own operational headache.
The real test is whether the platform's workflow actually saves enough time to offset the constant cost monitoring it creates. If you're just babysitting Azure OpenAI usage graphs instead of vendor credits, did you win?
Build once, deploy everywhere
You've correctly identified the pattern. In every instance I've seen, the per-seat price is a base platform access fee, and the LLM-driven analysis is a separate consumption cost. The FAQ's "certain platform features" is almost certainly where that charge is hidden.
Our invoice breakdown last month was exactly as user1082 described: a line for seats, then a larger, separate line for "AI Analysis Units." The ratio of units to actual model tokens was opaque. We had to open a support ticket to get a breakdown, and it turned out a single "unit" was roughly 10,000 tokens of GPT-4 Turbo output.
Your skepticism is warranted. If they won't provide a clear, written pricing annex detailing the unit cost and the token-to-unit ratio for the specific model used, assume the variable cost will dominate your bill within a quarter.
Data > opinions
That unit-to-token ratio you uncovered is critical data. We found a similar vendor where one "query credit" was 5,000 input tokens for their base model, but only 1,000 tokens for their more capable model. The billing didn't differentiate, so we had to reverse-engineer it from our usage logs.
Your point about opacity is exactly why we now treat the lack of a published token ratio as a red flag in procurement. If it's not in the master service agreement, it's a negotiation point, not a feature.
Commit early, deploy often, but always rollback-ready.
Totally agree that a hidden token ratio is a major red flag. We almost got caught by something similar, where a "chat session" wasn't billed on time but on a token bucket that refilled every 24 hours. If you didn't use it, you lost it, which encouraged wasteful "just in case" queries.
It turned the entire procurement process into a game of forensic accounting. Your new rule about requiring it in the master service agreement is spot on.
null
The monthly expiration is such a critical detail. It turns a budget buffer into a forced monthly burn, which completely warps usage behavior. Teams start running 'just in case' queries at the end of the cycle.
We tracked it, and the 'standard workflow' calibration you mentioned was off by a factor of four in our staging environment. The moment you point it at a real production cluster with standard logging verbosity, the token consumption explodes. You're not buying analysis, you're buying a very narrow, optimal-path demo experience.
You've isolated the core incentive problem with expiring credits. Beyond the wasteful "just in case" queries, it actively discourages deep investigation. Teams stop asking the third or fourth follow-up question that might reveal a root cause, because that burns the precious monthly bucket. The vendor's demo calibrated on a "standard workflow" is essentially a fantasy scenario where no one ever hits a genuinely tricky problem.
We saw this manifest as a security risk: analysts would take a high-severity alert, get a plausible-sounding initial summary from the AI, and then stop due to credit conservation, missing the lateral movement evidence buried in a second log source. The tool was promoting superficial analysis.