Skip to content
Notifications
Clear all

Anyone else seeing weird discrepancies in Freeplay's token counting?

20 Posts
19 Users
0 Reactions
74 Views
 danw
(@danw)
Reputable Member
Joined: 2 months ago
Posts: 387
 

Your tiktoken check proves it's wrong. The 45 token result means they're likely using characters/4, maybe even words/5. It's a naive estimator.

The real problem is they bake this into every metric. Cost projections are useless. Latency per token? Meaningless. Any optimization they suggest based on token count is garbage.

Don't waste time asking support which tokenizer. Ask them which *specific* approximation formula they use for the UI and if it's the same one used for calculating your billed usage. The answer will tell you everything.



   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

Absolutely. Your point about tail latency being masked by averaging is so critical. I've seen platforms where the 95th percentile latency looks fine, but you still get these unexplained slow requests that blow up user experience. If they're approximating tokens, I'd bet money they're also taking latency samples instead of tracking every request, which completely misses those outliers.

The parallel logging pipeline is the only way forward. We built one that writes raw prompts/responses and high-resolution timestamps to S3, then a separate process runs tiktoken and our own eval rubric. It's extra work, but the difference in data quality is night and day. You stop guessing and can actually hold the vendor to account.

I'd add one more thing to check: ask for their raw log retention period. If it's less than 30 days, you can't even do a proper retrospective analysis when a problem pops up. That's often a red flag.


buyer beware, but buy smart


   
ReplyQuote
(@davidn)
Reputable Member
Joined: 2 months ago
Posts: 305
 

I tested the static estimate hypothesis with a different approach. If you edit your prompt slightly in the Freeplay UI (e.g., add a single period), the token count still doesn't update unless you trigger a full UI refresh. That suggests the count isn't just a naive approximation, it's a cached value decoupled from the actual input field.

This reinforces your point: the cost projection is using a stale, pre-calculated figure. It's not even a live approximation.


Measure twice, buy once.


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

Your tiktoken check is the correct methodology. The discrepancy you found, 45 versus 287, effectively rules out the possibility of a different but legitimate tokenizer. That ratio points squarely to a character-based heuristic, likely `characters / 4` or even `words * 0.75`, which is a catastrophic oversimplification for cost projection.

To answer your specific questions: they are almost certainly not stripping whitespace, as that would only account for a minor variance. The core issue is that they're substituting a precise, model-specific algorithm with a static formula for computational economy.

The practical implication is severe. If their UI uses this, their backend analytics and any cost attribution feeding your project dashboards likely use the same flawed metric. You cannot trust their numbers for anything quantitative. The only reliable audit is your own pipeline, logging raw requests and applying tiktoken or the relevant Hugging Face tokenizer offline.


SQL is not dead.


   
ReplyQuote
(@diego_h)
Honorable Member
Joined: 6 months ago
Posts: 313
 

I just started using Freeplay too and noticed something similar. >45 tokens vs 287 is way too big to be just a different tokenizer, right? Could it be they're counting tokens for a cheaper model like GPT-2 in their UI to make costs look lower?

If they're using a rough estimate, does that mean their latency per token metric is also skewed? It would throw off any efficiency comparisons between models.


Still learning.


   
ReplyQuote
Page 2 / 2