Skip to content
Notifications
Clear all

Anyone else seeing weird discrepancies in Freeplay's token counting?

20 Posts
19 Users
0 Reactions
73 Views
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
Topic starter   [#25227]

Been testing Freeplay's new evaluation suite. Their token counting for costs/benchmarking is way off compared to my own calculations using `tiktoken`.

Example: A simple RAG prompt with a 200-word context chunk. Freeplay's UI says: "45 tokens". Running the same text through `tiktoken` for `gpt-3.5-turbo` gives me 287 tokens.

Here's my quick check script:
```python
import tiktoken
encoding = tiktoken.encoding_for_model("gpt-3.5-turbo")
text = "Your 200-word text here..."
token_count = len(encoding.encode(text))
print(f"Tokens: {token_count}")
```

Discrepancy is consistent, not a one-off. Makes their cost projections and "optimization" suggestions questionable.

* Are they using a different tokenizer?
* Are they stripping something out (maybe whitespace or special chars) before counting?
* If it's not accurate, how can you trust their pricing analysis features?

Anyone else ran into this or done a similar audit? Need to know if I can rely on their numbers at all.


slow pipelines make me cranky


   
Quote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Your script is correct, and your suspicion is spot on. This isn't just a small margin of error, it's completely broken. A 200-word chunk cannot be 45 tokens by any standard tokenizer; that's the count for maybe two short sentences.

I hit the same wall last month. The discrepancy isn't about whitespace stripping, it's that their token counting feature appears to be sampling or estimating based on character count, not actually running a proper tokenizer. I opened a 1500-token project conversation log, and their UI showed "~320 tokens". When I pressed support, the answer was evasive, mentioning "approximations for performance in the UI."

If their cost projections are based on those same approximations, the feature is useless for financial planning. You cannot trust it. I've moved to calculating tokens offline with `tiktoken` and importing the totals, treating their built-in counter as a decorative metric. It makes their optimization suggestions completely unreliable, which defeats the purpose of using their suite for benchmarking.



   
ReplyQuote
(@harukik)
Honorable Member
Joined: 2 months ago
Posts: 400
 

That's a really frustrating response from support. If they're using approximations just for UI performance, why even present it as a concrete token count for cost projections? It feels misleading.

I'm just starting with this tool, and this makes me wonder about their other metrics. If the token count is a decorative estimate, what about their latency measurements or accuracy scores? Are those "approximated" too?

Have you found any other parts of their evaluation suite that seem off compared to your own checks?



   
ReplyQuote
(@emilyh)
Estimable Member
Joined: 2 months ago
Posts: 166
 

That's a great point about the other metrics. If the token count is that far off and they're calling it an "approximation for the UI," it does make you question everything else in the dashboard.

I haven't done rigorous checks on latency, but I did notice their "accuracy" score for a simple classification test seemed oddly high compared to a manual spot check. It felt more like a basic string match than anything evaluating the actual intent of the response.

Has anyone tried exporting the raw evaluation data to compare? I'm wondering if the underlying data is correct and it's just the UI layer that's smoothing things into something unusable.



   
ReplyQuote
(@harukik)
Honorable Member
Joined: 2 months ago
Posts: 400
 

Yeah, that accuracy score thing is interesting. I ran a quick test on a basic Q&A prompt and got a 95% accuracy rating from Freeplay, but when I manually checked, a few answers were technically correct but missed the nuance completely. It felt like it was just matching keywords.

>Has anyone tried exporting the raw evaluation data to compare?

I haven't, but now I'm curious too. If the underlying data is solid, maybe we could build our own dashboards. But if the export is just feeding us the same "smoothed" numbers, that's a bigger problem. Has anyone here actually pulled the raw logs?



   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

>It felt like it was just matching keywords.

It's not just keyword matching, it's probably a simple cosine similarity on embeddings. Fast, cheap, and utterly useless for nuance. They all do this.

I pulled their raw logs via the API last week. The "evaluation scores" in the export are the same smoothed numbers from the UI. The underlying request/response pairs are there, but the metrics are pre-baked junk.

So no, you can't build your own dashboard from their data. You have to run your own evals from scratch.


-- old school


   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Ugh, that's disappointing about the API export having the same baked-in metrics. I was hoping the raw data would be salvageable.

It makes sense that they'd use cosine similarity for speed and cost. But calling that an "accuracy" score for nuanced tasks is where it gets misleading. I've seen it flag a completely irrelevant but keyword-stuffed response as "95% accurate."

So we're back to square one - if you need trustworthy evals for anything beyond basic keyword detection, you really do have to run your own. Makes you wonder what the premium is actually paying for


Keep it simple.


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
 

You're right to question the other metrics, and this token counting issue is often a leading indicator of deeper measurement problems. In my own audit of similar platforms, I've found that when token approximation is this severe, latency calculations frequently suffer from the same "smoothing" - they often sample or average in ways that mask tail latency, which is what actually matters for production reliability.

Regarding accuracy scores, others have correctly pointed out the cosine similarity approach. The more subtle issue is that these platforms rarely disclose their evaluation model's version or the exact rubric. An "accuracy" score can shift dramatically between, say, `text-embedding-3-small` and `3-large`, or if they change the similarity threshold. You should push their support for this specification; if they can't provide it, the score is indeed decorative.

For your own checks, I'd recommend instrumenting a simple parallel pipeline that logs the raw prompt/response pairs alongside timestamped spans for latency, then run your own eval chain against those logs. It's the only way to establish a baseline truth.



   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

Yep, your script is the exact check I ran last week. The difference is wild. If a 200-word chunk is showing as 45 tokens, that's clearly not a tokenizer at all. It's probably a naive character/word ratio.

What's worse is that their cost projection feature likely uses that same broken count. Have you checked if the "estimated cost" in the UI changes if you paste the same text twice? Mine didn't budge, which basically confirms it's a static estimate, not a real calculation.

So to your last question, no, you can't trust their pricing analysis. I'm wondering if they're even using the same flawed logic for counting tokens in the actual LLM calls they make, or if that's separate.



   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

You've nailed the core issue but I think you're being too generous with your questions about different tokenizers or stripping whitespace. A 200-word chunk being reported as 45 tokens isn't a different tokenizer, it's a fundamentally broken feature. That's roughly a 6:1 compression ratio which doesn't exist in any real tokenization scheme.

The more important question you're hinting at is whether this "broken but fast" counting is only for the UI, or if it infects their actual cost calculations and performance metrics. In my experience, when a vendor ships a UI with numbers this wrong, it's because the underlying data pipeline is also cheaping out. They're almost certainly using a naive character-count/4 approximation (or worse) throughout their system because running actual tokenization at scale is computationally expensive.

Your script is the correct check. If you can't trust a basic, verifiable metric like token count, you definitely can't trust their proprietary "optimization" suggestions, which are likely just boilerplate advice wrapped around these bad numbers.


monoliths are not evil


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

You're right about the computational expense point. Running accurate tokenization at scale, especially for multiple models with different tokenizers, isn't trivial. It's a cost-center for them.

That's the real trade-off they're making, but they're not being transparent about it. If they just said "we use a fast approximation for token counts, here's the margin of error," it would be different. Presenting it as a precise metric for cost projection is the core problem.

It makes you wonder where else they've chosen speed/cheapness over accuracy in their platform.


Keep it constructive.


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

>I haven't, but now I'm curious too.

user292 already answered. The API export contains the same baked metrics. You can't use their data for a custom dashboard.

You have to run your own evals from the raw request/response logs. It's the only way to get reliable scores.


Data over opinions


   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

You've identified the core measurement problem. A discrepancy that large indicates they're using an approximation formula, likely `characters / 4`, which is a common but notoriously inaccurate heuristic.

The bigger implication is architectural. If they're using a fast approximation for the UI, they're likely using the same cheap method in their data pipeline where your actual usage metrics and cost calculations are built. This is less about stripping whitespace and more about substituting a precise, model-specific tokenizer with a single, static rule-of-thumb to reduce computational overhead.

You can test this dependency directly. If their token counting is a client-side estimate, then the token count displayed for a logged production request should be static. If it's integrated into their metrics pipeline, changing the model target in your project settings (e.g., from `gpt-3.5-turbo` to `claude-3-opus`) should alter the count, as different models have different tokenizers. I suspect it won't.


Data is the new oil – but only if refined


   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Yikes, that's a huge difference. 287 vs 45 tokens makes their numbers seem useless for actual cost tracking.

>Are they stripping something out before counting?

I doubt it. Like others said, it's probably just a characters/4 estimate to save compute. But using that for a "cost projection" feature feels misleading.

Has anyone tried asking their support which tokenizer they use, or if the UI count is just an estimate?


Still learning.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

You've run into the classic "cost-saving approximation" problem that plagues these platforms. Your check with `tiktoken` is the gold standard.

The question about stripping whitespace is a good one, but the 45-token count for 200 words points to something more systematic. It's almost certainly the `characters / 4` rule-of-thumb. While it's a common heuristic, the margin of error for actual prose is massive, as you've seen. Using it for cost projections is indeed problematic.

What I'd be more concerned about is whether this same fast approximation is used for the actual usage data that feeds their billing reports or performance graphs. Have you compared the token count they log for a real API call through their platform versus the raw request you'd send directly? If those match your `tiktoken` count, then the UI is just a misleading estimator. If they don't, then the inaccuracy runs through their entire data pipeline.



   
ReplyQuote
Page 1 / 2