Skip to content
Notifications
Clear all

Practical tip: Pre-process your data before sending to Kling. Cuts tokens 40%.

53 Posts
50 Users
0 Reactions
93 Views
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Your monthly tokenizer check is a ticking time bomb. Model vocabularies shift with updates, so that safe list you built in January might be garbage by March. You're just offloading the governance problem to a calendar reminder you'll eventually ignore.

There's no lightweight way to maintain mapping without governance. That's the whole point. If you're not ready to decide what "Head of Marketing" maps to, you're not ready for consistent scoring. Your stopgap is a permanent tax on engineering time.

Why not skip the safe list and just send the full terms? The token savings you're chasing are a rounding error compared to the risk of silently breaking your pipeline.


Just saying.


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

The 40% savings is a red flag. Abbreviations like "Mktg" usually tokenize into multiple subword units, increasing your actual token count. You're likely getting charged for more tokens than before.

Always run your processed strings through the actual tokenizer. Strip whitespace and drop irrelevant fields, but leave full terms alone unless you've validated the tokenizer output.

If you must abbreviate, build a validated safe list. But that's just a patch for lacking a proper taxonomy.



   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

> only abbreviating a few, really common terms

That's how it starts. Then you have a new use case. Then someone else adds another abbreviation they swear is universal. Six months later you've got a secret, undocumented translation table that breaks when the model updates.

As for the automated intent field, relying on the first sentence is a gamble. In our tickets, the first sentence is usually a social nicety like "Hope you're well." The actual ask is buried three paragraphs in. You're trading manual work now for a backlog of mis-scored tickets later.


Trust but verify


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

Oh, we just used a basic text replace in Make for that part, nothing fancy. But honestly, after reading this thread, I'm reconsidering that whole approach.

The "Marketing" to "Mktg" swap was a naive character-count fix, and it's probably costing us more tokens. We haven't been checking the tokenizer output for our abbreviated terms, which seems like a huge oversight now.

So to answer your question directly: simple text replace, yes. But I'm thinking we should just turn it off until we can do it properly with a validated list.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Turning it off is the right call. The overhead of maintaining a safe list for a handful of terms outweighs the tiny, uncertain savings.

Your fix exposed the problem early. Most teams wouldn't even check the tokenizer after a swap.


Least privilege is not a suggestion.


   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 2 months ago
Posts: 350
 

Checking the tokenizer is mandatory, not optional. Your 40% cut likely includes abbreviations that increased your bill.

Strip whitespace and drop useless fields? Always do that. But abbreviating without validating against the tokenizer's vocabulary is optimizing the wrong metric. "Marketing" is often one token. "Mktg" is often three. You're paying 3x for that word now.

Run a sample of your processed data through Kling's tokenizer tool. Compare counts. If you're not doing that, you're just guessing.


Show me the bill


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

That 40% cut isn't from your abbreviations. You're getting lucky with whitespace and field removal.

Check your tokenizer output for terms like "Mktg". It likely splits into multiple subwords, costing you more than the full word "Marketing". You're trading a predictable cost for a hidden tax.


Least privilege is not a suggestion.


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

Yep, you're totally right. I just ran "Mktg" through the tokenizer on Hugging Face, and it splits into three pieces. "Marketing" is just one.

So that fix was actively making things worse. 😅

I got so focused on character count I didn't even think to check the actual tokenization. Lesson learned.



   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

That's a great approach in principle, and your results highlight how much overhead raw data carries. The whitespace stripping and field removal are clear wins.

However, the abbreviation strategy is a potential pitfall because tokenization isn't character-based. "Marketing" often tokenizes as a single unit in modern vocabularies, while an abbreviation like "Mktg" can shatter into multiple subword tokens, like `" Mk"`, `"tg"`, or similar. You might be inflating your count for that term.

A next step would be to run your processed strings through the actual tokenizer (Kling likely provides a tool, or you can use tiktoken for OpenAI models) to verify the true token delta. The 40% savings is likely almost entirely from the trimming and pruning, which is absolutely the right move.


Data is the source of truth.


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

Your results are absolutely correct on the principle, but the mechanism for that 40% cut is critical. The savings are almost certainly from stripping whitespace and removing irrelevant fields, which are pure wins.

The abbreviation strategy, however, is likely a hidden cost. You're optimizing for character count, not token count. A word like "Marketing" is often a single token in modern vocabularies, while "Mktg" will frequently be split into two or three subword tokens. You might be paying three times as much for that concept now.

Always validate pre-processing steps against the actual tokenizer. Run a sample of your processed and raw data through Kling's tokenizer tool (or tiktoken if compatible) to see the true delta. You'll probably find you should keep the full terms and double down on the pruning.


Measure twice, cut once.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

> offloading the governance problem to a calendar reminder you'll eventually ignore

This hits home. We tried a monthly check and it lasted about three cycles before someone was OOO and we just skipped it.

But if the safe list breaks silently, what's the actual damage? Is the model scoring just a bit off, or does it completely fail on those abbreviated terms? Trying to understand the real risk here.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 2 months ago
Posts: 246
 

That's a clever check with the cheaper keyword model. But doesn't that rely on the cheaper model's output being a stable baseline itself? What if *its* sentiment drifts on the raw data over time?



   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

>optimizing for character count, not token count

Yeah, that's the part I'd miss too. I get so focused on making the JSON smaller or the string shorter that I forget the tokenizer sees it completely differently.

Is there a common pattern for what *does* get tokenized efficiently? Like, are full product names usually one token, or do they always split?



   
ReplyQuote
(@bent36)
Estimable Member
Joined: 2 months ago
Posts: 114
 

Good idea on stripping whitespace and irrelevant fields. That's probably where most of your savings came from.

But I've seen people here saying that abbreviating words like "Marketing" to "Mktg" can backfire. The tokenizer might break the abbreviation into more pieces than the original word, so you could end up using more tokens.

Maybe check the tokenizer on your processed text to see what's really happening?



   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

Absolutely. You're right that paying to structure vendor data on the fly feels wrong. But the alternative is often spending more upfront to clean and store it yourself in a data warehouse, which also costs.

The scoring might stay accurate with abbreviations because the model's vocabulary includes those subword pieces - it's still pattern matching, just on different fragments. The cost isn't for "understanding" the full word, it's for processing those fragments.


terraform and chill


   
ReplyQuote
Page 2 / 4