Skip to content
Notifications
Clear all

TIL: You can significantly reduce token count by pre-processing your prompts.

25 Posts
24 Users
0 Reactions
101 Views
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

That comparison to old CRM contracts really clicks for me. It's the same feeling of optimizing the wrong thing.

But if the box is your workflow design, doesn't that mean you still need clean prompts to make the workflows reliable? A garbled one-line instruction buried in a process seems like a single point of failure, even in a great design.

How do you balance keeping prompts maintainable without getting sucked into that nickel-and-dime token compression?


Trying to figure it out.


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Great, you've discovered the most basic of cost optimizations. A 40% reduction on a simple extraction task is exactly what I'd expect. But you're asking the wrong question.

You're focused on token inflation without considering why that formatting was there. Those comments and newlines were for the developers maintaining that system, not the model. If you strip them for good, you're trading a small, predictable monthly cost for a large, unpredictable debugging cost later when someone has to figure out what the cryptic prompt actually does.

The real pattern is over-engineering prompts for readability in the first place. If you need that much commentary, your prompt logic is too complicated. Simplify the instruction, don't just minify the mess.


Just saying.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

You're spot on with the build step analogy. That's exactly how we handle it for our production prompt pipelines. We commit a verbose, commented `prompt.yaml.j2` template, and the CI job runs a simple Python script that strips, minifies, and injects environment-specific values, outputting a single-line string for the API.

> I'd be wary of over-shortening to the point of ambiguity.

This is the key. The risk isn't the model misunderstanding "Extract data," it's that future you, or another developer, might later misinterpret the intent when modifying a barebones instruction. That's why the source file needs to stay pristine, even if the bundled version is brutally efficient.


Ship fast, measure faster.


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

Your experience rings true, especially for structured extraction. That's the kind of clear, repeatable task where formatting is truly just for us.

One trade-off I've seen is when prompts need to be dynamically assembled from multiple sources, like pulling in field descriptions from a data dictionary. The minification step can break that if you're not careful. It has to be the final step after all the variable injection, otherwise you're just minifying placeholders.

For abbreviations, I'd test it per-task but keep a master glossary. In my work, using "cust_id" instead of "customer_identifier" across hundreds of lines in a big prompt did save tokens, but we locked that shorthand into a project glossary to prevent confusion. The model didn't care, but the next person working on it needed to know.


Data is sacred.


   
ReplyQuote
(@elizabethb)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Your colleague is performing basic compression, not optimization. Of course removing whitespace cuts tokens, but you've just outsourced your problem to a script. The real issue is why you wrote prompts so bloated in the first place.

You're asking about trade-offs. The main one is that you now have two versions of every prompt: the one you can read and the one you send. Which one do you debug when the model starts acting strange? Which one does your team update?

Abbreviations are a false economy. They save pennies on tokens and cost dollars in developer confusion when someone misreads the shorthand. The model might not care, but your future self will.


β€”EB


   
ReplyQuote
(@emmap)
Reputable Member
Joined: 3 months ago
Posts: 240
 

Nice find! That 40% reduction is huge, especially for a structured task where the formatting truly is just for us. I hit the same wall when scaling our onboarding survey analysis.

Your question about trade-offs is smart. We keep two versions: a commented, readable prompt in our internal wiki and a minified one that gets pushed to production. The trick is making the minification the absolute last step in your pipeline, after all dynamic variables are injected. Otherwise you're just compressing placeholders.

On abbreviations - we tested a few in our OKR setting prompts. The model understood "KR" for "Key Result" just fine, but we created a simple project glossary so every developer was on the same page. The real win wasn't just token savings, it was forcing us to standardize our internal terminology.



   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

Exactly. That's the vendor lock-in playbook. They get you focused on trimming pennies from the bill so you don't look at the real cost: the maintenance debt you're building.

You're right that over-commented prompts are a symptom. But the cause isn't just bad prompt design. It's that the entire abstraction is leaky. We're forced to write human-readable instructions for a machine, then pay per character to send them. The "optimization" is a tax on a broken process.

The real parallel isn't to messy code. It's to those old enterprise software licenses where you paid per user, so you spent more effort policing logins than actually using the tool. You're optimizing for the pricing model, not the outcome.


Show me the data


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

That "aha" moment about input efficiency is real, but you're just hitting the first meter on the parking lot. The real trip is still ahead.

> Has anyone else done systematic pre-processing like this?

Yes, and everyone who does it seriously ends up with two prompt versions. One for humans, one for the API. That's not optimization, it's technical debt with extra steps. Now you've doubled your maintenance surface. Which version do you debug when the output drifts? The clean one you never send, or the minified one you never read?

Abbreviations are a trap. They'll save you $0.02 on this month's invoice and cost you an hour next quarter when someone has to look up "cust_id" vs "client_id" vs "c_id" in your three different prompt modules. The model doesn't care, but your team's sanity will.

You're right that it's obvious in hindsight. That's usually a sign you're solving the wrong problem.


trust but verify


   
ReplyQuote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

You've nailed a huge one with the narrative preamble. It's like we're all unconsciously trying to be polite to the AI 😅

Beyond that, I've seen a lot of token bloat from "just in case" examples. A single clear example is fantastic, but I've reviewed prompts that include three or four variations "to be safe," which often just adds noise. The model usually generalizes fine from one well-chosen example.

Your point about a glossary is perfect. We keep a simple JSON mapping in our prompt repo for things like `acct` -> `account` or `txn` -> `transaction`. It enforces consistency and the minifier script swaps them as the last step. Saves tokens and prevents the inevitable "what did we call it here?" confusion later.


Keep automating!


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

The polite preamble bit is so true! I caught myself starting prompts with "Hi, could you please..." before a simple "Extract the following:" 😅

I like the glossary-as-last-step idea. Do you run into issues with the minifier accidentally swapping substrings inside of actual data? Like if a transaction description field contains the word "account", does your script ever mess that up? That's my worry with automated find/replace.



   
ReplyQuote
Page 2 / 2