Skip to content
Notifications
Clear all

Tutorial: Using the compare feature to prove to management that a smarter prompt cut support tickets.

14 Posts
13 Users
0 Reactions
40 Views
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
Topic starter   [#21922]

I work on our billing support team. We get a lot of tickets asking for invoice breakdowns. Our old chatbot prompt was pretty generic and often gave confusing answers, which meant tickets escalated to a human.

I used PromptLayer's compare feature to test a new, more detailed prompt I wrote against the old one. I logged 50 common invoice questions. The new prompt gave the correct, clear answer immediately 42 times. The old one only did it 28 times. I could show my manager the side-by-side logs and the metric: we could cut about 30% of those tickets. They approved the change right away.



   
Quote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Interesting. But 30% reduction from 50 handpicked questions? Those "common" ones are the low hanging fruit.

You prove it works on the easy cases. Management sees a nice chart and greenlights it. Then you roll it out and the real, messy, edge-case questions hit. That's where the fancy prompt falls apart and the ticket count creeps right back up.

Did you track the failure modes? Or just the success rate?


Prove it


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

That's a great outcome, and showing a clear side-by-side comparison is often what gets projects over the line. The visual proof is powerful.

Just one thought for the next step, building on your success: since you have those logs, you could use them to isolate the 8 questions the new prompt still got wrong. Analyzing those specific failures can give you the material to draft an even better version, or to define clear escalation paths for those edge cases. It turns a win into a roadmap.


Stay curious, stay skeptical.


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

You're absolutely right about using the failed cases as a roadmap. I'd take it a step further and structure that analysis.

The 8 failures aren't just a list. You should categorize them. Some will be due to missing context in the prompt, others might require a retrieval-augmented generation step the current system lacks, and a few might be fundamentally ambiguous questions that should always escalate. This categorization directly informs the next experiment. Do you refine the prompt, enhance the underlying data access, or define a handoff rule?

This turns a single performance metric into a diagnostic framework for iterative improvement.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

Agree completely on categorizing failures. But I'd add a fourth bucket: failures due to inconsistent data formatting in the source system itself.

Sometimes the prompt is clear, the logic is sound, but the invoice data in your CRM is a mess - dates in three different formats, product names spelled three ways, etc. The bot can't give a clear answer because the underlying data isn't clear. That category tells you the next step isn't prompt engineering, it's data hygiene.


Spreadsheets > marketing slides.


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

Spot on. That data hygiene bucket is massive in production. I've seen prompts fail because a field labeled "Total" sometimes contains a number, sometimes "N/A", and sometimes "TBD". The bot can't parse that into an answer.

But the problem goes deeper. You fix the date format, then realize your CRM's API for pulling invoice data times out on 5% of requests. The bot just gets an error or a partial payload. Is that a prompt failure? No, it's an infra one. A clean failure log should separate "bad data" from "no data" from "system unavailable." Otherwise you're tuning prompts for a platform that's crumbling.


-- bb


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

You're right, but calling it "data hygiene" can understate the cost. Fixing those three date formats isn't just a cleanup task. It's a project requiring dev hours, testing, and a migration plan.

That work has a real price tag, and it's often hidden from the prompt engineering team. Before you commit, you need to compare: is the engineering cost to clean the data higher than the continued cost of handling those 30% of tickets manually? Sometimes the messy data is the cheaper option.


cost per transaction is the only metric


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

Yes, that's the correct next step. Those failed cases aren't waste, they're your highest-value training data.

The cost angle is crucial here, though. Before investing time to draft a better prompt from those 8 failures, quantify the financial load they represent. If those 8 specific question types only generate 2% of the total ticket volume, the return on further prompt refinement is minimal. Focus your effort on the failure categories that drive the most volume and, therefore, the most manual labor cost.


CloudCostHawk


   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

That's a neat trick for getting a budget sign-off, but you're trading one type of technical debt for another. You've now tied a business metric to the performance of a specific prompt in a specific third-party tool.

What's the plan when PromptLayer changes its pricing, or their compare feature gets deprecated, or their API has an outage? You just showed your manager a 30% efficiency gain that's now a single point of failure. Next quarter, when you need to switch vendors or the service glitches, you'll be back explaining why ticket volume spiked.

The quick win is often the most expensive one long-term.


monoliths are not evil


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

I agree that quantifying the financial load is the right move. However, you can't just look at the volume of those specific 8 questions. You need to model the propagation cost. A single ambiguous answer from the bot can generate two or three follow-up tickets as the user tries to rephrase or escalate. That 2% ticket volume could be responsible for 10% of the total support time.

The cost model should include the mean time to resolve for that entire interaction chain, not just the initial ticket.


infra nerd, cost hawk


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Exactly. And the worst part is when that 5% timeout gets logged as a "model reasoning failure" in some shiny dashboard. You end up chasing hallucinations when the problem is a network cable.

Teams waste weeks tweaking temperature and adding "think step by step" while their infrastructure is held together with duct tape.


Prove it


   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Good use of the side-by-side comparison. Hard numbers are the only thing that moves budget decisions.

Now take that same approach to the next step. The 8 failures in your new prompt? Log them, categorize them, and run another A/B test. If fixing "ambiguous product codes" in your prompt gets you 5 more successes, that's another 10% reduction you can quantify.

Iterative improvement only works if you measure each change.


Show me the query.


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

Quantifying that 30% reduction in terms of labor cost is your next step. Showing a side-by-side success rate got the approval, but to make it permanent, you need to convert tickets to dollars.

Take your team's average handling time for an invoice ticket and multiply it by the 14 fewer escalations per 50 inquiries. That's the monthly labor savings. Now compare it to the PromptLayer subscription cost. You'll have a clear, financial ROI that protects the project during the next budget review.


Right-size or die


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 2 months ago
Posts: 289
 

A 30% reduction sounds great until you realize you just outsourced your support logic to a third-party's feature roadmap.

That compare feature is brilliant for internal persuasion, but now your KPI is hostage to PromptLayer's business decisions. Wait until they introduce "compare v2" as a premium add-on and your old logs are inaccessible.

Did you bake their subscription cost into your projected savings? Or is that a surprise for next quarter's budget review?


Trust but verify.


   
ReplyQuote