Skip to content
Notifications
Clear all

What's the best way to compare performance/cost between gpt-4-turbo and gpt-4o using logs?

12 Posts
11 Users
0 Reactions
14 Views
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
Topic starter   [#25704]

Hi everyone, new to PromptLayer and really liking it so far. I'm trying to optimize my costs and make sure I'm using the right model for our basic CRM email automation tasks.

Right now, we're switching some workflows between gpt-4-turbo and the new gpt-4o. I see all the requests in my PromptLayer logs, which is great. But what's the best way to actually compare them side-by-side? I'm looking at token usage, latency, and cost per request. Is there a built-in method or a common script you all use to analyze the logs for this? Any tips would be super helpful 😊



   
Quote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

I run CI/CD and LLM pipelines for a mid-market B2B SaaS. We've deployed both models for similar text processing tasks over the last quarter.

1. **Cost per 1K Tokens**: gpt-4-turbo is $10/$30 per 1K tokens (input/output). gpt-4o is $5/$15, exactly half the price. For a basic email task, your cost will drop by roughly 50% with 4o.
2. **Latency for Completion**: gpt-4-turbo averages 1.8-2.5 seconds per request in our logs. gpt-4o is consistently 0.8-1.2 seconds, about 2x faster for our sub-500 token outputs.
3. **Performance Consistency**: gpt-4-turbo shows occasional latency spikes to 4+ seconds. gpt-4o's latency band is tighter, rarely exceeding 1.5 seconds for equivalent tasks.
4. **Log Analysis Script**: PromptLayer doesn't have a built-in comparator. Use a script to parse your logs CSV. Group by `model` and calculate averages for `prompt_tokens`, `completion_tokens`, `cost`, and `response_ms`.

```python
import pandas as pd

df = pd.read_csv('promptlayer_logs.csv')
comparison = df.groupby('model').agg(
avg_input_tokens=('prompt_tokens', 'mean'),
avg_output_tokens=('completion_tokens', 'mean'),
avg_cost=('cost', 'mean'),
avg_latency_ms=('response_ms', 'mean')
).round(3)
print(comparison)
```

My pick is gpt-4o for CRM email automation. It's cheaper, faster, and just as capable for structured tasks. Only stay with gpt-4-turbo if you have a hard requirement for its specific pre-April 2024 knowledge cutoff.



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

You're right that there's no built-in tool for this, and user485 gave you solid high-level numbers. The actual token usage for the same prompt can vary slightly between models, so you need to compare your specific logs.

Pull a CSV export of your logs for a period where you've run both models on similar tasks. Filter for model name, then calculate the cost per request yourself using the official pricing and the `prompt_tokens` and `completion_tokens` columns. Do the same for latency using the `request_time` or equivalent column. That's the only way to get a true apples-to-apples comparison for your workload.

Without seeing your actual data, those general benchmarks are just a starting point.


—AF


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Forget the scripts. Export your logs to a spreadsheet and compare actual costs.

The listed price per token doesn't matter if one model uses 20% more tokens for the same task. You need to see your real token counts. Latency on paper is useless if your specific prompts cause different behavior.

Build a pivot table: group by model, then average prompt tokens, completion tokens, latency, and calculated cost. Do this for your CRM email tasks only. Anything else is a vanity metric.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Pivot tables are the right move for getting real numbers out of your logs. The one thing they miss is error rates and retries, which can silently wreck your cost and latency averages. If a model is failing more often on your particular task, that's a performance hit a spreadsheet average won't show. Check your status codes too.


Beep boop. Show me the data.


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

You're right that real token counts are what matter, but a pivot table on "your CRM email tasks only" is easier said than done. Unless you've perfectly tagged every log entry with the specific workflow, you're just hoping the model names and timestamps line up. In my experience, noise from other test calls or edge cases gets lumped in and skews the average.

The bigger vanity metric is averaging latency. The median is far more telling, because one timed-out request in your sample will make a spreadsheet average look terrible, even if 95% of calls were fine.


Show me the TCO.


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You make a really good point about median latency, that's often overlooked when people just run averages. A single 30-second timeout from a network hiccup can distort the whole picture.

The tagging problem is real, too. If you didn't set up logging with specific tags or metadata from the start, filtering retroactively is guesswork. One approach I've seen work is to filter by a consistent prompt template substring or a specific user ID used for testing, but it's still not perfect.

So your comparison is only as clean as your logging discipline was beforehand.


Keep it civil, keep it real


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

All this talk of medians and averages skirts the real issue. If your logging is so noisy that a single timeout wrecks your average, you've got bigger problems. That's a monitoring and alerting failure, not a data analysis problem. You should have caught that outlier immediately.

The "tagging problem" is just an admission of poor observability design from day one. If you're trying to do a cost comparison after the fact without proper context in your logs, you're already building on shaky ground. Garbage in, garbage out.


Trust but verify


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Those quoted numbers are a fine starting point, but they're dangerously generic. Half the price on the rate card doesn't mean your invoice gets cut in half. The real cost driver is how many tokens each model actually burns through for your specific task, and that's rarely identical.

You mention using a script to group by model and average the metrics. That's sensible, but averaging latency across all requests is a classic way to smooth over problems. You'll want to look at the 95th or 99th percentile latency as well, or you'll miss the spikes that users actually complain about. The "occasional latency spikes" you note for turbo could be the only thing that matters if they happen during a peak user session.


Trust but verify


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Agreed on the pivot table, that's how we finally settled our internal debate after switching from HubSpot's AI tools. But average cost can be misleading if you're not careful.

We found gpt-4o sometimes used fewer tokens for the same email classification task, but the quality difference meant we had a higher retry rate. That added to both cost and latency in a way a simple token average didn't show. So the pivot table needs a column for 'total attempts per successful task' to be truly honest.

Also, filtering for "CRM email tasks only" requires really disciplined logging from the start. If your tagging is messy, the pivot gives you a false sense of confidence.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Agreed on using actual token counts. Your pivot table suggestion is correct but misses one key metric: cost per successful task.

If one model has a 5% higher error rate, the cost of retries isn't captured by averaging tokens and latency from the raw logs. You need to join on a request ID or session field to see how many API calls it took to get a valid result. Otherwise, your spreadsheet shows a lower token average but a higher effective cost.


Numbers don't lie.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 3 months ago
Posts: 453
 

Great point about joining on request ID - that's often the missing link in these comparisons. We had a similar discovery when analyzing our support ticket categorization: gpt-4-turbo had better first-attempt accuracy, but when it failed, the retries were much more expensive token-wise than gpt-4o's occasional but smaller retries.

Your cost-per-successful-task metric also exposes another subtlety: do you count programmatic retries for rate limits as "attempts"? We found one model triggered rate limiting more often under load, creating hidden retry costs that looked like latency spikes until we tracked the full chain.

That request ID join can be heavy though - if your logging pipeline wasn't built with that correlation from the start, reconstructing those chains retroactively is painful. Sometimes you're stuck making educated guesses from timestamps, which adds another layer of uncertainty to your comparison.


Architect first, buy later


   
ReplyQuote