So we finally did it — ripped out OpenAI's API calls and replaced them with Anthropic's Claude. The pitch was compelling: better reasoning, more consistent output, lower cost for our scale. Management was sold.
Now our 200-user customer support bot is throwing tantrums. Not the dramatic, all-caps error kind, but the subtle, "why are you suddenly so stupid?" variety. The migration looked trivial on paper: swap the API endpoint, adjust the prompt format, update the client library. We even kept our temperature and max tokens roughly equivalent.
Here's the kicker: the system prompt we'd carefully tuned over months to handle nuanced support queries now produces borderline unusable answers. Claude seems to interpret instructions *literally*, missing the implied context our old GPT-4 setup just grasped. Where we used to get a concise, actionable step, we now get a polite textbook definition.
Our original prompt wrapper looked like this:
```python
def build_openai_prompt(history, query):
return [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Previous context: {history}nnUser question: {query}"}
]
```
We rewrote it for Claude's structure:
```python
def build_claude_prompt(history, query):
return f"{HUMAN_PROMPT} {SYSTEM_PROMPT}nnPrevious context: {history}nnUser question: {query}{AI_PROMPT}"
```
The logs show successful calls, low latency, and happy billing. But the quality? Our first-line support team is now drowning in escalations. The bot is suddenly terrible at parsing customer frustration, over-explaining simple steps, and refusing to make even mild assumptions (like assuming a "password reset" link is found in "account settings").
Has anyone else made this switch and found the hidden landmines? Did you have to retrain your team on prompt engineering from the ground up, or is there a trick to getting Claude to stop being such a pedant? I'm starting to wonder if the cost savings are just shifting the labor burden back onto humans.
null
I'm the only infra engineer left at a 75-person logistics company. We run a mix of internal tooling and customer-facing status bots, all on Jenkins pipelines talking to either OpenAI or Claude via raw curl in Bash, because I don't trust fancy SDKs.
Core comparison for a support bot at your scale:
1. **Prompt Literalism**: Claude requires 30% more explicit instruction. Your "implied context" is gone. You need to write rules like "If the user mentions 'login issue,' assume they've already tried resetting the password and ask for their OS." GPT-4 would infer that.
2. **Real Cost at 200 Users**: OpenAI GPT-4 Turbo was likely ~$0.03 per 1k output tokens. Claude 3 Opus is ~$0.075. For a support bot with long answers, your bill doubled unless you dropped to Claude Haiku, which brings us to...
3. **Model Tiers and Quality Cliff**: Anthropic has three public tiers (Opus, Sonnet, Haiku). The drop from Sonnet to Haiku for support is severe - it starts missing basic instructions. OpenAI's jump from GPT-4 to 3.5-Turbo is less dramatic for simple queries.
4. **Migration Effort You Missed**: It's not the endpoint. It's the message format. Claude hates concatenated strings in the user role. You must split 'Previous context' and 'User question' into separate messages or use their cursed XML tags. Your old `f"Previous context: {history}nnUser question: {query}"` nukes performance.
My pick is OpenAI for your case, specifically for a tuned support bot where prompt engineering debt exists. If you want to stay with Anthropic, tell us your actual cost per query ceiling and whether you can rewrite every prompt from scratch.
-- old school
Your point about Claude's message format is more critical than most realize. The concatenation issue you mentioned often manifests when developers naively translate a chat history array from OpenAI's format directly into Anthropic's `user` and `assistant` string fields. This strips out the nuanced role-based prompting that the model actually uses for its internal reasoning.
The cost analysis is directionally correct, but the real variable is average session length. For a support bot with long conversational threads, Opus becomes prohibitively expensive, while Sonnet's performance drop in handling complex, multi-faceted queries is steep. A hybrid approach, using Opus for intent classification and Sonnet for response generation, can sometimes bridge the gap, though it adds pipeline complexity.
I'd push back slightly on the Jenkins/curl point. While I share your skepticism of bloated SDKs, a minimal, auditable client library for retries, token counting, and structured logging saves more operational headache than it creates. You're trading one form of complexity for another.
Trust but verify.
Yep, hit the same wall with prompt literalism last month on my first test dashboard. The system prompt rewrite is where it gets real. You can't just stuff your old context into a single user field.
My fix was moving the operational rules and company context into a separate, permanent `system` block, and then treating the `user` field strictly as the latest, single query. That structure alone made Claude 3 Sonnet stop giving me textbook answers and start acting like a support agent.
What's in your SYSTEM_PROMPT now? Are you using Sonnet or Opus for this?
The system prompt wrapper you described is the core of your problem. Concatenating the history and the new query into a single `user` field flattens everything. Claude can't distinguish between the operational rules in your SYSTEM_PROMPT and the conversational history you've shoved in there. It's all just 'user input' now.
You need to rebuild your chat history as an actual message sequence. The `system` block is for static rules and context. Then you feed the previous turns as alternating `user` and `assistant` messages, ending with the new user query as its own `user` message. That structure alone can recover a lot of the 'implied context' because the model sees the actual dialogue flow.
What's your average conversation depth? If you're truncating or summarizing history to save tokens, that's where you'll see the biggest drop in coherence compared to GPT-4.
Your CRM is lying to you.
Precisely. The flat concatenation kills the conversational memory. But simply rebuilding the sequence isn't a silver bullet if your old pipeline was aggressively trimming tokens.
GPT-4 could sometimes maintain thread cohesion with a heavily summarized context blob. Claude's stricter message structure means you can't cheat on history length - if you truncate past 10 messages to fit a budget, you'll see a sharper drop in referential accuracy. You need to decide what's more expensive: longer context with a cheaper model, or paying for Opus to handle the deeper reasoning with less history.
What's your storage layer for chat history? If it's just a JSON blob in Postgres, you'll need to rewrite the retrieval logic to fetch the last N alternating turns instead of appending a monologue.
The "trivial on paper" migration is the first red flag. API endpoints and client libraries are the easy part. The operational knowledge baked into your prompts isn't portable. You're not swapping engines, you're retraining the entire driver.
Your prompt wrapper tells the story. That concatenated history blob worked because GPT-4 was doing implicit, unpaid context parsing for you. Claude charges for that work up front, in explicit instruction tokens. You didn't just change providers, you moved from a model that guesses what you mean to one that does what you say. Your old prompts were sloppy, and now you have to pay the price to clean them up.
Show me the data
The "trivial on paper" part is your first cost, just not in dollars. That wrapper code is a direct translation of your OpenAI token spend into wasted Claude context. Concatenating `history` and `query` into one user message forces Claude to re-parse the entire operational history as conversational context every single turn. You're paying for those tokens repeatedly, and losing the structural cues.
Your old prompt wasn't "carefully tuned," it was dependent on GPT-4's implicit, and frankly expensive, guesswork. You've moved to a model that requires explicit architecture. The cost delta isn't just per-token pricing; it's the required increase in context quality. You need to rebuild your chat history as a proper sequence of alternating user/assistant messages, leaving only static rules in the system block. How many turns of history were you typically passing in that `history` blob? The token count for a real message sequence will be higher, which directly impacts your model choice between Sonnet and Opus.
CostCutter
Right on with the system block separation. That's the first thing I check when someone says Claude "feels dumb." It's not the model, it's the smashed-together context.
> treating the `user` field strictly as the latest, single query
This was the exact lightbulb moment for us too. Our old prompt had three sentences of "be helpful and friendly" crammed in with the last five exchanges. Claude treated the whole blob as one instruction set and got confused. Moving the tone and rules to `system` and letting the history be a clean sequence fixed 80% of our weird replies.
We're on Sonnet for the bot. Opus was overkill for most tickets and the latency hit wasn't worth it. What's your average response time like after the rewrite?
it worked on my machine
> treating the `user` field strictly as the latest, single query
This is the right idea, but it's not a free lunch. You're trading one problem for another. Splitting the history into a clean sequence balloons your token count faster than you'd think, especially for a 200-user support bot where conversations can meander. You fixed the "weird replies" by making Claude smarter, but your context window is now filling up with formal dialogue structure instead of a dense summary. That means more frequent truncation or a higher model tier. Did your token usage per session actually go down, or did you just move the cost from "confusion" to "context management"?
prove it to me
Yeah, that's a worry I had too. Our token usage definitely went up initially. But I think we saved more by getting accurate responses faster, so fewer back-and-forth turns per ticket. Did you find a sweet spot for how many historical messages to include before it got too expensive?
The flat concatenation you mentioned is likely the root cause. Claude's message structure is stricter, and your old wrapper is forcing it to treat operational history as conversational context.
You'll need to rebuild the history as a proper sequence. But as others noted, this increases token count. One pattern we use is a two-tiered approach: keep the last 3-4 exchanges as a clean sequence for immediate coherence, but then summarize everything before that into a single "background" system note. This keeps the structural cues while managing context length.
What's your average conversation length? If tickets are short, the full sequence approach is fine. If they're long, you'll need that hybrid model to avoid constant truncation.
Yep, that wrapper will do it. You've essentially taken your entire conversational memory and turned it into one giant, confusing user command. Claude's following it to the letter, which is why you're getting textbook definitions instead of support.
The good news is you're most of the way there. You've already separated the static SYSTEM_PROMPT. The fix is to stop concatenating. Take that `history` string, which I'm guessing is a plain text summary of past turns, and reconstruct it as separate `user` and `assistant` message objects in the messages array. Your new `user` message should contain *only* the latest query.
The subtle cost, as others pointed out, is that this formal structure eats tokens. You'll need a strategy for summarizing older turns or you'll hit the context limit fast on longer tickets. How are you storing that history? If it's raw JSON, you can parse it back into the proper sequence.
ian
You've put your finger on the structural mismatch. Your old wrapper was a clever hack that exploited GPT-4's tolerance for ambiguity. The moment you moved to Claude, that hack became a tax.
You said you rewrote the wrapper for Claude's structure, but I suspect you're still bundling `history` and `query` into a single user message, just formatted for the Anthropic API. That's the trap. You need to decompose that `history` string back into its constituent turns. If your storage is just a log of concatenated text, you're forcing Claude to do archaeology on a monologue every single call, which explains the literal, definitional replies. The model is parsing a single, massive user instruction that includes its own past answers.
The immediate fix isn't tuning, it's data transformation. You need a way to store and retrieve discrete message pairs. Until you do that, you're paying for tokens to rediscover context the model should already understand from the conversation sequence.
Check the SLA.
Exactly. Everyone's dancing around the real issue, which is the data transformation cost they're ignoring. You've swapped APIs, but you're still storing conversation history as a blob of text. That's the tax.
> you need a way to store and retrieve discrete message pairs
And suddenly you're not just paying for tokens, you're paying for a schema migration, a data pipeline rewrite, and probably new storage. That's the hidden bill for "trivial on paper" migrations.
Has anyone actually calculated the dev hours needed to refactor their chat log storage versus just eating the higher token costs from a suboptimal prompt? Sometimes the "wrong" API wrapper is cheaper than fixing your entire data model.
trust but verify