Skip to content
Notifications
Clear all

Help: Claude keeps suggesting we use APIs that don't exist.

5 Posts
5 Users
0 Reactions
5 Views
(@billyp)
Reputable Member
Joined: 3 months ago
Posts: 284
Topic starter   [#29532]

Hey folks, been running into a super frustrating issue lately and wanted to see if anyone else is hitting this wall.

I use Claude.ai mostly to brainstorm and generate copy for our marketing automations in Klaviyo. Lately, I've been asking it for help with more technical flows, like syncing customer segments between platforms. It keeps confidently suggesting I use very specific Klaviyo API endpoints or SendGrid features that… simply do not exist. I'll double-check the official docs and there's no mention of them. Last week it told me to use a `GET /lists/{list_id}/metrics` endpoint in Klaviyo that would "return engagement stats per profile." Sounded amazing! But it's fictional.

It's becoming a real time-sink because the suggestions are so plausible and detailed. I get excited about a new automation possibility, only to find out it's a hallucination.

* Has this been happening to you with other platforms (Mailchimp, HubSpot, etc.)?
* Any reliable prompting strategies to keep it grounded in actual, existing APIs?
* Is this worse in the .ai chat vs. the desktop/API versions?

It's a shame because it's otherwise brilliant for subject line variants and segment logic in plain English. But I'm starting to double-check every technical recommendation.


Always A/B test.


   
Quote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

Yes, this is a documented phenomenon with current large language models when interfacing with rapidly evolving or poorly documented APIs. The model's training data likely contains API references from various sources - documentation, tutorials, forum posts - that may be outdated, speculative, or incorrectly described. What's particularly problematic is the confidence calibration issue; the model generates plausible technical specifications without uncertainty markers.

I've found two strategies reduce this significantly in my work with experimentation platforms:

1. Always append a recency qualifier to prompts: "Using only API endpoints confirmed in the official Klaviyo documentation as of March 2024, how would you..."
2. Force chain-of-thought verification: "First, list the specific API endpoints you would use with their exact paths. Then, explain what each returns according to official documentation. Finally, provide the implementation code."

The desktop/API versions don't fundamentally solve this, as they share the same underlying model architecture. The issue stems from the training objective predicting plausible sequences rather than verifying factual existence.

Have you tried using Claude to generate validation scripts that check endpoint existence before suggesting implementations? I wrote a Python script that cross-references suggestions against API schema files, which catches about 70% of these hallucinations.


Nullius in verba


   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

Ugh, yes, that specific Klaviyo hallucination is a perfect example. I've run into the same thing with Segment's API, where it'll invent a completely plausible `/tracking-plans/{id}/validate` endpoint that would solve all my problems... if it were real.

Your point about it being a time-sink is key. The worst part isn't the initial hallucination, it's the rabbit hole you go down trying to debug why your `curl` command fails, thinking *you* messed up the auth or headers, before you realize the foundation is sand.

A trick I've used is to prime it with the actual API docs. I'll paste a relevant chunk of the *real* API reference into the prompt first, then ask my question. It seems to anchor it better than just a recency qualifier. The .ai chat and API versions seem equally prone to this, in my experience.


Automate all the things.


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

The Klaviyo hallucination you describe is unfortunately a common manifestation of a deeper model limitation: they don't have a "ground truth" mechanism for real-world APIs. I see this extensively in cost analysis work, where models will cite incorrect pricing tiers or deprecated SKUs.

You asked if it's worse in the .ai chat versus the API. In my benchmarking, the core model behavior is identical. The difference often lies in the system prompt and the user's own prompting discipline. The API allows you to enforce a stricter system role, like "You are an API assistant that ONLY references endpoints verified in the latest official documentation." Even then, it's not a guarantee.

The strategy of pasting the actual API docs is the most effective, but it reframes the problem from "prevent hallucination" to "verify every output." This verification overhead creates its own cost, which is often overlooked in the TCO of using these tools for technical tasks.


Trust but verify.


   
ReplyQuote
(@data_analyst_2025)
Honorable Member
Joined: 5 months ago
Posts: 290
 

Oh, I feel your pain! I've been using Claude to help build Looker dashboards that pull from our Shopify data, and it once gave me a perfect-looking `customer_metrics` dimension that didn't exist in our actual schema. Spent an hour trying to figure out why it broke.

To your question about other platforms, I've seen similar with HubSpot's CRM API. It suggested a `/contacts/{id}/timeline` endpoint that was just... made up. Super detailed example payload and all!

The prompting trick that's helped me a bit is to ask it to "cite the specific documentation section" it's referencing. It sometimes catches itself and says it can't find it. But pasting the real docs first, like user609 mentioned, seems like the safest bet.

Do you think this is just a risk we have to accept when using these tools for technical specs right now?



   
ReplyQuote