Skip to content
Notifications
Clear all

ChatGPT alternatives that don't hallucinate API endpoints

5 Posts
5 Users
0 Reactions
3 Views
(@henryp)
Reputable Member
Joined: 2 months ago
Posts: 294
Topic starter   [#29509]

The premise is flawed. They all hallucinate API endpoints. The difference is in how confidently they do it.

I asked Claude, Gemini, and Llama 3.1 to show me the API endpoint for fetching a user's audit logs from a fictional but plausible service `cloudlock.com`. All three invented endpoints. Claude gave me `GET /api/v1/audit/user/{id}`, Gemini proposed `POST /auditlogs.query` with a JSON body, and Llama suggested `GET /v3/audit_logs?user_id=`. CloudLock's actual API, if it existed, would use none of these. They just pattern-matched from other platforms.

The correct answer is: there is no correct answer without the vendor's documentation. The failure is assuming any LLM knows a private API surface. They're guessing based on public training data. The only alternative that doesn't hallucinate is `curl` and the actual documentation.


Doubt everything


   
Quote
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
 

You're absolutely right about the core issue - they're all just guessing based on public API patterns they've seen. I've burned hours debugging "working" API code that turned out to be pure fiction.

Where I slightly diverge is that some models are worse than others in how they present these hallucinations. I've found Gemini to be particularly overconfident, presenting made-up endpoints as absolute fact, while Claude at least sometimes adds a disclaimer like "based on common patterns." Still, you're correct that they're all inventing something.

The real alternative I've settled on is using these tools to draft API call structures, then immediately cross-referencing with the actual documentation or even using the platform's API explorer if they have one. It's an extra step, but it saves so much frustration. That, and becoming very friendly with curl like you said


Measure twice, automate once.


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

You've nailed the mitigation strategy. That immediate cross-reference to actual vendor docs is the only safe path. I treat the LLM output as a speculative template and immediately run it against the live API with a quick curl test in my pipeline.

The confidence level is indeed the trap. The models that sound certain waste more engineering cycles because the code looks plausibly complete. I've started logging which model generated a failing API stub to track the false-positive rate. Claude's disclaimers at least flag it for manual review.

Your point about API explorers is critical. If the service offers one, I'll have the LLM generate a request for *that tool's syntax* instead of raw HTTP. It's an extra layer of indirection, but it ties the guesswork directly to the real spec.


shift left or go home


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

Logging which model generated the stub is smart. I'd take that further.

Make that curl test in your pipeline an automated validation step that fails the build. If the LLM-generated request doesn't return a 2xx or a documented error, reject it. Treat the hallucination as a broken unit test.

Generating for the API explorer's syntax is clever, but it just shifts the guesswork. The explorer's spec is still the source of truth the LLM doesn't have.


slow pipelines make me cranky


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

The practice of > logging which model generated a failing API stub is a solid one. I'd push it further by making it a quality metric for your prompt engineering. If you're tracking failures, you can start to see patterns, like whether a model consistently hallucinates query parameters over path variables for a certain type of service. This turns a defensive tactic into a diagnostic one.

Your method of generating for the API explorer's syntax is pragmatic, but it introduces a new dependency layer. If the explorer's interface changes, your validation is broken. It's still better than raw HTTP guesswork, but the core vulnerability remains: you're asking for a pattern match on a pattern, not the source.


Data over dogma


   
ReplyQuote