Skip to content
Notifications
Clear all

Help: Kling keeps 'hallucinating' API endpoints that don't exist in our docs.

11 Posts
10 Users
0 Reactions
32 Views
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
Topic starter   [#23876]

We've been piloting Kling for automated API integration and data ingestion from several internal microservices. While its natural language parsing for endpoint discovery is impressive, we're encountering a persistent and problematic pattern: the agent is confidently generating calls to API endpoints that are not present in our actual OpenAPI/Swagger documentation.

The issue manifests during the specification parsing and planning phase. For example, when given our `user-service` spec, which has a documented `GET /users/{id}` endpoint, Kling's execution plan will often include a call to `GET /users/{id}/preferences`, claiming it's for fetching user settings. No such endpoint exists in the spec file we provided. This causes the pipeline to fail at runtime with 404 errors.

We've tried the following to mitigate this, with limited success:
* Verified our OpenAPI 3.0 specs are valid and hosted at a stable URL.
* Explicitly provided the exact spec file path in the agent configuration.
* Used clear, imperative prompts like "Using only the endpoints defined in the provided specification, fetch data for user ID 456."

Our current configuration snippet is straightforward:
```yaml
agent:
source: kling
spec_url: "https://internal-api.company.com/specs/user-service.yaml"
instructions: "Ingest data from the user service API. Do not assume or invent endpoints."
```

The "hallucinated" endpoints are often plausible (e.g., `/preferences`, `/history`), suggesting it's drawing from a general pattern in its training data rather than our specific docs. This undermines reliability for automated workflows.

Has anyone else faced this? Is there a configuration parameter or prompt engineering technique to strictly bind the agent to the provided specification and suppress this extrapolative behavior? We need deterministic behavior based on the documented contract.

— DN


Data is the only truth.


   
Quote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Oh that's frustrating. We ran into something similar with a different tool last year. It kept inventing `/project/{id}/summary` endpoints because it had seen similar patterns in public API docs during training.

I wonder if Kling is doing pattern completion based on its training data, rather than strictly adhering to your provided spec. Have you checked if your prompts can include a stronger boundary instruction, like "Do not infer or create any endpoints not explicitly listed in the specification file"? Sometimes these agents need that extra, explicit guardrail.



   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Yeah, the stronger boundary prompt idea is a good one to try. I've found these tools can be surprisingly literal when you force them into that mode. Something like "You must only use endpoints exactly as they are named and described in the following spec. Do not deviate, assume, or create new endpoints" can sometimes shift the behavior.

But your core issue, where it's inventing `/preferences` routes, feels like it might be a deeper training bias issue, as user805 suggested. Kling's model might have been heavily trained on common API patterns, and it's defaulting to "completing" what it thinks a RESTful service *should* have, rather than what it *does* have. Have you reached out to their support team? A persistent pattern like this might need a model-side adjustment or a config flag they haven't documented yet.


Raise the signal, lower the noise.


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

The training bias point is spot on. I've seen this exact pattern with other agents - they try to be 'helpful' by filling in perceived gaps based on public API trends. It's less hallucination and more inappropriate extrapolation.

Your boundary prompt suggestion works in many cases, but it's brittle. If the model's training data strongly reinforces common patterns like `/preferences` or `/summary` subroutes, it can override even explicit instructions. That's when you need vendor support.

Have you tried providing Kling with only the spec, stripped of any natural language task description? Sometimes removing the 'why' from the prompt forces stricter adherence to the 'what'.



   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Ugh, this hits close to home. I had the exact same issue last year, but with a different vendor's "smart" API connector trying to infer `/{id}/history` endpoints from our sales object specs. It was so confident and wrong.

Everyone's suggestion about training bias is correct, but the boundary prompt alone might not stick. What finally worked for us was a two-part approach: first, that ultra-strict prompt you already tried, but second, and more importantly, we started feeding the spec *and then* immediately asking for a verification step. We'd prompt: "First, list every endpoint verb and path you see in the attached spec. Confirm that this is the complete list you will use." Making it explicitly enumerate its source material before planning seemed to ground it.

It's a tedious extra step, but it cut the phantom endpoint calls by about 90% for us. Have you tried forcing that kind of manual checkpoint in the flow?



   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
Topic starter  

The pattern completion hypothesis is likely correct. These models treat API path structures as a sequence prediction problem, and common suffixes like `/preferences` have high probability weights.

Have you examined the raw output from Kling's spec parsing stage? The intermediary representation might reveal where the extrapolation occurs. I'd instrument the pipeline to log the exact OpenAPI spec nodes it claims to be using for its plan generation.

If you can access that trace, compare it to the spec's actual paths object. The discrepancy often isn't in the final call but in an intermediate reasoning step where it "infers" related resources.


Data is the only truth.


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

Tracing intermediate steps is the only way to diagnose this, but you're assuming you'll get that level of transparency. Most of these tools are black boxes.

The real waste is engineering hours spent debugging their opaque reasoning. You instrument, you log, and you still have to file a support ticket.

Focus on what you can control. Run a validation layer before execution that compares planned calls against a known spec. Reject anything not on the list. Don't trust the tool to self-audit.


show me the bill


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Yeah, that's a solid prompt structure. I've had good luck with that literal instruction mode too, but you're right about the training bias being a tougher nut to crack.

Even with a strong prompt, if the underlying model is heavily tuned on public API patterns, it can sometimes "autocomplete" in its internal representation. I'd still try the strict boundary, but pairing it with the verification step user354 mentioned seems like the pragmatic next move before a support ticket.


data over opinions


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

That `/users/{id}/preferences` example is a textbook case of the pattern completion issue. Since you've already tried the clear imperative prompts and validated your spec, the next step really depends on your access to the tool's internals.

Can you check the logs for Kling's initial spec ingestion? Sometimes the issue isn't in the final plan, but in the agent's internal parsed representation of your API. If it's already added the inferred route there, all downstream steps will use it. If you can't see those intermediate steps, then implementing an external validation gate, as some suggested, might be your most reliable stopgap while you push the vendor for a fix. It's an extra layer, but it prevents the runtime 404s.


Architect first, buy later


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

The suggestion to strip out the natural language task description is a good one, and it often helps. But I've found it's not always enough when the underlying model has been heavily optimized for 'helpful' task completion.

In my tests, providing only a raw spec sometimes just shifts the problem. The agent still makes internal assumptions about what a 'complete' task would require, and if it decides a `/preferences` endpoint is logically necessary, it may still attempt to construct or infer it, treating the spec as an incomplete reference rather than an authoritative source. The pattern completion is baked into the reasoning, not just the prompt interpretation.

So while it's a necessary step for cleaner input, it's often a precursor to needing that external validation layer everyone's mentioning.


Data is the source of truth.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That's a good diagnostic step if you can get the data. I've had to do similar digging with other tools, but in my experience, the parsing trace often shows a sanitized, post-reasoning output, not the raw inference step. You'll see a clean list that matches the spec, because the internal 'completion' step happens before that log point.

If you can't see the actual chain-of-thought, you're stuck. The workaround is that external validation gate, like others said. Log the planned calls from the agent's final output and diff them against a programmatically loaded version of your spec. Reject the entire operation if there's a mismatch. It's redundant, but it's the only reliable control point.


Automate everything. Twice.


   
ReplyQuote