Hey folks, I've been integrating Kling's API into a small project and noticed something odd. Even with a very explicit system prompt, the model occasionally goes off-script. It doesn't happen all the time, which makes debugging tricky.
For example, I set up a prompt to enforce a strict JSON output format for a weather microservice. The system instruction was clear:
```json
{
"system_prompt": "You are a weather data assistant. You MUST ONLY respond with a valid JSON object containing 'temperature', 'unit', and 'conditions' keys. Do not include any additional text, explanations, or markdown formatting."
}
```
Yet, about 20% of the time, I'd get a response like:
```
Sure! Here's the weather information you requested.
{
"temperature": 22,
"unit": "celsius",
"conditions": "clear"
}
```
This breaks the parsing logic downstream. My initial thought was a context window issue, but the conversations are short.
Has anyone else run into this? I'm trying to figure out if it's:
* A token limit problem with the system prompt being pushed out?
* An inherent non-determinism in how the model weights system instructions versus user messages?
* Or maybe a need for a more forceful, repetitive instruction pattern?
I'm used to working with systems where the "contract" is strict, like a REST API or a message queue schema, so this unpredictability is a bit concerning for production use. Any insights or similar experiences would be helpful.
--builder
Latency is the enemy, but consistency is the goal.
I've had similar issues when working on a small SaaS project that needed clean API output. For me, adding a very specific instruction in the user prompt itself, repeating the format, helped reduce the failures. Something like "Remember: output only the JSON, no other text." in the actual query.
Could temperature or top_p settings be a factor? I've found that lowering them makes the model more obedient to the prompt structure.
Do you get these failures more often on the first try in a fresh session, or after several back-and-forth exchanges?
Your point about adding the instruction to the user prompt itself is a solid workaround, and I've seen it help too. It's like reinforcing the signal.
On your question about whether failures happen more on the first try or later, in my experience it's often the first reply in a session. Once the model starts a pattern in a conversation, it tends to stick with it. So if your first response is clean JSON, subsequent ones usually are too. That's why a flawed first reply is so frustrating.
Lowering temperature definitely increases compliance, but it can also make the output feel less natural for other use cases. It's a trade-off.
Stay grounded, stay skeptical.
Your observation about temperature and top_p is correct. Lowering them reduces stochastic behavior, which directly impacts adherence to formatting instructions. I typically set temperature below 0.3 for strict output tasks.
Adding the instruction to the user prompt is reinforcing the signal, as user814 said. I'd add a caveat: if you're using a chat completion endpoint with multiple messages, the placement of that reinforcing instruction matters. I've had better results putting it in the *last* user message before the expected model response, rather than the first message in the chain.
And on your last question - failures are definitely more common on the first turn in a session, before a pattern is established. It's one reason I often use a cheap, separate "priming" call to set the pattern before the actual production call.
—Anita
I've hit this exact issue when trying to get clean JSON for IaC template generation from an API. While temperature settings and prompt reinforcement are valid points from others, the underlying cause often comes down to how the model internally prioritizes instruction tokens against its own tendency to generate conversational completions.
You can experiment with the `system_prompt` weight parameter if the API offers it. Some providers allow you to increase the weight given to system instructions relative to the user message. If that's not available, a practical workaround is to use a two-stage approach: first, a cheap, fast model call to generate a structured response, then a second, stricter validation or reformatting pass with a different prompt. This adds latency but can drastically improve reliability for production microservices where parsing failures are costly.
Your suspicion about an "inherent non-determinism" in weighting is correct. The model is trained to be helpful, and adding a conversational preamble often feels more helpful than raw JSON, despite explicit instructions. This is a known tension in instruction tuning.