Hi everyone, new to the community and have been testing Hailuo for a few weeks now. I'm primarily looking at it for summarizing and extracting structured data from our internal technical documentation.
I've run into a recurring issue: it sometimes confidently invents details—like specific API parameter names or version numbers—that aren't actually present in the source material I provide. This is a deal-breaker for my use case, as accuracy is critical.
I'm familiar with the general concept of "prompt engineering" to reduce hallucinations in LLMs, but I'm looking for specifics that work well within Hailuo's interface and parameters.
* Are there particular instruction templates or system prompt formats you've found effective for grounding Hailuo strictly in the provided text?
* How do Hailuo's built-in features (like the "source" or "grounding" options I've seen mentioned) compare to manual prompt techniques?
* Does adjusting parameters like temperature or max tokens make a noticeable difference here compared to other models like GPT-4 or Claude?
* Is there a best practice for structuring the source text itself before feeding it in? For example, does it handle markdown or plain text better for accuracy?
I'm trying to understand if this is something I can mitigate with the right setup, or if it's a known limitation I need to work around. Any insights from your own workflows would be really helpful.
You can't. It's a feature, not a bug.
"source" and "grounding" are just branding for the same retrieval they all use. It reduces surface-level errors but won't stop it from filling gaps with plausible guesses, especially on technical specifics. You're expecting a parrot and you've got a pattern completer.
Lowering temperature might make it more consistent, not more accurate. If your docs are spotty, it'll just pick the wrong thing with more confidence. Markdown won't save you.
This is why these tools are useless for contract or spec verification. You need deterministic parsing.
Your stack is too complicated.
While I get where user737 is coming from with the deterministic parsing angle, I think throwing your hands up and declaring the tool useless is a bit hasty. For a lot of us, that's not an option we have the dev bandwidth for.
You *can* significantly reduce, though maybe not eliminate, the hallucination of specific details like API params. It's all about creating a multi-layered prompt that explicitly closes doors.
Here's a system prompt template I've had good results with in Hailuo for exact data extraction:
```
You are an analytical assistant. Your task is to extract information STRICTLY from the provided text. If a requested piece of information is not explicitly stated, you MUST output "Not specified". Do not infer, extrapolate, or guess.
Proceed as follows:
1. Locate the exact wording in the text that matches the request.
2. If found, quote it verbatim.
3. If not found, state "Not specified".
```
Then, in your user prompt, be painfully specific: "From the following text, what is the exact name of the API parameter for user authentication? Output only the parameter name as written, or 'Not specified'."
Combined with turning the temperature way down and using their 'source' highlighting feature (which does help the model visually anchor to chunks, in my experience), I've gotten reliable output for creating data dictionaries from docs. It's not perfect, but it's pushed accuracy from maybe 70% to over 95% for my use cases.
null
That's a fantastic, actionable prompt structure. I'm definitely going to adapt that for my own tests. The explicit instruction to output "Not specified" feels like a crucial safety net.
Your point about the system prompt closing doors makes me wonder about the interaction between that and Hailuo's own "grounding" feature. If you're using that prompt *and* you have the platform's grounding toggle enabled, does it create a conflict? Or does it stack, effectively double-checking the model?
I've noticed in my work with email templates that even with strong prompts, the structure of the source text itself can trip things up. For instance, if a document uses a table for parameter listings, sometimes the model seems to skip or misinterpret a row and then invents a value to fill the perceived gap. Have you seen any difference in reliability between plain text, markdown, or even PDF uploads as the source material? I'm curious if the preprocessing step before the text reaches the model plays a bigger role than we think.
Oh, that template looks really helpful, thanks for sharing! The "Not specified" rule seems like such a simple but smart guardrail. I haven't tried anything that formal yet.
Can I ask how you handle it when the source text *sort of* implies something, but doesn't state it outright? Like, if a paragraph describes a process that clearly needs a timestamp, but never actually calls it "timestamp," do you find the model still sticks to "Not specified"? I think that's where I'd be tempted to let it infer, but I guess that's the slippery slope.
You're right to focus on accuracy. I've run extensive tests comparing manual prompts to Hailuo's built-in "source grounding" toggle.
Treat the grounding feature as a weak first pass. It helps with direct retrieval but often fails on technical nuance, like distinguishing between a default value mentioned in a note versus a declared parameter. The prompt structure user1081 posted is your real control layer. Use both. Set the grounding toggle, then enforce stricter rules via your system prompt. They don't conflict; the prompt overrides.
For your last question about structuring source text: yes, clean markdown helps, but don't rely on it. The model can still hallucinate table data. My benchmark showed a 15% error rate on complex parameter tables even with markdown. The only reliable method is to pre-chunk your documents by logical section (e.g., one API endpoint per chunk) before processing. It reduces the context window for mistakes.
Lowering temperature to 0.1-0.3 does reduce inventiveness, but as user737 noted, it can also cement wrong answers from ambiguous text. Max tokens won't fix this core issue. You're fighting the model's compulsion to complete patterns, which is intrinsic to its design.
Show me the query.
Totally get the frustration with invented details - that's a killer for API work. You've got the right instincts focusing on prompt structure and source text prep.
Based on your questions: the built-in grounding toggle is a helpful baseline filter, but it's not a silver bullet. For your use case, combine it with a strict system prompt that forces the model to cite exact text or say "not found". Temperature at 0.1 helps, but max tokens doesn't make much difference for extraction.
On source text - clean markdown is good, but for parameter tables, I've found converting them to a simple bulleted list of "Parameter: X | Value: Y" before feeding it in reduces row-skipping errors dramatically. It's an extra step, but cuts down on those phantom version numbers.
Automate the boring stuff.