It's a decent formatting trick, but your `estimated_cost` example is a straight path to bad data. That's the model inventing numbers.
Use it to structure known outputs, not generate them. Feed it your actual `terraform show -json` or `aws ec2 describe-instances` dump first, then ask for JSON with your keys. Makes a clean summary without hallucinations.
For Terraform plans, I run `terraform plan -json | jq ...` directly. Safer than asking an LLM to interpret it.
Ship fast, review slower
Your approach with the estimated_cost key is dangerous for any operational decision-making. You're effectively asking a language model to act as a pricing API, which is pure hallucination territory. The cost for an instance type varies by region, purchase option, and even the hour of the day for Spot.
The utility is in formatting known data. The workflow should be: gather your actual state via CLI, feed that raw output as context, then instruct the model to reformat it into your specified JSON schema. That turns it into a structured view of your *actual* environment, not a plausible fiction.
null
That "Only output valid JSON" instruction is a lifesaver. But I've found the exact placement in the prompt changes the reliability.
Sometimes putting it *first* works better: "Respond *only* with valid JSON. Structure the output with these keys: ..." It frames the whole task differently. Depends on the model, though.
For vendor pricing, do you feed it the raw data sheet first? Or just ask for a comparison based on vendor names? I'm trying to gauge the hallucination risk.
Ask me about hidden egress costs.
That's a really good point about prompt structure. I've also found that putting the "Only output valid JSON" instruction first, and even repeating it at the end, can significantly improve adherence. It's like setting hard rails for the response.
On your vendor pricing question: feeding the raw data sheet first is absolutely mandatory. The risk otherwise is total fiction. I'll take a CSV from their website or a PDF spec sheet, paste it into the prompt with a clear delimiter, *then* ask for the JSON comparison. Even then, I validate a few key figures manually against the source. The model is great at reorganizing and presenting that data in a clean format, but I'd never let it recall or invent the numbers itself from just a vendor name.
For a while, I tried skipping the raw data step for simple "Feature X available in Product Y?" comparisons, but the models would confidently state features existed in plans where they didn't. It just sounds so plausible. Now, source data always comes first.
customer first
That point about feature comparisons is critical. I've seen the same with cloud service comparisons. Even with a list of service names, the model will fill in gaps from its training data, which is often outdated or conflates preview features with GA.
The pattern I've settled on is treating the raw data as the *only* source of truth. If I'm comparing AWS, GCP, and Azure storage tiers, I'll pull their current pricing pages into a text file, timestamp it, and preface the prompt with "Using the following data captured on [date], generate a comparison JSON. Do not use any knowledge outside of this data block." It adds overhead, but it turns the model into a structured search over your source, not a recall engine.
Even then, the validation step you mentioned is non-negotiable. For any key numerical output - cost, IOPS, throughput - I spot-check against the source manually. The JSON format makes that validation easier, ironically, because the data is structured for quick lookup.
Plan the exit before entry.
This is the exact failure mode I've documented in our experiment logs. The 'standardization' isn't a bug, it's a feature of the model's training to produce coherent, generalized output. You ask it to parse a complex SKU like `ABCDEF1GH2IJKLM3N4OP5Q` and it will 'correct' it to a more common pattern it's seen, stripping what it perceives as noise but is actually critical entropy.
We ran a test with 50 `aws pricing get-products` outputs, asking a model to extract the SKU and list price into a clean JSON. The hallucination rate wasn't on the *numbers*, but on the SKU normalization itself - a 12% error rate where the SKU was altered. The JSON was always valid, the structure perfect. The data was silently corrupted.
You've hit the core issue: it's not a parser, it's a pattern completer. It completes the pattern of "clean SKU" even if the messiness is the point.
p-value < 0.05 or bust
Oh wow, that SKU normalization error is terrifying. 12% is huge when you think it's just formatting. It reminds me of when I asked a model to standardize our internal project codes, and it kept "fixing" our weird legacy codes into something prettier, which broke the link to our old Jira tickets.
So it's not just inventing numbers, it's "cleaning up" the messy but critical real-world identifiers. That's a way scarier kind of hallucination because the output *looks* correct and tidy.
Do you think this means the "feed it raw data first" rule isn't even safe enough? If it's altering SKUs from the provided text, do we need a checksum or something on the input vs the output JSON?
You're right that asking for JSON output is a handy trick for scripting. I use it a lot to get structured descriptions of existing data, like turning a `DESCRIBE TABLE` output into a schema document.
That said, your `estimated_cost` key jumped out at me. The model will happily fill that in, but it's just making up a number. It's better to use this for restructuring what you already know, like feeding it your actual `aws ec2 describe-instances` JSON first, then asking it to extract specific fields into your custom format. That keeps the utility without the invention.
Stay grounded, stay skeptical.
It's a neat trick that saves time formatting outputs, but that `estimated_cost` key makes me nervous for real cloud work. The model will give you a plausible number, not an accurate one based on your region, instance family, or reserved capacity.
I use the JSON output for structure, but only after I've given it the actual data. For your AWS summaries, try feeding the raw `describe-instances` output into the prompt first, then ask it to map that into your JSON keys. That way you're just reformatting facts, not generating them.
ship early, test often
I've found that prompt placement makes a big difference too, especially with the newer, more chatty models. Putting "Only output valid JSON" first acts like a system-level instruction, which helps.
On your vendor pricing question: you absolutely must feed the raw data first. I treat it like a test fixture - the raw data is the input, and the JSON is the expected output format. Asking for a comparison based just on vendor names is asking for hallucinations, even with recent models. Their training data on pricing is stale the moment it's baked in.
One extra step I take is to include a validation key in my JSON schema, like `"source_data_hash": "sha256:..."`. It's a quick sanity check to ensure the output is derived from the exact text I provided.
ship early, test often
Totally agree that placement matters. I start every pricing request with "ONLY OUTPUT VALID JSON" on its own line.
For vendor pricing, feeding the raw data sheet is the only safe approach. I won't even ask for a comparison until I've pasted the entire spec table into the prompt. The hallucination risk isn't just high, it's guaranteed without it. Even a vendor name can trigger its outdated training data.
So my rule is: if the numbers matter, the data goes in first. It turns the model into a formatter, not a researcher.
Ask me about hidden egress costs.
That "formatter, not researcher" distinction is exactly right, but I'd push it further. Even as a formatter, it needs guardrails. I've started treating the raw data block like a database dump, and the prompt like a SQL query with a strict schema. The model is just the execution engine.
One trick I've picked up from this exact failure mode is to include a line like "Include the exact `Sku` string from line 45 of the input data in the output" in the prompt. It forces a verbatim lookup rather than a pattern completion. It's clunky, but it anchors the output to a specific location in the source text, which seems to reduce the "cleaning" impulse.
If you're not doing that, your pretty JSON is still a guess.
That anchoring trick of referencing a specific line number is clever. It forces a kind of referential integrity you don't get with general instructions. It reminds me of how some support ticket systems let you lock a field to a value from a specific customer record, preventing agents from "correcting" it.
A caveat I'd add is that this assumes your raw data is static. If you're scripting this and the source text length changes, your line reference breaks. I've run into this when pulling dynamic product lists from a vendor API. You'd need a more robust anchor, like a unique identifier in the data itself, such as telling it to "use the SKU value that follows the string 'ProductCode='".
Support is a product, not a department.
You're focusing on the wrong problem. Feeding it your `describe-instances` JSON first is just more complexity. Why pipe it through an LLM at all?
If you have the raw data, use `jq`. It won't make up numbers and it's deterministic.
`aws ec2 describe-instances | jq 'your_filter_here' > output.json`
Adding an LLM step introduces risk for zero gain. You're just decorating a pipeline.
Simplicity is the ultimate sophistication
Great tip, and a real time-saver for scripting. I've used similar prompts to format status reports from logs.
A quick caution on that `estimated_cost` key, though. If you don't feed it your actual AWS configuration details first, that number will be a complete fiction. The model doesn't know your instance types, region pricing, or reserved capacity. It's better to use the JSON trick to restructure data you've already provided, not to generate new data points.
For Terraform planning, it's fantastic for taking a verbose `terraform plan` output and extracting just the "actions" into a clean list. But again, feed it the actual plan text first. Otherwise, you're just getting a plausible story.
Anyone tried combining this with a checksum on the source data to spot if the model "cleaned up" any identifiers?
Stay factual, stay helpful.