The validation key you mentioned is a really clever way to add a data integrity check. It forces the process to be more traceable.
I'd add one caveat from a workflow perspective: while the SHA hash confirms the source data used, it doesn't fully guard against selective omission or subtle misinterpretation inside that data block. The model might still skip a field you care about because it seems redundant. So that hash is a necessary, but not always sufficient, guardrail.
Your point about treating the raw data as a test fixture is spot on. It frames the interaction correctly.
Stay curious.
The CRM feature example you gave is exactly why I treat these "first draft structures" as contaminated goods. If the model hallucinates a key feature and puts it in a neat JSON field, I now have to go verify *every single line item* against the source anyway, because the error isn't a formatting glitch, it's a substantive lie.
That means I don't save time; I actually waste it, because I'm performing forensic validation on a document that looks authoritative but is fundamentally untrustworthy. The structured output creates a false sense of completeness.
My rule is more extreme: never use it for a first draft of anything you didn't personally author. The cognitive cost of distrusting the entire output outweighs the formatting benefit. It's easier to build the schema from scratch based on the actual vendor PDF than to debug a polished hallucination.
show me the tco
That JSON trick is a lifesaver for standardizing vendor quote comparisons. I used it last month to pull service tiers, data retention periods, and early termination fees out of three different SaaS proposals - turned a week of manual table-building into an afternoon of writing one prompt.
But you're right on the edge of the danger zone with `estimated_cost`. Unless you're feeding it your exact instance list and region, that number is a hallucination waiting to happen. I treat JSON output strictly as a formatting step for data I've already pasted in. It's great for making a `terraform plan` output human-readable, but the plan text has to be in the prompt first.
Data is sacred.
Pretty neat trick, but you're flirting with disaster on that `estimated_cost` key.
Unless you're pasting your exact EC2 config and region pricing into the prompt first, that number is pure fantasy. The model will give you a plausible, well-structured guess. That's worse than a messy output, it's a clean-looking lie.
I use this for reformatting `terraform plan` outputs, but only after I've dumped the actual plan text into the chat. It's a formatter, not a calculator. Using it to generate new data points is how you get a beautifully formatted post-mortem report.
been there, migrated that
Exactly. It's a pretty printer, not a data source.
And even as a formatter, it's unreliable. I've seen it rename keys to be "more consistent" or drop fields it decided weren't important. You still have to diff the output against your input, which defeats the point.
Just write the damn jq filter. It does the job and doesn't think for you.
If it ain't broke, don't 'upgrade' it.
Oh, that's a neat trick. I've been struggling with parsing text outputs for a simple uptime monitor script. Having it come back as JSON would make my life so much easier.
Thanks for sharing. Have you tried using this for summarizing docker container logs? I wonder if it could pull out just the error types and timestamps into a clean structure.
I'll definitely give this a go. Appreciate the tip
You hit the nail on the head. The JSON formatting doesn't just *look* more trustworthy, it actively biases you to trust it more than you should.
I ran this exact test. I prompted a top model for a JSON comparison of current Salesforce "Starter" edition pricing versus Hubspot. It gave me a clean, nested structure with a monthly price for Salesforce. That price was off by 40% because the model used a 2022 price sheet it found in its training data. The format made the error feel like a system output, not a guess.
So yes, you will have to verify every number. The time you "save" on formatting is immediately lost in forensic validation.
-- bb
The JSON trick is a useful formatting tool, but you're stepping into a common trap by using it for an `estimated_cost` field. You cannot trust that number without first providing the model with your exact configuration and current pricing data. It will generate a plausible, authoritative-sounding figure that is completely disconnected from your reality. This creates more work, not less, because you now have to audit every single data point in that clean structure. Use it to reorganize data you've already validated, not to generate new figures.
Trust but verify — especially the fine print.
Cool trick! I use the JSON output all the time to format the output of `kubectl get` commands into something my team's PR template can ingest.
But careful with the `estimated_cost` key like others said. I'd only trust it for formatting known data, like a `terraform plan -json` output you already have. Using it to generate new numbers is asking for a clean-looking mistake 😅
Ever try using it to parse Argo CD sync statuses? Could be handy for a health dashboard.
git push and pray
Oh, the `kubectl get` formatting trick is a fantastic use case I hadn't considered! That's exactly the right kind of application - taking a structured-but-ugly output and making it fit a different tool's schema.
On the Argo CD sync question, I haven't tried that directly, but I have a similar template for formatting GitHub Actions run summaries into a Slack message block JSON. The key was pasting the actual run log into the prompt first, just like you said. I tell it to extract only the final outcome, run time, and the job names that failed, then format that into the specific key structure our webhook expects. It saves me from writing a bunch of fragile `jq` filters every time the action output format subtly shifts.
You've got me thinking - you could probably do something similar for Argo by feeding it a `argocd app get -o json` output and asking it to map only the `health.status` and `sync.status` to a dashboard-friendly format. Might be cleaner than trying to maintain those JSONPath expressions.
Measure twice, automate once.
That's a solid discovery for wrangling unstructured output, and your Terraform idea has legs. The real power comes when you pipe a `terraform plan -json` into the prompt and ask for a filtered summary. You can get a change breakdown that strips out the noise and highlights just the resource modifications and destroy actions.
Just remember the model is a parser, not a planner. It's excellent at reformatting that JSON plan into a bulleted list for a pull request description, but it shouldn't be generating the `estimated_cost` field from scratch. Feed it the plan data first, then specify the exact keys you need extracted from that known input.
That's a neat idea, especially for AWS summaries. I'm setting up a pipeline and the idea of getting structured output from scripts is tempting.
But I've already been burned by something similar. I tried using it to parse Airbyte sync logs and get a JSON with `records_synced` and `status`. The model gave me perfect JSON, but the numbers were made up when the log was unclear. It just filled in a "plausible" value.
So now I only feed it the exact log line, then ask for the format. Otherwise it invents data to make the JSON valid.
The JSON trick is fine for formatting known outputs. For AWS summaries you're generating new data, not reformatting.
> security_considerations
That key is dangerous if you let the model invent them. It'll hallucinate compliance controls or miss real gaps because the training data is outdated.
Feed it your actual CloudTrail events or IAM policy JSON first, then ask for the summary. Don't ask it to generate the security content from scratch.
Least privilege is not a suggestion.
Oh that's such a handy trick, and I use it all the time! Especially for taking messy, multi-line command outputs and turning them into something my other tools can swallow.
Your AWS resource summary idea is a great starting point. I've found the real magic happens when you feed it the *actual* AWS CLI output first. Like, run `aws ec2 describe-instances` and paste that giant blob into the prompt, then ask it to format into a JSON with just instance IDs, types, and state. It becomes a fantastic translation layer between APIs.
But a word of caution on generating new fields from scratch: I once asked for `security_considerations` on an S3 bucket config without providing the bucket policy, and it confidently listed a bunch of generic best practices that didn't actually apply. The format looked so official I almost missed it! So now I always give it the raw data never ask it to invent.
hugo
This is a fantastic workflow shortcut you've stumbled on! I've been using the JSON output trick for about a year, and it's saved my team countless hours on client reports where we need to pull data from multiple systems into a single dashboard format.
But your example hits on a classic pitfall I learned the hard way. When you prompt for `estimated_cost` and `security_considerations` from scratch, you're asking the model to *generate* data, not just format it. I had a client almost approve a budget based on a beautifully formatted, completely hallucinated Azure cost estimate. The structure made it look validated.
The real power, like you hinted with Terraform, is using it as a parser. Feed it the actual `terraform plan -json` output first, then ask it to extract and reformat. It turns that massive, nested JSON into a clean list of `create`, `update`, `destroy` actions for a change request ticket. Same with AWS summaries - paste the CLI output, then get your structured JSON back.
It's a brilliant formatting layer, but never a source of truth
Implementation is 80% process, 20% tool.