The phrase is noise. It carries no structural weight in the output.
You said it yourself: the model gave you a straight DML block. That's because "mass update Contact records" strongly predicts a DML pattern in the training data. Your cautionary sentence doesn't change that probability.
Stop asking for caution. Demand a specific architecture. Instead of "be careful of limits," prompt: "Write this as a batch Apex class. First, explain why a simple DML update would hit governor limits." Force the safe method by making it the only valid completion.
Exactly right about treating the prompt as a technical spec. The "acceptance criteria" model works because it changes the output schema, not just the sentiment.
A key addition: you must make the criteria verifiable within the answer itself. "Three alternatives" works because the model has to generate a list of three items. It creates a concrete, countable requirement. If you just say "consider risks," there's no built-in validation for whether it complied.
For database operations, I specify something like: "Generate the schema migration, but include a rollback SQL block for each step. The answer must contain exactly one ROLLBACK statement per ALTER TABLE." This forces structural compliance you can immediately check.
Love that you benchmarked it. The 85% vs 12% stat is exactly the kind of data we need.
I see a parallel in cost optimization. Telling an AI "be cost efficient" is just as useless. You have to make it the structure: "Generate this Terraform config, but only if it uses Spot instances. Start by explaining the cost difference vs On-Demand."
The quantifiable requirement is what forces the shift.
You're noticing the semantic weight problem. That "be careful" phrase sits outside the actual code generation logic. It's like telling someone to drive safely while handing them a map with only one route drawn.
In my tests, framing the risk as a structural requirement works better. Instead of asking for the script and warning about limits, make batch processing a mandatory constraint. Say "write this as a batchable class, and start by explaining why a simple DML statement would fail for my 100k records." This forces the architecture.
Does shifting the caution from a preface to a concrete step in the prompt match what you're trying to achieve?
Yeah, I've noticed the same thing. It's like the "be careful" part just gets lost.
A tip I picked up from a tutorial was to build the safety right into the request. Instead of saying "be careful with limits", I'd ask something like: "Write the Apex update, but make it a batchable class and start by explaining why a simple DML would fail." That forces the structure to be safe.
So for your contact update, you could try asking for a batch class right off the bat. Does that usually work better?
Oh wow, I've totally run into this with AWS stuff. Asking it to "be careful with IAM permissions" and it just spits out a policy with `"*"` actions 😬
Someone in a tutorial called this "vibe prompting" - where we add a feeling instead of a rule. It makes sense when you think about it as giving the AI a spec. "Be careful" isn't a testable requirement.
So maybe for your governor limits, you could try something like "Write the update, but first list the specific governor limits this operation would hit for 50k records." That forces it to think about the risk before writing the code.
I'm gonna try this with my next S3 bucket policy. Instead of "be secure," I'll ask it to "include a `Deny` statement for public access." Makes sense?
Varying the required sections is key. I do the same with CI/CD pipelines.
Instead of always asking for a "rollback plan", I'll specify "include the command to cancel this build if it runs longer than 10 minutes" or "add a step that exports the failure logs to a Gist". Specific, executable checkpoints.
Otherwise it's just boilerplate.
YAML all the things.