I've been neck-deep in a data migration for a new ERP, and part of the security review involved checking our prompts for PII scrubbing. Like many here, I used an assistant to help draft and validate the prompts. The results were **alarmingly confident and completely wrong** in a very specific way.
The task was to create a system prompt that would ensure the model never outputs raw credit card numbers from processed logs. The assistant suggested a prompt that was very detailed and sounded robust. It included lines like:
* "You must never, under any circumstances, output a full credit card number."
* "If you encounter a credit card number in the input, you must redact it, showing only the last four digits."
* "A credit card number is a 16-digit sequence conforming to the Luhn algorithm."
The assistant even provided example inputs and the expected redacted outputs, which it got perfectly. The confidence was sky-high. However, when we did independent testing with a more varied dataset, the failure was systematic.
**The Failure:** The prompt relied heavily on the model *recognizing* a credit card number as a concept. In practice, when given a log line with a string like `"CardNumber: 4111111111111111"`, it worked. But when the input was obfuscated or formatted differently, it failed spectacularly. Examples from our tests:
* Input: `"Customer provided card 4111-1111-1111-1111 for payment."`
* **Assistant Output:** `"Customer provided card 4111-1111-1111-1111 for payment."` (No redaction)
* **Correct Handling:** Should redact to show only last four: `"Customer provided card XXXX-XXXX-XXXX-1111 for payment."`
* Input: `"The number is four one one one one one one one one one one one one one one one."`
* **Assistant Output:** No redaction at all.
* **Correct Handling:** This is extremely tricky and requires a different strategy entirely—the prompt logic failed.
The core issue was the assistant's overconfidence in the model's pattern-matching abilities based on the simple examples we co-developed. It created a prompt that passed a basic sanity check but lacked the rigorous, regex-based pattern definitions needed for real-world security. We had to scrap it and build a validation suite with edge cases first, then craft the prompt around that.
This has made me incredibly cautious about using assistants for any security-adjacent prompt engineering. The overconfidence masks a lack of systemic, adversarial testing.
- h
Data is sacred.