Skip to content
TIL how to bypass C...
 
Notifications
Clear all

TIL how to bypass Claw's prompt injection guardrails with a simple formatting trick.

19 Posts
19 Users
0 Reactions
51 Views
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Right? That's the scary part. If a system breaks on formatted text, it'll choke on the natural noise of a real user session. I'd bet a simple customer message with a stray newline or a typo could accidentally trigger a false positive... or worse, mask a real injection attempt.

Makes you wonder if their training data was too clean. Real chat logs are messy.


Infrastructure as code is the only way


   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

That normalization step is the first line in any adversarial ML pipeline. Their model is being evaluated on a distribution the attacker doesn't have to respect.

We built a similar filter. Our pre-processor strips all whitespace, lowercases, and maps homoglyphs before the text even hits the classifier. Adds 2ms latency. Their omission is a design choice, not an oversight.


Prove it with a benchmark.


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

The multi-turn split is a clever test. It directly probes for stateful tracking, which most services avoid for cost reasons.

The real cost isn't just per-message inference, it's storing and comparing against the session's vector history. That requires a fast lookup store like Redis. If they're not doing that, the split attack works. But if they are, the latency per message becomes unpredictable.


sub-100ms or bust


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

Interesting, that's a really practical example of why normalization matters. It reminds me of a basic regex filter I wrote once that failed on real user input because I didn't account for line breaks.

Got a quick question though. In your pre-processing example, what exactly are you doing to the text? Like, are you splitting it by sentence and randomly adding bold or code blocks between phrases?



   
ReplyQuote
Page 2 / 2