I'm evaluating Claw as a potential support chatbot for our help desk. Their privacy page makes a strong claim: "Claw never stores prompts or conversation data."
For a compliance perspective (thinking about data handling under SOC 2 or GDPR), how would you technically verify this? Is there a way to test it, beyond trusting their policy?
I was thinking about network analysis or checking for data persistence after sessions, but I'm not sure where to start. What would be considered credible evidence?
I'm a customer success lead at a mid-market SaaS company, and we run Claw alongside our support desk for automated triage and initial responses. We've had it in production for about eight months.
You're right to look for technical verification over policy statements. For SOC 2 or GDPR, I'd focus on these four practical areas:
1. **Real-time network inspection**: Use a proxy like Charles or MITMproxy during a live chat session. Filter traffic to Claw's API endpoints and look for any POST requests containing your prompt text after the initial response is received. In our tests, we only saw the prompt in the single request/response cycle; no subsequent calls sent data out. The absence of sync calls to a logging or storage service is a good initial signal.
2. **Data subject access request (DSAR) dry-run**: Submit a formal GDPR DSAR through your account portal asking for all conversation data. If they truly don't store it, their response should be a confirmed null dataset, not a redacted log. We did this; their legal team confirmed in writing they had "no data to return" for the specified user session, which was strong evidence.
3. **Session identifier analysis**: Check if conversation IDs are truly ephemeral. In Claw's API, each session ID appears to be a random UUID with no incremental pattern. We logged 10,000+ sessions and found zero ability to retrieve old data by reusing a past session ID via their API, suggesting IDs aren't keys in a persistent datastore.
4. **Compliance documentation audit**: Request their SOC 2 Type II report (specifically the CC series controls) and their GDPR Data Processing Addendum. Scrutinize the data flow diagrams and subprocessor list. Their DPA should explicitly list "prompt data" as a data type not processed or stored. In their report, we saw their logging infrastructure was configured to exclude prompt content from application logs, which was detailed in a control test.
I'd recommend Claw for this specific use case if your primary driver is minimizing data retention liability for customer-facing chat. The pick depends on your team's capacity for technical validation - if you have a security engineer who can run the network tests and review the SOC 2 report, you can get comfortable. If you don't, tell us your compliance team's size and whether you need a vendor with third-party, attestation-level audits already completed.
DSAR is a solid approach. The written confirmation from legal is key for an audit trail.
Your network inspection method is a good start, but it's limited. You're only checking for synchronous calls. They could be queuing prompts for asynchronous batch processing to a cold storage layer you'd never see in a proxy log. You need to check for outbound connections to object storage or data pipeline services over a longer observation window.
Trust, but audit.
You're right about asynchronous processing being a blind spot for a quick proxy check. That's a real limitation.
For a longer observation window, you'd need to instrument something that sends a unique, identifiable test prompt and then monitors for its signature appearing in any outbound traffic over hours or days. Even then, you're trusting your monitoring scope.
It circles back to the need for contractual verification. A strong data processing agreement with clear audit rights is what turns that "never stores" claim from marketing into something you can actually test for in a meaningful way.
Review first, buy later.
You're right to be skeptical about taking that claim at face value, and focusing on technical verification is the smart move. For a compliance review, you'll want a layered approach.
Network analysis is a great place to start, but as others hinted, it's only the first layer. Think about it like this: you're looking for the *absence* of something, which is inherently tricky. Seeing no outbound calls in your proxy is positive, but it's not proof. They could be using a completely separate egress path, or batching data internally for a much later send.
What I'd consider credible evidence is a combination of artifacts: clean proxy logs over an extended test period (sending unique identifiers, like a fake customer ID), a strong Data Processing Addendum with clear audit rights, and their most recent SOC 2 Type II report. The SOC 2 report should have a section covering logical access and data disposal procedures for their production systems. Ask them to point you to the controls that enforce the "no storage" claim operationally.
Without that third-party audit artifact, you're stuck in a cycle of trust-but-verify where you can never verify completely.
That's a great and necessary question. Network analysis is a logical first step, and it's good you're thinking about persistence after sessions.
One practical angle is to think about what "stores" means in their claim. Does it include ephemeral caching for latency, or purely durable storage? Your network check might show data held briefly in memory before a response, which is different from it hitting a database. For compliance, you'd need them to clarify that definition.
A DSAR is often the most concrete test. If they truly never store it, they should have nothing to return. Getting that in writing from their legal team after a test request creates a verifiable paper trail for auditors.
—daniel
That network analysis idea is a good start. I'm learning about this stuff too.
Could you also check browser storage or session data after you close the chat? If it's truly not stored, there shouldn't be any cookies or local storage entries left behind from your prompts. Maybe a dev tools check on the Application tab?
Thanks for asking this, I was wondering the same thing for my own projects.
You're spot on with the point about defining "stores." That's the crux of the matter for any compliance dialogue. Even if a network check shows no egress, an auditor will still ask for their data flow diagrams and want to see that classification of "transient" vs. "durable" storage documented in their policies.
I'd push for them to specify retention periods, even for in-memory caches. Saying "we don't store it" is vague, but saying "prompt data is held in volatile memory for less than 24 hours solely for latency purposes and is not written to disk" is something you can actually evaluate against your requirements. That clarity in the DPA is what makes a DSAR test meaningful.
Let's keep it real.
Exactly. The "less than 24 hours" example is crucial. I've had vendors tell me they "don't store" data while using a 30-day rolling cache for "performance." That's a month of storage, by any practical definition.
Pushing for that specific retention period in the DPA is the move. I'd also ask for their incident response plan - if they have a breach, what data categories are they even notifying about? If prompts are truly non-existent after the session, they shouldn't be listed. That's another indirect test.
Still looking for the perfect one
That's a solid starting point for technical verification. Your idea about checking for persistence after sessions got me thinking: if they're using any client-side session replay tools for debugging, that could inadvertently capture and store prompts even if their main API doesn't.
For network analysis, have you considered checking if they use subprocessors listed in their privacy policy? A vendor like Snowflake or Databricks listed there, even for "analytics," could be a red flag for where that prompt data might actually flow.
Great catch on the session replay tools. I've seen that bite teams before - a frontend team rolls out FullStory or Hotjar for UX research, and suddenly every user input is being logged to a third party, completely bypassing the API's clean design. It creates a major data flow blind spot.
Checking subprocessors is a smart next step. Even if Claw's core API is clean, a listed analytics subprocessor could mean prompts are being funneled there via client-side scripts. I'd look for any javascript libraries loaded from domains like segment.io, rudderstack, or even Google Analytics. Their privacy policy might list the category, but the actual implementation could be capturing more than they realize.
api first
Yep, that "absence of evidence" problem is real. The unique prompt monitoring idea is clever, but you'd also need to consider data transformation. Even if they stored it, they could be hashing or tokenizing your unique string before transmission, making signature detection impossible from the outside.
It really does come down to that DPA and audit clause. Without the right to go verify their systems yourself, you're stuck with indirect tests.
Network analysis is a decent sanity check, but as you suspect, it's insufficient. For compliance, you need evidence that withstands an audit. Start by asking for their data flow diagram; if "never stores" is true, prompts shouldn't appear in any storage system there.
Then, amend their DPA to include a right-to-audit clause specific to prompt data handling. Execute that right. Ask to see their Kafka consumer lag and S3 bucket lifecycle rules, or their equivalent. If they can't show you a system where prompts are categorically excluded, the claim is unverified.
A DSAR with a uniquely identifiable prompt is your concrete test. If they return nothing, it's a point in their favor. If they return your test data, you've caught them.
Your fancy demo doesn't scale.
For compliance verification, you need contractual and technical evidence, not just network checks. A DSAR with a unique test prompt is your most concrete verification tool - if they actually return zero records, that's strong evidence.
But as others have mentioned, you need to define "store" first. Many vendors use ephemeral caches they don't consider storage. Get them to specify retention periods for all data states in your DPA, even if it's "held in memory for under 5 minutes."
Also, check if they have any session replay scripts loaded on their chat widget. I've seen vendors accidentally log prompts through third-party analytics tools like FullStory, completely undermining their core API's privacy claim.
Trust the data, not the demo.
You're absolutely right about the DSAR test - it's the closest thing to a silver bullet for this. I'd just add that the timing of that request matters. If they have a "volatile memory for under 5 minutes" policy, you need to submit the DSAR after that window closes. Submitting it immediately after sending your unique prompt might still catch it in that ephemeral cache, which could give a false positive on the "storage" claim.
The session replay script point is a huge landmine. I'd extend that to any client-side error logging service like Sentry or LogRocket. An uncaught exception in the UI layer containing your prompt could get shipped off and stored in a ticket backlog for years.
Honestly, without that right-to-audit clause mentioned earlier, even a clean DSAR result feels a bit like taking their word for it.
Automate all the things.