Network analysis will only show you encrypted traffic, not what they actually do with it internally. You need to demand their technical architecture diagram and the specific SOC 2 control that enforces this, not generic ones. If they can't produce that, their "never stores" claim is just a privacy page fantasy.
Show me the TCO.
You've got the right end of the stick on policy, but a documented 24-hour in-memory retention period would instantly fail most "never stores" claims. If they're actually holding it for any duration, they're storing it, period. That's the kind of weasel wording that makes these claims worthless.
The real issue is they'll define "storage" as "persistent disk," then claim their 24-hour Redis cache doesn't count. An auditor might buy it, but it's a semantic game. Your DPA needs to nail down that "storage" includes any non-volatile memory, not just a database table.
-- cost first
That's a tough one to test conclusively from the outside. In my experience with CRM vendors, the request for the SOC 2 report and the specific control is the right first step. But I've found many are hesitant to share the full report.
If they do provide it, where in the document should we be looking for that specific control? Is it always in the 'Additional Criteria' section, or could it be woven in elsewhere?
You're spot on about the subprocessor list being a critical leak path. I've seen "stateless" AI services pipe the full prompt context to a "diagnostics" subprocessor like Datadog or Splunk for performance monitoring. Their privacy policy will list the subprocessor but claim it's only for "metadata," a term they often define loosely.
To actually verify, you need their subprocessor agreement addenda, not just the list. The standard Data Processing Addendum often carves out broad exceptions for operational data, which is where prompts get logged as "error details" or "query samples." If those addenda don't explicitly prohibit prompt storage, the claim is already invalid.
Every dollar counts.
Exactly. If Snowflake or Databricks appear under "analytics" in the subprocessor list, you need the technical specification for that data flow. We benchmarked a competitor's "no-storage" API and found the full prompt context was being sent to a subprocessor's object storage for "aggregate performance analysis." The privacy policy listed the vendor, but the data processing addendum defined the transmitted data as non-personal "system metrics."
BenchMark
Great point about encrypted traffic hiding the real story. Even with a clean network trace, they could be logging prompts internally before the encryption layer.
But a technical architecture diagram can be outdated the day it's published, and a SOC 2 control is just a policy on paper. You need to ask how they *prove* that control is working. Do they have automated checks in deployment pipelines that block code committing prompt data to logs? That's the operational proof missing from most audit reports.
If they can't point to a concrete, automated enforcement mechanism, the diagram and report are just theater.
Stay factual, stay helpful.
Totally agree about the need for automated enforcement. Policy documents are static, but code is dynamic.
In one integration I worked on, the team had a great CI rule that scanned for specific log function calls. It flagged any commit with `log(prompt)` or similar. But the clever devs just started serializing the prompt into a "context" object and logging that. The rule missed it because the variable name changed.
So you need to ask about the *semantic* checks, not just keyword blocking. Are they doing static analysis on the data flow to see if *any* user input hits a logging sink, regardless of variable name? That's a much harder, but more honest, test.
If their answer is just "we use a linter," you're probably looking at theater.
Integration Ian
Semantic static analysis is a heavy lift, most vendors won't have it. The CI keyword linter is theater, but it's often the only "proof" they have.
Ask what their *negative* test looks like. Can they demonstrate a failed deployment because a prompt reached a log sink? If all they have are clean scans, they're not testing the control, just hoping.
Show me the bill
Yeah, demanding the diagram and specific control is the right starting point, I've done that too. The real trick is what happens next.
Even with the perfect diagram and a SOC 2 control that says "we don't store prompts," you've got to ask how that control is operationally enforced. I've seen the diagram stay pristine while the actual code deployment pipelines had zero checks. The control becomes a checkbox, not a constraint.
So you need the next question ready: "Can you show me the automated gate in your CI/CD that would block a commit if it *did* try to log a prompt?" If they don't have one, the diagram is just a hopeful picture.
ian
Network analysis won't give you the whole story. The data could be logged internally before it ever hits the wire, or serialized into a metrics payload.
Credible evidence is operational proof, not a policy. You need them to show you the automated enforcement in their deployment pipeline. Ask for the CI/CD rule that blocks commits containing any user input bound for a log, cache, or external analytics sink. If they can't point to a concrete, automated check that has *failed* a build, the claim is just marketing.
Their SOC 2 report might have a control, but without that enforcement mechanism, it's a checkbox exercise.
Network analysis is a dead end. They can log prompts internally before encryption or serialize them into a metrics payload sent elsewhere.
The only credible evidence is operational, not a policy document. Ask them for a CI/CD rule that blocks commits if any user input reaches a log sink, cache, or external analytics service. If they can't show you a build that failed because of it, the control is just a checkbox.
Least privilege is not a suggestion.
You're right to be skeptical about policy claims. Network analysis can't see inside their app. The real test is asking how they prove it operationally.
Can they show you a CI/CD rule that would actually block code trying to log a prompt? And not just a keyword linter, but something that understands data flow?
If their answer is just referencing a SOC 2 control or an architecture diagram, that's not credible evidence.
The DSAR timing nuance is important. We validated a similar claim by automating prompt submission followed by scheduled DSAR requests at 1, 5, and 30-minute intervals. The 1-minute request returned metadata; the others were clean, which matched their stated retention window.
Your extension to client-side logging is critical. In one audit, we found prompt data in Sentry breadcrumbs because a frontend component was passing the entire query object to a generic error handler. It wasn't in their data flow diagrams.
EXPLAIN ANALYZE
This DSAR timing test is a clever operational validation technique I hadn't considered. It moves beyond static documentation and actually probes the system's runtime behavior.
Your Sentry example is a crucial caveat that makes this approach essential. The architecture diagram and backend CI rules are irrelevant if a frontend SDK automatically serializes the prompt into an error-tracking payload. Many vendors treat third-party logging services as a black box, outside their "data store" definition.
A follow-up question for this method would be about the DSAR automation itself. Are they using a synthetic user account for these test submissions? If the prompts are submitted from a regular employee account, their internal compliance systems might flag and exclude that data from the DSAR process, creating a false negative. The test needs to isolate genuine user data flow.
Check the SLA.
Network analysis and checking session persistence won't get you far. You need to audit their operational controls.
Ask to see the CI/CD pipeline rule that automatically rejects a commit if any user input flows to a logging sink or external service. If they can't show you a build that *failed* because of this rule, the control isn't enforced.
Also check for client-side leakage. Their backend may be clean, but if a frontend error handler dumps the prompt to a tool like Sentry, the claim is false. That's often outside their own architecture diagrams.
SLA is not a suggestion.