Skip to content
Notifications
Clear all

Breaking: Security review of Speechify's data handling for confidential docs.

19 Posts
19 Users
0 Reactions
79 Views
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
Topic starter   [#23978]

I've been trialing Speechify for processing internal meeting transcripts and some client-facing documentation. The convenience is obvious, but I'm hitting a pause button before we approve it company-wide.

My main concern is around data retention and access control for uploaded documents. Their privacy policy mentions data is used to improve services, but it's vague on specifics for the enterprise tier. If I upload a confidential contract draft, who at Speechify can access that raw text? Is it ever reviewed manually? How long is the audio output stored on their servers after generation? I couldn't find clear answers in the admin panel.

Also, their SSO setup (which we'd need) seems to only be on their highest plan. That makes me wonder about their overall security maturity. Has anyone here done a deeper security assessment or requested a vendor security questionnaire from them? I'm particularly interested in their subprocessor list and where the actual text processing occurs.



   
Quote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

Yeah, that SSO setup being locked to the top tier is a red flag to me too. It makes the whole pricing feel like security is an upsell, not a standard feature.

Have you tried directly asking their sales team for the subprocessor list? I had to do that with a different vendor recently and it was... illuminating. They sent a PDF that basically listed every major cloud provider, which just created more questions about where our data actually touched down.



   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

Their privacy policy being "vague on specifics for the enterprise tier" is the point. That's where the real pricing starts.

You won't get clear answers on manual review or retention from the admin panel. Those are contract terms. Expect a separate data processing agreement (DPA) that costs extra and defines everything. SSO on the top plan proves it.

Ask for their security questionnaire now. The subprocessor list is usually massive, but the real cost is the internal hours you'll spend mapping it.


always ask for a multi-year discount


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

You're absolutely right that the real terms are hidden in a DPA, but the cost isn't always a direct line item. In my experience, the negotiation of those terms - specific retention windows, audit rights, breach notification timelines - becomes the leverage they use to avoid discounting the base enterprise license.

That subprocessor mapping exercise is crucial. A massive list often signals heavy use of serverless functions or third-party tooling for internal analytics, which creates more data egress points than a simpler, more contained architecture. Requesting their data flow diagram alongside the list can reveal if your document text transits multiple environments before a final audio file is generated.


Plan the exit before entry.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Data flow diagrams are often outdated marketing collateral. They'll show a neat box for "audio processing" that's actually a dozen different Lambda functions and third-party APIs.

The subprocessor list is what matters, but you need the *purpose* for each one. If they list AWS S3 for "temporary storage," ask for the IAM policy that restricts access to that bucket. That's more revealing than a diagram.

Negotiate audit rights that let you verify the bucket policies and IAM roles, not just review a static diagram.


Least privilege is not a suggestion.


   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Agreed on the subprocessor list's purpose being key. While a data flow diagram can hint at complexity, the actual technical safeguards are in the cloud permissions. For instance, a vendor listing "Amazon S3 for temporary storage" should also provide the bucket policy showing that access is restricted to specific IAM roles with short-lived credentials, not a broad internal service account.

Negotiating for audit rights to verify these policies is essential, but in my experience, many vendors will only agree to a third-party audit report (like SOC 2) rather than allowing direct access. The more critical term to push for is the right to receive evidence of these controls, such as a screenshot of the S3 bucket policy configuration at a point in time, as part of your annual review.


CPU cycles matter


   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Good point about pushing for evidence vs direct access. That screenshot can be powerful.

Just remember, a static screenshot from a year ago is only good if their change management is locked down. I'd try to add a clause about being notified of any material changes to those bucket policies during the term, not just during an annual review.

Otherwise, you're trusting their internal ticketing system as much as their S3 config.


Trust the trial period.


   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Exactly. That's the contract lever to pull - notification of material changes. I've traded a slightly higher per-seat cost for that clause before.

The tricky part is defining "material" in a way that isn't just "any change to our security policies," which they'll never agree to. Aim for something like "any change affecting the encryption, access, or geographic storage location of customer data." It gives you a real trigger.



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

Your instinct to pause is exactly right. Those specific questions about manual review and retention windows won't be answered publicly. They're reserved for the negotiation phase with an NDA in place.

When you request their security questionnaire, ask directly about "access logging and alerting for internal employees." That gets to your question of who can see the raw text. Their answer should detail if access requires an approved Jira ticket, is logged to a SIEM, and triggers alerts. A vendor with real maturity will have those controls in place, regardless of SSO tiering.

Push for a DPA that includes the right to receive evidence of those specific access controls, not just a policy statement.


Review first, buy later.


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

That's a solid point about access logging and SIEM alerts. In my last vendor review, I asked that exact question and got back a policy document full of "should" statements instead of technical specifics.

You really need to ask for an example alert rule or a sanitized log entry. A mature process will have those ready to share under NDA. If they can't produce one, it's a strong indicator the control isn't actually operational.



   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

That's the right technical control to ask for, but requesting an example alert rule or log under NDA often hits a legal wall. They'll claim it's part of their proprietary detection logic.

A more practical question is asking for their mean time to acknowledge (MTTA) and mean time to resolve (MTTR) for employee access alerts. If they can't quote you a target SLA from their runbooks, the SIEM alert is just noise.


—davidr


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 2 months ago
Posts: 294
 

MTTA and MTTR targets are meaningless without knowing the alert volume and false positive rate. They could meet a five-minute SLA by having an alert that never fires.

What if they quote you a two-hour resolution target, but their logs show manual access to raw text data happens a dozen times a day? The speed of response is less important than the frequency of the event you're responding to.


Doubt everything


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

You're right to question the privacy policy vagueness and the SSO tiering. It's a classic signal: features that should be table stakes for enterprise security are gated behind premium plans because their underlying architecture isn't built with zero-trust principles from the ground up.

Your specific questions about manual review and retention won't be answered in public docs. When you request their vendor questionnaire, demand the data retention schedule tied to their log lifecycle policies. If they claim "audio output is deleted after 30 days," ask for the S3 lifecycle rule or blob storage policy that enforces it. A verbal or policy statement is worthless without the automation that executes it.

Regarding the subprocessor list, don't just get the names. You need the *data classification* each one handles. If a contract draft goes to Google Cloud for speech-to-text, that's different than it going to a human review subcontractor. That distinction is what you need to pressure-test in their DPA.



   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

That's a sharp distinction on data classification for subprocessors. In our last vendor review, we found they'd listed "Google Cloud Platform" as a single line item, but the attached data classification matrix showed that only metadata passed through GCP Pub/Sub, while the actual audio content was sent to a different, named AI provider for processing. That's the level of granularity you need.

Pushing for the S3 lifecycle rule as evidence is correct, but you should also ask for the execution logs. I've seen a lifecycle rule exist in Terraform but fail silently for months because the IAM role attached to the bucket lacked the necessary permissions. The proof is in the CloudTrail event showing the actual deletion API calls.


Latency is a liability


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Agree completely. Asking for a sanitized log or alert rule cuts through the policy fog. I've had vendors provide them, and it's revealing. One shared a Splunk alert that only triggered on five failed login attempts from a single IP within an hour, which completely missed the insider threat scenario everyone was worried about.

The counterpoint is that a mature team might still refuse to share actual detection logic, considering it a security control itself. In those cases, I've accepted a walk-through of a dummy alert in their staging SIEM environment, demonstrating the workflow from log generation to ticket creation. It proves the pipeline is live and monitored.

Without that, you're right, it's likely theater. An operational control leaves operational evidence.


--perf


   
ReplyQuote
Page 1 / 2