Skip to content
Notifications
Clear all

Thoughts on the privacy policy? They say they train on your data. Concerned.

4 Posts
4 Users
0 Reactions
37 Views
(@kubernetes_tinker_99)
Estimable Member
Joined: 7 months ago
Posts: 56
Topic starter   [#9865]

Hey folks, been tinkering with SciSpace (formerly Typeset.io) for a few weeks to organize research papers. The tool itself is pretty slick for summarization and Q&A across PDFs.

I was going through their docs before potentially rolling it out to my team, and their privacy policy gave me pause. Specifically, the part about using your uploaded content to train and improve their models. Here's the snippet that caught my eye:

> "We may use the Content to... train our models and improve our Services."

This is common for many AI-powered SaaS tools, but the ambiguity is the concern. As someone who operates on a "trust but verify" principle with cloud services, I'm left with questions:

* **What exactly is "training"?** Is it fine-tuning a core model, generating synthetic data, or something else?
* **Does this apply to all data?** What about proprietary research drafts or papers under peer-review embargo?
* **Is there a data retention policy for this training data?** Can you opt-out?

In our K8s/Argo CD world, we'd define this in a ConfigMap or a policy-as-code rule! But here, it feels a bit like a black box.

Has anyone dug deeper or reached out to their support? I'm trying to weigh if:
- This is a standard clause we accept for the functionality.
- If we need a paid plan for clearer data handling terms.
- Or if it's a dealbreaker for any non-public content.

Would love to hear your experiences or if you've found a solid alternative that's more transparent about data pipelines.


#k8s


   
Quote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Classic. You're right to be wary.

> "We may use the Content to... train our models"

That's the catch-all. Unless their DPA or a support ticket says otherwise, assume "train" means exactly that, and it applies to everything you upload. Embargoed draft? Now it's part of their training corpus.

Opt-out? Good luck. If it's core to their service, they won't let you. Their "improvement" loop is fueled by your data. In our world, you'd set a retention policy. In theirs, it's a value extraction feature.

Slick tools always have a price. Your proprietary research is likely part of it.


Just my two cents.


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 4 months ago
Posts: 723
 

That snippet is standard boilerplate, but the devil is in the detail they don't provide.

Your questions are valid, but you won't get answers from a public policy. You need to pressure support or legal for their Data Processing Addendum (DPA). The DPA is where the actual carve-outs are, if they exist.

Look for:
* Explicit opt-out mechanisms
* Definitions of "Training Data" vs. "Service Data"
* Specific retention windows for processed content

If they won't provide a DPA or it's just as vague, you have your answer. Assume all uploads are training fodder.


Benchmarks don't lie.


   
ReplyQuote
(@latency_king)
Trusted Member
Joined: 6 months ago
Posts: 44
 

Pushing for a DPA is sound advice, but in my experience with these platforms, the data pipeline's velocity often outstrips contractual assurances. Even with a DPA clause stating content is isolated for "live service only," you must examine the actual request flow.

Where does the inference call go? Is there a separate, gated endpoint for model training that your data could be routed to post-processing? The network latency between the service's frontend pod and its training cluster is the real tell. If they're colocated in the same VPC with sub-millisecond hops, the architectural separation is likely non-existent. A true silo would introduce noticeable, added latency for the training path.

Ask support not just for the DPA, but for a high-level data flow diagram. If they can't provide one, the operational model is probably "everything goes everywhere," and your verification is complete.


Every microsecond counts.


   
ReplyQuote