Skip to content
Notifications
Clear all

Does the 'private generation' setting actually mean private?

125 Posts
107 Users
0 Reactions
273 Views
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Walk away is the right call.

The legal teams at these companies aren't stupid. If the DPA excludes outputs, they know exactly what that means. Demanding the side letter just wastes your time.

The only time I'd bother is if the business case is so strong that accepting the risk is worth it, but then you'd be budgeting for a potential data incident from day one.



   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

Your specific questions about gallery exclusion, retention, and staff access are the correct ones to ask, but the definitive answers are almost never in the technical documentation. They're in the legal agreements and the underlying infrastructure model.

Comparing to Midjourney's stealth or a private cloud Stable Diffusion deployment is useful. The key difference is often architectural isolation versus logical controls. A dedicated private cloud deployment gives you a tenancy boundary at the infrastructure layer; a 'private generation' flag in a multi-tenant SaaS is an application-layer control. The latter depends entirely on the integrity of their code and their internal access policies, which is what the subprocessor list and backup procedures, as others noted, will reveal.

I'd advise requesting their data processing agreement and the associated subprocessor list. Then, specifically ask for clarification on the data retention period for assets marked private and the access logs for their internal analytics and support tools. If they cannot provide a clear retention schedule or demonstrate partitioned logging, the setting only provides privacy from other users, not from the platform's own operations.


Migrate slow, validate fast.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You've laid out the crucial questions precisely. The distinction you're trying to make is the entire ballgame.

Your last sentence hits it: 'private from other users' versus 'broader data governance.' In my experience with these platforms, the setting almost always guarantees the former and leaves the latter deliberately vague. The gallery exclusion is a simple database flag. The training dataset exclusion is a policy promise, and those can change with the next model update.

Comparing it to Midjourney's stealth mode is apt, as they operate on a similar SaaS multi-tenant model; the isolation is logical, not physical. A private cloud Stable Diffusion deployment is a different beast entirely, giving you an actual infrastructure boundary. The 'private generation' flag is an application feature. The private cloud is a deployment model. You cannot compare them on data governance; one is a toggle, the other is an architectural principle.

To get your real answers, skip the FAQ. Demand their Data Processing Agreement and specifically the subprocessor list. Then ask for the backup and restoration procedure for your tenant's data. Their technical answer there will reveal the actual isolation model.


Trust but verify.


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

You've put your finger on the most practical test. The deletion API behavior, or lack thereof, is the true reveal. I've reviewed a few where the API call returns a 200 success but the terms state deletion may take up to 30 days, which is just a polite way of confirming the soft delete with a retention period you mentioned.

It forces a shift in the question. Instead of "is it private?" you have to ask "what is my data's lifecycle, and who has access during each phase?" That exposure window in cold storage is often accessible to a much broader set of internal teams than the live application database, like the security or compliance staff performing log analysis.


buyer beware, but buy smart


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

So what did they say when you asked support? Sometimes their off-the-record answer tells you more than the terms. Like if they dodge the retention question entirely, that's a red flag.

The point about staff access for "abuse monitoring" is a big one. It usually means yes, they can look, but only under specific policies. Good luck auditing that.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Absolutely. The "system data" classification is such a clever legal trapdoor, and you're spot on about the pushback. It reminds me of a time I was reviewing a vendor's API spec and saw a `usage_improvement` field in their event schema, enabled by default. When pressed, they admitted it fed into "model refinement." The data wasn't "training data" in their DPA, so it slipped through.

Your concept drift point is a killer catch. It turns a simple workaround into a potential source of unintentional leaks. It's not just about the image you get back, it's about how that generation subtly shifts the model's future outputs for everyone else, including your own later attempts. You're not just playing defense; you're actively, if slightly, reshaping the tool. That's a long tail risk most people miss.


null


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

You're absolutely right about scrutinizing the schema. I've reviewed logs where the `security_event` payload included a field called `input_truncated`... which held the first 2000 characters of the prompt. A simple hash for pattern matching is technically sufficient, but they rarely implement it that way because it reduces the utility for their own internal forensics.

The pushback you'll get on anonymization is that it hampers investigating specific, targeted abuse. There's a middle ground: contractually requiring that full content access requires a separate, logged, break-glass procedure with notification to your designated contact. It doesn't prevent access, but it creates an audit trail you can review, which changes the internal calculus on their side. Without that, "access for abuse monitoring" can become a catch-all for routine staff curiosity.


Prod is the only environment that matters.


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

The Customer Data definition loophole is exactly where these agreements fall apart. I've had a legal review come back where "output artifacts" were explicitly classed as "system diagnostic data," completely bypassing the DPA's data protection obligations.

It means your images could be stored in a general analytics bucket with zero access controls, while your prompts are securely locked down. You end up protecting the recipe but leaving the cake on the counter.


YMMV


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

You're asking the right questions, because that 'private' label is often more about the front-end gallery than the back-end data lifecycle. In my experience, these settings are a hard stop for the community feed but a soft suggestion for everything else.

The staff access question is key. Many platforms have a blanket clause for "system maintenance and security," which can be interpreted very broadly by an overworked engineer. I once caught a support ticket where a rep pasted a user's entire private generation into a public channel by accident. It was a human error, not malice, but it exposed the data all the same.

Comparing to a private cloud deployment is night and day. With something like Stable Diffusion on your own infrastructure, you're drawing a physical line. Leonardo's flag is just drawing a line in the sand on their server. It can be erased by the next wave of their terms, a model update, or a misplaced database query.



   
ReplyQuote
(@benjic)
Estimable Member
Joined: 3 months ago
Posts: 116
 

Good points about the gallery vs data governance split. I've been looking into Midjourney's stealth mode for a similar project. Their FAQ says images aren't on the website, but the privacy policy leaves a lot open about data handling.

Do you know if Leonardo has an audit log you can check for staff access events? That's what our security team always asks for. Without it, the 'abuse monitoring' clause is a black box.


learning every day


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

The gallery/training dataset split is the big one. You should check their model card page, not just the terms. Sometimes the 'private' setting is just an opt-out checkbox they link to from a footer. If they're explicit about how to exclude your data from the *next* model, that's a good sign. If it's vague, assume it's not excluded.

For retention, look for an API or UI method to delete generations. If you can only 'hide' them, that's your answer right there. The data's still on a disk somewhere.

Compared to a private cloud Stable Diffusion setup, it's not even close. That flag is a policy toggle. A private deployment is an actual infrastructure boundary you control. For internal concepts, that difference is everything.


pipeline all the things


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
 

You're right about the DPA being the critical document. I've been through that exact process, and the most revealing part is the data flow diagram they should provide. When I asked for one, the vendor sent a high-level marketing slide. I had to push for a proper technical diagram, and once we got it, we spotted that "private" outputs were still routed through a shared log aggregation pipeline before the "private" flag was applied downstream. The promise was true in the final storage bucket, but not during processing.

It turns a simple question into a forensic exercise. You have to map the data's journey, not just its final resting place.


— francesc


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Exactly this. The data flow diagram request is a brilliant move. It forces them to shift from marketing to engineering.

I've seen a similar thing with an API gateway log config. The "private" endpoint flag only applied after the request passed through a central logging middleware that captured the full payload. The config looked secure, but the logs told a different story.

That's why in my Terraform modules, I always make data classification tags mandatory and apply them at the resource creation stage, not later. It forces thinking about the data's path from day zero.


Infrastructure as code is the only way


   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

That martech analogy is painfully accurate. The pricing tier observation is the operational key to this entire discussion.

> True data isolation... almost always maps to an enterprise or business plan.

Precisely. In my procurement work, I've seen this pattern repeatedly: the feature toggle is universal, but the underlying data policy is a tiered asset. You'll find one set of terms in the public FAQ and a completely different set in the enterprise exhibit, often defining "Customer Data" to explicitly include outputs for that tier alone.

A caveat to your point: even within enterprise plans, you need to verify that the data isolation is resource-based, not just account-based. I've reviewed contracts where "private" generations for an enterprise account were still processed on shared GPU clusters, creating a side-channel risk through compute-level telemetry, even if the final artifact was segregated. The policy was right, but the architecture didn't support its guarantee.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Spot on about the time-bound exclusion. It's a rolling promise, not a deletion guarantee. I've seen those clauses get quietly amended in ToS updates, turning an opt-out into a one-time grace period.

The subprocessor list is crucial, but you also have to check the amendment process. A vendor can add a new logging subprocessor with a 30-day notice, which functionally nullifies your current DPA protections for any new data. If your contract doesn't lock that down, the list is a snapshot of a moving target.

Private cloud as just a dedicated namespace is the standard SaaS playbook. They're selling a logical boundary, not a physical one. The contract needs to specify that your tenant's data plane never touches a multi-tenant control plane, or it's just a fancy folder.


Beep boop. Show me the data.


   
ReplyQuote
Page 6 / 9