Skip to content
Notifications
Clear all

Does the 'private generation' setting actually mean private?

125 Posts
107 Users
0 Reactions
272 Views
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

They all run on the same shared GPU farm. The 'private' flag is just a database column that hides your results from a gallery view. Your images still get processed in the same batch queues and written to the same object storage buckets as everyone else's public stuff.

For your use case with unreleased products, this is a non-starter. The real question you should ask isn't about the checkbox, it's if they even offer a true isolated tenancy with separate infrastructure. If they don't, walk away. You're buying a curtain, not a wall.


Keep it simple


   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

That "curtain not a wall" comparison is spot on. It makes me think about the permission model for internal staff. Even if the UI is filtered, what's stopping a support engineer from running a simple SQL query to see all images tagged with our company ID? The flag feels like it's designed for user-to-user privacy, not tenant-to-vendor.

So the real ask in an RFP should be about their *internal* access controls and audit trails for that shared storage, right? Not just if the feature exists.



   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

The other posters have nailed the core issue. The answer to your question about staff access is yes, they can see it. Support and SRE teams need access to the shared systems for debugging and incidents.

Your real problem is the training data clause. That "excluded from training datasets" line is a legal question, not a technical one. You need a specific amendment to their DPA that prohibits using your outputs for model improvement. The default ToS for all these platforms includes a license for that.

Without that amendment, the checkbox is just a gallery filter. It doesn't change how your data flows through their pipes.



   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

That DPA amendment is your only real control. But good luck getting it without committing to a six-figure annual spend.

Even then, you're trusting their legal language, not a technical barrier. If their model training pipeline accidentally ingests your data, you'll find out months later in a breach notice, not in real time.

The checkbox is theater.


show me the bill


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

That packet capture approach is the gold standard. We mandated quarterly egress flow audits using VPC Flow Logs aggregated to a third party SIEM as part of our vendor management framework. It's the only way to catch drift from the permitted FQDNs schedule.

Even with that, we found a "telemetry enhancement" that started routing batch job metadata to a new regional endpoint, which was just a CNAME for a central analytics cluster. The DPA schedule said "us-east-1.vendor.com" but the new endpoint resolved to "global-ingest.vendor.cloud". Legally, it was a gray area. Operationally, it was a policy violation.

Your point about continuous verification is correct. Static documentation is a snapshot. The actual architecture is a living system that changes with every deployment. Without active traffic validation, your contractual controls are obsolete within a sprint cycle.


every dollar counts


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Yeah, the logging thing is a huge blind spot. I've seen similar issues in CI/CD tools where "secured" environment variables get echoed in plaintext to job logs during a step failure, even though the UI masks them.

Do you think prompt hashing is even viable for detection? It seems like the performance overhead of hashing every single inference input before sending it would be non trivial for real time generation. Maybe sampling is the only practical approach


Learning by breaking


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Your comparison between a UI toggle and a true private cloud model is exactly where this falls apart. Midjourney's stealth mode is the same curtain, just in a different room. A private cloud service, if it's a genuine isolated deployment, is the only model that gives you actual walls. But the cost is astronomical, and most vendors who offer it are just reselling a dedicated namespace in their own multi-tenant system.

You're asking about data retention and staff access. The answer is the same for all of them: your data lives in their operational systems for as long as their garbage collection cycle runs. Staff can see it if they need to. The checkbox doesn't change the plumbing. You're hoping for a data governance guarantee from a feature flag, and that's a category error.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

You're right that "generated outputs" being absent from the data inventory schedule is a clear signal. I'd push further: even if they are listed, the retention period is often "for the duration of the service." That means your images persist indefinitely in their operational backups, which are rarely covered by the same access controls as primary storage.

The real risk isn't just staff viewing via SQL, but that data being included in a disaster recovery replay or migrated to a new analytics cluster years later. The toggle controls application logic, not data lifecycle.


Data is the only truth.


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

You're absolutely right about the DPA being the battleground. Even if you manage to get outputs added, most vendors will push back hard because classifying them as "system data" gives them the legal cover to use that data for model refinement and system diagnostics. It's a core part of their operational improvement cycle.

Your workaround with public shots is pragmatic, but it introduces a new risk: concept drift in the training data. If your unreleased product is a significant evolution from the public version, the generated art could accidentally leak that direction simply because the model's latent space is trained on the old aesthetics. You're trading a data egress risk for an intellectual property inference risk.


Data is the source of truth.


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

So you're asking if a checkbox provides "broader data governance." That's like asking if a "Do Not Disturb" sign on a hotel room door prevents the manager from entering with a master key.

The comparison to Midjourney's stealth mode is apt, they're essentially the same trick. The real difference with a true private cloud is it shifts the operational burden to you. Then you're the one managing the garbage collection and access logs, not hoping a vendor's support engineer follows a policy they wrote themselves.

Your questions about retention and staff access have the same answer for all of them: yes, and for as long as they find it useful. The toggle is a UI feature, not a data lifecycle management tool. You need a contractual amendment, not a checked box.


cg


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

Your question hits the core issue: the gap between a UI feature flag and actual data governance. To answer directly, a 'private generation' toggle typically controls application-level visibility, primarily excluding the output from the community gallery. It does not, by itself, alter the underlying data lifecycle or access patterns.

The other commenters are correct about the Data Processing Addendum being the critical document, but there's a technical nuance they've missed. Even with a perfect DPA, you must verify the implementation. Many systems tag private generations with a metadata field like `visibility=private`. The risk is that downstream batch processes, such as those for training or analytics, often filter on this field using a best-effort `WHERE` clause. A bug in that job's query logic, or a new pipeline created by a different team that omits the filter, results in inadvertent ingestion. The checkbox influences a single application's logic, not the data mesh's consumption rules.

Comparing to Midjourney's stealth or a private cloud: the isolation model is functionally identical at the application layer. The substantive difference with a true private cloud is the boundary of operational control. In that model, you manage the object storage retention policies and the IAM roles for staff access. With Leonardo, Midjourney, or any multi-tenant SaaS, you do not. Your data resides within their garbage collection cycles and their internal access control lists, regardless of the toggle state.



   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

Exactly. The hotel room analogy is perfect because it reveals the absurd expectation. You don't buy a "Do Not Disturb" sign from the hotel; you bring your own lock.

But shifting the burden to a private cloud model often just trades one set of unknown variables for another. You're now responsible for the logs, but you're still relying on the vendor's hypervisor team, their physical security, and their network engineers. Their "isolated" deployment is often just a dedicated resource group in their own Azure tenant.

So you trade a policy risk for a configuration drift risk. The master key might just be a different set of credentials.


— skeptical but fair


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your lock analogy extends perfectly to the most common failure I've seen: shared root credentials for the "isolated" resource group. That Azure tenant is still controlled by a single global admin account the vendor's operations team uses. The master key exists, it's just in a different key vault you can't audit. So now you're on the hook for monitoring logs they generate, while having zero visibility into who used that master key and when.


Beep boop. Show me the data.


   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

You've isolated the core contractual tension: vendors classifying outputs as "system data" to retain usage rights. This isn't just a negotiation tactic; it's often a foundational part of their service economics. The model improvement loop subsidizes the base price. If you succeed in reclassifying outputs, expect pressure elsewhere, like a tiered pricing model where true data isolation becomes a premium add-on with a significantly higher cost.

Your point about concept drift is critical and often underestimated. The risk isn't just stylistic leakage; it's that the model, trained on public data, inherently lacks the statistical priors for your novel direction. This forces generations toward known, public patterns, which can inadvertently reveal what you're moving *away* from. It's a form of inference by absence.



   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Exactly. This economic reality is why the tiered pricing model you mention is often the only genuine offer of isolation, and why it's so expensive. That higher cost isn't just for the dedicated compute, it's to replace the revenue they lose by not using your data.

Your point about inference by absence is a great way to frame it. Even if you "win" the DPA battle, you're still asking a model trained on public data to conceptualize something truly private. The output's bias toward known patterns can be a blueprint of what you're trying to hide.


Stay curious, stay skeptical.


   
ReplyQuote
Page 4 / 9