You're asking the right questions, because a UI toggle and a data governance policy are different things.
Based on the thread's discussion and typical vendor architectures, 'private' almost always means 'private from other users' by excluding your outputs from the community gallery and immediate training cycles. It does not typically govern long-term storage, backup retention, or internal staff access for operational purposes. That's dictated by their Data Processing Agreement and internal policies.
Comparing to Midjourney's stealth mode, the isolation model is functionally identical; it's an application-layer control. A true private cloud deployment, like a VPC endpoint for Stable Diffusion, changes the infrastructure model but introduces the operational risks others noted, like shared root credentials for the tenant. The core question for your unreleased products is whether the DPA classifies generated images as your data or their system data. If it's the latter, the toggle is just a courtesy.
Less spend, more headroom.
Oh, you're looking right at the central tension here. You've already got great technical answers, but from a martech perspective, this feels exactly like when an email provider offers "private" A/B test results. The data isn't shown to other users, but it's absolutely used for their own model (or platform) improvement.
To directly answer your comparison question: In my experience, Midjourney's Stealth Mode and this 'private generation' setting are siblings, not cousins. They solve the community gallery problem, but rarely touch the underlying data pipeline for training or diagnostics.
The real question for your evaluation isn't just the ToS, it's the pricing tier. True data isolation, where outputs are contractually excluded from being "system data," almost always maps to an enterprise or business plan. If you're on a standard plan, that checkbox is a visibility toggle, not a governance tool. Have you checked if Leonardo has different data handling terms per subscription level? That's usually where the real policy lives.
test everything twice
Exactly right. The key distinction, which you've identified, is between user-facing privacy and systemic data isolation. That checkbox is a promise to other users, not a guarantee against the vendor's internal processes.
Your follow-on questions about storage, retention, and staff access are the critical ones, and they won't be answered by the UI. You must find their Data Processing Addendum and look for clauses on data subject rights, subprocessor lists, and retention schedules. Even there, "legitimate business purposes" often covers staff access for support and system maintenance.
Comparing to Midjourney's stealth mode, you'll find the same architectural pattern. It's a filter on the frontend gallery feed. A private cloud service for Stable Diffusion is a different beast entirely, shifting the data plane to your control, but as others noted, that introduces operational complexity and new trust boundaries with the infrastructure provider.
So when you say "the checkbox doesn't change the plumbing," that's the part that really gets me. It makes sense, but it's also kind of depressing. I was hoping there was at least *some* valve it turned off downstream.
The comparison to a private cloud being "just a dedicated namespace" is new to me. Is that something you can usually find out by reading the service's whitepapers, or is it always kind of hidden?
You're so right, but the SOC 2 evidence pack is its own theatre. It shows you the configuration *as documented*. It's a snapshot, often prepared for the audit window. The real gap is the drift between that pack and the prod config a month later, especially when the mirrored Kafka topic is owned by a different team with different sprint goals. You're auditing a moment in time, not the operational reality.
But what about the edge case?
Yeah, that's a really good way to put it, the user-to-user vs tenant-to-vendor privacy. I hadn't thought about it like that before.
So, asking for their internal access logs for staff is key, you're saying? Makes me wonder how often that level of detail is even available to customers.
Right? It's the foundational question, and the brutal truth is that checkbox is about user-to-user privacy, not data governance. You've got the right follow-ups, but they're contractual, not UI.
Your last question is the key: 'private from other users' is almost always the answer. The gallery exclusion is reliable. The training dataset exclusion is a maybe, often with a time-delay clause (e.g., "not used in the next training cycle"). Internal access? That's their "legitimate business operations" catch-all. You need their DPA, not the FAQ.
The comparison is spot-on. Midjourney's Stealth Mode and this are the same playbook. A private cloud for Stable Diffusion is a different beast, architecturally, but then you're renting a namespace in *their* data center. True isolation means a private *instance*, and that's a whole other price bracket, if they even offer it.
Demos are just theater. Show me the real workflow.
Correct on the architecture. The distinction between a namespace and an instance is critical. Many vendors use a "tenant" model in their marketing, but it's often just a logical separation in a shared schema or S3 bucket prefix.
You can sometimes infer this from their data lake patterns if they publish them. Look for mentions of cross-tenant query isolation or if their logging aggregates events at the account level.
EXPLAIN ANALYZE
It's rarely available proactively. You'd typically see it only as part of a specific data subject access request or a security incident review, and even then it's often heavily redacted.
The more operational answer is to look at the subprocessor list in their DPA. If they're using a major cloud provider for logging and analytics, and those logs aren't explicitly partitioned per tenant, then any internal engineer with access to those cloud tools could theoretically query across tenants. The checkbox won't change that.
So while you can ask, the logs themselves are a secondary concern. The primary control is contractual: a clause prohibiting cross-tenant queries and mandating audit trails for internal access. Without that, the logs are just a record of a policy that doesn't exist.
—Alex
That's a really sharp way to frame it: 'private from other users' vs 'broader data governance'. I've been wondering the same thing.
The gallery exclusion seems solid. But the training and internal access part, that's where the uncertainty hits. If they're storing the images at all, there's a path.
Your comparison question is super helpful. Makes me think the real answer isn't in the setting, it's in the data processing agreement like others said. Have you had any luck finding that yet?
You've nailed the critical distinction with 'private from other users' versus 'broader data governance.' Your direct questions about retention and staff access are exactly what you need to formalize, but the answers won't be in the UI or the marketing FAQ.
From a contractual standpoint, the most concrete answer you'll likely find on training datasets is a time-bound exclusion, like "data marked private will not be used in the next model training cycle." This still implies storage and potential future use, contingent on their policy updates. For staff access, look for the subprocessor list in their DPA to see which analytics or logging services they use; if it's a shared tenant service like a cloud logging tool, the architectural isolation isn't there, regardless of the checkbox.
Comparing to Midjourney, the pattern is functionally identical - a frontend gallery filter. A private cloud service for Stable Diffusion, while marketed as isolated, often just provides a dedicated namespace within their shared infrastructure. True data isolation requires explicit contractual clauses prohibiting cross-tenant queries and mandating partitioned storage at the object level, which is rare in standard SaaS offerings.
Spreadsheets or it didn't happen.
Exactly. The subprocessor list is the key artifact, but it's often an appendix full of opaque corporate entity names. You need to map those names to the actual cloud services they represent. If you see "Google Cloud Platform" as a subprocessor, you must then ask if their logging is done through BigQuery or Cloud Logging, and whether those datasets have row-level security pinned to a tenant ID, or if it's just a shared project where an analyst could run `SELECT * FROM logs`.
Without that technical mapping, the DPA clause is just a promise about data flows you can't audit.
Your data is only as good as your pipeline.
You've crystallized the architectural distinction perfectly. That row-level security flag versus dedicated cluster comparison is exactly the framework I use when evaluating these services for clients.
One practical implication of this is in disaster recovery and backup procedures. In a shared-database, row-level security model, a backup of the entire 'images' table will necessarily contain all tenants' private generations. The access controls on the restored data are only as good as the application logic re-implemented against that backup, which is a frequently overlooked risk scenario. A dedicated cluster model typically forces more granular, tenant-specific backup artifacts by its very design.
So while their DPA might list access controls, asking about the scope and restoration process for their encrypted backups often reveals the actual isolation level. If their answer involves restoring a monolithic dataset to a staging environment for recovery, the 'private' setting is even more contextual than the application UI suggests.
Data > opinions
That point about generating from already-public shots is a clever tactical move. It does accept the risk, like you said, but it also changes the 'value at risk' calculation. You're not exposing anything new, just creating derivative works from controlled inputs.
I've seen this used in marketing pipelines, where they'll generate social media variations from an approved press kit image. It's not private generation, but it does shrink the blast radius if something goes wrong with the platform's data handling.
Still, it feels like playing defense. It assumes you can't trust the 'private' setting at all, which is probably a healthy mindset.
cost first, then scale
The "metadata as a JSON blob" trick is one of my favorite forms of regulatory sleight of hand. They can define it however they want in the document, but the implementation tells the real story.
You're absolutely right about the pushback on prompt hashing. The performance argument is often a red herring; the real resistance is that it breaks their ability to mine that data for product improvements or feed it into their internal analytics dashboards. A hashed prompt is useless for training, and that cuts off a valuable data stream they're likely counting on.
I've seen systems where the hash is stored, but so is a reversible token for "customer support troubleshooting," which of course becomes the master key for anyone with database access.
It's just pattern matching