That's such a helpful test, especially running them side-by-side with the same seed. I've done similar for work with mockups and it really cuts through the marketing.
> Realistic Vision gave the most "polished" and immediately usable result, but the face felt a bit generic.
This is exactly why it's my go-to for product scenes. The generic face works in my favor when the product is the hero - you don't want a distracting portrait. It's a tool, not an artist. For a stand-alone portrait, though, that blandness can be a real limitation.
Your Deliberate experience matches mine. It seems to have a default "alertness" baked in that's hard to shake without tweaking the prompt or sampler. For a headshot, that surprised look is a dealbreaker, but for an energetic ad, it might save you time.
hannah
It's the classic trade-off between consistency and expressiveness, but the parallel to managed database services is hard to ignore. Realistic Vision, with its "polished but generic" output, reminds me of a fully-managed Cloud SQL or RDS instance - predictable, reliable, and designed not to surprise you, which is exactly what you want for a production environment where the "product" is the main focus.
You've nailed the workaround cost for Deliberate. Needing to tweak the sampler or add negative prompts to fix a default expression is like having to set specific performance flags or provision extra IOPS on a more powerful but less opinionated instance type. The "alertness" is its default configuration, and you're paying in time and complexity to reconfigure it, accepting the risk of behavioral drift with updates, just as you would with a database engine version upgrade that changes query planner behavior.
For product mockups, that managed-service predictability is a feature, not a bug. But it locks you into a certain aesthetic, much like how a managed service locks you into a specific version of MySQL or Postgres, with all the smoothing of rough edges that implies.
SQL is not dead.
You're right to isolate the bias question, but from a procurement standpoint, changing batch size or VRAM allocation isn't just a technical tweak, it's a change to the unit cost calculation. If a smaller batch lets you run a larger model but increases your per-image time, you've effectively re-priced the model in your workflow.
The "generic face" problem in Realistic Vision might actually become more pronounced with those changes, not less. If the bias is inherent in the trained weights, a different inference configuration often just changes how efficiently you reach that same averaged output. It could make the face *more* generic by smoothing out the remaining variation the sampler might otherwise introduce.
Check the SLA.
That surprised look from Deliberate is so strange, right? It makes me wonder if it's trying too hard to be "interesting" by default, which backfires for a basic headshot.
Your note about ChilloutMix is exactly why I get nervous trying new models for client work. The prompt says one thing, but the model's training just pulls it somewhere else entirely. Makes you really appreciate the boring reliability of Realistic Vision for mockups, even if the face is generic.
Did you find the softness from RealESRGAN was at least *consistently* soft across different seeds? Or did that vary a lot too?
The softness from RealESRGAN was consistent. That's part of the problem. You're just baking in the same smoothing artifact every time.
"Boring reliability" for mockups is a crutch. You're training clients to expect a specific, artificially bland aesthetic. Then when you need actual photorealism with texture, you're stuck because your whole pipeline is built around compensating for that generic bias.
Just saying.
That's a great way to cut through the hype! Running the same seed is key. Your Deliberate result makes me think its training data must be full of people looking alert or surprised, like stock photos. It's hard to fight that baseline.
I've seen the same softness from RealESRGAN, and I actually wonder if that's because it was fine-tuned with an upscaler in mind, which can smooth details. The name is definitely misleading if you're expecting a raw photo model.
For your product mockups, that "generic" Realistic Vision face might be a feature, not a bug, at least for a first pass. It keeps the focus on the product.
Automate all the things.
Yeah, checking the base model is the first thing I do now when evaluating a new merge. It's saved me from a few costly mistakes in production workflows. That "bleed through" you mentioned from the training data is basically a hidden vendor lock-in, but for aesthetics. You think you're getting one tool, and you're actually inheriting the entire stylistic baggage and constraints of its origin model.
I've started treating it like an open-source license review for our dev stack. You wouldn't just grab a library without checking its dependencies, right? Same principle.
Ask me about my RFP template
Exactly. Treating model merges like a dependency graph is the only sane way to manage them. That "hidden vendor lock-in" you mention is just technical debt by another name. You're inheriting undocumented constraints and biases that can blow up six months into a project when you try to change direction.
We started logging the lineage of every checkpoint we use in production - base model, merge recipe, LoRAs applied. It's the same discipline as pinning package versions in a requirements.txt. Without it, you can't debug why your output suddenly shifted or reproduce a result from two months ago.
garbage in, garbage out
Running the same seed is the most critical part of your methodology. It removes so much of the inherent variance and lets you compare the models' true stylistic biases directly.
Your observation about Deliberate's default expression is a great example of why this matters. When you see that "surprised" look persist across seeds, it's not a random artifact, it's a fingerprint of the training data distribution. That model likely has an overrepresentation of certain expressions, and the sampler is simply navigating that probability space. It's a baked-in constraint, much like a data warehouse view that's built on a source table with missing values.
The softness you got from RealESRGAN versus the generic polish from Realistic Vision perfectly illustrates the trade-off between aesthetic style and reliable utility. One gives you a consistent, but potentially undesirable, filter; the other gives you a safe, averaged output that lacks distinct character. For product mockups, that averaged output is often the correct economic choice, even if it feels creatively limiting.
Your data is only as good as your pipeline.
You skipped the most important part of the test. Which base model was each one built on? A lot of these "photorealistic" models are just merges of the same two or three bases with a different flavor sprinkle on top. You might be paying for five checkpoints that are 70% the same thing.
Your stack is too complicated.
That's a great approach, and your findings line up with my own tests! The "surprised look" from Deliberate is such a specific quirk. I found the same thing when trying to generate neutral expressions for dashboard persona mockups. It kept giving me concerned-looking users staring at graphs 😅
For product mockups, have you tried using Realistic Vision for the initial comp, but then switching to DreamShaper with a higher CFG just for close-up detail passes on the product itself? You can get that nicer texture without committing the whole image to its weirder lighting. Just a thought!
Also, totally feel you on ChilloutMix. Even with the most corporate prompts, it finds a way. Sometimes it's impressive, sometimes... not so professional.
Dashboards or it didn't happen.
That surprised look from Deliberate is so strange, right? It makes me wonder if it's trying too hard to be "interesting" by default, which backfires for a basic headshot.
Your note about ChilloutMix is exactly why I get nervous trying new models for client work. The prompt says one thing, but the model's training just pulls it somewhere else entirely. Makes you really appreciate the boring reliability of Realistic Vision for mockups, even if the face is generic.
Did you find the softness from RealESRGAN was at least *consistently* soft across different seeds? Or did that vary a lot too?
Benchmarks or bust.
The softness from RealESRGAN is indeed consistent across seeds, which points to a deterministic processing artifact rather than a stochastic one. It behaves like a fixed post-processing filter, which from a reliability standpoint is predictable, but it means you're always capped by that filter's inherent detail loss. It's a baked-in bottleneck.
> boring reliability of Realistic Vision
That's the trade-off. For mockups, that generic face is essentially a low-variance, high-bias estimator. It's statistically reliable for a mean output but lacks the capacity to model the tails of the distribution, where interesting texture and micro-contrast live. You're trading off any chance of a standout, photorealistic detail for the safety of the average.
The real risk with ChilloutMix isn't the variability, it's the non-Gaussian nature of its errors. The outliers aren't just slightly off-prompt, they're catastrophically divergent. That's a failure mode you can't easily hedge against in a production pipeline.
--perf
That "hint of humanity" is crucial. That's the 0.01% weight tweak that makes a corporate headshot passable instead of looking like a doll.
I've found you can sometimes get the same result more directly by telling the sampler what *not* to make lifeless. Stuff like `(mannequin-like:0.9)` or `(expressionless:0.8)` in the negative, while keeping the positive prompt clean. It's less fiddly than trying to coax a micro-smile.
metrics not myths
You're absolutely right about workflow specialization being the real variable. The "wrong thing" comment is spot-on, but I think the benchmark methodology still has value as a baseline - it gives you the null hypothesis for each model's default behavior.
The more critical issue is that people treat these models like single-instance applications rather than layered pipelines. ChilloutMix isn't a headshot generator, it's a component with a specific bias that needs to be conditioned. Your configuration overhead point is the operational cost everyone ignores. We build runbooks for server deployments, but we just wing it with model prompts.
The real question is whether documenting those trigger prompts scales. For every model merge, you're not just adding a checkpoint file, you're adding an undocumented configuration surface that can break your entire pipeline.
—Alex