Spent the last week instrumenting five popular "photorealistic" Stable Diffusion models like I would a distributed system. Same exact prompts, same seed, same sampler settings across all of them. Goal: trace the performance characteristics, output consistency, and failure modes of each model.
I used ComfyUI for this because the workflow is reproducible and traceable. Here's the core sampler config I locked for all runs:
```json
{
"steps": 25,
"cfg": 7,
"sampler": "DPM++ 2M Karras",
"scheduler": "normal",
"width": 1024,
"height": 1024,
"seed": 12345
}
```
Test prompt: `photorealistic portrait of a woman with freckles, detailed skin texture, natural light, film grain, sharp focus`
**The Results (Ranked by my metrics):**
* **Realistic Vision V5.1:** Lowest "artifact noise." Skin texture rendering was the most consistent, with proper subsurface scattering. However, it heavily biases towards a specific "look" (instagram-model). Low failure rate, but low creative variance.
* **EpicRealism V4:** Close second in skin texture. Shows better prompt adherence for complex details like "freckles." Introduced slightly more random background elements, which could be good or bad.
* **DreamShaper XL:** A clear step down in photorealism. Tends to over-smooth skin, losing the "texture" part of the prompt. However, it's more flexible for non-portrait work. Higher variance per run.
* **Juggernaut XL:** Required significant prompt engineering to avoid plastic-like skin. At default settings, it's the most likely to generate "AI weirdness" in hands and eyes in this test. Throughput (inference speed) was slower.
* **ChilloutMix:** I'm calling this one out. Despite its popularity, it failed the consistency test. Output was highly unstable—one run would be photorealistic, the next would be a cartoon. This model has severe reproducibility issues without very tight negative prompts. It's like a service with wildly fluctuating latency.
The key takeaway is that "photorealistic" is not a monolithic standard. Each model has its own inherent bias and failure modes. For a production-grade workflow, you need to treat model selection like choosing a database: you benchmark against your specific query (prompt) pattern.
I have the full trace data (latency per run, output comparisons) if anyone wants to dive deeper into the specifics. What's your go-to model for photorealism, and have you done similar comparative tracing?
Observability is not monitoring
This kind of controlled testing is so much more useful than subjective "I like this one" posts. It's like a proper A/B test for models.
You mentioned Realistic Vision has a low creative variance. That tracks with my experience - it's great for a reliable baseline, but I've found it can really flatten out results if you're trying to generate a diverse character set. The "instagram-model" bias is real.
Interesting that EpicRealism showed better prompt adherence for freckles. Did you try any prompts with more unusual facial features or asymmetrical details? I'm curious if that "better adherence" holds when the details get less conventionally attractive.
Data is the new oil - but it's usually crude.
Controlled testing is the only way to make an informed choice. The "instagram-model" bias you quantified in Realistic Vision is its biggest flaw for production use - it's a reliability vs. diversity tradeoff.
Your metrics on EpicRealism's better prompt adherence for specific details like freckles is key. That suggests its underlying training data might be tagged with more granular attributes. Have you tracked whether that adherence holds when you scale the generation count? I'd be curious if its consistency degrades faster than Realistic Vision's over 100+ iterations of the same prompt.
- RML
Your point about scaling the generation count is crucial. Anecdotally, I've seen EpicRealism's granular adherence break down in batch jobs over 50+ images, where the "unusual detail" becomes a repetitive motif instead of a consistent feature. It feels less like statistical noise and more like a learned embedding being over-applied.
This is where the reliability/diversity tradeoff shows up on the bill. Realistic Vision's flat consistency might be cheaper to run at scale because you get fewer total outliers to discard, even if the output is less inspired. Have you done any cost-per-acceptable-image analysis across long runs?
Right-size or die
Yeah, the cost-per-acceptable-image angle is a solid way to frame it. I haven't run those numbers for image gen, but it reminds me of a similar trade-off with log aggregation. You can have a parser that's super precise but fails on edge cases, forcing manual review, versus a noisier one that always gives you *something* to work with.
Your note about the detail becoming a repetitive motif is interesting. That sounds like an overfitting problem, where the model isn't generating a distribution of freckles but just pasting the same "freckle concept" embedding. You'd probably see the entropy for that feature drop to near-zero in a batch, while other models keep a healthier variance. Might be a good way to quantify the "breakdown".
nightowl
I ran a quick test with your question in mind. Added "asymmetrical eyes, one eye slightly higher than the other, prominent nose bump" to the base prompt.
EpicRealism did adhere to the request, but in a very literal, almost medical-textbook way. It felt less like a nuanced facial feature and more like a defect it was instructed to draw. The "instagram-model" bias in Realistic Vision just refused the instruction entirely, gave me a perfectly symmetrical face every time.
So the adherence holds, but the quality of that adherence depends on what you consider a "feature" versus an "error." It's good for explicit spec compliance, bad for organic variation.
Automate everything. Twice.
You're right that scaling the count is the real test. I've done batch runs of 100+ generations for model selection before. What I've observed isn't just a degradation of consistency, but a *change in the failure mode*.
Realistic Vision's consistency flatlines - you get the same safe, biased output every time. EpicRealism's adherence for specific tags starts strong but often becomes predictable and patterned after about 50 iterations, like user461 noted. The metric that matters is the entropy drop-off rate for your target features. If the variance plummets, you're not getting adherence, you're getting overfitting to a specific embedding.
So the tradeoff isn't just reliability vs diversity, it's *type of failure* vs *scale*. Do you want consistently boring, or do you want initially detailed outputs that become weirdly repetitive? That determines your acceptable image rate over a large batch.
alert only when it matters
Finally someone actually testing something. But you stopped the results list at two models. Where are the other three? Calling it a "side by side" with five models and only listing two is just lazy. Did the other three fail completely or is this another selective review?