I've been conducting a rigorous comparison of SD 1.5 and 2.1 base models for a synthetic image generation pipeline, and the results are concerning. Specifically, the facial outputs from the SD 2.1 base model (`stable-diffusion-2-1`) are consistently subpar—distorted features, asymmetrical eyes, and a general "uncanny valley" effect that 1.5 rarely produces at equivalent settings.
My hypothesis is that this isn't just about the model checkpoint, but the interplay between model architecture, the negative prompt, and the specific VAE. In my controlled tests, I've used the following parameters as a baseline:
```yaml
checkpoint: v2-1_768-ema-pruned.ckpt
sampler: Euler a
steps: 25
cfg scale: 7
size: 768x768
prompt: "portrait photo of a woman, detailed eyes, professional photography"
negative prompt: "deformed, blurry, bad anatomy"
```
Even with these, the facial coherence is markedly worse. I've observed:
* Pronounced "melting" or merging of facial structures (e.g., cheekbone into hair).
* Inconsistent rendering of fine details like eyelashes or teeth.
* A higher propensity for grotesque distortions compared to 1.5, which seems more robust.
Has anyone else replicated this in their workflows? More importantly, has anyone found a reliable configuration pattern or necessary preprocessing step to mitigate this with the 2.1 base? I'm considering whether this is a known regression that necessitates always defaulting to a dedicated portrait fine-tune (like `Realistic_Vision`) when faces are a priority, which adds complexity to the pipeline.
--crusader
Commit early, deploy often, but always rollback-ready.
You're definitely not alone on this. I've noticed the same thing when we were stress-testing 2.1 for some of our internal content guidelines.
One thing that made a significant difference for me was switching to a different sampler for portraits. Euler a can be a bit aggressive sometimes. I've had better luck with DPM++ 2M Karras or even the simple DDIM for those finer facial details at a higher step count, say 35-40. It seems to smooth out that "melting" effect you described.
Also, have you tried explicitly adding something about face symmetry to your negative prompt? I know you have "deformed," but I found adding "asymmetrical eyes, uneven face" gives the model a more specific nudge away from those common distortions. It's not a perfect fix, but it helps reduce the failure rate.
Keep it constructive.
I've been down this exact rabbit hole for a project generating synthetic profile images, and your controlled parameters mirror my own initial tests. You're right to suspect the interplay of components.
The VAE is a huge factor often overlooked. The 2.1 base model uses a different VAE than 1.5, and it's notoriously weaker on faces. Swapping in the vae-ft-mse-840000-ema-pruned VAE from SD 1.5 into your 2.1 pipeline can stabilize features significantly, reducing that "melting" artifact. It's a kludge, but it works.
Also, the training data shift for 2.1 can't be ignored. The filters removed a lot of "unsafe" content but also much of the nuanced facial data. Your pipeline might need an explicit fine-tuned model on top of the base checkpoint for consistent portraits, treating 2.1 as a worse starting point than 1.5 for this specific domain.
Extract, transform, trust
You're spot on about the training data shift, it's a key piece of the puzzle. I'd treat the VAE swap as a necessary step for any face-centric work with 2.1, not just a kludge. It's become standard practice in our workflows.
That said, I've found the "worse starting point" idea cuts both ways. For a corporate project needing very generic, neutral faces, the 2.1 base's limitations after the filtering actually made our job easier by reducing unwanted "style." We just had to accept a higher initial discard rate. It's a trade-off.
Keep it constructive.
The core issue isn't your prompt or sampler. It's the dataset. The aggressive NSFW filtering for 2.1 removed huge swaths of portrait data, crippling its understanding of facial geometry. You can't prompt your way out of a fundamentally broken training set.
Your "rigorous comparison" is correct. 2.1 base is objectively worse for faces. Treating it as a starting point for a production pipeline is a mistake unless you're fine-tuning it on a dedicated face dataset first. The VAE swap mentioned later is a band-aid on a bullet wound.
Stop using 2.1 base for portraits. Use a model derived from 1.5 or accept that you're working with a compromised tool. No amount of negative prompting fixes missing data.
— geo
Your controlled test matches what we've seen in the moderation queue for projects using the 2.1 base. It's a common frustration.
I think the bigger question your tests highlight is whether the base model should be the go-to for a production pipeline, or if the expectation needs to shift. For community guidelines, we often advise treating 2.1 as a specialized tool rather than a universal upgrade. It's good for some things, but for reliable, consistent faces, the consensus here is leaning towards using a model fine-tuned for portraits from the start, as others have said.
Have you considered running your pipeline with a 1.5-based portrait model as a direct A/B test against your current setup? The difference in moderation overhead for odd faces might be significant.
Raise the signal, lower the noise.
That's an interesting point about treating it as a specialized tool. But if the base model is so compromised on a fundamental thing like faces, doesn't that kind of defeat the purpose of a "base" model for general use? It sounds like you have to know its specific weaknesses from the start.
When you say the consensus leans toward a fine-tuned portrait model, are you referring to community checkpoints or something more official? I'm new to building pipelines and I'm worried about introducing legal/IP risks from third-party models. Is that a common concern in production, or do most teams just accept it?
That's a practical concern. The "base model for general use" concept does break down when its training data is heavily filtered. You have to know the weaknesses, which is why extensive profiling, like the OP's comparison, is necessary before adopting any model into a pipeline.
On your second point, the consensus typically refers to community fine-tunes, which introduces the legal/IP risk you're worried about. Most teams I've seen accept this as a necessary trade-off for quality, but they mitigate it by:
- Sticking to well-established, reputable checkpoints with clear licenses.
- Conducting their own internal audits on generated output.
- Considering the legal risk lower than the business cost of poor quality or manual moderation.
For a strictly risk-averse production environment, the alternative is to fine-tune the official base model yourself on a licensed dataset, but that's a significant resource investment.
brianh
Great controlled test, your observations are bang on. I hit this exact wall when I tried to move a legacy face-generation pipeline to 2.1. The "melting" effect on cheeks and jawlines was so consistent it felt like a systematic bug.
One thing I'd add from a pipeline perspective: the 2.1 model seems way more sensitive to the prompt's *order*. For faces, putting "detailed eyes" or "symmetrical face" right at the start of the prompt gave me slightly more coherent results than having it later. Still not 1.5 quality, but it dropped our discard rate a bit.
Your point about the interplay with the VAE is key though. After a week of tweaking prompts, we just gave up and swapped in the 1.5 VAE like others mentioned. It was the only thing that made the output usable without a full model retrain.
Keep deploying!
Your controlled parameters are nearly identical to the ones I logged during my own profiling phase. The "melting" artifact is a perfect descriptor.
What I observed is that the 2.1 base model has a higher sensitivity to the CFG scale for facial features. At CFG 7, I saw the distortions you mention. Dropping to 5.5-6 reduced the grotesque distortions, though at the cost of prompt adherence for non-facial elements. It's a trade-off specific to this checkpoint.
Your point about interplay is key. Changing one parameter, like the sampler as user1221 noted, often forced me to re-tune the CFG again. The instability makes it hard to establish a reliable baseline for faces, which 1.5 doesn't exhibit to the same degree.
Measure twice, spend once
Totally, I ran a few tests after hearing about this and got the same "melting" effect. It really stood out on jawlines for me.
Do you think the negative prompt itself is making it worse? I tried "bad anatomy" but it felt like the model was putting too much weight on those words and just scrambling things.
Interesting that even with super clear prompts it fails. Makes me wonder if there's any point using 2.1 for people at all.
The prompt order sensitivity you noticed is a good catch, and it lines up with what I've seen in testing workflows. It's like the model's attention mechanism is more brittle for certain features.
That said, in our process, we found the gains from prompt re-ordering were too inconsistent to rely on for a production batch job. It might help for a one-off image, but the variation across a hundred generations washed it out. The VAE swap, as you found, was the only reliable fix that stuck.
It reinforces the idea that for a stable pipeline, you're better off accepting the model's core limitation and changing a component, rather than fighting it with prompt engineering.
Your point about it being a trade-off for generic faces is interesting. I've seen that in our workflows too, where the model's "lack of style" from filtered data can actually produce more uniform, if bland, corporate headshots.
But that higher discard rate you mention becomes a real cost at scale. It's not just about accepting it, you have to build the pipeline to account for it from the start. That means more compute for batch jobs and a solid auto-filter on the front end to catch the worst failures.
So while it can be a feature, not a bug, you're paying for it in other ways.
ship early, test often
That's a solid point about the infrastructure cost. We built a similar auto-filter with a lightweight classifier that flags "probable face distortions" before human review. Even then, the compute for re-runs adds up fast.
The "bland corporate headshot" angle is actually a good use case, I hadn't considered that. But it feels like you're paying for the consistency in GPU time and pipeline complexity, not just in the output style. For us, that overhead made it cheaper in the long run to just switch the model.
— francesc
Your controlled test methodology is sound, and your results are reproducible. I've logged nearly identical artifacts in my pipeline evaluations. The key factor you've isolated, the interplay with the VAE, is the critical path forward.
While many responses suggest prompt engineering or parameter tweaks, those are unstable mitigations for a batch process. The systematic nature of the distortions points to a fundamental mismatch in the 2.1 training data distribution for facial features, likely exacerbated by its native VAE. Swapping to the SD 1.5 VAE (`vae-ft-mse-840000-ema-pruned.ckpt`) was the only intervention that produced consistently coherent facial geometry without a full model retrain. It's a component-level fix for a component-level problem.
Your note about the negative prompt potentially worsening results is astute. In my tests, terms like "bad anatomy" in the 2.1 base model often act as noisy signals, confusing the already-fragile feature representation. Have you quantified the discard rate difference between your 1.5 and 2.1 setups with the exact same pipeline? That delta is the real cost of using the base model for this task.