Okay, so I've been deep in NightCafe for a few months now, mostly using it to generate abstract art for presentation decks and blog headers. But recently, I decided to really push for photorealistic human portraits. You know, the kind that don't look like a wax museum nightmare or have that weird, soulless stare.
The big unlock for me wasn't just the model choice (though that's crucial). It was treating the prompt like a detailed CRM data record. Instead of just "a woman smiling," I build a full profile. I specify the lighting source and quality ("studio lighting with softbox, subtle rim light"), the lens type ("85mm portrait lens, f/1.8, shallow depth of field"), and even the environment ("neutral gray backdrop, studio setting"). This is like ensuring your lead source and campaign details are clean—it gives the AI the right context to work with.
I've found that **NightCafe's "Stable Diffusion" model** with the "Realistic Vision" VAE is my go-to starting point. The real magic happens in the negative prompt. This is where you banish the uncanny valley. My standard negative prompt includes: `disfigured, deformed, plastic, doll-like, mannequin, wax figure, unnatural skin texture, asymmetric eyes, blurry, out of focus, bad anatomy`. It's a deny list for creepy outcomes, similar to how we set validation rules to block bad data entry.
Finally, don't sleep on the **"Portrait"** preset under Advanced Settings. It adjusts the aspect ratio and some under-the-hood weights perfectly for faces. And just like iterating on a sales workflow, generation is iterative. I'll generate a batch, pick the one with the best base structure, then use that as an init image for another pass, maybe tweaking the prompt to add "detailed pores, natural skin texture, subtle eye moisture." That second pass often brings in the convincing human details. It's a process, but the results are getting scarily good.
Absolutely love the CRM data record analogy. It's spot on - you're giving the system structured fields to populate instead of a vague text search.
I'd add that for the negative prompt, I've had good luck with very specific technical flaws. Things like "chromatic aberration", "lens flare", "motion blur", and "poor focus" seem to steer it away from that cheap CGI look. Also, mentioning "subsurface scattering" in the positive prompt can sometimes help with that natural skin glow.
What's your experience with the CFG scale on these? I find I have to crank it down lower than usual for portraits, maybe around 5-T, to avoid that over-processed, hyper-detailed stiffness.
api first