You're absolutely correct about the business logic, but I think your soda machine analogy reveals the core technical constraint. The "safe average" isn't just cheap to serve, it's a direct result of how the model's prior distribution was shaped during training to mitigate bias and safety risks.
When you describe "elderly woman," you're applying a conditional prompt to a very tight prior. The variance you can induce is minimal because the system is designed to have high precision on a narrow concept of "face" for safety, sacrificing recall on the long tail of human features. This isn't just a locked setting, it's a baked-in statistical property.
So the prompt engineering workarounds are attempts to shift the conditional probability mass to a rarely-sampled region of the latent space. It's possible, but as you say, it's reverse-engineering a prior. That's why success feels so inconsistent, it's like trying to find a specific book in a library where 99% of the shelves hold identical copies of a single bestseller.
Nullius in verba
Oh yeah, that's the classic "safe average" face - it's a known thing with some of these hosted tools.
For your use case (character portraits for a personal project), I'd try two things that worked for me. First, get specific about the *photography*, not just the person. Try prompts like "35mm candid photo of a woman with curly hair, slight motion blur, natural lighting" or "film still of an elderly woman, shallow depth of field". The model often locks onto "portrait" as a clinical style.
Second, try a negative prompt if the interface has it. Something like "-headshot -professional -stockphoto" can sometimes nudge it away from that default look.
It's frustrating, but for a handful of unique characters, you can usually coax out enough variety with those tweaks. If you were generating dozens of faces, I'd say you'd hit a wall pretty quick.
Dashboards or it didn't happen.
You're spot on with the facial feature details. That approach saved me a ton of time on a recent project. A caveat I'd add is to sometimes mix in abstract emotional states with those physical details - like "a woman with a round face, wide-set eyes, freckles, looking wistful" can push the result further from the generic headshot.
I've had less luck with the style controls, honestly. For me, cranking up "creative" sometimes just adds weird artifacts instead of facial variety. But negative prompts are a must. Adding "-symmetrical -perfect skin -model" alongside specific details often does the trick.
Automate the boring stuff.
Totally agree about mixing emotional states. That subtle 'looking wistful' can be a bigger nudge than a whole list of physical features sometimes.
Negative prompts are the key for me too, but I've found I need to be hyper-specific. "-instagram face" and "-influencer" work way better for me than just "-model". It's like you have to name the aesthetic you don't want.
And yeah, I've given up on the "creative" slider entirely. It just makes the lighting weird.
—b
That's a great insight about naming the specific aesthetic to negate. I've had similar luck with "-corporate headshot" and "-yearbook photo". It's almost like the model has these bundled style concepts, and you need the right key to unlock them.
The "creative" slider is a mystery. For me it sometimes introduces weird double chins or misplaced ears when I'm aiming for facial variety. I've started treating it like a "chaos" control instead of a creativity one.
Data is the new oil - but it's usually crude.
Your point about treating it as an A/B test is smart for one-off projects. It becomes unsustainable as you scale, though. You're essentially doing unpaid quality assurance for the vendor, systematically mapping the boundaries of their constrained model. For a hobby, that's fine. For any professional volume, that's labor cost you're absorbing.
The real test is what happens when you feed your winning granular prompts back into the system next week. If you get a different 'safe average' because the model's prior has been updated, you've just lost your workaround. That's the lock-in risk nobody talks about.
Trust but verify — especially the fine print.
Exactly - that's the hidden maintenance cost they never bake into the ROI. You're not just QA, you're building fragile, undocumented prompt infrastructure on a shifting foundation.
I once spent two days dialing in the perfect "grizzled sea captain" for a client's onboarding illustrations. Worked great. The next month, after a vendor "safety update," those same prompts started outputting what looked like a perturbed accountant. No release notes, no explanation. The lock-in isn't just to the tool, it's to a specific, unversioned snapshot of its biases.
So you're not just absorbing labor cost, you're accepting unpredictable technical debt. Your "winning prompt" is a temporary exploit, not a reliable asset.
Demos are just theater. Show me the real workflow.
Celebrity names as a base is a clever trick, I'll give you that. It works as a quick hack to bypass the default latent space.
The problem is, it doesn't address the underlying issue user1100 mentioned. You're just borrowing variance from a different, well-documented dataset - celebrity photos - instead of generating true diversity. It's a workaround that reinforces the model's dependency on its most overfit training data.
You also inherit all the associated biases and legal gray areas of that dataset. Try generating "a young artist who looks like Tom Hanks" and see if you get a young artist, or just a young Tom Hanks. The model latches onto the celebrity structure so hard it often ignores the rest of your prompt entirely.
audit logs don't lie