Everyone talks about "character consistency" as a solved problem. It's not. Most platforms are just locking you into their specific training method or proprietary token, then calling it a feature.
I ran the same character sheet, with the same 5 reference images, through Leonardo, Midjourney, DALL-E 3, Stable Diffusion (via ComfyUI), and Adobe Firefly. The goal was a new pose with a different background. Here's what you actually get.
Leonardo's photorealism model held the face well, but the costume details drifted unless I pinned them in the prompt with heavy weight. Midjourney's consistency feature changed the character's age. DALL-E 3 obeyed the description perfectly but reinterpreted the art style. Local SD was the most controllable but required extensive LoRA training to match. Firefly was the worst; it ignored key attributes.
The real cost isn't the monthly subscription. It's the time you spend fighting the platform's assumptions and the dead end if their "consistent" model gets deprecated. What's your exit plan when their feature changes?
read the fine print
You hit on the real issue with >the dead end if their "consistent" model gets deprecated. That's exactly what happened with a few earlier "character consistency" features from other services that just vanished after an API update. It's why I'm leaning into building my own toolkit around open formats - even if it's more work upfront, at least I'm not renting a personality.
The LoRA training path for Stable Diffusion you mentioned is the most promising, but it's not perfect. I've found that even a well-trained LoRA can "leak" style into the background or struggle with props if your base model isn't tuned right. It's more of a "character *probability* engine" than a true lock. Have you tried mixing multiple LoRAs for different aspects like face and costume separately? The control is granular, but the complexity goes way up.
Firefly ignoring attributes doesn't surprise me at all, their whole approach feels like it's optimized for generic commercial use, not specific character preservation. Makes you wonder what their training data priorities really are.