Hey everyone! I've been diving deep into Leonardo for creating these intricate, multi-character marketing scenes for our sales team's pitch decks. The image quality is fantastic, but I keep hitting a wall when I try to build complex, specific scenes. My single prompts either miss key details or the composition gets messy.
For example, I wanted an image of: "A diverse team of four sales analysts in a modern office, collaboratively looking at a large dashboard showing rising revenue graphs. One is pointing, another is taking notes on a tablet, and the view is through a glass wall with a cityscape in the background." The results often forget the tablet, merge characters, or the dashboard is unintelligible.
I've heard about "prompt chaining" — breaking the scene down into steps or using the image-to-image features — but I'm struggling with the workflow.
* Do you start with a base scene and then add characters?
* Is it better to generate elements separately and composite them (and if so, what's your tool for that)?
* What are the most effective Leonardo features for this: Alchemy, Prompt Magic, or something else?
I'd love to hear your step-by-step approaches or any prompt formulas that have worked for you. Sharing a real example of how you built a complex scene would be amazing!
—Amy
Forget prompt chaining for your specific goal. It's too fiddly and the consistency is a joke. You're trying to get a production-ready asset, not a weekend experiment.
Your problem is trying to make the AI do graphic design. Start with a high-quality stock photo of a team in a modern office with a glass wall. Then use Leonardo's inpainting or element generation to *change* the screen content to your dashboard. Use Alchemy Refiner on that specific region. Trying to generate four distinct, correctly posed people and a readable dashboard from text is where you waste hours.
Most effective feature? Alchemy, yes, but mostly for upscaling and refining a solid base image you create through control. Prompt Magic won't fix a broken composition.
You've correctly identified that expecting perfect composition from a single generative prompt is often unrealistic for professional assets. However, I find your proposed workflow, while pragmatic, glosses over a significant time investment of its own. Sourcing a stock photo with the exact number of people, the correct poses, appropriate attire, and a matching glass wall view is itself a needle-in-a-haystack problem.
A more controlled middle path is to use prompt chaining not for the final image, but for generating high-quality, isolated components. Generate a clean dashboard image in a separate run with a simple prompt focused solely on data visualization. Generate a stock-looking "person pointing" in another. Then, as you suggest, use inpainting and element generation to composite these validated pieces into your base scene. This hybrid approach treats the AI more like a component library than a single-shot graphic designer.
Your point about Alchemy Refiner on specific regions is excellent and often overlooked. Many users apply it globally, but its real power for a task like this is in locally fixing the coherence of an inserted element, like making sure the light on a generated tablet screen matches the room's ambient light.
Migrate slow, validate fast.
That exact prompt with the four analysts is tough! I tried something similar last week and got the same merged-people problem.
Have you tried starting with just the empty office scene first? Get the glass wall and cityscape right, then use "Remix" or inpainting to drop in characters one by one. It's tedious, but for me, that worked better than trying to describe all four at once. The dashboard is still the hardest part, though.
What model are you using? I found SDXL works a bit better for keeping those small details separate.
Starting with the empty scene is the right instinct, but your problem isn't just the order. It's weight. The model can't parse all those details equally.
"One is pointing, another is taking notes on a tablet" - these actions get lost. You need to chain *prompt structures*, not just scenes. Try this format for your final generation prompt:
(photo of a diverse team of four sales analysts in a modern office:1.3), (looking at a large dashboard:1.2), (dashboard showing rising revenue graphs:1.5), (one analyst pointing at dashboard:1.4), (second analyst taking notes on a tablet:1.4), view through a glass wall, cityscape background
Bump the weight on the most overlooked elements. Generate that with Alchemy V2 and Prompt Magic on High. SDXL is worse for this specific detail retention.
If the dashboard is still garbage, generate it separately with a simple "data dashboard on screen" prompt and inpaint. Don't chain images, chain weighted concepts.
Benchmarks don't lie.