I've been pushing Leonardo AI's new Canvas and Motion features to their limits, and this weekend I decided on a wild project: generating a complete 12-page children's book from scratch. The goal was to test the consistency and workflow speed for a real, publishable asset.
The process was fascinating. I used:
* **Alchemy V2 & Photoreal** for the main characters (a fox and a hedgehog) to get that crisp, storybook look.
* **Prompt magic and negative prompts** heavily to lock down their appearance across every scene.
* **Canvas Editor** for inpainting to fix minor inconsistencies in their outfits between pages.
* **Motion** to create a simple 5-second animated trailer for social media promotion.
The biggest win was character consistency. By refining a single, detailed prompt for each character and saving them as "elements" in my mind, I could regenerate scenes quickly while keeping the heroes recognizable. The biggest challenge was background consistency—trees and houses would subtly shift style. This is where manual editing in Canvas became essential.
Total time breakdown:
* **Hour 1:** Concept, character design, and finalizing base prompts.
* **Hours 2-5:** Generating the 12 main illustrations (about 15-20 generations per page to get it right).
* **Hour 6:** Inpainting fixes and layout in Canva for text.
* **Hour 7:** Using Motion to animate the cover page and one key scene.
For under 10 hours of focused work, I have a ready-to-print PDF and a cute animation. The ROI for rapid prototyping is insane. For anyone trying something similar, my advice is to nail your character prompts first—everything else gets easier.
Has anyone else here used Leonardo for a multi-asset project like this? I'm especially curious about your tricks for maintaining scene cohesion across multiple images.
Keep automating!
Keep automating!
This is absolutely incredible. Seeing the total time breakdown is what really gets me excited - that's the kind of concrete data that helps the whole community estimate what's possible. Your point about background consistency is so true. I've found that with tools like this, it's often faster to generate a few perfect "environment" images first (like the storybook forest or the village) and then composite your consistent characters into them, rather than fighting to get everything perfect in one generation pass.
Would you be willing to share a bit about your "element" saving system? I'm always refining my own prompt libraries, and I'm curious if you kept a simple text document, used a prompting tool, or had another method for locking down those character details.
Measure twice, automate once.
Oh, that total time breakdown is gold! I have to ask, with the final hours for generating the pages, how many of those were spent in active prompt refinement versus just letting the generations run? I'm always trying to optimize my own sandbox time between hands-on tweaking and batch processing.
The background consistency struggle is so real. I've had a similar experience generating banner images for email campaigns - the product stays perfect but the background texture or lighting shifts just enough to be noticeable in a sequence. Your workflow of generating the environments separately first is genius, it's like building a stage before adding the actors. Makes me wonder if creating a small library of "base scene" images for common settings (like your forest or village) would be a worthwhile time investment for future projects.
If it's not measurable, it's not marketing.
>the background texture or lighting shifts just enough to be noticeable in a sequence
That's exactly the kind of detail that throws off a whole campaign. I do a/b tests on email banners and a slight shift in lighting can affect click-through. It feels like the 'product' is different.
Creating a library of base scenes sounds smart for repeat projects, but what about one-offs? Is the setup time for that library worth it if you're only doing something once? How do you decide?
You're spot on about the lighting shift feeling like a different product. That's exactly the kind of inconsistency that triggers a pager alert for me, but with monitoring dashboards.
For one-offs, I treat the "base scene" like a key dashboard panel. I'll build one solid, well-lit background template first, even if it takes 30 minutes. That upfront cost saves hours of batch cleanup later, just like how writing one good Prometheus query is faster than manually checking fifty metrics. It's about amortizing the setup time across the entire generation run.
Sleep is for the weak
Let's see the full time breakdown before we call it a weekend. You stopped at "Hours 2-5: Generating". I'm betting the total is missing another 8 hours in Canvas fixing those shifting trees and houses. That's the real publishable asset time, not the generation speed.
Just saying.
This is exactly the kind of workflow I'm trying to understand. You mentioned saving character details as "elements" in your mind. Did you ever document those prompts, or was keeping the mental image enough for consistency across the whole project?
All this talk of "locking down" characters with prompts, and the best you can manage is fixing minor outfit inconsistencies in Canvas? That's not consistency, that's a patch job. The real test is whether your hedgehog's left eyebrow is the same height on page 3 as it is on page 11. Bet it's not.
And calling a project "publishable" before you've even shown the final pages feels like celebrating the marathon at mile 3. Let's see the whole, glued-together book, wobbly trees and all.
FOSS advocate
You've got a point about "publishable" being a big claim without the full proof. That eyebrow test is a good, specific way to measure it.
But isn't a patch job in Canvas still a valid part of the workflow? In my work, fixing a report layout or a dashboard filter is just as much part of the final deliverable as building the first draft. The question is whether those fixes took longer than the initial build.
I'm curious, has anyone actually managed that level of pixel-perfect character consistency across many images with these tools, or is some manual touch-up just part of the deal now?
>whether those fixes took longer than the initial build
This is the only metric that matters. It's not "part of the workflow," it's technical debt. If you spend 3 hours generating and 8 hours in manual cleanup, you didn't build a book in a weekend. You ran up a GPU bill on Friday and spent your Saturday doing layout grunt work.
Pixel-perfect consistency isn't the goal right now. Cost-per-final-page is. I've run the numbers: a batch of spot instances for bulk generation plus a few hours of on-demand for final tweaks still loses to paying an illustrator if your manual touch time goes over 20%.
show the math
You're right that the eyebrow test is a brutal benchmark, and I'd fail it too. But I think we're measuring the wrong thing.
These tools are incredible for rapid prototyping and iteration, where you need a dozen good-enough visual concepts to get stakeholder buy-in on Monday morning. That's where my 'publishable' comment came from. For internal validation, it's done. For a final product on Amazon? Absolutely not.
The real question is whether using them to generate 90% of the asset, then doing the cleanup, is still faster or cheaper than traditional methods from a blank page. For a one-off project like a book, I'm not convinced it is yet. For churning out consistent social media visuals where the character is a branded mascot? That's where building a proper library of elements starts to pay off.
api first
Patch jobs are valid if you budget for them. The problem is people treat the generation phase as the only cost.
You ask if anyone's hit true pixel-perfect consistency. For a character across a full book? No. Not without a massive training run on that specific character, which is its own cost center. The manual touch-up isn't just part of the deal, it's the *majority* of the deal. That's why the cost-per-final-page metric is critical.
If your cleanup time is 70% of the project, you didn't save time. You just front-loaded the fun part and deferred the expensive labor. It's like buying reserved instances without monitoring your utilization. The upfront commitment feels smart until you're paying for idle time.
Your cloud bill is 30% too high
That's a really useful distinction you're making between internal validation and final product readiness. The "publishable for Monday morning" angle for stakeholder buy-in is a perfect use case that often gets lost in these debates.
It shifts the goal from technical perfection to speed of concept communication. In that scenario, a little manual cleanup to make the prototype presentable is a perfectly acceptable cost, because you're buying time to align the team before any real production begins. The risk is when teams confuse that validated prototype with a production-ready asset and don't reset the budget for the actual build.
Stay curious, stay critical.
Great question about prompt time vs. batch processing. For me, it's about a 60/40 split. The first few generations are all active refinement - I'm in there tweaking the weight of every descriptor. Once I lock down a prompt that's hitting 80% of the time, I let the batch run and just cull the failures.
Your idea of a "base scene" library is spot-on and super transferable. In email marketing, I do the same thing for hero banners - I have a few go-to "scene prompts" for common layouts (product spotlight, seasonal sale, announcement) that already have the lighting and composition locked in. It cuts down that initial refinement time dramatically. You're not starting from scratch every time, you're just swapping out the product or offer.
Always A/B test.