I've seen enough "magic AI" demos that work on stock photos and fail on real product catalogs. My team's budget isn't for playing with shiny toys. We need to generate consistent, on-brand marketing assets, and for that, we need a model that understands our specific products.
I'm evaluating Leonardo for this. The documentation talks about fine-tuning and custom models, but I need the blunt, practical steps from someone who's actually done it.
My scenario:
* We have ~500 high-quality, studio-shot product photos (various angles, on white background).
* Our products have distinct, consistent design features (think a specific logo placement, unique shape).
* Goal: Generate new background scenes, lifestyle contexts, or slight variations while keeping the product itself 100% accurate.
What I need to know:
* **Dataset Prep:** What's the *real* minimum number of images that actually works? Is 500 enough, or am I wasting my time? Do I need to manually tag/caption every single image with specific keywords?
* **Training Process in Leonardo:** Which training option do I use? "Fine-tune" vs. "Custom Model" – what's the operational difference for my use case? What settings (epochs, resolution) gave you usable, non-overfitted results?
* **Output Reality Check:** After training, does the model *actually* retain precise product details, or does it just generate "something similar"? If I ask for "our product on a beach," will the logo become a blurry mess?
* **Pitfalls:** What are the common failure modes? Does it start generating products with melted parts or extra limbs if you overtrain?
I don't need theory. I need a workflow that delivers a model I can plug into a pipeline without manual cleanup for every image.