I’ve been tasked with evaluating Leonardo AI for our marketing asset pipeline over the last quarter. My team handles everything from blog illustrations and ad variations to social media thumbnails. The pitch was appealing: a cost-effective, high-quality image generator that could integrate into our workflows. After a three-month deep dive, here’s my blunt, data-driven assessment.
**The Good: Performance and Consistency**
* **API Reliability:** Their API uptime was 99.94% for us. Latency was consistently between 1.2-1.8 seconds for a 512x512 generation, which is acceptable for batch jobs.
* **Prompt Adherence:** For straightforward descriptive prompts ("a modern office desk with a laptop and a potted plant, sunny window"), Leonardo outperforms several competitors. The output is predictable, which is critical for operational workflows.
* **Cost Structure:** The pricing is transparent. You pay for tokens, and the consumption is predictable. For high-volume, low-complexity asset generation, it can be cheaper than some alternatives.
**The Bad: Where It Falls Apart in Real Workflows**
* **Lack of True Determinism:** This is a major operational flaw. You cannot seed a generation reliably. We built a pipeline to regenerate updated versions of assets (e.g., change the product color on a model), and the background, composition, and lighting would shift dramatically even with the same prompt and settings. This makes versioning impossible.
* **Fine-Tuning is a Black Box:** Their "fine-tuning" (training your own model) is marketed as a solution for brand consistency. In practice, the results are erratic. You need a massive, perfectly curated dataset (think 100+ images), and even then, the output quality varies wildly. The documentation is vague on what "training strength" or "dataset diversity" actually *does* under the hood.
* **Observability Gaps:** The API and dashboard give you nearly zero insight into *why* an image failed or looked strange. No token consumption breakdown per element, no guidance on prompt conflicts. It's a black box. For ops, this is a deal-breaker. We need logs and metrics, not just a finished image.
**Technical Implementation Notes & Pitfalls**
We attempted to integrate it via their API for an automated banner ad variation system. Here's a snippet of our generation logic and the issue we hit:
```python
# Example of our batch generation call
payload = {
"prompt": "professional photo of a {product} on a {background}, clean studio lighting",
"modelId": "leonardo-1.0",
"width": 1024,
"height": 512,
"num_images": 4,
"guidance_scale": 7,
"seed": 42 # This seed is IGNORED in subsequent identical calls. No consistency.
}
response = requests.post(f"{API_URL}/generations", json=payload, headers=headers)
```
The `seed` parameter does not guarantee reproducible results across API sessions. This meant our A/B test comparisons were invalidated because we couldn't regenerate the "A" variant reliably.
**Verdict:** For casual, one-off image creation, Leonardo is competent and cost-effective. For any serious marketing *operations* workflow requiring consistency, reproducibility, and observability, it is currently unfit. The inability to deterministically reproduce or tweak images creates massive overhead in asset management and version control. We are continuing our evaluation with other platforms that offer more granular control and better logging.
We'll be piloting a different solution next quarter, focusing on systems that provide proper seed reproducibility and detailed generation metadata. The cost savings aren't worth the operational chaos.
—DL
Benchmarks or bust