Alright, who else is in the trenches with Leonardo's API? We built a slick integration for dynamic asset generation on our e-comm platform, leaning hard on their photorealism. The image quality is mostly there, but the API latency is now our single biggest bottleneck in the user journey.
We're seeing P95 response times creeping above 12 seconds for a simple image generation call, even with optimized prompts and fixed dimensions. That's before any retries, which happen more often than I'd like. Our conversion tests show a nasty drop-off on any page that calls the API synchronously. Moving it async helps the UI feel faster, but it kills the "instant customization" magic we were selling.
We're running a dedicated instance, so it's not a plan issue. Their support blames "elevated demand" and suggests batching, but our use case is real-time, per-user. Not a fit.
Before I rip this whole thing out and start over, has anyone actually moved workloads off Leonardo successfully? I'm looking at:
- Midjourney's API (when it finally lands)?
- Stable Diffusion via some managed service (like Replicate or Banana) but with a quality hit?
- Something else that doesn't require training our own models from scratch?
The core need is high-quality, consistent, photorealistic output via API, with latency under 3 seconds P95. Willing to pay for performance, obviously.
Or am I dreaming, and this is just the state of gen AI APIs right now?
Data over dogma.
Ugh, I feel your pain. We benchmarked this exact scenario. That P95 over 12 seconds is a conversion killer, full stop.
Your shortlist is solid. From our tests, Midjourney's vibe is different, less photorealistic for product-style assets, so that's a risk. The managed Stable Diffusion route is worth a deep look. We saw Replicate's SDXL offerings get within striking distance of Leonardo's quality for many use cases, with significantly lower latency. The key is prompt engineering - you'll need to invest time there to close the gap.
One alternative you didn't mention: RunwayML. Their Gen-2 has gotten shockingly good, and their API performance was more consistent in our load tests. Might be worth a sprint to prototype.
Did you try any aggressive client-side caching of common generations? It's a band-aid, but it can preserve some of that "instant" feeling while you evaluate a switch.
RunwayML's performance is indeed better, but you're trading one form of lock-in for another, potentially on a faster treadmill. Their pricing is opaque, and their model versioning is aggressive. You'll get that latency win, then find yourself rewriting prompts and retraining workflows every six months when Gen-3 inevitably drops and they start deprecating old endpoints.
The real devil in the details with Replicate isn't just prompt engineering. It's the cold start problem on their auto-scaling containers. Your first request after a lull can still hit 10+ seconds, which is just as damaging to that "instant" feeling. They all have a catch.
Beware of free tiers
Yeah, the cold start on Replicate is brutal and unpredictable. We ended up implementing a "warm-up cron" that pings our primary model endpoint every minute during business hours. It's a ridiculous workaround, but it kept our P95 under 3 seconds for active periods. Feels like paying for compute you aren't using just to avoid a penalty, which sums up this whole space.
You're dead on about the treadmill with RunwayML and others. The real cost isn't just the API call, it's the constant re-engineering of your integration layer and prompt logic every time they decide to push a major version and sunset the old one. That's a massive hidden tax on engineering time.
Automate everything. Twice.
That warm-up cron is a clever, if frustrating, hack. It really does feel like you're paying for idle time just to avoid a performance cliff.
Your point about the "hidden tax on engineering time" is spot on and often the biggest cost. Even if a new provider has better latency today, their roadmap can turn your integration into a permanent maintenance sink. I think that's a strong argument for baking migration agility into your own architecture from the start, even if it's more upfront work. It turns a crisis into a planned project.
Stay curious, stay skeptical.
Oof, 12+ seconds P95 is rough. That warm-up cron hack someone mentioned is painful but might be your only stopgap while you evaluate.
Have you considered treating the API as an unreliable external service in your gitops flow? We set up a pattern where our app checks a config map (managed via git) for the active provider endpoint and API key. Lets us shift traffic between, say, Leonardo and a Replicate SDXL deployment with a merged PR, not a redeploy. Cuts the migration agony way down.
The quality hit with SDXL is real, but maybe you can run a dual-write experiment for a week? Generate with both, log the results, see if the latency win outweighs the quality loss for your use case before a full rewrite.
git push and pray
The gitops config map trick is really clever. I'm taking notes for our own setup.
That dual-write experiment idea though, that's where I get stuck on the specifics. How do you actually run that in production without doubling your costs or messing up your analytics? Do you send the same prompt to both backends silently, pick one to actually serve to the user, and just log the other for comparison?
And for the logging part, are you just storing image outputs and timestamps somewhere, or do you have a manual review process for the quality? I love the concept, I just can't picture the execution without it becoming a huge side project.
Great point about abstracting the provider! That gitops config map trick is basically treating your image API like a feature flag, which is smart.
But for the dual-write, we found the simplest version was to do it silently for a small percentage of traffic, like 5%. We'd still serve the user from our primary provider, but fire off an async call to the new one in the background. Log the outputs and latency to S3, and later we'd run a script to generate a simple collage for side-by-side manual review. It's not a full analytics suite, but it gives you a gut-check on quality trade-offs without a huge engineering lift.
The real trick is picking that sample percentage low enough that the cost is negligible, but high enough to get meaningful data in a week.
That 12-second P95 is a brutal spot to be in, especially for an e-comm flow where you're selling that instant feeling. Been there.
I don't blame you for looking past batching when the magic is in real-time per-user creation. The async band-aid solves the timeout but wrecks the vibe, like you said.
You mentioned not wanting to train your own model, which is totally fair. But have you thought about a hybrid approach as a stepping stone? Something like using a fast, decent-quality model from Replicate or Banana for the initial instant preview, and then queuing a Leonardo job in the background to 'upscale' or refine that image for the final, high-quality asset the user gets emailed or sees in their account later? It keeps the interface snappy but still delivers that photorealism you built the experience around. It adds complexity, sure, but it might let you salvage the core integration while you trial other providers in the background.
hugo
That hybrid approach suggestion from user1533 feels like the most practical way forward to me. Keeps your promise of an "instant" feel, even if the final asset is slightly delayed.
I'm curious though, have you done any internal tests to see if a faster, lower-quality preview would actually satisfy users in the moment, as long as they know a better version is coming? Might be worth a quick hallway test with some mockups before you build it.
Good luck with this, rooting for you
Testing the faster preview with users is a solid idea, but I'd keep it to a smoke test. The psychology changes when they know it's just a mockup versus when it's actually generating in the app.
We tried a similar hybrid flow and the real issue was version drift. If your background job fails or the high-quality model returns something wildly different, you've now promised two assets that don't match. You need a solid reconciliation process and error handling baked in from the start.
It adds complexity, but it's better than a 12-second wait.
Ship it, but test it first
The config map abstraction is a smart architectural move. It treats the volatility of third-party AI APIs as a known risk, which it absolutely is.
However, the dual-write experiment concept needs careful design to avoid confounding variables. Simply logging outputs and timestamps isn't sufficient for a valid quality comparison. You must also log the exact random seed used for generation, otherwise any observed quality differences could be due to randomness, not the model. Most APIs allow passing a seed parameter; if they don't, your experiment's internal validity is compromised.
The logging approach also assumes a manual review process, which introduces its own bias. For a more scalable analysis, you could use automated metrics like CLIP similarity scores against the prompt or a dedicated image quality assessment model, though those come with their own interpretative overhead.
Nullius in verba
That hybrid approach is a very practical middle ground, and your point about version drift is the key one to nail. We've seen teams get stuck in support hell when the 'preview' and 'final' images diverge too much, and users feel misled.
If you go down this route, I'd suggest making the state of the process crystal clear in the UI - something like "Preview generated, final HD version processing" - and building a dead-simple way for a user to flag if the final image is unexpectedly different. It turns a potential bug into a feedback opportunity.
That 5% silent sampling trick is exactly the kind of low-risk approach I'm looking for. My worry is always the logging part becoming a hidden cost monster.
When you log to S3, are you storing the full image output, or just a path to it? If it's the full image, have you run into issues with storage costs or S3 list performance when you go to analyze a week's worth? I'm trying to think of a way to keep it simple but also not create a new data swamp.
That 12-second wait is tough, especially when you're selling that instant feel. I've been testing Stable Diffusion on Replicate for some internal dashboards, and the latency is consistently under 3 seconds. The quality isn't quite Leonardo-level for photorealism, but it's close.
Have you considered using it just for the initial preview? You could trigger the Leonardo job in the background and swap in the final image when it's ready. It might keep the UI feeling snappy.
How much of a quality drop would you say is acceptable for that first preview? Is it more about having *something* fast, or does it need to be near-perfect?