That point about hardware routing is really interesting, I hadn't considered that. So even if they keep the model version the same, my API request could land on different physical chips from one minute to the next?
It makes sense from their side, using a pool for efficiency. But it's a nightmare if you're building a pipeline that depends on a consistent result. The self-hosted server solution sounds like the only real fix, but that's a huge jump in complexity from just calling an API. I'm not there yet, skill-wise. 😅
So for someone using a service like Replicate or Fal, is there any way to even check if that's happening? Or is it totally opaque?
rookie
Your point about dynamic hardware routing is a critical, often undocumented variable. I've seen this happen even within a single data center, where a service's autoscaling group mixes different generations of the same GPU family, like A100 40GB and A100 80GB cards. The underlying memory bandwidth difference was enough to introduce floating point divergence after thousands of operations.
The move to self-hosting is the nuclear option, but even then, you trade one set of variables for another. You have to start pinning specific CUDA driver versions and disabling GPU power management features that can cause clock speed throttling. It's a deep rabbit hole.
—at
Hold up. You say you're using "identical generation parameters" but your payload snippet is literally missing them. The seed is just one variable. Where's your steps, guidance scale, and scheduler in that JSON?
If they're not there, the API fills them with defaults. Those defaults are the first place variance creeps in, even before you worry about hardware. You've gotta lock the whole thing down, not just the seed.
Trust but verify.
You're right that my payload snippet is missing those parameters. It's a bad habit from testing in the UI where those fields have defaults, and I assumed the API would lock them in if I didn't provide them.
Looking at the other replies, it seems like even if I add `steps`, `guidance_scale`, and `scheduler`, I'm still up against potential hardware variance in a shared API service. For my documentation pipeline, I don't think I need pixel-perfect matches, but the two test images I got were different enough that a logo was missing in one. That's a functional problem, not just a visual variance.
So, if I fully specify the payload and still get meaningful differences, is that a sign the service is fundamentally non-deterministic for my use case, and I should look elsewhere? Or is there a way to request more consistent hardware routing, maybe through a different API tier?