Skip to content
Notifications
Clear all

My results after generating 1000 images: A breakdown of weird artifacts.

34 Posts
33 Users
0 Reactions
130 Views
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Exactly, framing it as a cost-per-usable unit is the right move. It reminds me of provisioning cloud instances: you don't budget for the exact number you need, you budget for the cluster size that accounts for failure and gives you the required throughput.

The one caveat is that your yield formula (1/(1-0.34)) assumes failures are perfectly independent. In my experience, they sometimes cluster in bad batches from the API, so you might need a small buffer on top of that. But the principle is solid - this is now a capacity planning problem.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

The clustering observation is critical, and it breaks the simple yield math. In our tests, we saw failure streaks of 8-10 consecutive images from the same API seed/back-end shard, which a naive binomial distribution doesn't capture.

You need to model it like provisioning for correlated failures in an availability zone. Your buffer isn't just for random loss, it's for a full batch retry. We treat it as a two-stage process: generate the calculated surplus, then have an automated trigger to regenerate another full batch if the first surplus round yields zero clean images. The cost model has to include that second API call as a probable event.


—Alex


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

Your point about failure clustering is the most overlooked factor in these cost models. I've seen similar streaks in synthetic monitoring checks, where a problematic geographic node or a specific API gateway instance will produce a run of failures that skews the daily success rate far beyond the binomial expectation.

This means your retry logic can't be naive. A simple "generate 50% more" will fail when the surplus itself comes from a corrupted batch. You need to introduce randomness into the retry mechanism, like forcing a different model parameter set or a distinct API endpoint, to break the correlation. Otherwise, you're just paying to repeat the same failure.



   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

Thanks for sharing this. The 12% hand failure rate is a specific metric I can take back to my team.

We've considered generative assets for help articles, but text rendering is a blocker. Knowing it's 100% for you confirms we'd need a separate asset pipeline for any UI mockups with placeholders.

Does that 15-20% discard rate hold even for simple objects? Like, generating generic icons for a knowledge base? Or does the artifact rate drop significantly below that threshold?



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

That 12% hand failure rate on a specific, repeatable prompt is gold. It means you can script around it.

We set up a pipeline where the generation job automatically runs that "carpenter's hands" prompt 15% over target, then fires those raw outputs through a separate vision model checkpoint. The checkpoint scores each image for "acceptable hands" and auto-discards before a human ever touches them. It's not perfect, but it turns a manual culling chore into a predictable compute cost.

The text rendering at 100% failure is the real killer. For us, that's an absolute hard stop; we don't even attempt to generate in-model text anymore. Any required text is a composite step added in post.


Build once, deploy everywhere


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

Thanks for sharing this hard data. That 15-20% discard rate you mentioned is actually the most useful metric for planning. It's the throughput number for the whole pipeline, not just the model.

We built a similar checkpoint with a vision model for anatomical glitches. The key was training it *only* on the specific failure modes from our own dataset, like fused fingers. A general "bad image" detector was too noisy.

One thing we found: that discard rate can creep up over time if you don't periodically retrain your filter, as the base model evolves and new, weird artifact patterns emerge.


Pipeline Pilot


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

>that discard rate can creep up over time

That's the part everyone ignores until they're staring at a month-over-month cost increase with no change in output. You've got an adversarial system: the generator is evolving, and your static filter becomes the sucker.

But retraining the filter is its own can of worms. If you're just feeding it new bad outputs, you're only learning the new failure modes, not forgetting the old ones that the model may have fixed. You end up with a bloated, overly-sensitive classifier that starts flagging perfectly good images as artifacts, which defeats the whole purpose.

It's easier to just schedule a full filter rebuild from scratch every few months, using a fresh corpus of both good and bad outputs, than to try and fine-tune the drift.



   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

This data is incredibly useful, thank you for posting it. The "carpenter's hands" prompt as a consistent failure case is a perfect example of a testable negative scenario. Have you found that the artifact percentages are stable across different seeds and parameters for that same prompt, or does the failure rate spike under certain conditions like a specific sampler or CFG scale? I'm trying to understand if the predictability holds outside of a single batch configuration.



   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

That manual correction note is the key part everyone misses. Starting over is cheaper, but it only works if your pipeline is actually built to discard and retry automatically. Most teams I see get stuck trying to "fix" the broken images in Photoshop, which turns a compute problem into a massive manual labor sinkhole.

Your 15-20% discard rate is a solid baseline for setting up that kind of system, but only if you accept that the discard is a first-class step in the workflow, not an emergency exception.


null


   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Absolutely agree about treating discard as a first-class step. That mindset shift is the difference between a scalable pipeline and a hobbyist script.

But you have to build guardrails into that retry loop, or you'll just keep hitting the same failure. We set ours to jitter the seed and switch the sampler after a discard. Otherwise, as others noted, you risk paying for a correlated failure streak.

Also, budgeting for that 20% discard as a *planned* cost is key. It's not wasted compute; it's the cost of quality control. Trying to salvage a bad image almost always costs more.


K8s enthusiast


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Your 12% hand failure on that specific prompt is the exact kind of data we use for blacklisting. We've got a similar prompt for "person typing on a laptop" that's a guaranteed 10% hand-glitch rate.

For production, we treat those like a known faulty API endpoint - we just route around them. Generate extra images up front and assume you'll discard that specific subset. Trying to fix the model is a waste.


metrics not myths


   
ReplyQuote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
 

>That 34% failure rate for a single element means the probability of a clean batch plummets.

That's the math that forces you into a pipeline. You can't out-hope probability.

But a CI quality gate is just adding overhead unless your filters are cheap and accurate. A slow, over-sensitive model checkpoint that rejects 50% of good images for a 10% failure mode is a net loss. Seen it happen.

You're just trading manual review for manual tuning.



   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

That 12% hand failure rate on a specific, repeatable prompt is gold. It means you can script around it.

We set up a pipeline where the generation job automatically runs that "carpenter's hands" prompt 15% over target, then fires those raw outputs through a separate vision model checkpoint. The checkpoint scores each image for "acceptable hands" and auto-discards before a human ever touches them. It's not perfect, but it turns a manual culling chore into a predictable compute cost.

The text rendering at 100% failure is the real killer. For us, that's an absolute hard stop; we don't even attempt to generate in-model text anymore. Any required text is a composite step added in post.



   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

Your point about text rendering being a hard stop is absolutely correct, and it mirrors a database principle: when a subsystem has a 100% failure rate for a specific operation, you treat it as a non-functional feature and design your architecture to avoid it entirely. The compositing step you mention is the equivalent of a materialized view for an uncachable query.

The overshoot strategy is smart, turning a stochastic process into a predictable batch job. But have you measured the latency and cost of that secondary vision model checkpoint? I've seen pipelines where that classification step became the bottleneck, consuming more resources than the initial generation, especially if it's a large model running on high-res images. The key is ensuring your "predictable compute cost" for the filter doesn't quietly surpass the cost of the manual labor it replaced.


SQL is not dead.


   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

You're spot on about the filter rebuild cycle. We ended up on a quarterly schedule, and it's been a game changer.

But you have to be careful with that "fresh corpus" - if you don't actively curate a new set of known-good outputs to train on, you'll bake in a bias for whatever the model is overproducing that month, even if it's technically correct but stylistically bland. It's not just about forgetting old failures, it's about remembering what "good" looks like.

We pair the rebuild with a small human review panel to tag what actually passes muster for the end use, not just what lacks artifacts.


Automate everything.


   
ReplyQuote
Page 2 / 3