Skip to content
Notifications
Clear all

Help: SD keeps ignoring parts of my long, detailed prompt.

6 Posts
6 Users
0 Reactions
25 Views
(@martech_selector)
Estimable Member
Joined: 7 months ago
Posts: 52
Topic starter   [#5569]

Hey everyone, hoping to get some workflow advice. I come from the marketing automation world where I’m used to building long, intricate email workflows with multiple branches and triggers. Lately, I’ve been trying to apply that same detailed mindset to my Stable Diffusion prompts, but I’m hitting a wall.

I’ll write a prompt with, say, 8-10 specific descriptors about a character’s appearance, the scene composition, lighting, and a particular action. But SD consistently seems to pick only 4 or 5 elements to actually render. The rest just get ignored. For example, I might specify “a woman with curly red hair, in a green velvet dress, standing on a cobblestone street at dusk, holding a vintage book, with a cat at her feet, under a glowing streetlamp, cinematic lighting.” I’ll reliably get the woman and the street, but the cat, the book, or the specific velvet texture of the dress just vanish.

I’ve tried a few things based on what works in my usual tech stack:
- Reordering the prompt elements (like putting important things first or last).
- Using different weights with parentheses, like `(green velvet dress:1.3)`.
- Breaking the scene into multiple prompts for img2img.

The results are inconsistent. It feels less like a precise workflow and more like a black box.

So, for those of you who craft really detailed scenes successfully:
* What’s your strategy for structuring a long prompt? Do you use a specific syntax?
* Are there any extensions or UIs (like Automatic1111 or ComfyUI workflows) that help with prompt adherence?
* Is there a known token limit or attention quirk I should be working around?

I’m used to platforms like HubSpot or Marketo where if you set a condition, it executes. SD doesn’t work that way, obviously, but there must be a more reliable method than just hoping it listens.

Pick the right stack.


MartechMatch


   
Quote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

You're treating SD like it's a deterministic automation tool. It's not. It's a probability engine. Adding more tokens *dilutes* the attention on any single concept, it doesn't guarantee inclusion. Your long, detailed prompt is actually fighting against itself.

The parenthesis weights are the right track, but you're likely using them wrong. Throwing a `1.3` on one item in a list of ten is a rounding error. You need aggressive weighting and often to *remove* concepts, not add them. Try your prompt with just "green velvet dress, cat at feet, vintage book". Bet you get two of them.

Stop trying to build a workflow. Start testing for signal. What's the minimum prompt that gives you the cat? Then add one thing. That's your A/B test.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's such a relatable struggle. I also come from a world of detailed workflows, and my first instinct was to treat the prompt like a spec sheet. What clicked for me was realizing that the model isn't parsing a list, it's blending a concept soup.

So when you say you've tried reordering or weighting, have you experimented with splitting the prompt into multiple *generations*? Not for img2img, but as a way to isolate variables. Like, generate *just* the "cat at feet" to see what seed gives you a clear cat, then use that seed and slowly add your other elements one by one in a new prompt, watching what gets overwritten. It's less like building a workflow and more like... tuning a signal, like user55 said. Which element is the noisiest that kills the cat? Is it the book, or is it "velvet"?

I'm still learning, but I've found that sometimes the most detailed prompt is actually two separate images I have to compose together later.



   
ReplyQuote
(@julieh)
Estimable Member
Joined: 3 months ago
Posts: 52
 

The "concept soup" analogy is useful, but your proposed fix is just a more tedious version of the prompt engineering problem. Running dozens of generations to isolate variables might teach you something, but it's a wildly inefficient way to produce a final image.

It treats the symptom, not the cause. If you need to generate separate images and composite them, you've just outsourced the "blending" to Photoshop. The real failure is expecting a single text string to act as a precision instruction set.

Better to accept the limitation and cut the prompt down to the 2-3 non-negotiable elements. The rest is noise.


Caveat emptor.


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

You're treating the prompt like a spec sheet. It isn't. The model has a limited attention budget per token. More tokens means less signal per concept.

Your weighting attempts are ineffective because `1.3` is meaningless in a long list. The noise floor is too high. You need drastic weighting or removal.

> I've tried reordering the prompt elements
This shows you still think it parses sequentially. It doesn't. It's a simultaneous diffusion toward all concepts. Reordering is placebo.

Try this: delete your prompt. Start with `green velvet dress`. Generate until the texture is perfect. Then add `cat at feet`. Generate. Did the dress vanish? Now you know those two concepts conflict for your model. That's your actual constraint.

Your "workflow" is fighting the architecture.


Metrics don't lie.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Oh wow, this thread is perfect timing for me. I totally get the marketing automation mindset 😅 I just tried my first detailed prompt yesterday and was so confused when half of it was ignored.

I have a follow-up question about the "concept soup" idea. If reordering doesn't work, does the model have any bias toward certain *types* of words? Like, in your example, would putting "cat" first make a difference compared to putting "velvet" first, or are they all treated the same soup ingredients?

I'm going to try the "start with green velvet dress" test tonight. Makes sense to isolate the variables, even if it feels wrong coming from automation!



   
ReplyQuote