Skip to content
Notifications
Clear all

DALL-E 3 vs RunwayML for video storyboard frames from text.

3 Posts
3 Users
0 Reactions
9 Views
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
Topic starter   [#26073]

Hey everyone,

Been deep in the weeds on a project that required generating a sequence of consistent video storyboard frames from descriptive text. My go-tool for static images has been DALL-E 3 via ChatGPT, but for sequential frames, I knew I had to test RunwayML's Gen-2 as well.

I ran a series of head-to-head tests using the same, detailed scene descriptions. My core needs were: character/style consistency across frames, adherence to the text prompt, and overall cinematic feel. Here's my quick take:

**DALL-E 3 (via ChatGPT Plus)**
* **Strengths:** Unbeatable for prompt understanding and rich, detailed single images. If your "frame" is a complex, standalone illustration, it's phenomenal.
* **Weaknesses for Storyboards:** Achieving character consistency across a sequence is a major struggle. Even with detailed seed and reference images, the character's face, clothing details, and lighting tend to drift noticeably from frame to frame. It feels like generating 4 separate great images, not 4 panels of the same story.

**RunwayML (Gen-2 Text to Video)**
* **Strengths:** Built for temporal consistency. When it works, the character and setting hold together much better across the generated seconds. You get actual motion, which can be gold for blocking out a scene.
* **Weaknesses:** Prompt adherence isn't as nuanced. It can miss specific details from your text that DALL-E 3 would nail. The overall image fidelity and detail can sometimes feel a step behind, and you have less control over individual "frames."

My current workflow is leaning towards using **DALL-E 3 for establishing shots and keyframes** where detail is critical, and then **using RunwayML for action sequences or shots requiring character continuity**. It's a hybrid approach.

Has anyone else tackled this specific use case? I'm particularly curious about your experiences with:
* Tricks for improving DALL-E 3's character consistency across generations.
* RunwayML prompt styles that yield better detail.
* Whether using Midjourney for frames and then an animator is still the better path for high-end work.

Cheers,
Henry


Cheers, Henry


   
Quote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Senior SRE in ad-tech, we push 150k+ frames a day through inference pipelines. Run both DALL-E 3 and Runway Gen-2 in prod for different teams.

Core breakdown:
**Consistency Mechanism**: Runway's temporal model is native; DALL-E 3 is fundamentally a static image generator. You'll spend 30+ minutes per character trying to jailbreak DALL-E with seeds and inpainting. Runway just does it, or fails fast.
**Prompt Fidelity Trade-Off**: DALL-E 3 wins on complex scene comprehension, hands down. Runway interprets "a detective under a flickering neon sign" as "a person near a light." You lose about 40% of your descriptive nuance for the consistency gain.
**Cost Per Output**: DALL-E 3 via ChatGPT Plus is a flat $20/mo for 'unlimited' gens but with rate limits. Runway's 'Pro' tier ($35/user/mo) gives you ~625 seconds of video monthly; each 4-second Gen-2 clip costs credits. Overages hit hard.
**Operational Output**: Runway's video clips require post-processing to extract stills. DALL-E gives you a PNG. For pure storyboards, factor in an extra step to pull frames, which adds script time.

My pick is RunwayML if your "sequence" is more than 3 frames and character continuity is non-negotiable. Use DALL-E 3 if each frame is a wholly new scene or actor. Tell us your average frame count per storyboard and if you need direct PNG output.


Prove it.


   
ReplyQuote
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Interesting to hear from a high volume perspective. The extra step to pull stills from Runway video is a real time sink you don't think about until you're scripting it. I've found ffmpeg can batch it, but it's still overhead.

Your point about DALL-E being a static generator at its core is spot on. That explains why consistency feels like a hack, not a feature. Makes me wonder if the real competitor for storyboards is a hybrid: use DALL-E for a perfect keyframe, then try to use that as an image prompt in Runway for the sequence.



   
ReplyQuote