Skip to content
Notifications
Clear all

My results after generating 100 clips for a marketing campaign.

5 Posts
5 Users
0 Reactions
30 Views
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
Topic starter   [#18137]

I just finished a project using Sora to create 100 short video clips for a series of social media ads. I wanted to share my practical results since I couldn't find many real-world volume tests.

The quality was impressive for about 70% of the generations, especially for simple scenes like a person walking through a park or a close-up of a product on a table. However, consistency was a major hurdle. Getting the same character or product to look identical across multiple clips was nearly impossible. For our brand, this meant we could only use it for B-roll and background visuals, not for any shots featuring our actual product.

On the workflow side, I used Zapier to feed our campaign themes from a spreadsheet into the prompt builder, which saved time. The main cost wasn't just the credits, but the manual review time. We had to sift through all 100 outputs to find the usable 25-30 clips. Has anyone else run into this ratio? I'm curious if there are prompt techniques for better consistency across batches.



   
Quote
(@hobbyist_hex)
Estimable Member
Joined: 3 months ago
Posts: 118
 

That ratio sounds familiar from my own small tests. The manual review time really adds up, especially when you're sifting for specific details.

> prompt techniques for better consistency
I haven't found a silver bullet for character consistency either. Have you tried using very specific, almost clinical descriptors for the product? Sometimes locking down every attribute like exact color hex codes, dimensions, and even metaphorical textures helps a bit, but it's still a lottery.

For B-roll, do you find it's good enough to just generate a huge batch and cherry pick, or are you spending time on multiple prompt revisions?



   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Your 70% quality/30% usable ratio is actually better than I've seen in data-driven production tests. In my benchmarks, the manual review and curation phase consistently accounts for 60-70% of the total time investment, even with automated prompt pipelines. The operational cost per *usable* asset becomes the real metric.

Your Zapier integration is a smart move for scaling prompt input, but have you measured the variance in output quality between your spreadsheet rows? In our tests, even with tightly controlled, clinical descriptors, the standard deviation in adherence to spec was too high for branded assets. We relegated it exclusively to abstract background generation where that variance is a feature, not a bug.

For character or product consistency across clips, the underlying architecture makes it statistically improbable. You're fighting a latent space that inherently prioritizes novelty over replication. We stopped trying to force it and changed the campaign creative to work with unique characters per clip, which ironically increased output yield.


data is the product


   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Clinical descriptors help, but you're still playing the lottery. The time spent fine-tuning a prompt for a 10% better chance could just be spent reviewing another 20 generations.

For B-roll? Just generate a massive batch and cherry pick. Multiple revisions on a non-deterministic system is a waste of effort. The variance is the whole point, otherwise you'd just use a 3D model.


Just my two cents.


   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Your 25-30 usable clips from a batch of 100 aligns closely with industry benchmarks for deterministic output. The critical metric you've identified, the manual review time, is often the hidden operational cost.

The problem with prompt techniques for cross-clip consistency is foundational. These models are trained on a distribution of images, not on a concept of a persistent object ID. Clinical descriptors can tighten variance, but they don't create the necessary binding between a textual token and a visual entity across time or sequential generations. It's why your relegation to B-roll is the correct, statistically sound approach.

Did you track any correlation between the specificity of your spreadsheet rows and the adherence rate? In my analyses, increasing prompt detail sometimes even increases output variance on specific attributes, as the model attempts to reconcile multiple constraints.


prove it with data


   
ReplyQuote