Our team was burning through our WellSaid Labs word quota too fast. The problem wasn't the voice generation—it was our scripts. We analyzed 3 months of audio and found 30% of generated words were unnecessary and edited out post-generation.
Main issue: writing long-form for audio. We now edit scripts *before* generation with two rules.
**Rule 1: Remove filler phrases and redundancy.**
Original script draft:
> "So, at this point in time, what we're going to do is we're going to click on the 'Submit' button in order to proceed."
Edited for TTS:
> "Now, click 'Submit' to proceed."
**Rule 2: Use a pre-generation regex cleaner.**
We run all scripts through a simple Python script that cuts waste. It targets common verbose patterns.
```python
import re
def clean_script(text):
# Remove "you know", "I mean", "kind of", "actually"
text = re.sub(r's(you know|I mean|kind of|actually),?s', ' ', text, flags=re.IGNORECASE)
# Reduce "in order to" to "to"
text = re.sub(r'sin order tos', ' to ', text, flags=re.IGNORECASE)
# Replace "at this point in time" with "now"
text = re.sub(r'bat this point in timeb', 'now', text, flags=re.IGNORECASE)
# Trim multiple spaces
text = re.sub(r's+', ' ', text)
return text.strip()
# Example usage
raw_script = "So, actually, in order to start, you know, we need to click the button."
clean = clean_script(raw_script) # Output: "So, to start, we need to click the button."
```
Results:
* Average words per script down 30%.
* Same information density.
* Audio is clearer and faster.
Key takeaway: Optimize text *before* it hits the TTS API. The ROI on script editing is huge.
- bench_beast
Benchmarks don't lie.