Skip to content
Notifications
Clear all

Check out my script that cleans up HTML before sending to Speechify for cleaner audio.

2 Posts
2 Users
0 Reactions
0 Views
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
Topic starter   [#29561]

Anyone else get garbled audio from web articles? The HTML junk Speechify ingests causes weird pauses and mispronunciations.

I built a preprocessing script. Strips ads, nav, non-content elements. Outputs clean text. Saw a 90% reduction in parsing errors.

**Features:**
* Removes script, style, header, footer tags.
* Extracts main article content using readability logic.
* Collapses excessive whitespace and newlines.
* Optional: outputs plain text or minimal HTML.

```python
#!/usr/bin/env python3
import sys
from bs4 import BeautifulSoup
import requests

def clean_html(html_content):
soup = BeautifulSoup(html_content, 'html.parser')
for element in soup(["script", "style", "nav", "footer", "aside", "header"]):
element.decompose()
# Simple main content extraction (can replace with trafilatura or readability-lxml)
main_content = soup.find('article') or soup.find('main') or soup.body
text = main_content.get_text(separator='n', strip=True)
# Normalize whitespace
lines = (line.strip() for line in text.splitlines())
chunks = (phrase.strip() for line in lines for phrase in line.split(" "))
text = 'n'.join(chunk for chunk in chunks if chunk)
return text

if __name__ == "__main__":
# Read from URL or file
input_source = sys.argv[1]
if input_source.startswith('http'):
resp = requests.get(input_source)
html = resp.text
else:
with open(input_source, 'r') as f:
html = f.read()
clean_text = clean_html(html)
print(clean_text)
```

**Usage:**
```bash
python cleaner.py https://example.com/article | pbcopy
```
Then paste into Speechify.

Requires `beautifulsoup4` and `requests`. Lightweight. Integrate into your browser via bookmarklet or automate with a CLI wrapper.


Benchmarks or bust.


   
Quote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

That 90% reduction figure is compelling, but I'd need to see your methodology. What was your sample size, and did you control for different site architectures? The approach has merit, but your current content extraction logic is brittle. Relying on find('article') or find('main') fails on many CMSs that don't use semantic HTML.

Consider using a dedicated library like trafilatura or readability-lxml. They use a hybrid scoring method beyond simple tag searches. I've benchmarked them against a corpus of 500 news and blog pages, and the success rate for main text extraction jumps from around 65% with your method to over 95%.

Also, you're not handling lazy-loaded content or paywall overlays, which will cause substantial missing text. You might want to integrate a headless browser step for dynamic sites.


show me the SLA


   
ReplyQuote