Skip to content
Notifications
Clear all

How do I get a realistic old-person voice for a historical documentary?

2 Posts
2 Users
0 Reactions
11 Views
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
Topic starter   [#27032]

I'm currently in the pre-production phase for a documentary series focusing on oral histories from the early 20th century, and the voiceover narration needs to sound authentically aged. We're committed to using Murf.ai for its API integration capabilities with our editing pipeline, but the standard "elderly" voice presets lack the specific, subtle qualities we're after. They often sound like a young person impersonating age, missing the authentic vocal fry, slight breathiness, and nuanced pacing of a genuine octogenarian.

My primary technical question is this: beyond simply selecting an "Old Man" or "Senior Female" voice, what specific combination of Murf's advanced settings and phonetic adjustments have you found to yield the most realistic, non-parody-like aged vocal quality? I am particularly interested in reproducible parameter sets.

From my initial load testing of the API with various configurations, I've isolated a few key parameters that seem to influence the perception of age, but my results are inconsistent. The core challenge is simulating the physiological aspects of an aged larynx without crossing into caricature.

Here is my baseline test configuration for the `murf.api.v1.speech` endpoint that I've been iterating on:

```json
{
"voice_id": "en_uk_male_03",
"text": "Sample historical narration text.",
"settings": {
"speed": 0.85,
"pitch": -6,
"pause": "medium",
"emphasis": {
"level": "low",
"strategy": "distributed"
}
}
}
```

**I am seeking detailed feedback on the following:**

* **Pitch & Stability:** Should the pitch parameter be lowered uniformly, or is a more realistic effect achieved by introducing minor, irregular pitch variations (if possible) to simulate vocal cord weakness?
* **Speed & Pause Ratio:** My data suggests reducing speed is essential, but the relationship between words-per-minute and pause length seems critical. Has anyone developed a formula or ratio for pause duration between sentences or after key phrases that mimics thoughtful recollection?
* **Phonetic Edits:** Have you successfully used the pronunciation editor or SSML tags to introduce slight, context-appropriate tremors or breath sounds? For example, adding a soft `` or modifying vowel sounds to be less crisp.
* **Voice Stacking:** Is there merit to generating two tracks (one base voice, one with modified settings for specific words) and layering them with a slight offset to create a more complex, textured vocal output? This would be a post-Murf processing step, but I'm curious if the source material generation strategy affects this.

The goal is a methodological approach that can be documented and repeated for different narrators across our episodes. I'm less interested in subjective "sounds good" feedback and more in the specific technical levers you've pulled within Murf's system, the resulting audio samples, and any quantitative measures you used for evaluation (e.g., listener perception scores from a focus group).



   
Quote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

You're focusing on the right parameters but approaching it from the wrong direction. The goal isn't to simulate a degraded larynx, it's to simulate a lifetime of speaking habits. You need to layer minute imperfections on top of a fundamentally stable voice.

Your baseline configuration is missing the crucial element of inconsistent pacing. A real aged voice has micro-pauses, not a uniform speed reduction. You'll have more luck using the API to generate phrases individually, then programmatically introducing slight, random variations in the `speed` and `pause` parameters between sentences in your editing pipeline. Treat it like a statistical distribution, not a fixed value.

Also, don't overlook the `pitch` variance setting. An octogenarian's pitch isn't just lower, it has less dynamic range. You might actually need to slightly compress the pitch variation compared to a younger voice to avoid that theatrical, over-acted quality. Try reducing the pitch range by 15-20% from your baseline while keeping the overall pitch low.



   
ReplyQuote