You're right to notice the unnatural quality. What you're hearing is likely a post-processed adult voice, not a model trained on genuine child speech patterns. The prosody and cadence are wrong because the underlying data is wrong.
Testing their flagship adult voice with your script is the correct next step. It's a separate pipeline. If the core voice is solid, you've just identified a bad feature to avoid. If it's also poor, then you have a platform-wide quality issue.
The marketing for children's audiobooks is, frankly, a serious misjudgment. It reveals a gap between product checklist and user experience.
every dollar counts
It's not just you. The issue is they're all basically the same low-effort post-processing trick, so letting it shake your faith in their main catalog is a bit naive.
The diagnostic step isn't to just *question* the other voices, it's to *test* them. Run your exact script through their standard adult conversational voice. If that also sounds robotic, then you've got a core quality problem. If it's fine, you've just learned that their "child" feature is a checkbox they can't actually deliver on, which is its own kind of useful data about the vendor.
You've nailed it with the "good reminder to test the specific voice you need." That's the cornerstone of solid validation.
I'd add that this is where a structured QA checklist is invaluable. We treat voice selection like any other feature test, with a clear scoring rubric for naturalness, emotional fit, and listener fatigue. It prevents the team from getting charmed by a demo and missing the mismatch with the actual script.
Your workaround suggestion for the "young adult" preset is exactly right. It often gets closer to the warmth you need without hitting the uncanny valley. Finding those functional alternatives is a much better use of time than trying to make a broken feature work.