Skip to content
Notifications
Clear all

Anyone else's "instrumental" tracks still have weird vocal artifacts?

2 Posts
2 Users
0 Reactions
10 Views
(@liamr)
Trusted Member
Joined: 3 months ago
Posts: 30
Topic starter   [#5553]

Hey folks, been putting Suno through its paces for a few weeks now, mostly generating background music for some web app demos. I'm loving the creative potential, but I've hit a consistent snag.

When I generate a track with the "Instrumental" style tag selected, I *still* get these faint, weird vocal artifacts. They're not full lyrics, more like mumbled syllables or ghostly "oohs" and "aahs" buried in the mix. It happens maybe 60% of the time, even when I'm super specific in the prompt (e.g., "purely instrumental synthwave track, no vocals").

A few examples from my last session:
* A "corporate presentation loop" had what sounded like a distorted, distant voice saying "hey" every few bars.
* A "cinematic orchestral piece" had a choral pad that distinctly shaped itself into a vowel sound.

It's not a dealbreaker, but it makes the tracks unusable for true instrumental purposes. I have to either re-roll multiple times or clean it up in Audacity.

**My questions:**
* Is anyone else running into this consistently?
* Could it be a style tag conflict? I sometimes use "Catchy" or "Emotional" alongside "Instrumental" – maybe that's confusing the model?
* Any workarounds or prompt phrasing that reliably gives you clean results?

I'm really curious if this is a common training data bleed-through issue, or just my bad luck with the generations. Would love to compare notes!

liam


If it can be automated, it will be.


   
Quote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh yeah, I've had that happen. It's like the model has a hard time letting go of the human voice as an instrument. I ran into it making background music for my home lab's monitoring dashboard - generated a nice ambient track and kept hearing what sounded like a whispered "status" in the background. Had a good laugh, but it ruined the track.

Try pushing the prompt further away from anything vocal-adjacent. Instead of just "no vocals," I've had better luck with something like "electronic ambiance, only synthetic textures, completely inhuman soundscape." It seems to help if you actively forbid the tools it might use to make those artifacts, like pads or choirs, even if you want a similar feel.

> Could it be a style tag conflict?
That's probably part of it. "Emotional" or "Catchy" likely pulls from training data full of songs that use vocal elements to achieve that, so the bleed-over makes sense. I'd try "Instrumental" with purely descriptive mood tags like "serene" or "mechanical" and see if the ghost choir fades away.


it worked on my machine


   
ReplyQuote