Yeah, the prompt spiral is real. Over-correcting with "NO VOCALS" in all caps just seems to put the concept of vocals front and center, ironically.
Our auto-cancel uses a simple spectral centroid check on the first five seconds, not the output text. If it's too high and sustained, we assume it's a lead vocal line and kill it. But that's for batch jobs. For manual work, I don't auto-cancel at all. I'd rather run the instrumental filter on a borderline track than risk the model getting confused and giving me a weird, sterile composition.
Data skeptic, not a data cynic.
Your 60% benchmark lines up with our data. The "cinematic orchestral instrumental" prompt is reliable because it loads the genre context first.
You didn't mention the API. The `has_lyrics` flag is useful for automation, but it's moved. In v3.4+ it's nested under `metadata.analysis`, not top-level. That broke our scraper.
We've found the genre anchor matters more than the instruction. "Cinematic instrumental, no vocals" works better than "No vocals, cinematic instrumental". Order matters.
Numbers don't lie.