I've been systematically testing Udio's 'Genre Blender' feature over the last week, using a standardized set of prompts to evaluate its consistency and output quality. My initial hypothesis was that it would interpolate stylistic elements in a meaningful way, perhaps blending instrumentation or rhythmic patterns. However, my results suggest a significant failure mode where the output often becomes a structurally incoherent pastiche rather than a coherent fusion.
For example, using the prompt "A melancholic ballad blending synthwave and bluegrass" yielded a track where:
* The introductory 8 bars established a plausible synthwave pad and slow 4/4 beat.
* At the 9-second mark, a banjo sample was abruptly introduced with no harmonic or rhythmic integration.
* The vocal melody, when it entered, switched time signatures erratically, creating a disorienting listening experience.
This wasn't an isolated case. My test matrix included blends like "baroque pop with drum and bass breakdowns" and "doom metal meets bossa nova," all with similar issues. The core problem appears to be a lack of hierarchical understanding; the model seems to be performing a simple, surface-level feature swap (e.g., inserting a genre-stereotypical instrument) without reconciling the underlying musical grammar.
I'm curious if others have undertaken similar structured tests. Specifically:
* Have you found any prompt engineering techniques that mitigate this, or does the issue seem fundamental to the model's architecture?
* Is the output more coherent when blending closely related genres (e.g., folk and country) versus distant ones?
* Could this be a training data issue, where the model lacks sufficient examples of true, well-executed genre hybrids?
prove it with data