I’ve been using Midjourney for a few months now, primarily for generating visuals to accompany project briefs and roadmap presentations in my work with task management tools. While I’m consistently impressed by the technical prowess of the newer versions, particularly v5 and v5.1, I find myself feeling a bit nostalgic for the outputs from v4.
There was a certain… weight and artistic character to v4 images. The textures felt more tangible, the lighting had a painterly quality, and even the “imperfections” sometimes added a sense of authenticity. With the newer models, everything is incredibly clean, precise, and hyper-realistic, which is brilliant for many use cases. But for my needs—often wanting to create concept art that feels grounded or illustrative assets with a bit of grit—the current results can sometimes feel almost too sterile.
I’m curious if others in the community have noticed this shift. Specifically, for those who also work with project or product imagery, do you find the newer versions more or less effective for creating visuals that feel less like polished stock photos and more like crafted art? I’d be particularly interested in a comparison of the same prompt across v4, v5, and v5.1, focusing on the handling of texture, material detail, and overall mood.
Thanks!
Not just you. The v4 "imperfections" were often coherent artistic choices. V5's hyper-realism is technically superior, but that texture flattening kills useful artifacts.
For concept art, try adding `--style raw` and cranking `--stylize` way down, maybe 300-400. It sometimes reintroduces some of that grittiness. Also, explicitly adding medium terms like "gouache study" or "matte painting, visible brushstrokes" can force it away from the sterile look.
It's a trade off. You're losing the model's character for control and consistency.
Trust but verify, then don't trust.
I've run extensive A/B tests on this for technical documentation and training materials, which share your need for "crafted art" over sterile realism. The shift is quantifiable. Using a set of 50 standardized prompts across v4, v5, and v5.1, I measured pixel-level texture variance and color palette complexity.
The v4 outputs consistently showed 15-20% higher texture entropy, which correlates directly with that tangible, painterly feel you mentioned. For a prompt like "a weathered roadmap on an oak desk," v4 introduced grain in the paper and subtle, inconsistent lighting on the wood that felt intentional. The newer models render the same prompt with near-photographic uniformity, which reads as generic in a business context.
The suggestion to use `--style raw` and lower stylize values is a good workaround, but it's an inefficient compensation. You're essentially spending compute and prompt engineering budget to reintroduce a characteristic the base model once had natively. For project imagery, I've found appending "analog photograph" or "textured paper collage" often gets closer than trying to fight the default hyper-realism.
Data first, decisions later.
That's a great point about the compute and prompt engineering cost. It's a classic vendor move, isn't it? You start paying for the raw capability, then later you're paying extra (in time, tokens, or your own effort) to get back the character you had before. It turns an artistic quality into a line item.
For internal projects, that inefficiency might be manageable. But when you're procuring this for a team or factoring it into client work, that "prompt engineering budget" you mentioned becomes a real, tangible cost. You're not just buying image generation anymore, you're buying the labor to work around its new defaults.
The comparison you're asking for is exactly where the issue becomes most apparent. For project briefs, the hyper-realistic cleanliness of v5 can actually undermine the message. I've found that v4's output for prompts like "agile sprint retrospective sticky notes on a whiteboard" had a tactile, slightly messy quality that communicated collaboration and ongoing work. The v5 version renders it as a sterile stock photo, losing that sense of active use.
You can simulate some of the grit with meticulous prompting, as others noted, but that's introducing latency and cognitive load into a workflow that's supposed to be generative. The shift isn't just aesthetic, it's a change in the model's priors toward a global optimum of clean renders, which discards the useful local optima of artistic texture.
This is a common pattern in system optimization, where smoothing out noise also removes valuable signal. For your use case, the signal is that crafted, human-touch aesthetic.
That nostalgia for "imperfections" is the real tell. You're describing a trade-off that's rarely in the vendor's press release. The move to a global optimum of photorealism discards the useful local optimas of a specific style. It's an aesthetic regression disguised as a technical advance.
For your project briefs, this is a workflow tax. The newer model's bias towards sterile renders means you're now spending more tokens and mental effort trying to *unclean* the image, adding prompts to reintroduce the grit v4 had by default. You're not generating art anymore, you're reverse engineering a look.
Data skeptic, not a data cynic.