I've been conducting a series of standardized aesthetic output evaluations on the v6.1 model since its release, specifically targeting the impact of the `--style raw` parameter. My methodology involves generating controlled prompt sets across multiple categories (portraiture, landscape, architectural, abstract concept) with and without the raw flag, using identical seeds for direct comparison.
The initial hypothesis was that `--style raw` would produce more predictable, less opinionated outputs by reducing Midjourney's default 'stylization' layer, theoretically leading to greater user prompt fidelity. However, the observed variance suggests the opposite; it appears to introduce significant, non-deterministic instability in the output characteristics.
My benchmark runs indicate the following inconsistencies:
* **Prompt Adherence Volatility:** For a technical prompt like `aeronautical blueprint of a coaxial rotor system, isometric view, technical drawing`, the standard v6.1 produces consistently schematic-style images. With `--style raw`, the results randomly oscillate between a pure technical drawing, a photorealistic image of a helicopter, and an artistic sketch, despite identical seed values.
* **Stylistic Incoherence:** In portrait benchmarks, using a prompt such as `young woman with curly hair, studio portrait, neutral background`, the non-raw outputs are uniformly stylized in the recognizable Midjourney v6.1 'photo' style. With `--style raw` enabled, the results are a statistically random mix:
* ~40%: hyper-realistic, near-photographic output.
* ~30%: heavily painterly, reminiscent of older model styles.
* ~20%: bizarre, desaturated, and oddly flat compositions.
* ~10%: nearly identical to the non-raw version.
* **Parameter Interaction Conflicts:** The `--stylize` parameter (`--s`) seems to have an inverted or chaotic relationship with `--style raw`. Lower stylize values sometimes yield *more* dramatic stylization in raw mode, which contradicts the documented intent.
This behavior points not to a simple reduction of style, but to what seems like an incomplete or buggy implementation of the style removal process. It's as if the system is randomly selecting from a disjointed set of underlying legacy stylization kernels instead of cleanly subtracting a single layer.
I propose a community benchmarking effort to gather more data. If others are willing to run controlled tests, please use the following standardized prompt format and report back with your seed and observed output style category.
```markdown
**Prompt:** [Your test prompt]
**Seed:** [Your seed]
**v6.1 Default:** [Description of output]
**v6.1 + --style raw:** [Description of output]
**Stylization Category:** (e.g., Photorealistic, Painterly, Default-like, Flat/Desaturated, Other)
```
Without reproducible and consistent behavior, the `--style raw` parameter cannot be considered a reliable tool for a controlled workflow. It currently functions as an uncontrollable randomizer, which is frankly unacceptable for any professional use case requiring predictable outputs.
Has anyone else performed systematic tests and quantified similar variance? I am particularly interested in data regarding the interaction between `--style raw` and the `--weird` parameter.
-- bb42
-- bb42