I've been running a systematic benchmark of NightCafe's image-to-image transformations, specifically testing all "Artistic" filters (Coherent, Expressionist, etc.) against a standardized set of 50 source images. My goal was to identify the optimal filter for different input types.
The initial results reveal a significant issue: **convergence**. Regardless of the source image—whether a portrait, a landscape, or an architectural photo—the output from the Artistic filter set exhibits a strikingly similar texture and color palette. The distinct "handwriting" one expects from different artistic styles is largely absent.
**Key observations:**
* The "painterly" overlay applied often dominates the source content, reducing detail diversity.
* Color temperature tends to shift towards a warm, ochre-dominated spectrum across multiple filters.
* Brushstroke texture, while different in scale between (e.g.) "Coherent" and "Expressionist," follows a similar underlying pattern algorithm.
Here's a data snippet from my test log for a single input image:
```
Input: studio_portrait_01.jpg
Filter | Output Similarity Score* | Dominant Hue Shift
----------------|--------------------------|-------------------
Coherent | 0.78 | +22° (warmer)
Expressionist | 0.75 | +18° (warmer)
Scream | 0.82 | +25° (warmer)
Fantasy | 0.71 | +15° (warmer)
(* vs. median output of all Artistic filters)
```
This suggests the underlying model for the "Artistic" category is heavily biased towards a specific aesthetic, limiting its utility for generating truly distinct stylistic variations. For my workflow, I've gotten more diverse results by using the "General" or "Abstract" filter sets and then applying a secondary prompt for style.
Has anyone else performed comparative analysis and found parameters or combinations that break this pattern? I'm particularly interested in `--style_weight` or `--cfg_scale` adjustments that might mitigate the homogeneity.
Numbers don't lie
Interesting methodology, and your "convergence" finding matches my experience when I tried to script batch conversions for a digital asset pipeline. The underlying model seems to have a surprisingly narrow latent space for that entire filter category.
Have you tried adjusting the **influence slider** alongside your filter tests? I've found that setting it below 50% sometimes breaks that texture dominance, allowing more source image structure to come through. It doesn't solve the hue shift, but it might introduce more variance in your similarity scores.
Your test log snippet makes me think you could treat this as a CI problem. Define a baseline "difference threshold" for the output of different filters against the same source, and fail the pipeline if they're too similar.
Commit early, deploy often, but always rollback-ready.
Not sure a benchmark is the right tool here. You're measuring a subjective style output like it's a performance test. The "convergence" you're seeing is probably the entire point - they're just slapping a single branded 'artistic' look on things. It's a feature, not a bug. Did you expect actual Van Gogh vs. Monet algorithms?
Trust but verify.