After conducting an extensive comparative analysis of Sudowrite against other AI-assisted writing platforms, I have reached a preliminary conclusion regarding its flagship "Rewrite" feature: it exhibits a consistent failure to preserve authorial voice across multiple genres and stylistic inputs. This is not a minor stylistic quibble, but a fundamental failure in the model's ability to perform style transfer, a task that is central to its marketed value proposition.
My methodology involved feeding identical source passages—each with a distinct, pre-analyzed voice profile—into Sudowrite's rewrite engine, a competing general-purpose LLM via API with a specific style-preservation prompt, and an open-source model fine-tuned on literary data. The control prompt in Sudowrite was consistently "Improve this passage." The results were quantitatively and qualitatively assessed.
* **Quantitative (Lexical):** Sudowrite's outputs showed the highest deviation from source material in metrics like:
* Type-Token Ratio (TTR) variance.
* Average sentence length shift.
* Profanity/qualifier injection where none existed previously.
* **Qualitative (Stylistic):** The engine demonstrated a pronounced tendency to overwrite the input with a homogenized, commercially "polished" voice. A terse, hard-boiled noir paragraph would be rewritten with superfluous adjectives and a softening of cadence. A academic, passive-voice laden section would be transformed into inappropriately direct and simplified prose.
Consider this brief illustrative example. Source passage with a detached, technical voice:
```plaintext
The mechanism failed. Diagnostic codes 45 and 78 were logged. The system entered a failsafe state. No user data was compromised.
```
Sudowrite's typical rewrite output:
```plaintext
Unfortunately, the mechanism encountered a failure. We recorded diagnostic codes 45 and 78, which prompted the system to seamlessly enter a protective failsafe state. Rest assured, all user data remained completely secure and uncompromised throughout this event.
```
The transformation is clear: active to passive voice, injection of empathetic qualifiers ("Unfortunately," "Rest assured"), and a shift from factual reporting to customer-facing reassurance. The original voice is eradicated.
This suggests one of several underlying issues: the model may be heavily biased towards a single "optimized" corporate or marketing voice due to its fine-tuning dataset, or the "Rewrite" function may be a complex, non-transparent chain-of-thought prompt that prioritizes expansion and "improvement" over fidelity. From a privacy and ethics perspective, this also raises questions about the training data's diversity of style and whether the tool is designed to subtly conform user output to a standardized, inoffensive median.
I am soliciting further data points from the community. Have you conducted similar tests? Are there specific trigger words or prompt engineering techniques you've found that mitigate this voice-erasure effect, or does the underlying architecture seem inherently biased against stylistic preservation?
- labrat
Your point about profanity and qualifier injection is something I've seen in my own testing, though I focused more on contract and RFP language. The system seems to have a default "voice" it reverts to, which is often oddly conversational or tries to add emphasis where none is needed. It overcorrects.
Have you experimented with using the "Describe your voice" feature in the Story Bible? I found its effectiveness to be highly genre-dependent. It worked moderately well for maintaining a technical, dry tone in procedural documents, but completely failed to replicate a more nuanced, persuasive sales voice. The engine appears to prioritize grammatical "improvement" over stylistic fidelity.
Your quantitative approach is interesting. Did you track whether the deviation was consistent? In my notes, the shifts weren't random - they always trended toward a more standardized, middle-of-the-road readability score, which suggests an underlying optimization goal that conflicts with voice preservation.
Buyer beware, but with a spreadsheet.
The "middle-of-the-road readability score" you're seeing is the giveaway. It's the same pathology I see in SaaS sales contracts: a vendor's boilerplate is always optimized for their risk, not your specific context. The model is tuned for a generalized, inoffensive, corporate "voice" because that's the safest default from a liability and mass-appeal standpoint.
That's why the Story Bible feature is a band-aid. You're giving it a style guide, but the underlying engine's reward function is still prioritizing grammatical conformity and a vanilla tone. It'll absorb your "voice" notes as just another set of parameters to be balanced against its core objective, which is to produce a clean, universally palatable output. It explains the genre-dependency you noted; technical prose is already close to that vanilla standard, so the deviation is minimal.
So we're not really paying for style preservation. We're paying for a product whose primary optimization goal is directly at odds with its marketing claim. The contract for this service probably has a wonderfully broad disclaimer about "results varying," doesn't it?
The small print is where the fun is.