I've been testing the 'improve' command on my existing blog drafts. It makes the text more polished, but the output starts to sound generic, like any other AI content.
What specific steps or prompt phrasing do you use to keep your original tone intact? I need practical methods, not just "add more details." How do you handle this before committing to a long-term plan?
I run the content team for a 500-person B2B SaaS company, and we've used the Claude API for blog and documentation drafting for about a year now.
The core issue is preserving authorial voice, which is a tuning and workflow problem. Here's how we break down the approaches we tested:
1. **Fine-tuning Cost vs. Control**: A custom fine-tune on ~50-100 of your existing posts works, but it's a resource sink. At my last shop, the project cost roughly $2-3k in dev and API time and needs a refresh every 6-12 months. It's for teams with a large, consistent back catalog and a dedicated budget.
2. **Prompt Engineering Investment**: This is free but demands rigorous effort. We built a "voice specification" document we paste into every session. It's not just style notes; it includes 5-10 concrete examples of your "before" and desired "after" text. Maintaining and iterating on this document takes 1-2 hours a week.
3. **Workflow Guardrails (The Editing Layer)**: Never use 'improve' on a full draft. We only apply it to specific paragraphs flagged as "awkward" or "unclear." The final step is always a human editor reading aloud against the original. This adds about 15% more time to our editing cycle but cuts AI-generic output by about 80% in our experience.
4. **Tool Choice and Context Limits**: Using the API (Claude, GPT-4) directly beats any built-in "improve" button, because you control the system prompt. The 100k+ context windows now let you feed the entire voice spec plus 2-3 example posts. The limitation is cost: thorough editing this way can run $0.10-$0.25 per post, which adds up.
My pick is prompt engineering with a strict paragraph-by-paragraph workflow. It's the most deterministic and cost-effective for a solo blogger or small team. If you have the budget and volume (50+ posts a month), then a custom fine-tune starts to make sense. Tell us your monthly output and if you're using an API or a tool's built-in button.
Trust but verify — especially the fine print.
Your problem isn't the 'improve' command, it's the premise. Asking an LLM to improve your draft is asking it to impose its own standard of polish, which is inherently generic.
Stop using 'improve'. Use 'rewrite this in my voice' and attach a bullet list of your actual voice traits. Give it examples of sentences you've written and tell it to mimic the structure and cadence, not the content. If you can't describe your own voice in concrete terms, you shouldn't be automating its editing.
The long-term plan is figuring out if you need polish or just a proofreader. Most early drafts need tightening, not a full AI makeover.
Trust but verify.
"Improve" commands prioritize readability metrics over your voice. They're built to converge toward a standard average.
You're asking the wrong question. Your workflow is the problem. Don't polish the whole draft with AI. Feed it only the sentences that are objectively broken - awkward phrasing, unclear antecedents. Ask it to fix just those, using the surrounding text as the voice guide.
If you can't identify which specific sentences need correction, you need a human editor, not a better prompt.
Least privilege is not a suggestion.
Forget prompts. Benchmark it.
Take a single paragraph. Run it through the 'improve' command three times with these instructions:
1. "Improve for clarity and flow"
2. "Improve for clarity and flow. Maintain the original author's casual, first-person tone."
3. "Rewrite for clarity, matching the sentence length and vocabulary in this sample: [paste another one of your paragraphs here]"
Compare the outputs. You'll see the generic voice comes from the default, not your request. Option 3 works if you can provide a strong sample.
If the third output still sounds wrong, your workflow is the issue, not the tool. You need stronger reference text.
Benchmark or bust