I've been advising a small SaaS startup on their website copy, and we hit a snag on the "About Us" page. The founder wanted something that balanced professionalism with their unique team culture. To get a feel for the current landscape, I ran the exact same prompt through five popular AI writing tools.
The prompt was straightforward:
> "Write an 'About Us' page for a B2B SaaS company named 'NexusFlow'. We provide intelligent workflow automation for mid-sized financial firms. Our tone should be professional yet approachable. Key points to include: founded by ex-industry experts who saw a gap in affordable, intuitive tools; our platform turns complex processes into simple, visual workflows; we believe technology should empower teams, not replace them; and we value deep partnerships with our clients. Aim for 250-300 words."
The outputs were fascinating in their divergence. One tool produced a bloated, 500-word essay full of generic phrases like "leverage synergistic solutions." Another gave a crisp 200-word draft that nailed the key points but felt a bit sterile. A third surprisingly injected a subtle, welcoming warmth that aligned perfectly with "approachable," while a fourth got the facts right but the voice sounded like a stiff press release.
The main differences weren't in factual coverage—they all hit the bullet points—but in **voice, conciseness, and how they framed the core message**. The one that performed best required the least editing to match the desired tone. The others needed significant rewrites, either to cut jargon, inject more humanity, or tighten the structure.
It's a useful reminder that the tool choice isn't just about generating text; it's about which one starts you closer to your specific voice and audience. The "best" tool completely depends on the brand's personality.
—daniel
That's a really interesting experiment. I've seen similar variation when using AI to generate config files or CI/CD pipeline YAML. You give the same spec to GitHub Copilot, ChatGPT, and a CLI tool, and you get three different structures that all technically work.
The "bloated, 500-word essay full of generic phrases" sounds painfully familiar. It's like when a tool over-engineers a simple deployment script, adding unnecessary complexity.
Which tool gave you the sterile but crisp draft? That's the kind I'd probably end up using as a first draft, then trying to manually inject the "approachable" tone afterward. I wonder if the results would be more consistent if you added a single, very specific instruction like "avoid the phrase 'leverage' and use 'use' instead."
Learning by breaking
Absolutely. That "sterile but crisp" one came from the tool that's basically a GPT wrapper with a minimalist UI - it tends to do exactly what you ask, then stops. No extra fluff. It's useful as a skeleton, but you're right, you then have to spend time adding the warmth back in.
Your point about a single, specific instruction is spot on and something I preach to clients. We did a similar test with marketing email copy and found that adding "replace all instances of 'leverage' with 'use'" dramatically cut down on the corporate jargon across every tool. The consistency improved, but the *style* still varied wildly.
It's the classic automation paradox: sometimes the tweak needed to make the output perfect is more work than writing the first draft from scratch. Been there with workflow builders too - you can spend an hour debugging a flow that would've taken 10 minutes to do manually.
Implementation is 80% process, 20% tool.
Your comparison to config files is so on point. That's exactly it, the tools are generating valid output, but the *architecture* of the sentences is wildly different, just like one tool will write a flat YAML structure and another nests everything.
And yes, the "avoid 'leverage'" trick is a lifesaver. I've found you need to get hyper-specific, almost like writing a style guide prompt. My addition is to also ban "empower" and "solution." It cuts the fluff, but you're right, the stylistic bones still come from the tool's own default voice. Sometimes that sterile first draft is perfect for a technical spec page, just not an 'About Us.'
Spreadsheets > marketing slides.
I've hit that automation paradox so many times building Zaps. You can get a 90% solution in minutes, then burn an hour trying to fine-tune the last 10% with filters and formatters. It's exactly like you said - sometimes you just need to ditch the tool and write the last mile yourself.
Your style guide prompt idea is key, though. I treat it like configuring an API connector: you have to set the default parameters. I have a saved prompt for client work that starts with "Avoid: leverage, empower, solution, disruptive, robust, synergy." It's a night and day difference, even if the underlying structure still varies.
api first
The automation paradox is real, but I think the deeper issue is we're conflating tools for structured code with tools for subjective writing. You can write a style guide prompt, but the variance in the underlying model's training data and its interpretation of "approachable" will always be a black box. It's not like a linter where you get deterministic output.
That sterile draft you mention as a skeleton is the only part with any real value, honestly. It's the minimum viable text. Everything after that is just manual labor with extra steps, which begs the question: why not just write the warm, human part from the start? You're trading the 10 minutes of manual work for an hour of prompt engineering and editing, which feels like a classic case of over-optimizing the wrong layer. We see this all the time in infra - a team will spend a week automating a 15-minute monthly task.
Your k8s cluster is 40% idle.
Your experiment highlights a core inconsistency in generative AI that mirrors issues in distributed systems. The variance you observed isn't just about "style"; it's a symptom of non-deterministic output from models with different fine-tuning and sampling parameters. You're essentially getting five different consensus outcomes from the same initial state, without a reliable commit log.
The >"bloated, 500-word essay full of generic phrases"< output is the equivalent of a system with poor backpressure, generating data without regard for the word limit constraint. The crisp 200-word draft is a system that honors the boundary condition but fails to deliver on the qualitative "approachable" metric, treating it as a soft requirement.
This is why treating these tools as deterministic template engines is flawed. You wouldn't run a database replication protocol without specifying RPO and RTO. Similarly, using these tools effectively requires defining, and more importantly, testing for, concrete acceptance criteria beyond word count. For instance, you could run a secondary check on the output for a "jargon density" score before even evaluating tone.
Yep, that was Copyforge. Gives you the barebones structure and nothing else. A decent starting point if you treat it like a wireframe.
Your 'avoid "leverage"' trick works, but only up to a point. The underlying model still decides what "approachable" sounds like. I find you need to get even more granular - "use contractions like 'we're', avoid passive voice, write for an 8th grade reading level." Even then, the 'voice' is a toss-up.
Exactly. Getting down to that level of granularity with instructions is the only way I've found to get consistent results. It's like building a complex filter in Make or Zapier, you have to account for the edge cases.
But even with a bulletproof style prompt, you're still at the mercy of the tool's baseline. I tried the same approach for help desk auto-replies, specifying everything from sentence length to which transition words to avoid. Two different platforms gave me outputs that technically followed every rule, but one still felt like a friendly colleague and the other like a policy manual.
That's the part you can't fully prompt away - the model's inherent "voice." Sometimes the sterile wireframe is the best you can hope for, and you just have to accept that the final human touch is non-negotiable.
api first
Your experiment highlights the core challenge of using non-deterministic systems for what is essentially a branding and messaging problem. The variance you're seeing isn't a bug in the tools, it's a fundamental property of using LLMs without a tightly controlled schema for the output.
Think of it like generating JSON from a prompt. Without a strict, predefined schema, you'll get valid JSON, but the field names, nesting, and even data types can vary. Your prompt lacked that enforced schema for tone and style. The key points were your required fields, but "professional yet approachable" is an unconstrained enum that each model resolves differently based on its training bias.
The most reliable method I've found is to move the creative burden upstream. Instead of hoping the model interprets "approachable" correctly, you feed it explicit examples. Provide three sample sentences that embody the exact tone you want, instruct the model to analyze their linguistic patterns, and *then* generate the new copy adhering to those patterns. It's essentially creating a style transfer function, which is far more deterministic than relying on ambiguous adjectives.
—BJ
Yep, saving that style ban list is a must-have. It's like a pre-flight checklist. I use a similar one, but I also find I have to go beyond banned words sometimes and dictate positive replacements. "Use 'help' instead of 'empower'" or "say 'tool' not 'solution'".
But you're right, it still can't fix the base voice. I had the same thing happen drafting a support SLA. Two tools followed every rule, but one sounded like a robot reading a policy and the other sounded like a person explaining it. For the final 10%, you're just editing a draft that's already cost you an hour to prompt engineer.
The move from a simple ban list to prescribing positive replacements is critical. It's the difference between a firewall that blocks bad traffic and a routing policy that guarantees good traffic gets through.
But your SLA example nails the hidden cost. You're paying for that "prompt engineering hour" upfront, just like committing to a reserved instance. You're locking in a discount on the initial draft, but the final human edit is that unpredictable variable cost. It can be worth it if you're generating at massive scale, but for a single page, the math rarely works out.
That final 10% of human editing is where the real FinOps principle applies - you're optimizing for total cost, not just the cost of the first draft.
Less spend, more headroom.
So which tool gave you the welcoming, warm one? That's the interesting result. The rest sound like they're just shuffling the same corporate word bank.
It feels like finding that variance is the real test. If a tool can consistently hit a specific voice beyond just the keywords, that's valuable. The sterile one is probably fine for a placeholder, but you'd still have to rewrite it completely to get any real personality in there.
That welcoming, warm one is the outlier that'd actually be useful for a real "About Us" page. The sterile drafts are predictable and easy to generate, but they're just generic scaffolding you have to rebuild anyway.
The real question is whether that warm output was a fluke or repeatable. You'd need to run a few more variations of the same prompt through that specific tool to see if it consistently produces that voice. If it does, you've found something more valuable than a text generator, you've found a tool with a usable *default tone*.
Otherwise, you're back to square one, trying to engineer that warmth into a prompt, which is where everyone else in the thread is hitting the wall.
Latency is the enemy, but consistency is the goal.
Agreed. It's infrastructure automation all over again.
You automate the boilerplate, not the unique bit. That sterile skeleton *is* the boilerplate. The prompt engineering to get warmth is trying to automate the unique business logic, which never scales.
The real risk is wasting cycles on the wrong layer. You'll see teams building elaborate style-guide-as-code rigs when one person with a clue could write the final draft in twenty minutes.
Least privilege is not a suggestion.