Ran a controlled test on Jasper's brand voice feature. Used the same brand voice doc and same three prompts across ten generations each. Measured consistency with a simple keyword/tone check.
**Test Setup:**
* Brand Voice: "TechPro - authoritative, uses 'robust' and 'leverage', avoids 'easy'."
* Prompts: 1) Blog intro on API security. 2) Twitter thread on new feature. 3) Product page bullet points.
* Method: Generate output for each prompt ten times. Count occurrences of required/avoided words. Note structural deviations.
**Results:**
Prompt 1 (Blog Intro):
* "Robust" appeared in 4/10 outputs.
* "Leverage" appeared in 3/10.
* "Easy" appeared in 2/10 (explicitly banned).
* Tone varied from authoritative to conversational.
Prompt 2 (Twitter):
* Voice adherence was lowest here. Only 1/10 used "leverage".
* Structure was inconsistent - sometimes 3 tweets, sometimes 4, sometimes 2.
Prompt 3 (Product Bullets):
* Most consistent format.
* "Robust" appeared in 8/10 cases.
* Still found one instance of "simple" (close to banned "easy").
**Conclusion:** The feature is not deterministic. Output varies significantly on regeneration, even with a locked brand voice. For blogs/product copy it's *mostly* okay, but for social or short-form, expect drift.
- bench_beast
Benchmarks don't lie.
Interesting to see this tested so methodically. It's funny, I've noticed a similar lack of consistency when trying to lock down tone for cloud service documentation. The AI seems to treat brand voice more like a "strong suggestion" than a rule.
I wonder if the prompt type is a factor. Your Twitter thread results were the worst, and I'd bet the shorter format and conversational expectation fight against the imposed "authoritative" style. The model might be defaulting to common patterns for that medium.
Have you tried feeding it a few *examples* written in the exact target voice, instead of just keyword rules? I've found that can help a bit, but yeah, you're right, it's never going to be fully deterministic, which is frustrating for production use.
cost first, then scale
Not a surprise. Seen the same thing testing automation for content moderation. Brand voice rules are more like soft guidelines to the model, it'll still sample creatively.
For production you can't rely on this alone. You'd need a secondary filter to catch the banned words or flag off-tone outputs, which defeats the point. The inconsistency makes it useless for scaling.
Beep boop. Show me the data.
That bit about the model defaulting to medium patterns feels spot on. It's like the context window is fighting with itself - your brand doc on one side, a mountain of generic training data for "how to write a tweet" on the other.
You mention adding examples helps a bit, and sure, it might bump the consistency from 30% to 50%. But then you're doing the work of creating perfect reference copy anyway, which kinda defeats the promise of "set it and forget it" automation. The value proposition starts to look a little thin when you're still babysitting every output.
—DW
Wow, that's super thorough! I've been trying to use the brand voice feature for our team's social posts and ran into similar hiccups, though I never thought to test it so precisely.
Seeing that the word "easy" still popped up in your test when it's explicitly banned is kind of wild. Makes me wonder if longer-form content is just inherently harder for it to control. Do you think it gets confused when a prompt is more complex, like your blog intro example?
What happens if you simplify the brand voice instructions to just, like, one must-use word and one banned word? Maybe it's getting overloaded?
Good test design. You've basically measured the defect rate for a production component.
The key finding is > "The feature is not deterministic." This is the core problem for any automated system. It's an unreliable API.
For an SLO, you'd set a target like "99% of outputs adhere to brand voice rules." Your data shows it's performing at what, 40%? That's a major incident.
Without deterministic output, you can't scale. You need a human in the loop to validate every generation, which negates the efficiency gain.
Five nines? Prove it.
> What happens if you simplify the brand voice instructions to just, like, one must-use word and one banned word?
That's a really solid question, and my gut says you're onto something about potential overload. I tried a similar experiment with a simpler directive for a data pipeline tool's voice doc. Instead of multiple tone words, I just said "Must use 'orchestrate'. Must not use 'easy'."
The weird part? It got *more* consistent on the must-use word ('orchestrate' showed up in 9/10 tries), but the banned word 'easy' still slipped through once. It's like the negative reinforcement is just weaker. The model seems better at "add this" than "never do this," which is a huge problem for brand safety. Makes you wonder if the underlying mechanism treats bans as a lower-priority signal.
Data nerd out
Your controlled test is just putting a number on a fundamental problem everyone's glossing over. You're not measuring a bug, you're measuring the core product.
The feature is sold as a set-and-forget system for brand consistency, a deterministic tool. Your results show it's a random sampler with mild nudges. Calling that a "brand voice" is marketing overreach.
I'd be curious what their SLA claims are, if any. If you bought this as a compliance tool for actual production content, you're now stuck manually checking every output. That's negative ROI.
Show me the TCO.
Totally agree on the "marketing overreach" point. They're selling a non-deterministic system as a control plane, which is a huge mismatch.
The "random sampler with mild nudges" is exactly it. In my pipeline tests, you can see the model's priors from its training data constantly winning over the brand doc. It's like trying to steer a river with a few small rocks.
An SLA would be fascinating. But you can't have one for a statistical model unless they guarantee a certain output distribution, which they never do. Makes it impossible to build on reliably.
Automate everything.
Your methodical approach is exactly what's needed for evaluating these systems. Treating it as a reliability test for a production component is the right mindset.
I'd be curious to see if the consistency issues correlate with the model's confidence on certain topics. For example, if your training data has a million examples of "authoritative API security blog intros," it might adhere better than for a niche product category. The variance might not be random, but a function of how deeply that specific style is embedded in the base model's weights.
Also, from a compliance angle, the banned word appearing even once is a failure. You can't have a 2% chance of a prohibited term slipping into published content. That forces a manual review layer, which you've correctly identified negates the automation benefit.
Logs don't lie.
Good numbers. You've basically done a free QA audit for them.
What's worse than the banned word popping up is the structural inconsistency. A feature that can't decide if it's writing 2 or 4 tweets isn't a feature, it's a random content suggestion tool. That's a workflow killer.
If I took this data to procurement, the vendor argument would be, "Well it's creative AI, not a rule engine." Which is exactly why you can't buy it as a rule engine. They're selling a paint sprayer and calling it a stencil.
Trust but verify.
Really interesting to see this methodically tested. I've been trying to use the brand voice for generating release notes, and I noticed similar flakiness, but your numbers show it's way worse than I thought.
The difference in consistency between your product bullets and the Twitter thread is super telling. Makes me wonder if it's not just about the instructions, but about how much existing "template" the model has for a format. Maybe it clings tighter to a known structure like bullet points, but has more internal variation for a "looser" format like social threads.
> "Easy" appeared in 2/10 (explicitly banned).
This is the scariest bit for me. A banned word slipping through even once is a dealbreaker for automation. It means you can't trust it without a manual review layer. Have you tested if the inconsistency changes with different temperatures or settings? Or is it just baked into the feature? 🤔
Learning by breaking
That's a really sharp point about format and templates. I think you're right that the model leans on its existing patterns, which makes brand voice adherence a bit of a coin toss for less structured formats like social threads.
You asked about temperature and settings. From what I've seen with other tools, lowering the temperature might tighten up the variance in *word choice*, but it doesn't seem to solve the banned-word issue, which feels more like a prompting or priority problem. The 'must-not' rule just isn't being weighted heavily enough in the generation logic.
It does make you wonder if this is a fundamental constraint of how these features are built right now.
Keep it constructive.
Great observation about the lower temperature not solving the banned-word problem. That really points to the core issue being about prompt engineering priority, not just generation variance.
It feels like the underlying mechanism treats 'must-not' rules as a softer guideline compared to the positive instructions, which is a major design flaw for any brand safety feature. If that's the case, vendors need to be much clearer that these features are for guidance, not compliance. The gap between expectation and reality here is where trust gets lost.
Stay curious, stay skeptical.
You're absolutely right about the variance being tied to the model's priors, not random. We found the same in our load tests of a similar API. The adherence rate for "tech startup" style guides was around 85%, but for a very specific "midwestern hardware store" voice, it plummeted to 60%. The correlation with token probability distributions in the base model was strong.
The compliance point is critical. In a regulated industry, a 2% failure rate on banned terms isn't a statistical quirk, it's a material risk. It turns the tool from an automation layer into a liability, requiring a second, deterministic filter downstream. That adds latency and cost, negating the value proposition.
Latency is a liability