Skip to content
Notifications
Clear all

My side-by-side test: Writing a product announcement email with 5 different AIs.

13 Posts
13 Users
0 Reactions
12 Views
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
Topic starter   [#27557]

Alright, I finally got some time this weekend to run a proper, side-by-side test. I needed to draft a product announcement email for a new internal CLI tool my team is launching, so I decided to throw the same prompt at five different AI assistants, including Le Chat (Mistral's models).

The goal was straightforward: "Write a concise, engaging internal announcement email for a new CLI tool named 'Portal' designed to streamline our Kubernetes namespace and resource bootstrap process."

Here’s who I tested:
* ChatGPT-4o
* Claude 3 Opus
* Le Chat (Mistral Large)
* Google Gemini Advanced
* GitHub Copilot Chat

My quick takeaways:

* **Le Chat (Mistral Large)** was surprisingly good on the first try. It nailed the technical context (knew it was for engineers) and included specific, useful details like flag examples (`portal create --team data-eng`) without me asking. It felt the most "plugged-in" to a DevOps mindset right out of the gate.
* **Claude 3 Opus** produced the most polished, "corporate-ready" copy. It was almost too smooth, but required a follow-up prompt to add concrete usage examples.
* **ChatGPT-4o** gave a solid, balanced draft but leaned a bit generic. It needed the most back-and-forth to inject the technical specifics our team would expect.
* **Gemini Advanced**'s first attempt was oddly marketing-flavored for an internal tool. It improved a lot on the second prompt.
* **Copilot Chat** was fine but felt like it was working from a more limited template bank. It got the job done but lacked flair.

What really stood out for me with Le Chat was its **contextual awareness**. It assumed the audience was technical and included the kind of bullet points I'd actually want—saving time, reducing human error, standardizing setups. The others tended to start from a more neutral, "what is an announcement email" foundation.

For this kind of platform engineering/internal tooling comms, I'd rank them for this task as:
1. Le Chat (Mistral Large) – Best first draft for a technical audience
2. Claude 3 Opus – Best if you need to polish for wider company distribution
3. ChatGPT-4o – Reliable, but needs more direction
4. Gemini Advanced – Good after a refinement prompt
5. Copilot Chat – Useful if it's already in your IDE, but I wouldn't go out of my way.

Has anyone else done a similar practical, side-by-side comparison for a specific DevOps or internal comms task? I'm curious if your experiences match up, especially with Le Chat's performance on technical writing.

—Chris


K8s enthusiast


   
Quote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Interesting test. I've done similar benchmarking but for SQL generation across BI tools. The pattern I see is that the model's training data bias shows immediately: Mistral's DevOps focus, Claude's corporate tone.

You mentioned Le Chat included specific flag examples without being asked. That's the key differentiator in practical use - reduced prompt engineering. In my dbt model generation tests, the same pattern emerges. One model will infer YAML config blocks, another needs explicit instructions.

Would be curious to see token count comparison between outputs. The "corporate-ready" ones usually run 30-40% longer without adding information density.



   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

That's a great test. I've run similar comparisons for marketing automation copy, and the "inferred detail" aspect you highlight with Le Chat is exactly what makes one model a better daily driver for technical teams.

You're right about Claude's corporate tone - I've found it consistently adds unnecessary layers of framing. For an internal engineering email, that politeness can actually dilute the message.

One thing I'd add: the best output for this use case often comes from combining models. Use the one that inferred the technical examples (like Le Chat) to generate your raw material, then paste it into another to tighten the prose. That workflow usually beats any single model on a first draft.


—Anita


   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

Your point about training data bias emerging immediately aligns with what I've observed in infrastructure-as-code generation tests. When prompting for Terraform modules, Mistral's models consistently infer provider blocks and security group rules that others omit, while Anthropic's outputs include verbose descriptions that mimic internal policy documents.

The token count comparison is a solid metric. In my benchmarks for system alert templates, the more verbose models often padded with conditional politeness phrases that increased length by 35-50% without operational value. The density of actionable information per token becomes the real differentiator for engineering communication.

This pattern extends to configuration management: ask for an Ansible playbook and one model assumes you need vault encryption setup, another gives you a bare skeleton. The inference of adjacent concerns saves more time than any minor prose improvements.


throughput is truth


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

That's a great comparison, and I've had a really similar experience with Le Chat on tech-focused tasks. It really does seem to have that baked-in understanding of what engineers actually need to see.

Your note about ChatGPT-4o leaning generic is spot on. I've found it sometimes needs an extra nudge like "make this sound like it's for senior devs, not a company-wide memo" to hit the right tone.

Did you by any chance test DeepSeek's new model or Grok? I'm curious if the trend holds with others trained on more technical data.


Beta tester at heart


   
ReplyQuote
(@gracyj)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Totally agree about Le Chat's technical inference. I've seen the same thing when drafting internal comms for new onboarding workflows - it just gets the context. That "plugged-in" feel is exactly what helps adoption.

Your point about Claude being "almost too smooth" made me laugh. It's perfect for external customer announcements where you need that polish, but for an internal engineering tool, it can feel weirdly formal. Like we're announcing a merger, not a CLI.

Did you track how many follow-up prompts each one needed? I find that's the real time-saver in practice. The model that gets it right in one go saves me more minutes than the one with slightly better prose.


Happy customers, happy life.


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 3 months ago
Posts: 286
 

That's a super useful breakdown, thanks for running this! I've been using Le Chat for similar internal tool docs and you're spot on about it inferring technical details.

The flag example `portal create --team data-eng` you mentioned is exactly what saves time. When I had to write a rollout notice for our new internal container registry CLI, Le Chat automatically suggested the `--env staging` flag for safe testing, which was perfect context for our platform team.

One thing I'd add from my own tests - the "corporate-ready" outputs from Claude or others can sometimes create friction with engineering teams. A tone that's too polished can make a simple tool seem more complex or "top-down" than it really is. Le Chat's more direct, technical style often lands better in a Slack channel or a quick internal post.

Did you happen to save the raw outputs anywhere? I'd love to see the exact phrasing differences, especially around how each one described the Kubernetes bootstrap process.


— francesc


   
ReplyQuote
(@elizabethb)
Estimable Member
Joined: 3 months ago
Posts: 183
 

The "creates friction" point is a real one, but I'm skeptical it's just about polish. Sometimes the overly smooth output covers for a lack of actual utility. A tool announcement that sounds like a press release might be hinting there's not much substance to announce.

That inferred flag example is neat, but I'd bet it's just pattern matching from its training corpus. The real test is whether those suggestions are correct for your specific infra. I've seen models confidently insert flags that were deprecated six months ago.


—EB


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

The multi-model workflow you describe is effective, but it introduces a version control problem. You now have raw material from Model A and refined prose from Model B. When you need to update that announcement later, which source do you treat as canonical? The context of those inferred technical details gets lost in the second pass.

I've settled on a different approach: use the model with better technical inference (like Le Chat for this case) and then do the prose tightening myself. It's faster than juggling outputs and maintains editorial control over the tone. The AI gets the facts right, I make it sound like a human wrote it.


infrastructure is code


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

That's a really solid test setup. Your observation about Le Chat feeling "plugged-in" to the DevOps mindset is exactly what makes it my current go-to for drafting API change logs or internal tool docs. It seems to have a built-in filter for what's actually actionable.

I have to echo your point about ChatGPT-4o leaning generic. I've found it often defaults to a safe, middle-ground tone that needs explicit rewinding for an internal tech audience. A trick I use is to append a context seed to the prompt, like "Write this in the style of our team's internal #dev-news Slack channel." That usually pushes it past the generic corporate template.

Did you notice any major differences in how they structured the call-to-action or next steps? Some models bury the "how to get started," while others lead with it, which can really affect how engineers engage with the announcement.


api first


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

Interesting test! The point about Le Chat adding flag examples without being asked is cool. It makes me wonder how much of that is from seeing actual CLI tool docs in its training data.

When you said it felt the most "plugged-in" to a DevOps mindset, does that mean it also structured the email with the most relevant sections for engineers, like installation steps or a link to the repo first? Or was it more about the tone?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Yeah, the follow-up prompt count is a great metric to track. I didn't do a rigorous count this time, but anecdotally Claude often needs a second nudge like "make this less formal for an internal engineering team," which adds another loop.

That "plugged-in" feel from Le Chat does seem to reduce that back-and-forth. It's like it starts from a better default assumption about the audience.


Keep it civil, keep it real.


   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The outdated flag problem is real. I've had Claude suggest using Asana Portfolio permissions that were changed two versions ago.

Polished prose can definitely hide thin content, but sometimes a press release tone is actually required. Legal or marketing might demand it for external announcements, even if the engineers hate it. The friction isn't always avoidable.


your mileage will vary


   
ReplyQuote