Just tried the new Expert Mode toggle on a few technical drafts. Initial impression: it's longer, but not consistently more accurate.
Ran two tests:
* Asked it to outline a CI/CD pipeline using GitHub Actions for a containerized app. The basic output was vague but usable. Expert Mode added unnecessary complexity—suggested a multi-stage Docker build when the prompt didn't specify it.
* Asked for a comparison of two infrastructure-as-code tools. Standard mode gave a decent feature list. Expert Mode produced a more structured table but introduced a minor factual error about a core feature.
My take: It seems to prioritize adding structural detail over verifying factual depth. For boilerplate or structuring known info, it might help. For true expert content, you still need to fact-check everything.
Has anyone run benchmarks on technical accuracy? I'm more interested in error rates than word count.
Proof in production.
Exactly. They're mistaking verbosity for expertise. Your multi-stage Docker example is perfect. It's adding assumed complexity instead of responding to the actual prompt.
I haven't seen accuracy benchmarks, but I've noticed the same pattern with CRM workflow prompts. It'll add more conditional logic steps, but sometimes misstates the native fields available in the platform. Structure over substance.
Until they publish error rates, assume it's just a formatting toggle.