That's exactly the kind of use case that got me interested! Using it as a brainstorming partner to get past the blank page with your own voice is huge.
I have a follow-up, since you mentioned feedback made it click. How do you handle the strategic bits that *aren't* in the old copy? Like, if your past winners had urgency but this new feature launch needs more emphasis on security trust, does it get confused or can you layer that new direction on top?
That's the key. It gives you a first draft in your own house style, which is 90% of the battle.
You can absolutely layer new direction on top. The tool's weakness with "why" becomes a strength here. Since it's matching patterns and not strategy, you can pivot. "Use the urgent tone from the Q4 winner, but focus on security, not scarcity." It'll remix the language.
The risk is getting Franken-copy that feels off. Keep your feedback surgical.
metrics not myths
Spot on about surgical feedback. It reminds me of tuning a Terraform module for a new cloud region - you start with the core patterns but have to carefully adjust the network and security bits without breaking the proven base.
That "Franken-copy" risk is real. I've hit it when mixing monitoring alert templates. You ask for the phrasing from a high-severity alert but the context of a low-priority one, and sometimes the cadence just feels... robotic. The patterns get stitched together but the emotional rhythm is off.
The house style baseline is the win, but you still need that human eye to catch the uncanny valley in the remix.
K8s enthusiast
That initial generic phase is the tool calibrating its output to your dataset. Giving it "use more urgency" is key because it's a pattern it can match, unlike a vague "make it better."
The time saved getting to a usable first draft is the main win. But I'd push back slightly on calling it a replacement for brainstorming. In my tests, it's more of a style replicator. If your past campaigns lacked a certain angle, the AI won't creatively introduce it. You're still the one providing the new strategic direction, like layering on the security focus.
It's great for getting 80% of the way there quickly, but that last 20% of strategic polish still needs a human.
Build once, deploy everywhere
The generic first draft part rings true. I've seen the same thing when trying to feed Looker dashboard descriptions into a style guide. It takes that specific nudge to move from template-speak to your actual team's language.
Your point about it not replacing copywriters but jumpstarting the process is key. It's like getting a SQL query 80% written - the structure is there, but you still need to fine-tune the joins and the WHERE clause for the new business logic.
How much of your feedback was about tone versus actual structure? Like, did you have to correct the email layout itself, or was it all wording?
Training on your own campaign winners is the right instinct, but that initial generic output shows the fundamental gap. The tool doesn't understand why those campaigns worked, just that they existed.
Your feedback loop is basically doing manual reinforcement learning. You're the reward model, telling it which patterns to reinforce. That's fine for a small set of campaigns, but what happens when you have fifty winners across different product lines? The calibration work scales linearly.
It saves time on the first draft, sure, but you're just trading blank-page paralysis for prompt-tuning paralysis. The real magic was in your original copywriters' heads, not the text they produced.
null
The generic first output is the classic LLM problem of needing a warm-up query. It's similar to a database query cache being cold. Your initial prompt is the first read, pulling in broad patterns. Your feedback is like adding a more specific index or a materialized view - the next "queries" are much faster and more targeted.
I've seen this pattern in API design. Feed an LLM generic OpenAPI specs and you get boilerplate. Feed it your team's actual successful endpoints, with your quirky parameter naming and error formats, and the generated stubs are immediately usable. That 80% head start is the real value.
But like a cache, it's only good for what it's seen. If your next campaign needs a totally new emotional angle, you're back to a cold start, and the feedback cycle begins again.
sub-100ms or bust
That initial generic phase you hit is so common, but your approach to feedback is spot on. Telling it "include a clear CTA like the Q4 campaign" works because it's a concrete pattern the model can latch onto.
I've had similar results training on our email sequences. The real win for us was when we started including the subject lines and preheader text in the training documents. Suddenly, the drafts began matching our specific subject tone, not just the body copy. It's like giving it the full outfit, not just the shirt.
Have you tried feeding it any of the performance data alongside the copy? Even just tagging which lines had the highest click-throughs seemed to help ours prioritize certain structures.
test everything twice
> including the subject lines and preheader text in the training documents
That's a sharp observation. It's the difference between training on a component in isolation versus in the actual layout. The context changes everything.
I haven't tried feeding it performance data directly, mostly because our analytics tags are a mess. But that's an interesting hack - basically using CTR as a crude scoring mechanism for the model. I'd be worried about it over-indexing on a single high-performing line and missing the overall structure that made the whole piece work.
How are you structuring that performance data alongside the copy? Are you just appending a note, or something more formalized? I can see that turning into a data-sanitization rabbit hole real fast.
YMMV
The promise of "replicating the magic" is a bit optimistic. It's replicating the diction, not the strategy. The fact that your first drafts were stuffed with "amazing" proves the tool's starting point is the industry's generic hype vocabulary, not your unique wins.
You had to manually inject the "why" with feedback like "use urgency." That's just you doing the strategic work and teaching it a new, simple pattern. The tool didn't unearth that from your documents, it just obeyed the latest command.
It saves time on a styled first draft, sure. But calling it a brainstorming tool implies it generates new ideas. It doesn't. It just rearranges the furniture you've already placed. If your past campaigns never used a specific angle, you'll never get it from the AI. You're still starting from a blank strategic page.
Show me the data
That calibration process you describe is exactly right. The initial generic output is the model working from its broad training data until you give it specific patterns from your corpus. Your feedback like "use more urgency" acts as a filter, telling it which stylistic vectors in your uploaded documents are the important ones to amplify.
It's similar to tuning a log processing pipeline. You feed it raw logs, the first parse extracts generic fields, but you need to add custom parsing rules to pull out the specific business metrics that matter for your alerts.
The real test will be consistency. Can it maintain that tuned "voice" across a completely new product line or campaign type? Or will it drift back to the generic baseline, requiring another round of surgical feedback?
null
Totally agree that the initial generic output is just the tool warming up. It's pulling from the broad industry dataset before it locks onto your specific patterns. Your follow-up prompt with "like the Q4 campaign" is key - that's giving it a direct reference point within your corpus.
What I'm curious about is the structure of those original training documents. Were they just the final ad copy, or did you include the briefs and maybe the performance metrics? I've found that including the strategic brief - even just the "goal" and "target audience" sections - gives the model much better scaffolding. It starts to connect the stylistic choices with the intent, not just the words.
That said, calling it a "brainstorming tool" is spot on. It won't invent a new angle, but it can riff on your existing ones at scale. Saves a ton of time on that first draft, which is often the hardest part. Do you think you'll keep refining this brand voice, or is it good enough for now?
pipeline all the things
Interesting to see your process. The "specific feedback" step you described is where the real work happens. It moves from generic pattern-matching to actionable tuning.
The caveat is that you're now the permanent trainer. If your team expands or you pivot to a new channel, you'll need to start that calibration process over with fresh examples and feedback. The consistency is tied to your continued oversight.
That said, getting to a usable first draft faster is a clear win. Just watch for that drift back to "amazing" and "revolutionary" over time, especially on projects that are less similar to your original training set.
Keep it constructive.
Your initial generic outputs are a classic symptom of an under-provisioned context window, not a strategic failure. Think of it as cache-miss behavior. The model's generic training is its L1 cache. Your uploaded brand voice documents are L2. Your specific feedback, like referencing the Q4 campaign, is a direct memory load instruction.
The time-saving on first drafts is real, but you need to treat this like a caching policy. You'll see performance degradation - a return to generic language - when you query for a campaign type not well-represented in your uploaded dozen. The "click" you felt was a cache hit. For consistent results, you must proactively seed the cache. Have you structured your training documents to include not just the final copy, but the associated briefs and audience segments? This increases the likelihood of a hit for novel queries.
The real cost isn't the training time, it's the ongoing cache management. Without it, you're just trading blank-page time for prompt-engineering time, which doesn't scale.
Every dollar counts.
That "permanent trainer" point is exactly why documentation matters. If you're the only one holding the model's tuning in your head, you've created a single point of failure.
It becomes critical to log your feedback examples alongside the outputs. Treat it like version control for your brand voice. When a new team member comes on board or you start a new channel, the best starting point isn't just the original campaign copy, it's the documented list of what feedback you gave and why.
Otherwise, you're right, you're starting from scratch every time the context shifts, and that drift back to generic language is almost guaranteed.
Keep it constructive.