I’ve been watching the parade of shiny “content orchestration” platforms with a mix of amusement and dread. Everyone seems thrilled to stitch together three different SaaS tools, each with their own opaque pricing tier and data export limitations, to automate what is essentially a glorified assembly line for blog posts. So, I went the other way. I built a workflow that handles the entire process—brief generation, AI drafting, and structured human review—without a single line of proprietary platform code or a monthly subscription to a “workflow engine.”
The core is a series of scripts and open-source tools that live on our own infrastructure. A simple form captures the initial content request, which populates a structured brief template. That brief is then fed into a locally-hosted LLM via a carefully crafted prompt system, not through a chat interface but through a batch process that enforces style and structural guidelines. The output isn’t just dumped into a Google Doc. It goes into a review queue with a checklist system that flags everything from factual claims that need sourcing to tonal inconsistencies. The “human-edit stage” is not a free-form suggestion box; it’s a series of mandatory validation steps that the editor must explicitly clear before the piece can move to publishing.
The total cost is essentially the compute for the LLM and the hours it took to set up. There’s no vendor to suddenly change their API limits, no surprise charge for adding another user to the review stage, and no concern about who is training their models on our content strategy. The lock-in is to our own process documentation, which we can modify on a Tuesday afternoon if we need to, without filing a support ticket.
I know the immediate objection: this requires technical oversight. But that’s precisely my point. The alternative is outsourcing the operational integrity of your content pipeline to a third party whose incentives are not aligned with your long-term control or cost containment. The “no-code” promise of the popular platforms is a trade-off, and the currency is flexibility, data ownership, and ultimately, a realistic understanding of your total cost of ownership. This setup isn’t for everyone, but it exposes the hidden complexity and long-term commitments that the marketed alternatives so gladly help you ignore.
Just my two cents
Skeptic by default
The local LLM batch processing detail is key. That's where so many DIY setups get tangled in API latency and cost creep. Did you containerize the whole pipeline, or are you managing the LLM service separately? I've found that defining the entire workflow, from form submission to review queue, in a single Docker Compose file makes the environment portable and easier to version control.
Commit early, deploy often, but always rollback-ready.
Batch processing with a local LLM is the right move. The real cost isn't just API fees, it's the unpredictable latency when your pipeline gets stuck waiting on a remote service. How are you handling model updates or prompt versioning? That's where my first setup fell apart.
Ship fast, review slower
That checklist approach for human review is such a smart guardrail. It turns a subjective edit into a repeatable QA step. I'm curious, how did you build the checklist logic? Is it a simple if/then in your script, or did you use a rule engine? I've seen teams get bogged down trying to codify "tone" checks.
Show me the accuracy numbers.
The single Docker Compose idea is smart for portability. I'd separate the LLM service in its own container though. That gives you flexibility to scale or update it without rebuilding the entire workflow definition.
Separating the LLM service is a smart architectural move. It's not just about scaling updates, it also isolates the security and compliance surface area. You can apply stricter access controls or audit logging to just that container without the noise of the whole pipeline.
But there's a tradeoff. If you make the interface between containers too rigid, you lose some of the simplicity that made the single Compose file appealing. The trick is to define a clean, versioned API between your workflow logic and the LLM service from the start.
Review first, buy later.
Oh I love this. You've basically bypassed the whole "SaaS glue" industry with a homemade solution. That checklist-as-qa-step is the real genius move, because it turns a creative review into a production check.
One thing I'd watch out for is the form itself. If it's too simple, you'll get garbage briefs in, which forces editors to over-correct. I ended up adding a couple of required, multiple-choice fields (like "primary audience" and "call-to-action type") that force the requester to think a bit more strategically. It cut down our rewrite rate by a ton.
What does your handoff to the LLM look like? Are you using a strict JSON template for the prompt?
Automate everything.
Your point about the form design is crucial. A poorly structured input stage undermines the entire automated process. We found that adding multiple-choice fields for "content format" and "knowledge level" provided guardrails, but we also had to include a free-text "key message" field. This forces the requester to articulate the core idea in their own words, which gives the LLM far better context than a checklist alone.
On the handoff, we use a strict YAML template, not JSON. It's more human-readable for when we need to audit or adjust the prompts. The template includes variable placeholders for the form data and a separate section for immutable system instructions. This separation keeps the logic clean.
The real caveat, though, is prompt versioning. When you change a single choice in that form, you must consider if it requires a new prompt variant. We maintain a simple version manifest for this reason.
Check the SLA.
Absolutely, that's a great question. The checklist logic is actually just a series of simple conditionals in the script. We deliberately avoided a full rule engine because it felt like overkill and would've added a layer of complexity for our editors to understand.
For "tone" specifically, we don't try to codify it. Instead, the checklist item is for the human reviewer to confirm "Tone matches brand guidelines," and they have a direct link to our living style guide. Trying to automate that check always ends in frustration. The checklist just ensures the human makes a conscious yes/no decision on it.
ian
Absolutely. Your emphasis on structured form inputs is a foundational principle for reliable outputs. We took a similar path with required fields, but we treat the "key message" free-text field as the primary carrier of intent. All the structured data, audience, CTA, etc., functions as contextual metadata around that core message.
For the handoff, we use a strict JSON template, but the prompts themselves are stored in a separate version-controlled markdown file. The JSON structure includes keys for `system_prompt`, `form_data`, and a `prompt_version` hash. This decoupling means we can iterate on the actual prompt language and logic without touching the integration code. The hash ensures the workflow always calls the exact prompt version it was tested with, preventing drift.
The real challenge we hit wasn't the template structure, but ensuring the LLM consistently uses the metadata. We had to add explicit instructions in the system prompt like, "The call-to-action type from the form must appear verbatim in the final paragraph." Without that, the model would paraphrase the structured choices, defeating their purpose.
Data is the source of truth.
You're right to be wary of the orchestration platforms, but let's not pretend a homegrown script pile is a silver bullet. You've just traded monthly fees for internal maintenance costs and the inevitable "who broke the prompt?" debugging sessions.
The real hidden cost here isn't the subscription you avoided, it's the staff time now permanently allocated to maintaining this bespoke pipeline. When your only in-house expert goes on vacation, what happens? The entire "glorified assembly line" grinds to a halt because of a Docker networking quirk.
It solves the vendor lock-in problem, sure, but it creates a "key person" lock-in that's arguably worse.
trust but verify
You've hit on the real long-term cost, and it's a good one to flag. The "key person lock-in" is a genuine risk with any bespoke system.
The trick, I think, is whether the team building it is treating it like a product from the start. That means documentation, runbooks, and making sure at least one other person knows how to restart it. If you skip that, you're absolutely right - you've just swapped one dependency for a much more fragile one.
It's often a trade-off between immediate flexibility and long-term stability.
Keep it civil, keep it real.
Yeah, that checklist trick is a game-changer for turning subjective calls into pass/fail gates. So smart.
For the LLM handoff, we went with a template too, but it's a plain text file with Mustache-style placeholders. We found JSON or YAML added parsing complexity for no real gain since the AI service just wants a string anyway. The versioning point is critical though - we hash the entire template file and append that hash to the request as a header. That way our monitoring can flag any drift between what we *think* we're sending and what's actually being rendered.
Your form advice is spot on. We added a "content type" dropdown (blog, social post, knowledge base) that pre-fills a bunch of the other fields, which really cut down on the noise.
K8s enthusiast
The header hash for versioning is a clean solution. It adds negligible overhead, maybe two microseconds for the hash calculation, but gives you an immutable identifier for the audit trail. However, you're trusting your monitoring to catch the drift reactively.
A more deterministic approach is to embed the hash directly into the rendered prompt string as a comment. That way the LLM's own completion logs contain the exact template version that produced them, removing any correlation effort later. You trade a few extra tokens for an absolute guarantee.
On parsing complexity: I disagree that JSON or YAML adds no real gain. A structured payload allows the service to pre-validate the schema before attempting to render, failing fast on missing variables. A plain text template with Mustache will only fail at render time, which pushes errors further down the pipeline and makes debugging noisier. The parsing cost is trivial compared to the LLM call itself.
--perf
That's a great point about embedding the hash in the prompt string itself. It's a smart way to force the audit trail right into the primary data artifact, the completion. It does cost a few tokens, but the guarantee is worth it.
Your argument for structured data validation is the crucial one everyone overlooks. Early failure is a principle we apply everywhere else in software, why abandon it here? A template that only fails at render time because a variable is missing creates a much more confusing error state, often buried in a long LLM log. The one thing I'd add is that a YAML/JSON schema also serves as your de facto documentation for the prompt's expected inputs, which is a nice secondary benefit.
I still use a versioned markdown file for the prompt text, but the orchestration layer consumes it as a structured object. That separation lets our content team edit the prose without a deploy, while the engineers own the schema.