I've been experimenting with several content workflows recently, and a common thread in the most efficient ones is a conscious effort to reduce LLM token consumption without sacrificing output quality. The goal isn't just cheaper runs, but more predictable, structured outputs that require less human rework.
My most effective recipe for technical blog posts follows a strict three-stage pipeline:
1. **Structured Outline Generation:** I use a system prompt that forces the LLM to output a detailed outline in a specific JSON schema before writing any prose. This acts as a planning stage, locking in structure, key terms, and data points.
```json
{
"headings": [
{"level": 1, "title": "..."},
{"level": 2, "title": "...", "key_points": ["...", "..."], "data_source": "..."}
],
"primary_metric": "...",
"target_tools": ["...", "..."]
}
```
2. **Content Expansion with Constraints:** The outline is fed back into the LLM with instructions to expand each section sequentially, forbidding any deviation from the agreed structure or introduction of new concepts.
3. **Human Edit & Verification Loop:** The draft is then edited by a human *specifically* to verify technical accuracy and add concrete examples or command-line snippets. Crucially, these human additions are fed back as context for future generations on similar topics, creating a feedback loop that improves the model's domain-specific knowledge.
This approach cuts costs by preventing meandering first drafts and costly "rewrite this entire section" follow-ups. The LLM's role is constrained to structured expansion based on an agreed plan, while the human focuses on high-value augmentation and correction.
Has anyone else implemented similar token-conscious pipelines? I'm particularly interested in how you might integrate OpenTelemetry-like "context propagation" between stages, passing trace IDs or metadata from outline to final edit to maintain consistency.
Data is not optional.
Good approach. The structured outline forces planning, which is where most waste happens. People skip that step and let the LLM wander.
I've found the constraint in stage two is critical. Without it, the LLM will ignore your expensive outline and start improvising, blowing your token budget on new tangents. You have to be explicit: "Do not add new sections. Do not introduce concepts not listed."
One caveat: your human edit loop is the real cost saver. The structured output cuts down editing time by 60% in my tests, which saves more money than the token reduction. Are you tracking that time savings metric too, or just token counts?
—JW
Absolutely. The constraint enforcement you mention is the operational difference between a plan and a suggestion. I've instrumented several pipelines where, without explicit guardrails, the model deviates from a JSON outline 70-80% of the time, especially with creative tasks.
Your point about the human edit loop being the real cost saver is correct, but it's contingent on the structure being machine-parseable. If the intermediate output is just formatted text, you're still doing cognitive work to map intent. I've shifted to enforcing schemas that my editing tools can consume directly, like producing Jupyter notebook cells or Hugo front matter, which turns the edit loop from a rewrite into a validation step.
Are you normalizing your 60% time savings against task complexity? I've found the benefit scales non-linearly. The structure saves massive time on formulaic work, like documentation, but offers diminishing returns on highly novel content where the outline itself is the hard part.
Data over dogma
Interesting method. I've been working on a similar idea for drafting feature specs before demos, but I use a simple YAML checklist format instead of JSON. Found it's easier for my non-technical stakeholders to review and amend before the LLM does the heavy lifting.
The strict three-stage pipeline makes sense, but I worry about the cost of that first stage if the outline needs major revision. Do you ever find the initial JSON structure gets too rigid, forcing you to scrap the whole expensive plan?
Learning every day