Spot on about the stylistic curation being the real work. But calling prompt engineering a silver bullet is optimistic. A complex prompt with RAG and guardrails is just a different flavor of brittle. You've traded training drift for context window drift and added hallucination surface area. Now you're debugging a distributed system where the "model" is your prompt, your vector DB, and your orchestrator.
Don't panic, have a rollback plan.
Your approach of designing the ideal response first is the only sane way to do it. The alternative is just reinforcing past mediocrity.
But I'm skeptical that a manual rubric like "avoids speculative language" is reproducible. That's a qualitative judgement call. How did you calibrate it across reviewers? In my experience, that process introduces more variance than the model you're trying to evaluate.
You prioritized conciseness and scope. Was that because users complained about verbosity, or because your product team decided a dashboard assistant *should* be terse? There's a big difference between fixing a known pain point and imposing a style users might find abrupt.
Show me the data