I've been experimenting with a specific use case for system prompts: turning ChatGPT into a tool for critical technical evaluation. As someone who constantly benchmarks frameworks, I've found its default tone to be overly optimistic and generic. However, you can programmatically force a more skeptical, evidence-first mindset.
Here's a system prompt I've used to analyze new API tooling announcements:
```
You are a senior backend engineer with a deeply skeptical, evidence-driven approach. You default to questioning claims and require concrete benchmarks, comparable alternatives, and a clear analysis of trade-offs before forming an opinion. Your responses should:
* Immediately identify the claimed benefits.
* List potential hidden costs or drawbacks (performance overhead, complexity, vendor lock-in).
* Demand specific, measurable evidence for performance claims.
* Compare to at least two established alternatives.
* Never conclude with unqualified praise.
```
This transforms the interaction. For example, when I fed it a press release for a new "high-performance GraphQL orchestrator," the output structure was markedly different:
* It first deconstructed the marketing language into testable claims (e.g., "50% faster resolver execution").
* It hypothesized the likely technical trade-off (e.g., "This likely pre-warms connections at the cost of higher baseline memory consumption, similar to Apollo Server's `persistedQueries` trade-off").
* It generated a list of specific metrics I should verify (P95 latency at scale, cold-start performance, introspection overhead).
* It suggested a comparative testing matrix against Apollo Server and Yoga.
The key isn't just getting a list of pros and cons; it's about forcing the model to adopt a methodology that prioritizes falsifiability and operational cost. This is far more valuable for vetting tools than the standard "Here are some great features..." response.
The limitation, of course, is that the model's "skepticism" is based on patterns in its training data, not novel insight. It won't identify a flaw that isn't represented in its corpus. But as a structured prompt for a first-pass technical review, it significantly raises the bar for the quality of analysis.
benchmark or bust
benchmark or bust
Oh wow, that's such a clever idea. I've definitely noticed the default "enthusiastic helper" vibe can be kind of misleading when I'm trying to evaluate something for my shop. I always end up having to double-check everything it says anyway.
So, to be clear, you're basically giving it a different personality or job title before you ask your real question? That's way simpler than I thought it would be. Have you tried this with non-technical stuff, like for reviewing new project management tools or email services? I'd love a skeptical second opinion on some of those sales pitches I get.
That's a solid prompt structure. I've found the same default optimism when asking about new SaaS vendors. The "never conclude with unqualified praise" directive is key - it cuts off the boilerplate niceties that aren't helpful for a real decision.
A related tactic I use is to explicitly instruct it to role-play a specific stakeholder, like a cautious procurement officer or an internal auditor focused on compliance risks. It forces the model to surface considerations like data residency or contractual SLAs that the "enthusiastic engineer" persona might gloss over. Have you noticed any diminishing returns on skepticism, where the output becomes overly negative or impractical?
Review first, buy later.
The specific directive to "demand specific, measurable evidence for performance claims" is crucial. I apply a similar principle when reviewing analytics or A/B testing platform announcements, where claims like "50% faster query performance" are common.
I've found it's also effective to add a structural requirement for the response, like forcing a comparison table. For example, when evaluating a new cohort analysis feature, my prompt instructs the model to create a table comparing the new tool against two others across dimensions like calculation granularity, historical data limits, and export functionality. This moves the output from abstract skepticism to a concrete, side-by-side evaluation.
One caveat I've observed is that you need to provide the model with the actual alternatives to compare against, or it will sometimes default to overly generic or popular choices that aren't relevant to your specific stack. The skepticism prompt improves the tone, but the quality of the comparison still depends on the inputs you feed it.
Data > opinions
Love this prompt. It's a great template. I've used a similar trick for email marketing service comparisons, where the hype is real.
Mine usually starts with "You are a cynical email deliverability consultant who's seen every vendor promise 'game-changing' inbox placement." It forces the model to ask about things like actual seed list results, reputation history with major ISPs, and contract exit fees - stuff the sales page obviously won't mention.
The one thing I'd add? Sometimes the skepticism can go too far and it'll start inventing drawbacks that aren't really industry-standard concerns. You almost need a final line in the prompt like "but keep critiques grounded in common vendor pitfalls, not theoretical extremes."
Automate the boring stuff.
Ha, the "inventing drawbacks" bit is so true. I've had it warn me about "potential IPv6 packet fragmentation issues" for an internal container registry setup. Like, c'mon buddy, we're not running a global anycast CDN here.
It's a bit like when you first get someone on your team to do code reviews and they go full "what if the CPU's L1 cache is poisoned" on a simple config change. You gotta rein them back into reality.
Your line about grounding it in common pitfalls is spot on. I might steal that for when I'm poking at new monitoring SaaS offerings. "Yes, I know distributed tracing adds overhead, but we're not Google. Focus on the vendor lock-in and data egress costs, please."
it worked on my machine
Yeah, you nailed it. I get the same issue when asking about potential Kubernetes operators or service meshes. It'll start outlining Byzantine fault tolerance requirements for a proof-of-concept app that gets three internal users. The model seems to have a "textbook perfect" architecture in its head and assumes that's always the goal.
The trick I've settled on is to explicitly state the operational context in the prompt itself. Instead of just "be skeptical," I'll add "for a team of three maintaining a legacy monolith with no dedicated SRE support." That forces the critique to be relevant to things like documentation quality and upgrade paths, not theoretical scalability limits. It cuts out maybe 80% of the noise.
Automate everything. Twice.
This is a brilliant foundational prompt for our line of work. I've adapted a similar structure specifically for data pipeline tools, and I've found the demand for `concrete benchmarks` is often the hardest to satisfy without providing the model with a concrete dataset.
For example, when I ask it to be skeptical about a new stream processing framework's latency claims, it correctly asks for the benchmark's workload shape (e.g., event size, partition skew). Since that's rarely in the announcement, I've added a step to my prompt: "If specific benchmark details are absent, propose a canonical, reproducible test scenario that would validate the claim." This forces it to outline what good evidence *would* look like, which is often more useful than just pointing out its absence.
Your point about transforming the output structure is spot on. It moves the model from generating a summary to performing an analysis.
Data is the source of truth.
It's absolutely a personality or job title override, yes. The default mode seems to be a mid-level engineer who's read the marketing copy but hasn't been burned by it yet. Giving it a jaded persona like "senior staff engineer who's seen three vendor acquisitions in this space" is basically installing the institutional memory your team lacks.
I've used it for procurement stuff. It works until it doesn't. Tell it to be a skeptical reviewer of a new project management tool and it'll correctly flag things like API rate limits and data export formats. But it has no real sense of the *social* costs, which are the actual deal-breakers. It won't tell you that switching to tool X will cause a mutiny from the design team because the UI is universally hated, or that the "seamless Jira integration" requires a consultant to set up.
So it gives you a false sense of a thorough review. You get a nice list of technical gotchas while missing the organizational ones that'll sink you.
That prompt is fantastic, and I've been using a near-identical structure for sales tools, especially CRMs with new "AI" lead scoring. It's perfect for cutting through the buzzwords.
My variation adds one line about cost-per-seat creep: "Question the scalability of per-user pricing for role-based workflows (e.g., a sales manager vs. an SDR)." Because inevitably, the shiny new feature gets rolled out and you need five more licenses just for the rev ops team to configure it.
The biggest win for me is the "never conclude with unqualified praise" directive. It forces the output to end with actionable questions or conditions, which is exactly what I need for a vendor evaluation doc. Instead of "This looks great!", I get "Proceed only if they can provide a trial instance with your own historical data to validate the scoring accuracy claims." That's gold.
hannah
Cost-per-seat creep is the real killer, especially with role-based pricing. That's where your TCO explodes.
Your point about ending with actionable conditions is key. I do the same for infrastructure tools, forcing it to produce a trial acceptance checklist. Things like "require a performance test against a snapshot of our current prod load" or "verify the SLA includes data export downtime."
The model still misses hidden operational costs, though. It won't flag that the new CRM's "AI" feature needs a full-time marketing ops person to tune and maintain it. You have to add that context yourself.
Show me the bill
Totally! It's exactly like giving it a different job title. I'm new to a RevOps role and I've been trying this trick with CRM add-ons and sales engagement platforms.
For the non-technical stuff like email services, it's hit or miss. I told it to act like a skeptical marketing ops manager and ask about things like warmup requirements and support SLAs. It was great at poking holes in the sales page, but like user1166 mentioned, it completely missed the hidden costs - like needing to hire a specialist just to manage the new tool's "smart" features.
Do you find you have to give it a really specific persona? I tried "be skeptical" for a project management tool demo and it just gave me generic questions about user limits.
Yes! That addition about proposing a canonical test is a fantastic upgrade to the basic skeptic prompt. It moves from just being a critic to being a constructive reviewer.
I've found this works really well for database or vector store benchmarks, where everyone uses a different "standard" dataset. Asking it to outline a reproducible scenario forces it to think about the variables that actually matter for your use case, not just the ones in the marketing gloss.
My only tweak: I sometimes have to add "...using publicly available datasets" to keep it practical. Otherwise it might design a perfect test that requires proprietary data you'd never have.
Show me the accuracy numbers.
The prompt structure is solid, but you're still missing the single biggest cost vector in API tooling: data transfer. That "high-performance GraphQL orchestrator" is probably just an expensive egress funnel.
Your line about hidden costs should explicitly call out cross-region and cross-cloud data flows. I've seen teams architect a "perfect" API layer that triples their AWS bill because it pulls from a GCP database per request. The model won't surface that unless you tell it to audit the data path.
cost optimization, not cost cutting
>keep critiques grounded in common vendor pitfalls, not theoretical extremes
Exactly. That's the fine line to walk with any prompt for a cost review. Tell it to be a skeptical infra auditor and it'll start warning you about cross-AZ data transfer costs for a static website. You have to bound the cynicism with reality.
For cloud services, I always add a constraint like "...focus on the 20% of resource types that typically drive 80% of the bill, like provisioned IOPS storage or NAT gateways." Otherwise you get a dissertation on rare edge-case pricing while it misses the obvious Reserved Instance mismatch that's bleeding cash every month.
pay for what you use, not what you reserve