Been using Elicit to speed-read academic papers for my team's lit reviews. The default quality prompts are okay, but I needed something more tailored for systems research, especially around k8s and edge computing.
I built a custom prompt that checks for the specific red flags I care about. It's basically a checklist I run against any paper's abstract and conclusion before diving in. Saves me hours.
Here's the core of my prompt setup in Elicit:
```
For the paper "[${title}]", assess the study quality for systems/engineering research. Focus on:
1. Methodology: Is there a clear, reproducible experimental setup? (e.g., cluster specs, workload details)
2. Artifact Availability: Are code, configs, or datasets provided? (GitHub link, Docker image)
3. Baseline Comparison: Are comparisons made against established benchmarks or state-of-the-art?
4. Scale & Realism: Is the evaluation done at a meaningful scale? (node count, request load)
5. Limitations: Does the paper explicitly discuss weaknesses or edge cases?
For each point, provide a short verdict: Strong/Moderate/Weak/Unmentioned.
```
This forces a structured response. If "Artifact Availability" comes back as "Unmentioned," I'm immediately skeptical. If "Scale & Realism" is "Weak" (e.g., "tested on a 2-node minikube cluster"), I know the claims might not hold up in production.
I've tuned this over time. For example, I added the "limitations" check after getting burned by a few papers that made bold claims about service mesh performance but didn't mention the network overhead.
It's not perfect, but it gives me a consistent first-pass filter. Anyone else crafting custom prompts for technical paper reviews? Curious what criteria you're using.
yaml all the things
That's a solid prompt for filtering papers. I apply a similar structured check but with a latency-first lens. For systems research, I'd add a point on measurement fidelity.
Specifically: Are latency distributions (p50, p90, p99) reported, or just averages? Averages in performance studies are often misleading. If they only report mean latency, that's a yellow flag for me - it can hide tail behavior that's critical in distributed systems.
You might also consider adding a check for whether the evaluation environment's noise floor is characterized. Virtualized/cloud instances can add variable jitter.
sub-10ms or bust