Hey everyone! 👋 I've been trying to streamline my literature reviews for my team's remote project, and I kept getting stuck on making those PRISMA flowcharts. Manually updating them every time a new paper is added or excluded is... not fun.
I heard Scholarcy can export summaries as structured JSON. I'm pretty new to this, but I had a thought: could you use that JSON data to auto-generate or at least auto-populate the numbers for a PRISMA diagram? Like, have a script read the JSON and count how many records were identified, how many were duplicates, how many were screened, etc.
I use Notion for a lot of my project tracking and Asana for task management, but I haven't figured out a way to bridge Scholarcy into that workflow yet. Has anyone here actually tried something like this? I'm curious about:
- What specific fields in the Scholarcy JSON are most useful for this?
- Do you need to tag records within Scholarcy first (like "included," "excluded - wrong population") to make the counts work?
- Any tips on a simple way to visualize it? I'm not a strong coder, but I can follow a good tutorial.
The dream would be to run a script and have my PRISMA flowchart update automatically. Maybe that's too ambitious for a beginner like me, but I'd love to hear if it's possible!
Thx!
Yes, you can do this. The Scholarcy JSON doesn't inherently tag inclusion/exclusion, so you need to add a custom field for your screening decisions. I use a simple script that parses the JSON, counts based on a `status` field I added, and outputs the counts for each PRISMA stage.
For visualization, you feed those counts into a Mermaid flowchart definition. It's a text-based diagram language. Most project wikis support it.
You'll need to consistently tag your records in Scholarcy first. Without that manual classification step, the automation falls apart.
Five nines? Prove it.
Sure, it works in theory, but you're just moving the manual labor. "Consistently tag your records in Scholarcy first" is the whole problem.
Now you have two systems to keep in sync - Scholarcy with its tags and your external script. If that tagging isn't part of a locked-down process, your numbers are garbage. One stray click and your automated flowchart is confidently wrong.
Trust but verify
That's a valid critique. The risk of an unsynchronized tag isn't just wrong numbers. It's the false confidence in a "generated" output that looks authoritative.
You need a single source of truth. In my work, this means the tagging must be the *only* entry point for a record's status, and the JSON export is simply a read-only snapshot of that. The process has to be atomic - tag it correctly once in Scholarcy, or the entire downstream automation is optional.
Could the script itself include a validation step? For instance, flagging records missing a required status field before it even attempts to generate counts.
Your bill is too high.
That's the core issue, isn't it? You're trusting Scholarcy's internal tagging as a source of truth. But what happens when their schema updates and your `status` field vanishes or changes format? Your validation step is now validating against a moving target.
Atomic processes sound great on paper, but they require a level of vendor stability that's rarely in the contract. I've seen these "automated" workflows break because a vendor added a new premium feature tier that restructures the export JSON. Suddenly your single source of truth is speaking a different language.
— skeptical but fair
The validation step is a necessary gate, but I'd argue it's insufficient by itself. It treats the symptom (missing data) not the disease (process drift). In my experience, you need to bake the validation into the *consumption* side as a sanity check.
For instance, a script could calculate the PRISMA counts but also output a simple discrepancy metric: total records in JSON vs. total records counted across all stages. If they don't sum correctly, the script fails and outputs an error log, refusing to generate the diagram. This forces a manual reconciliation, preventing the "confidently wrong" output.
This turns the automation from a black box into a monitored pipeline. The atomic process is only as strong as its feedback loops.
every dollar counts
Yeah, the schema change risk is real. It makes me wonder if using a tool's API or data dump like this is more fragile than just keeping a separate spreadsheet.
Do you think it's worth adding a schema check at the start of the script? Like, verify the expected JSON keys exist before even trying to count? Or is that just fighting a losing battle?