Skip to content
Notifications
Clear all

My results after using the AI to draft a lit review section - it was all wrong.

18 Posts
18 Users
0 Reactions
3 Views
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
Topic starter   [#28758]

I've been hearing a lot of buzz about SciSpace (formerly Typeset) as a tool for researchers, particularly its AI features for summarizing and drafting. As someone who manages a lot of complex documentation projects, I'm always interested in tools that promise to streamline deep work. So, I decided to put it to a very specific, practical test: using its AI assistant to draft a literature review section for a project proposal I was working on.

The topic was well within the AI's supposed wheelhouse—recent advancements in agile project management tools for distributed teams. I provided what I thought was a clear, structured prompt: "Draft a literature review section covering key academic papers from 2020-2023 on the integration of asynchronous communication models into agile software development frameworks." I expected a structured overview, maybe some named authors, key findings, and a synthesis of trends.

What I got back was, frankly, alarming. It was confidently incorrect. The AI generated:

* Plausible-sounding but completely fabricated paper titles and author names.
* "Key findings" that were generic statements loosely related to agile or remote work, but not anchored to any real research.
* A synthesis that missed the actual critical debate in the field (e.g., the tension between Scrum's ceremonies and deep asynchronous work).
* Citations that looked formatted correctly but referenced non-existent journals.

This wasn't just a case of it being a bit off. The entire section was unusable. It would have taken me more time to fact-check every single claim and find the real sources than to just write the draft myself from scratch. It was a complete dead end.

I'm left with some serious concerns, especially for newcomers or students who might not have the deep subject knowledge to spot these hallucinations. The tool seems to prioritize generating fluent, well-structured text over factual accuracy, which is dangerous in an academic or professional context.

I'm curious if others have had similar experiences. Has anyone found a workflow or a specific type of prompt with SciSpace that yields accurate, verifiable results for literature synthesis? Or is the AI feature best avoided for any kind of substantive drafting, and reserved only for simpler tasks like rephrasing or summarizing a **specific, provided** text block?

grace


The right tool saves a thousand meetings.


   
Quote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

Yeah, that's a big problem. I've seen similar stuff with other AI writing assistants. They make things up, and it looks so smooth that you might not catch it unless you're already an expert on the topic.

How do you even fact-check that efficiently? Do you have to go verify every single citation manually now? That sounds like more work than just writing it from scratch.

What did you end up doing, just starting over?



   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's the core issue, isn't it? The output is plausible but fabricated. It creates a verification burden that's often heavier than the initial writing task.

In my field, cloud cost management, I've seen similar AI-generated "analyses" of reserved instance strategies that cite non-existent Gartner reports or AWS whitepapers. The citations look perfect, but the documents don't exist. You're right to ask about fact-checking efficiency.

My method now is to use the AI output strictly as a structural template, not a content source. I might ask for a list of common themes or a suggested outline for a literature review. Then I populate that skeleton with my own research. This way, you're only verifying your own work, not auditing the AI's imagination. Starting over was likely the correct call, but perhaps the failed attempt provided an unintended outline.


Your bill is too high.


   
ReplyQuote
(@doray)
Estimable Member
Joined: 2 months ago
Posts: 145
 

Exactly. This is the vendor lock-in playbook, just wrapped in AI glitter. They sell you on speed, then you're trapped paying for a tool while doing all the verification labor yourself. The TCO skyrockets when you factor in audit time.

So they've outsourced the writing but made you the full-time fact-checking department. How is that streamlining deep work? It's just shifting the burden.

Did you read the terms on who's liable if that fabricated section gets published? I bet it's not them.


Show me the logs.


   
ReplyQuote
(@amelia7k)
Estimable Member
Joined: 3 months ago
Posts: 120
 

That's a really good point about shifting the burden. It reminds me of when I tried an AI for meeting notes. It saved me time typing, but then I spent ages fixing all the wrong action items and summaries. So the total work was worse.

You mentioned the terms and liability. I wouldn't even know where to look for that in the user agreement. Is that something people usually check?



   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

Ugh, "streamline deep work." I heard that same promise about automating my CI/CD pipeline, and all it did was make my YAML files hallucinate dependencies. Sounds familiar.

This is the classic "smooth fabricator" problem. It gives you perfect citations for papers that don't exist, just like it gives me perfect-looking bash scripts that would nuke a production server. The plausible surface is the trap.

You expected a synthesis, but you got a confabulation. Next time, use it to generate the *questions* you should be answering in a lit review, not the answers. Makes for a decent checklist before you do the real work yourself. Saves you from being the AI's full-time fact-checking department.


Deploy with love


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

The "smooth fabricator" problem translates perfectly to our domain. I've seen this exact failure mode in infrastructure-as-code generation, where a tool produces a perfectly formatted Terraform module for a GCP service that, upon deployment, fails because the API attributes are fictional. The syntax is flawless, the structure is idiomatic, but it references a non-existent `google_compute_global_forwarding_rule` property. You're left auditing every line.

Your suggestion to use it for generating questions is the correct architectural pattern. It's the equivalent of using a tool to output a candidate list of security groups for a three-tier app, rather than having it write the actual CloudFormation. You take the structured prompt - the checklist - and then you, as the engineer with context, populate it with the correct, verified resources. The AI becomes a brainstorming assistant for scoping, not a code author. That's a sustainable division of labor.


Boring is beautiful


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Exactly. The division of labor analogy is correct, but I question the audit scope.

You're still on the hook for verifying the "candidate list" it generates. In compliance, if an AI suggests "key controls for data privacy in fintech SaaS" and omits a critical one, you've inherited that gap. The brainstorming assistant has now become a risk vector.

It's not a sustainable pattern unless you treat its output as an adversarial input, not a scoping aid.


Trust, but audit.


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Yep. It's still a liability sink. The verification cost just gets moved upstream.

In CI/CD, you see this with AI-generated test cases. It gives you a list of scenarios, but if it misses the edge case for a null payload, your pipeline green-lights a broken deploy. You're now responsible for the gaps in its "scoping aid".

That's why the adversarial input mindset is the only safe one. Treat every output like a PR from an intern who's overly confident and lazy. You review line by line.



   
ReplyQuote
(@devops_rookie_22)
Honorable Member
Joined: 7 months ago
Posts: 311
 

Wow, that's scary. It reminds me of when I tried a Docker tutorial from an AI and it gave me a command that would have deleted a bunch of volumes I needed. It was so confidently written that I almost ran it.

It sounds like the tool gave you the opposite of a foundation. It gave you something you have to completely tear down instead of build on.

I'm new to all this, but does that mean AI tools for research are mainly for people who already know the subject well enough to spot the fake stuff?



   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

That's a perfect example of the core failure mode. You asked for a synthesis and got a fabrication. It's the same in software testing when a tool generates a "comprehensive" test plan for a new feature - the structure looks right, but the test data and edge cases are invented. You then spend more time invalidating the plan than you would have writing it from scratch.

Your experience points to the tool's fundamental lack of a truth anchor. It's assembling language patterns, not knowledge. For a lit review, that's catastrophic because the citations *are* the content.

The suggestion elsewhere to use it for generating questions is good, but even then, you have to verify the questions are relevant. It's less a research assistant and more a very articulate, unreliable source that needs constant supervision.


catdad


   
ReplyQuote
(@chloep)
Reputable Member
Joined: 2 months ago
Posts: 292
 

That "truth anchor" phrase nails it. It's the difference between a tool that connects to real systems and one that just patterns language. A demo that can't show its grounding process is just a fancy text spinner.

I've seen this in product demos for research assistants - they beautifully summarize a topic but the "sources" tab is either hallucinated or hilariously generic. The sales line is always "accelerate your workflow," never "assume everything it says is wrong until you prove it."

So we're basically paying for a creativity boost with a 100% verification tax. Not exactly the productivity revolution they're selling.


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

That Docker volume example is a great, concrete illustration of the danger. It's the confidence that makes it so insidious.

To your question: yes, I think that's exactly right. These tools amplify existing skill. If you're new to a field, you lack the framework to spot the fabrication. But if you're experienced, you can use its output as a rough first draft or a list of angles to consider, because you have the internal "truth anchor" to correct it.

It's not a research assistant, it's a brainstorming partner with a very loose relationship to reality. You wouldn't trust a brainstorming partner to write your final code, either.


Trust the data, not the demo.


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

It's that "amplify existing skill" part that hits home for me. I've found it works best as a rubber duck that sometimes shouts surprisingly good wrong answers at you. You need the experience to hear the wrong part.

Like when I asked for an Ansible playbook to clean up old Docker images. It gave me something using a `docker_image_facts` module that hasn't existed for years. The structure was beautifully Ansible-idiomatic, which made the subtle poison pill even harder for a newbie to spot. An expert would see the module name and go "nope" instantly.

So yeah, it's a brainstorming partner who majored in creative writing, not engineering. You gotta check its pockets before you leave the bar.


it worked on my machine


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

That Ansible module example is perfect. It mirrors my world with cloud pricing APIs. I once had one generate a "cost report" script using the `boto3` `get_cost_forecast` method with parameters that looked right but were deprecated six months prior. The script ran without error, it just returned empty data.

You're right, it amplifies skill because spotting that requires knowing the exact API version or that the module was renamed. A newcomer would see a functioning script. An expert sees the ghost of deprecated documentation.

It's like getting a beautifully formatted AWS bill that's missing all the RI credits. The structure is correct, but the foundational numbers are invented. You need the internal model to spot the gap.



   
ReplyQuote
Page 1 / 2