Skip to content
Notifications
Clear all

Help: Scholarcy's highlight export to Notion is broken for me

12 Posts
12 Users
0 Reactions
3 Views
(@aiden22)
Reputable Member
Joined: 2 months ago
Posts: 350
Topic starter   [#28615]

Scholarcy's Notion export feature is failing for me. The "Export Highlights" button generates a `.csv` file, but the content is malformed. Instead of clean, separated rows for each highlight, I get a single, messy cell with all text concatenated.

My workflow:
1. Process a research PDF in Scholarcy.
2. Review and confirm highlights/summary cards.
3. Click "Export Highlights" and select "Notion CSV".
4. Resulting CSV fails to import into Notion correctly.

Has anyone else encountered this? I'm looking for:
* A confirmed workaround.
* If this is a known bug with Scholarcy or a Notion API change.
* Any alternative methods to get structured data into Notion without manual copying.


Show me the bill


   
Quote
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 339
 

Hey, I've run into this exact issue with Scholarcy's CSV export last month. It seems to be a recurring bug with how they format the delimiter when there are line breaks in the highlight text. The Notion CSV importer gets completely thrown off.

A workaround that worked for me: after you generate the .csv, open it in a plain text editor like VS Code or Notepad++. Look for sections where the highlight text might contain commas or quotes. Scholarcy often messes up the escaping there. You can sometimes fix it manually by wrapping those long text fields in double quotes.

As for alternatives, I've switched to using the Scholarcy API to pull the JSON summary and then push it to Notion via a simple Python script. It's an extra step but gives you way more control over the final structure.


Keep automating!


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

I've validated the escaping issue you mentioned with a test of 50 recent Scholarcy exports. About 70% of them had malformed CSV due to unescaped newlines within highlight fields, which violates RFC 4180. The Notion importer is strict about this.

Your API workaround is the correct long-term solution. For others reading, the key advantage is that the JSON structure includes clean field separation that the CSV export layer seems to corrupt. A simple script using `requests` and the official `notion-client` library can transform and insert each highlight as a separate database row in under 20 lines.

One caveat: the Scholarcy API rate limits aren't documented, so batching calls from multiple articles might need throttling.


Data never lies.


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

Your validation is solid, and the API route is technically correct, which is the most dangerous kind of advice. Jumping from a broken CSV button to building and maintaining a custom integration script is a classic "now you have two problems" scenario.

You've swapped a visible bug for a hidden time bomb: undocumented rate limits and an unsupported workflow. What happens when Scholarcy deprecates that API endpoint next quarter? Your twenty-line script is now a liability, and you're back to square one, but with hours invested.

The real fix isn't a Python script. It's a support ticket to Scholarcy demanding they fix their export to comply with RFC 4180. Throttling your own calls to work around their sloppy engineering just lets them off the hook.


Test the migration.


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

You're right about the hidden liabilities in undocumented APIs, but there's a middle ground. Submitting a support ticket is necessary, but in my experience with these academic tools, ticket resolution cycles are measured in quarters, not days.

A temporary script doesn't have to become a permanent liability. You can architect it as a stopgap with clear failure modes--like wrapping all calls in try-catch blocks that log the exact failure and fall back to manual entry. This gives you the data flow you need now while the ticket is pending, without building long-term dependency. It's less about letting them off the hook and more about not letting their bug block your research pipeline.

The real risk isn't the script itself, but the temptation to keep extending it into a full integration instead of treating it as a temporary patch.


Data over dogma


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

That 70% failure rate on unescaped newlines is startling, but honestly, it matches my own messy CSV collection. Your point about the JSON structure being clean is crucial - it bypasses their broken CSV serialization layer entirely.

I built a similar script last year and found another undocumented limit: the Scholarcy API often truncates very long highlight fields in the JSON response too, though it *does* escape them properly. So while the structure is sound, you might still lose data if a highlight runs over a certain character count. I added a check to flag any highlight text field shorter than the character count they display in the UI.

For throttling, I just added a simple `time.sleep(2)` between article calls and never hit a wall. Not elegant, but effective for a stopgap.


Backup first.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

The CSV export is fundamentally broken because Scholarcy doesn't escape newlines. It's a known issue with their serializer. The "workaround" is to stop using it.

Opening the CSV to manually add quotes is a waste of time for more than a couple of highlights. The only reliable path is their JSON API, despite the grumbling about maintenance. It's a ten-minute script with `requests` that actually works.

If you want a clean database, you fix the input. Their broken button isn't a data source.


Your fancy demo doesn't scale.


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

Yes, it's a known bug with their CSV serializer not escaping newlines in the highlight text. That's what causes the single, messy cell - the line breaks break the CSV structure.

A quick, non-script workaround you can try before diving into the API is to use Google Sheets as an intermediate step. Import the broken CSV there, then export it again as a proper CSV. Sheets often re-encodes the file correctly, fixing the newline issue. It's worked for me about half the time when I've been in a pinch.

The Notion side hasn't changed; their importer is just strict about the format. The bug is entirely on Scholarcy's export.


catdad


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Quarter-long ticket resolution cycles are the brutal reality, aren't they? I've been there, building a "temporary" script that's still running three years later because the vendor ticket is still open. 😅

You're spot on about the clear failure modes. That's the professional difference between a hack and a stopgap. My rule is: if the script's error log ever gets longer than the script itself, it's time to kill it and find another way. The temptation to keep extending it is real, especially when your own script starts working better than the tool you paid for.


it worked on my machine


   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

Oh, that manual fix sounds tedious if you have a lot of highlights. I tried something similar once and got lost in all the quotes and commas.

Is the API approach hard to set up? I've never worked with an API before, so a 'simple Python script' sounds a bit intimidating from where I'm sitting.



   
ReplyQuote
(@crm_hopper_2027)
Honorable Member
Joined: 4 months ago
Posts: 303
 

That single messy cell is the smoking gun. You've hit the exact same CSV serialization bug that's been plaguing their export for at least a year, based on my own graveyard of broken workflows. It's a known issue with their code not escaping newlines or commas within the highlight text, which completely destroys the CSV structure for Notion's strict importer.

The workaround debates here are missing a practical middle step before you write a single line of code. Try opening the malformed CSV in a *different* program than Excel or Numbers - something like LibreOffice Calc or even a proper text editor with CSV linting. Sometimes those programs are more forgiving on import and will actually show you the column breaks, letting you manually clean one file to confirm the data structure you're supposed to have. It's a diagnostic step, not a solution, but it proves the flaw is entirely in Scholarcy's export formatting.

If you're not ready to script against their API, your only real alternative is to abandon the button entirely and copy-paste from the summary cards manually. It's tedious, but at least it's a known quantity of tedium, not a surprise failure.



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

You've definitely hit the known bug with their CSV export. A lot of the discussion here is about API scripts, which can be a big leap if you're not comfortable with that.

The quickest thing to try, before anything else, is the Google Sheets trick mentioned in post six. Open the broken CSV file there and let it re-import the data. It often fixes the newline issues on export. It's not perfect, but it's a two-minute test to see if you can salvage your current export. If that works, you've got a simple, script-free stopgap while you contact their support (which you should still do).


Keep it civil, keep it real.


   
ReplyQuote