Scraping the HTML is often *more* stable for features a vendor considers a display layer, as the UI tends to change less frequently than their internal API contracts. The risk is that you're still building a parser for a structure they never intended for export, which can be brittle.
Given the core issue is LaTeX compatibility, the browser extension route would at least give you access to the raw text before any CSV export mangles it. You could write a local script that ingests the scraped JSON and runs it through a dedicated LaTeX escaping function, preserving your `cite{}` and Greek symbols. This decouples the data extraction from the format translation, which is a cleaner architecture than trying to fix a broken CSV after the fact.
—Alex
Yeah, the CSV export mangling special characters is a classic encoding trap. Before you dive into a full script, check if you can force the CSV export to UTF-8-BOM or try opening it in a plain text editor like Notepad++ to see the raw encoding - sometimes a quick fix there saves you.
But you're right, the flat structure is the real killer. Pandoc won't help with that hierarchy. If their API is sparse, the browser extension scrape idea that came up later is actually a solid next step for investigation - you might get the raw data structure before they flatten it for the CSV.
Automate the boring stuff.
Totally feel your pain, the promise of a unified system falls apart the second you need to get your own data out cleanly. Been through similar with other platforms.
> Using their API directly
From my experience, when the export options are this bad, the API often treats those features as second-class citizens too. Sparse docs usually mean they haven't built it for external integration, so you're likely to hit weird limits or missing fields. I'd test a quick API call for a single annotated paper before investing any real time there.
The browser extension scrape idea that came up is actually the most promising for your immediate problem. It grabs the raw text before their export pipeline mangles it. Pair that with a local script to handle LaTeX escaping, and you might salvage your `cite{}` tags and Greek letters without rebuilding everything from a flat CSV.
Still looking for the perfect one
Agree on the API being second-class. If the docs are sparse, there's probably no versioning or deprecation policy either. A script against that endpoint is a ticking time bomb.
The scrape-to-raw-text idea is the right direction. You can isolate the encoding/escaping logic for LaTeX as a separate step, which is way easier to maintain than trying to reverse-engineer a broken CSV later.
Just make sure your scraper logs the exact element selectors it uses. When the frontend changes, you'll need to know what broke.
—cp