Hey everyone! I'm new to the monitoring world, but I just automated something that felt pretty cool. I keep having to compare pricing PDFs from different cloud monitoring vendors (Datadog, New Relic, etc.) and it was eating up so much time.
I built a simple workflow using `curl` and `pdftotext` (from poppler-utils) to scrape and extract text from a list of PDF URLs, then used `grep` to pull out key pricing lines into a single report. It's not fancy, but it works!
```bash
#!/bin/bash
# List of PDF URLs
PDF_LIST="pricing_urls.txt"
OUTPUT_FILE="price_comparison.txt"
echo "### Pricing Comparison $(date)" > $OUTPUT_FILE
while read url; do
echo "Processing $url..."
FILENAME=$(basename "$url")
curl -s "$url" -o "/tmp/$FILENAME"
pdftotext "/tmp/$FILENAME" "/tmp/${FILENAME}.txt"
echo "--- $FILENAME ---" >> $OUTPUT_FILE
grep -i -E "per host|per GB|per month|billing" "/tmp/${FILENAME}.txt" | head -5 >> $OUTPUT_FILE
echo "" >> $OUTPUT_FILE
done < "$PDF_LIST"
echo "Report saved to $OUTPUT_FILE"
```
Has anyone else tried something like this? I'm sure there are better tools (maybe even something in Prometheus/Grafana for tracking changes?), but I was happy to get this running 😅. Would love to hear how you handle doc comparison.
So you're scraping their pricing pages with `curl`. Did you check the robots.txt on any of those vendor sites first? Or consider what happens when they change their PDF structure next month? That grep for "per host" is going to miss any regional pricing tables or private offers.
Also, you're leaving PDFs in /tmp for anyone on that box to read. Not a huge deal for public data, but you should at least `curl -s "$url" | pdftotext - -` and pipe it instead.
You're tracking costs, which is good. But you're missing the real nightmare: commitment discounts, egress fees, and support add-ons. Those are usually buried in separate docs.
- Nina
That's really clever! I was just looking at Datadog's pricing the other day and it was such a headache. I never thought about trying to automate it like that.
When you run your grep for things like "per host," are you catching all the different plan names too? I got lost between their "Pro" and "Enterprise" columns.
Also, do you have this running somewhere regularly, or just when you need a fresh snapshot?
That's a really interesting approach. I've been looking at vendor pricing PDFs too, and I've found they sometimes list the same thing in multiple sections with slightly different wording. Your grep for "per host" might miss something like "monthly cost per monitored host."
Have you thought about adding a step to look for common units like "GB," "TB," or "hour"?
Good point about the wording variations. That's why regex alone usually fails after the first minor PDF update.
Adding unit searches helps, but you'll still miss the real cost drivers. You need to parse the actual pricing tables, not just lines with "GB." Vendors love burying minimum monthly fees in the footnotes of those tables.
And none of this catches the annual commit discounts, which often aren't even in the main pricing PDF. You're just automating the wrong 20% of the problem.
show me the bill
You're absolutely right about the annual commits. I've had to call sales reps directly for those discounted rate cards more times than I can count. They're intentionally kept off the public PDFs.
The footnote point is a killer too. A table might show $0.10 per GB, but the tiny print says "minimum $500 monthly per product." You can grep all day and still miss that.
It's a good first script for the baseline, but you're automating a moving target. The real number is often a conversation.
Keep it civil, keep it real.
Your pipeline approach is correct, but `pdftotext` will lose table structure, which is where most pricing details live. I'd suggest piping the output through a tool like `camelot-py` or `tabula` first to extract tables into CSV, then process those.
You're also missing error handling. Add a check for the HTTP status code from curl, or you'll silently process empty files when a PDF URL changes.
Here's a quick addition for the table extraction step if you want to stay in bash:
```bash
pdftotext -layout "/tmp/$FILENAME" - | sed -n '/^ *$[0-9]/p'
```
The `-layout` flag tries to preserve column positions, and the sed filter catches lines starting with dollar amounts. It's brittle, but better than pure line grepping.
benchmark or bust
You're spot on about `pdftotext -layout` for preserving some structure. I've had decent luck combining it with `awk` to reconstruct simple columns, but anything with merged cells or weird formatting falls apart fast.
The error handling point is critical too. That script will happily churn on a 404 page if a vendor changes a URL. I usually wrap the curl call in something like:
curl -f -s "$url" || { echo "Failed: $url" >&2; exit 1; }
That `-f` flag makes curl fail silently on server errors. At least then you know something broke.
But for real table extraction, I gave up on pure bash and wrote a small Python wrapper around Camelot. It's the only thing that reliably catches those footnoted minimums inside table cells.
K8s enthusiast
That's a really smart way to start tackling a tedious task. I'm in a similar boat with comparing email service provider contracts.
Your script gave me an idea for my own use case. But I'm curious - how are you handling different column layouts across PDFs? For instance, one vendor might have "per host" in the first column of a table and another might have it in a header row. Does your grep still pull the correct adjacent price, or are you getting a lot of mismatched data?
You're right that Datadog's PDF is a particular headache with all those plan names. But even if you catch "Pro" and "Enterprise" with your grep, you're still just scraping their marketing copy.
The real question is: when was the last time someone at your company paid the *listed* price for an enterprise plan? There's always a "custom" column that exists only in a sales rep's email. Running this regularly just gives you a false sense of having current data.
Data skeptic, not a data cynic.
The `-layout` flag helps, but it's still a crapshoot with multi-page tables or anything that uses cell shading. I've seen it collapse columns that just happened to align badly.
Error handling for curl is basic hygiene. I'd add a timeout flag too - some vendors' sites hang forever on HEAD requests. Something like:
curl -f -s --max-time 10 -o "/tmp/$FILENAME" "$URL"
And you're right about tools like Camelot. The problem is it's a heavy dependency for a pipeline. If you're already in Python for other steps, go for it. But if this is part of a lightweight cron job on a bastion host, you're stuck with the bash hackery.
shift left or go home
You've put your finger on the exact tension point. The "lightweight cron job on a bastion host" scenario is the real-world constraint that makes these beautiful Python solutions non-starters for a lot of us.
I've settled on a two-tier approach because of that dependency problem.
- A scheduled bash script using `pdftotext -layout` and your `--max-time` flag does the regular, automated scraping. It's brittle but it's a watchdog for major changes.
- Then, for actual procurement evaluations, I run a separate, manual Python process with Camelot on my laptop against the PDFs the bash script has flagged as updated. That's where I catch the shaded cells and footnotes.
It's not elegant, but it accepts that full automation for this specific task is a mirage. The bash job's main value is telling me *when* I need to do the manual deep dive.
null
That exact moment when you get a script to run and actually output *something* is the best feeling, right? Been there!
You've hit on the classic entry point for this task. I started almost the same way. The `grep -i -E` for those key terms is solid for a first pass to see if anything's changed at all.
But you'll run into the column alignment problem others mentioned really fast. Your grep will pull the line with "per host," but the actual dollar amount might be three columns over, and `pdftotext` without `-layout` just strings it all together. I'd at least swap to `pdftotext -layout "/tmp/$FILENAME" -` and pipe it directly into your grep. It keeps some spatial awareness, so the price has a better chance of being on the same line your pattern catches.
Also, heads up on the `head -5` - sometimes the first five matches are all from the cover page or a FAQ section. I found adding a context flag like `grep -B1 -A1` (shows one line before and after) gave me a much better clue about what I was actually looking at.
Have you thought about how you'll track changes over time? Like, does your report append or overwrite? That was my next headache after getting the first snapshot.
Yeah, the cover page thing is real. I tried grep -B1 -A1 too, but then I'd get the page footer mixed in. I ended up piping through a quick `sed '/^$/d'` first to drop blank lines and make the context a bit cleaner.
Good question on tracking changes. I'm just dumping to a dated CSV for now. It's messy, but at least I can diff them. How do you handle the actual comparison once you have two snapshots? Do you have a script for that, or do you just eyeball the diffs?