Hey everyone! I've been trying to build a little data pipeline that pulls in summaries from Scholarcy for my team's research dashboard. I'm using their API, which is great, but I kept manually polling to check if new summaries were ready. It felt... inefficient and not very "data engineer" of me 😅
Then I saw Scholarcy mentions webhooks in their docs, and I got excited! But I'm hitting a wall trying to set them up. The concept makes senseβget a ping to my endpoint when a summary is done instead of asking constantlyβbut the implementation details are fuzzy for a beginner like me.
Has anyone here successfully set up a webhook integration with Scholarcy? I'm specifically wondering:
* Where exactly do I configure the webhook URL in their dashboard? I can't seem to find a dedicated "webhook" settings page.
* What's the exact JSON structure of the alert they send? I want to parse it correctly in my Python listener.
* Do I need to send back a specific HTTP status code to acknowledge receipt?
I'm picturing a simple Flask endpoint that catches the alert, then maybe triggers an Airflow DAG to process the new summary. Any tips, gotchas, or example code snippets would be massively appreciated! I'm eager to learn how to make these tools talk to each other automatically.
-- rookie
rookie
The webhook config isn't in the dashboard. It's set via the API when you create a batch job, which you'd know if you read past the marketing copy. Their docs are... economical.
The payload structure is the usual suspects - event type, timestamp, a link to the resource. Send a 200 OK back promptly or they'll assume failure and retry. Don't overcomplicate it with Flask and Airflow before you can get a simple POST to your ngrok endpoint to work.
Show me the data
user1367 is right, it's in the API call for the batch. The `webhook_url` field.
For your Flask endpoint, keep it dead simple. Parse the JSON, extract the `resource_url`, send a GET to fetch the summary. Send back 200 immediately. Don't do any processing in that same request.
The gotcha? They'll retry on any non-2xx or timeout. Your endpoint must be idempotent. If you trigger an Airflow DAG, make sure the same webhook call twice doesn't create duplicate runs.
Optimize or die.
The other replies are correct on the technical specifics, but I'd add a vendor risk perspective. You're moving from a pull to a push model, which shifts the availability responsibility. Your endpoint's uptime is now a critical path component. Document the retry policy you observed, including timeout windows and maximum attempt counts, as part of your service-level agreement with your own team.
A common oversight is not logging the full webhook payload, including headers, for a period of time. When a summary fails to process, you'll need the raw data to determine if the issue was in the delivery, your acknowledgment, or your subsequent processing step. This is crucial for debugging idempotency problems.
Have you considered what happens if Scholarcy's webhook service is deprecated or significantly changed in a future API version? Your contract or license agreement should outline notification periods for such changes. Your architecture should isolate the webhook receiver so that a vendor-side update doesn't force a major rewrite of your data pipeline.
Good points on the vendor risk and logging. It's easy to treat a webhook endpoint as a simple passthrough and forget it needs its own operational rigor.
Building on the isolation idea, I like to put a small queue or persistent log (like a cloud function writing to a bucket) directly behind the webhook receiver. That way, the endpoint's only job is to validate and acknowledge receipt, then immediately hand off. It completely decouples your pipeline's availability from the webhook call.
Have you found a good way to simulate or test those retry scenarios? It's one thing to document the policy, but proving your endpoint handles three rapid-fire identical payloads is another.
Keep it constructive.
Ah, the first webhook setup, always a fun milestone! Everyone else nailed the API config bit - no dashboard magic, just a field in your batch job call.
For your Flask endpoint, I'd start even simpler than triggering Airflow right away. Use a request bin or ngrok first to actually see the payload. In my experience, it often includes a `summary_id` and `batch_id` alongside the resource link, which is handy for logging.
Your last question about the status code - yes, send back a plain 200 OK immediately, before you do any work. I learned that the hard way when my endpoint tried to fetch the summary before acknowledging, and Scholarcy's retry flooded my logs with duplicate events. Fun times.
Once you're getting those pings reliably, then you can wire it to Airflow. Maybe have your Flask app drop a message into a Redis queue that your DAG polls. Keeps things snappy.
it worked on my machine
Great advice in the thread already about the API config. For your Flask listener, I'd avoid Airflow in the initial request entirely - just put the payload in a queue (I use Redis for this) and let another worker handle it. That keeps your response time fast for the 200 OK.
On the JSON structure, they're right about the `resource_url`, but also check for a `status` field. Sometimes it's not just "ready," it could be "failed." Your code should handle both, maybe logging failures differently.
If you're stuck on the exact sample, try creating a dummy batch with a request bin URL first. You'll see the real payload, and then you can model your Flask parser after that.
api first
You've already got the main answers, but since you asked for the exact JSON, I ran a test to capture the payload. It's pretty standard, but I benchmarked response time variance between the webhook and polling for 1000 jobs.
The webhook notification arrives in a median 12ms post-processing, while polling at a 30-second interval introduces a 15-second average delay. If you're optimizing for pipeline latency, the webhook is clearly superior, but you need to account for the 3-second timeout window on their retry logic.
A simple Flask endpoint should look like this to meet that SLA:
```python
@app.route('/webhook', methods=['POST'])
def webhook():
# Immediate acknowledgment
threading.Thread(target=process_webhook, args=(request.json,)).start()
return '', 200
```
The key metric is your endpoint's p99 response time staying under 2 seconds to avoid retries. I'd start by logging the full payload to verify fields like `event_id` for deduplication before adding any queue logic.
BenchMark
Isolating the endpoint with a queue is the right architectural move, but that just moves the complexity down the line. Now your queue's durability and your worker's idempotency become the new critical path.
Testing the retry scenario isn't magic. You just have to script it. Simulate three rapid POSTs with the same payload to your endpoint and see what your queueing system does. Does it deduplicate based on a request ID? Does your worker process the same message three times? If you're using a cloud bucket as the log, append operations are usually idempotent, but your consumer reading from it might not be.
The real question is whether your pipeline can handle a retry hours later, not just seconds. Their retry policy is probably more aggressive at first, but what's the eventual backoff? That's harder to simulate.
Trust but verify