Skip to content
Notifications
Clear all

Just built a simple webhook listener to track data sync errors in real time during cutover.

4 Posts
4 Users
0 Reactions
1 Views
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
Topic starter   [#29499]

Just finished a three-month CRM migration from a legacy on-prem monster to a SaaS platform, and the single biggest tactical win wasn't the Terraform modules or the fancy K8s ingress setup—it was a dead-simple, home-rolled webhook listener we stood up two weeks before cutover. Everyone plans for the big data migration, but the real pain lives in the silent failures during the cutover window: records that look fine in the pre-validation but fail on the final sync because of some last-minute data entry, or API rate limits you didn't hit during testing because you weren't running at full production volume.

We configured the new CRM's integration layer to send all sync events—successes and failures—to a webhook endpoint. Instead of relying on batch logs or hoping the integration platform's UI would alert us, we had real-time visibility. The listener was just a small Flask app in a Kubernetes pod, backed by a Redis stream for durability. Its job was to parse the payload, filter for anything with an error status, and immediately push a formatted alert into our observability stack (Grafana Loki for logs, Prometheus for error counts, and PagerDuty for critical ones). The crucial part was enriching the alert with the source record ID and the exact failure reason from the provider's API response.

Here's the core of the listener logic (sanitized). The simplicity is the point.

```python
from flask import Flask, request
import redis
import json
import logging

app = Flask(__name__)
redis_client = redis.Redis(host='redis-stream', port=6379, decode_responses=True)
logger = logging.getLogger(__name__)

@app.route('/webhook/crm-sync', methods=['POST'])
def handle_webhook():
event = request.json
# The provider's payload included 'status', 'recordId', 'message', 'entityType'
if event.get('status') not in ['SUCCESS', 'SYNCED']:
error_payload = {
'record_id': event.get('recordId'),
'error': event.get('message'),
'entity': event.get('entityType'),
'timestamp': event.get('timestamp')
}
# Push to Redis stream for async processing/durability
redis_client.xadd('crm_sync_errors', error_payload)
# Also log immediately for Loki collection
logger.error(f"CRM Sync Failure: {error_payload}")
return '', 202
```

A downstream processor consumed the Redis stream, aggregated errors by type, and bumped Prometheus counters. We had a dashboard up that showed, in real time, error rates by entity (Contacts, Opportunities, etc.). This let us identify two critical issues we would have missed for hours:
1. A specific custom field mapping for a subset of legacy accounts was failing because of an unexpected character set. It was only 0.1% of records, but they were high-value.
2. Our rate limit buffer was too thin; we saw the error spike instantly and could trigger a manual back-off before the provider blacklisted our IP.

What I wish I had known before other migrations: Your integration platform's "successful sync" metric is often a lie. It might only mean the message was queued, not that the record was written. You need the destination system to tell you it's happy. Building this took half a day and saved us at least 40 hours of post-cutover forensic debugging and data repair. The hard truth is if you're not listening to the destination's real-time event feed during cutover, you're flying blind through the most critical phase.

---


Been there, migrated that


   
Quote
(@ava23)
Honorable Member
Joined: 2 months ago
Posts: 435
 

Yeah, the tactical win always comes from the dumb little thing you built yourself, doesn't it? Everyone's so busy buying the shiny integration platform with the "real-time monitoring dashboard" that they forget you need something that just *screams* when it's broken.

That said, I hope you're ready to sunset that Flask app soon. My experience is these brilliant, home-rolled monitors have a habit of becoming permanent, undocumented production infrastructure. You'll be debugging a Redis connection pool issue two years from now because "it just works" and nobody dares touch it.

Still, gotta love beating the vendor at their own game. Their UI would've shown a tiny, spinning loading icon while your entire pipeline melted.


Trust but verify.


   
ReplyQuote
(@clarak2)
Estimable Member
Joined: 2 months ago
Posts: 143
 

Totally agree about the tactical power of a simple, direct feedback loop. It's so easy to get lost in the "observability stack" and forget you just need a fast, loud signal.

What I liked about your setup was filtering for error status and pushing to Prometheus/Loki *first*. That's the key - it becomes a real metric you can alert on, not just another log to grep. I've seen teams dump the raw webhook to a log file and then wonder why they missed the incident. Having that filter at the ingress point is genius.

And you're right, the rate limit surprises are the worst. Testing never quite matches the chaos of the real cutover. Having that live view must have been a lifesaver


Docs save time


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You're absolutely right about the "temporary" tool becoming permanent. The operational cost of that undocumented Flask app over two years will almost certainly eclipse the one-time license fee for the vendor's proper monitoring dashboard.

The trick is to make sunsetting part of the design. On the last day of cutover, we scheduled a lambda function to send a kill switch SQS message to the listener's ASG, scaling it to zero. The CloudFormation stack was tagged with a hard deletion date. It still took a manual push to delete it, but the friction was lowered.

Without those baked-in constraints, it would still be running today.


Less spend, more headroom.


   
ReplyQuote