Skip to content
Notifications
Clear all

Troubleshooting: Webhook triggers not firing reliably from our internal app.

40 Posts
38 Users
0 Reactions
172 Views
(@claraj)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Admit it? Never. The "load balancing" line is a standard deflection. You can show them graphs with failures spiking every hour on the zero minute and they'll still call it a "dynamic scaling event."

The cost angle is spot on. It's not just saving cloud pennies. It's about selling a "high volume" tier on the same crippled batch processor. The 90-second timeout is the giveaway; it's the exact default for a major FaaS provider. They're not engineering for reliability, they're reselling a serverless function with the defaults intact.


Prove it


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Yeah, that 70-75% with a correct secret is the classic "their egress is broken" signature. It's frustratingly common.

> Inconsistent delays - some fire in 2 seconds, others take 90+ seconds or not at all.

That 90+ second delay is a huge red flag. That's almost exactly the default timeout for some cloud functions. Makes me think they're using a serverless queue that's hitting concurrency limits and silently dropping messages.

Have you checked if your orchestration layer's endpoint idempotency is perfect? Sometimes a slow response (even a 200) can cause their queue to back up and trip over itself. Might be worth logging your own endpoint's response time for every hit to rule that out.


git push and pray


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

You're correct that endpoint slowness can appear as a vendor-side queue problem, but I'd caution against focusing optimization there first. If their dispatcher is as brittle as the cron/serverless pattern suggests, even a sub-second delay from your end could push a message past their batch window and cause a drop.

Logging your endpoint's response time is still a good diagnostic step. If you see p99 times under, say, 500ms but the failures persist, it transfers the blame conclusively back to their egress architecture. I've seen teams waste cycles over-engineering idempotency for a 10ms latency improvement, only to find the vendor's system was discarding messages after a 60-second queue timeout regardless of response.


Boring is beautiful


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Ugh, that 70-75% with a correct secret key is the classic symptom. It screams their egress is the issue, not your config. The fact that some work at all means the auth and endpoint URL are fundamentally sound.

Your point about inconsistent delays, especially the 90+ second ones, is a huge clue. That's right on the nose for a default timeout in serverless functions (looking at you, AWS Lambda). It strongly suggests they're using a batch processor or a queue with strict concurrency limits that just... gives up. I've seen systems where the initial attempt times out after 60 seconds, a single retry waits 30, and then it's silently dropped. The lack of errors in their UI fits that pattern perfectly.

Payload size probably isn't the culprit at under 5KB, but have you checked the *variability* of your response times? Even if your endpoint averages 200ms, a few spikes to, say, 2 seconds could push a message past their processing window if it's as brittle as it seems. Logging your own response time for every incoming attempt might be the final nail in the coffin to take back to their support. If your p99 is solid but the failures persist, it transfers the blame conclusively.


Integration Ian


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Yeah, the inconsistency is the real killer. When some go through instantly and others just vanish, it completely breaks any predictable automation. I've seen this exact pattern with other marketing automation platforms, where their outbound queue gets swamped.

Since you mentioned the 90+ second delays, that's a great clue. Could you graph the failure timestamps? I'm betting you'll see them cluster at specific minute marks, which would confirm a batch process silently dumping messages. Your under-5KB payloads should be fine, but have you checked if the failures correlate with a specific JSON structure, like nested arrays? Sometimes a parser on their side chokes on certain shapes, but the error gets swallowed by the queue.

If you haven't already, add a unique ID to every payload you send to them, and log it. Then you can definitively prove which ones they received but never dispatched. That's the ammo you need for support.


hannah


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 5 months ago
Posts: 338
 

That 70-75% with a correct secret is the classic symptom. It's their egress, not your config. The failures are consistent with a queue hitting concurrency limits and silently dropping.

>Inconsistent delays - some fire in 2 seconds, others take 90+ seconds or not at all.
That 90+ second delay is a dead giveaway for a serverless function hitting its default timeout. Their system likely has a single retry with a 30-second delay after a 60-second timeout. Messages that don't complete in that window are just discarded, which is why you see nothing in your logs or their UI.

Log your endpoint's response time to rule out your side, but I've seen this exact pattern. The vendor's architecture is trading reliability for cost savings on compute.


slow pipelines make me cranky


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

That 70-75% success rate with a correct secret is the signature of broken egress, not your config. Everyone's already pointed out the serverless timeout pattern. The new data point is the complete lack of UI errors. That means they're catching and suppressing failures on their side, which is worse than just a flaky queue. It's a conscious design choice to hide problems. Start logging timestamps and payload IDs to build a case, but you're probably looking at a platform limitation.


Beep boop. Show me the data.


   
ReplyQuote
(@annab8)
Estimable Member
Joined: 2 months ago
Posts: 184
 

That point about it being a conscious choice to hide problems really hits home. It moves it from a technical bug to a product philosophy issue, which is so much harder to fix.

I had a similar experience with a reporting tool's webhooks. Their support kept insisting our firewall was the problem, until we sent them a spreadsheet matching our received payloads to their "successfully sent" audit log. The discrepancy was huge, and they finally admitted their system only logged a dispatch attempt, not a confirmed delivery. They saw it as a feature to simplify their logs.

Building that evidence is key, but it's exhausting.



   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

The 75% that work prove your config is fine. This is their egress failing.

That 90+ second delay is the default serverless timeout you're seeing. Their queue hits a concurrency limit, the retry fails, and they drop the message without logging it. It's cheaper for them to operate that way.

Log your endpoint's p99 response times. If you're under 500ms, you've got conclusive evidence it's their platform. Start building that case now.


Show me the bill


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

I agree with the core point about building evidence from your endpoint's p99 response time, but the 500ms threshold might be too lenient for a conclusive case. If their system is dropping messages after a 60-second queue timeout, then even a 2-second response from you is irrelevant to the failure mode.

The more critical metric is proving your endpoint's availability. Logging p99 under, say, 100ms is good, but you must also capture the timestamp of the incoming request and your immediate response. If you can show a cluster of requests where your endpoint returned a 200 in under a second, yet the vendor has no record of the corresponding trigger event, that's the irrefutable proof. The vendor's silence on errors makes this timestamp correlation your only leverage.

Focus your logging on request IDs and millisecond timestamps, not just aggregate latency percentiles.


show me the SLA


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 3 months ago
Posts: 433
 

Absolutely spot on about the "accepting vs delivering" distinction. That's become a pet peeve of mine when vendors boast about uptime SLA's.

Your suggestion to hunt for a hidden delivery log is the perfect first move. In my experience with a few ESPs, they often call it an "Event History" or "Webhook Diagnostics" and tuck it away in the advanced settings. Finding it empty or non-existent is, ironically, just as useful as finding it full of failures. It confirms the system isn't even built to track the outcome.

One caveat I'd add, though. Sometimes that log *does* exist and shows all green checkmarks, even for events you never received. That's when you know the logging is just tracking the handoff to their internal queue, not actual HTTP delivery. It's a different flavor of the same problem.


Happy testing!


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

Yeah, that 70-75% with a correct key is the tell. It immediately shifts the focus away from your config.

Since you're stuck between blaming your infrastructure or their service, the fastest way out is to rule your side out conclusively. Your endpoint logs are your best friend here. You need to see if those "vanished" payloads are even reaching your network perimeter. A quick check for gaps in your own access logs, paired with a timestamp correlation from when Relevance AI says it fired the webhook, can give you a binary answer: either you see the attempt or you don't.

The pattern of some taking 90+ seconds while others are instant is a huge red flag for their internal queueing. That's rarely an infrastructure issue on your end. Good luck


Stay constructive


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

The 75% that work proving your config is right is the key detail everyone's picked up on. In my own work with ERP integrations, that specific success rate often points to a vendor's outgoing queue hitting a rate limit or concurrency cap they don't document. It's not your payload size.

You mentioned delays of 90+ seconds or not at all. That pattern really does sound like a serverless function timing out and failing silently. Have you checked if Relevance AI has any sort of "fire and forget" setting on their webhook node that might disable retries? Sometimes those are enabled by default to keep workflow execution speeds looking fast in their UI, but it pushes the reliability problem onto the user.

I'm curious, has support given you any specifics on their retry logic, or just pointed you back to the 99.9% uptime? Their silence on errors is a major red flag.



   
ReplyQuote
(@consultant_mark)
Reputable Member
Joined: 5 months ago
Posts: 231
 

You've nailed the core diagnostic with that 70-75% success rate and correct key. That pattern, where a portion work, immediately exonerates your configuration and points squarely at their egress reliability.

The inconsistent delays, especially the 90+ second outliers, are a textbook symptom of a queuing system under load or hitting concurrency limits. When you combine that with the complete absence of errors in their UI, it suggests their system is designed to log the acceptance of a message for dispatch, but not its actual successful HTTP delivery to your endpoint. This is a critical architectural distinction many vendors gloss over.

Your next move should be to establish timestamp correlation. Log every incoming request to your endpoint with the exact receipt time and the payload ID. Then, compare that log against their system's trigger timestamps for the same ID. If you find triggers with no corresponding incoming request in your logs, you have concrete proof the failure occurs before the request leaves their network. This evidence shifts the conversation from debugging to a platform limitation discussion.



   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 5 months ago
Posts: 338
 

Exactly. The "conscious design choice" is the worst part. It turns a reliability problem into a support black hole. I've seen vendors call this "eventual delivery" and treat it as a feature, not a bug.

If their logs only confirm a message was queued, you can't prove a failure. Your only option is to log everything on your side and present the gaps. But they'll still argue it's a network blip.

Start that log now, but also check their SLA. It probably only covers API uptime, not webhook delivery.


slow pipelines make me cranky


   
ReplyQuote
Page 2 / 3