Just spent three days chasing down a notification delay that made our alerting system about as useful as a chocolate teapot. Turns out Grok’s webhook deliveries were batching up and arriving 2-7 hours late, completely at random.
Our setup is standard: event triggers a Grok workflow, which fires a webhook to our internal dashboard. No errors in Grok's logs, all marked "delivered" instantly. The receiving end? Nothing. Then a sudden flood of notifications hours later, usually at 3 AM. Fantastic for the blood pressure.
* Verified it’s not our endpoint: same infrastructure handles other webhooks with sub-second latency.
* No rate limiting warnings or quota issues visible in the console.
* Pattern seems completely arbitrary – affects maybe 30% of our production events.
Is this just us, or has anyone else seen Grok hold onto webhooks like they’re precious secrets? More importantly, has anyone found a workaround besides building a redundant polling layer (which defeats the entire point)?
The lack of a visible queue or delay metric in the UI feels like a hidden tax on reliability. We’re now evaluating if we need to budget for a secondary, real-time notification service just to cover for these silent delays, which frankly shouldn’t be necessary.
Cloud costs are not destiny.