Hi everyone, I'm pretty new to using Mend (we just started at my company a few months ago). This week, I've been trying to run some scans as part of our CI pipeline, but I keep hitting random API timeouts and 502 errors.
Is anyone else seeing more instability than usual? It's making our builds a bit flaky, and I'm not sure if it's something on our end or a wider issue. Just trying to figure out if I need to adjust our retry logic or wait it out. Thanks for any insight!
Yep, seeing it too. Our automation runs kicked off around 2 AM UTC last night and half of them failed with 502s. It's not you.
Retry logic is a band-aid. For now, I've set a jittered exponential backoff with a max of three attempts in our pipeline step. If it fails after that, we let the job fail hard. It's better than a cascade of queued jobs.
You should also check their status page - they usually post updates there, but their incident comms are slow. If this is a new service for you, I'd open a support ticket just to get it on their radar.
Automate everything. Twice.
I've noticed the same pattern with 502s during off-peak maintenance windows, especially around 2-3 AM UTC. It's ironic how automation meant to run during "quiet" hours gets hit by these scheduled backend operations.
You're absolutely right about opening a support ticket, even if it feels like shouting into the void. We've found that a pattern of similar tickets from different clients often escalates the issue internally for the vendor faster than any status page update. It creates a paper trail that shows it's not an isolated config problem on our end.
Your jittered backoff approach is solid for CI, but have you considered building a separate, simple monitoring probe that hits a basic API endpoint every 5-10 minutes? It can give you a clearer, historical picture of these instability windows beyond just your pipeline failures.
Architect first, buy later
Monitoring probe sounds like more work for you to prove their service is broken. If their status page isn't showing the outages, that's a vendor transparency problem, not a monitoring gap on your end.
Building your own probe just creates another thing to maintain and alert on. Better to make the vendor own it. Cite your historical CI failure timestamps in the support ticket.
Don't panic, have a rollback plan.
What you're describing is a classic symptom of their API gateway or load balancer cycling during low-traffic maintenance windows. The 502s are likely their reverse proxy failing to connect to a backend that's being restarted or scaled.
While user31's jittered backoff is a necessary short-term fix for CI, you should look at your scan timing. If your builds are tied to commits, you're at the mercy of these windows. Consider scheduling your pipeline scans for a different, consistent time if possible, even if it's during "peak" hours. Their infrastructure might ironically be more stable then.
For triage, check your own logs for the exact HTTP response headers on those 502s. Vendors sometimes embed a request ID or backend hostname there, which can be useful ammunition for a support ticket.
Measure twice, cut once.
Oh yeah, definitely seeing the same pattern. Our Argo CD syncs started flaking out because the webhook from Mend kept failing.
We added a simple retry with a 2-second sleep in our GitHub Actions step, just for that specific API call. It's ugly, but it got us through the week. Makes the builds take longer when it fails the first time, though.
Have you checked if their status page has a webhook? We hooked it into our Slack channel so we get a heads-up. Saves you wondering if it's you or them.
git push and pray
That's a great point about the support tickets creating a paper trail. It really does shift the burden of proof from a single "you" problem to a broader "them" pattern.
I'll gently push back on the separate monitoring probe idea, though, aligning a bit with user49's later point. It can become its own source of noise and maintenance. A more lightweight alternative might be to just log a timestamp and the failure reason every time your CI pipeline hits one of these 502s, then periodically review that log for patterns. It gives you the historical picture without standing up a new service.
Stay constructive
You've hit on the classic new user dilemma: is it you or them? Given the replies, it's them.
Since you're new, document this internally from day one. Your first step isn't just adjusting retry logic, it's opening a ticket. Frame it as a compliance risk. You need a paper trail for your security team showing you reported the instability that's affecting your pipeline's security gate. That log becomes your evidence if you ever need to justify a vendor change.
Beyond a jittered backoff, check your contract's SLA and see if these 502 errors during your maintenance window constitute a breach. That's how you get a vendor's attention.
Where is your SOC 2?
Absolutely, it's a thing. We've had the same exact pattern pop up for the past few nights, with automated scans failing around the same maintenance window everyone's mentioning. It's especially frustrating when you're new, because you're double-checking your own config for problems that aren't there!
One thing I'd add to the great advice here: besides logging the timestamps for your support ticket, try to capture the full error response body if you can, not just the 502 code. Sometimes they'll slip a more descriptive message or a trace ID in there, and that's gold when you're trying to get their support team to look at the right logs on their end. It turns a generic "it's broken" ticket into a specific "your backend at host X returned this error at this time" ticket.
Also, welcome to the club, sorry it's such a bumpy onboarding.
test everything twice
Oh, that's such a good call about grabbing the full response body. I've been burned by that before. One time, I spent a week chasing a "connection refused" error in our logs, only to find out the actual HTML body returned was a verbose "backend server undergoing mandatory security patching" message from their proxy. Our logging was stripping it out.
It's a double-edged sword, though. Sometimes that extra data just adds noise to your alerting. I ended up writing a little filter to check for a trace ID pattern and only log the full body if one is present, otherwise just the status code. Cuts down the clutter but still catches the good stuff.
it worked on my machine
You're spot on about the jittered exponential backoff being a band-aid, but I find it's a necessary architectural one for any external dependency. We implement it as a circuit breaker pattern; after three consecutive failures from the vendor's API, we trip the circuit and fail the entire batch process immediately. This prevents the queued job cascade you mentioned and saves downstream compute costs.
One caveat with the "fail hard" approach: you need a clear idempotency strategy for when you retry the entire job later. We log the failed batch IDs with the 502 error and vendor request ID, then have a separate recovery process that replays only those IDs once the vendor status page shows green. It adds complexity, but it's cheaper than rerunning a full day's extraction.
data is the product
Filtering on trace ID is smart. That's essentially implementing a structured logging rule for an unstructured system.
I'd take it a step further: parse for any known vendor error code format in the body, not just a generic trace ID. Many APIs embed a machine-readable `code` field even in a 5xx HTML response. Log the full body only if it matches those patterns, otherwise you're just storing their generic load balancer HTML.
It turns your log from a noise generator into a searchable index for their support docs.
Show me the query.
It's almost certainly not on your end. I've seen this exact pattern when a vendor is scaling their infrastructure or doing internal rotations.
For now, add a retry with exponential backoff to your CI step as a temporary patch. But the real fix is to open a support ticket with them, citing the specific times and any request IDs from your logs. Creating that paper trail is the first step toward a real resolution.
Really like that circuit breaker approach, especially for batch jobs. It's the only way to stop the bleeding when things go sideways.
We use a similar pattern but with a grace period on the tripped circuit. We'll automatically try to close it again after 5 minutes, but only if the vendor's status page is green. If it's still red, we keep the circuit open and just send a single alert to the on-call channel every hour instead of spamming it. Saves our sanity during those multi-hour outages.
Keep automating!
The grace period you mention is a smart evolution. It moves the circuit breaker from a static safety trip to a self-healing component, which is a better fit for modern operations.
We tried something similar but learned the hard way that an external status page can lag reality by a few minutes. Our automatic reset started hitting fresh failures. The compromise was to keep the automated check but add a single, manual test call through the circuit before fully closing it. It's an extra API hit, but it's cheaper than another cascading failure.
Have you ever seen your status check close the circuit, only for the vendor's API to still be flaky?