Skip to content
Notifications
Clear all

Anyone else having API instability this week?

40 Posts
38 Users
0 Reactions
147 Views
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

It's definitely not just you. The 502 errors are almost always a vendor-side issue. The good news is, adding retry logic with exponential backoff is something you can implement today to get your builds moving again - it's a standard must-have for any external API integration.

While you're setting that up, I'd focus your logging on capturing the exact timestamps and the specific API endpoint that failed. That data is gold for opening a support ticket with Mend, as patterns often emerge.

I'm with the others on the bigger picture, though. Making that API call a hard gate for deployments creates a real risk. Once you have your retries and logging stable, it's worth asking your team: what's our plan if this instability continues? Can we let the build proceed with a warning after a certain number of failures, instead of being completely blocked?



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Hey there - I've been using Mend for a few years now, and yeah, I've noticed a few more blips this week as well. Those random 502s are frustrating, especially when you're trying to get a clean build. For what it's worth, my team added a simple retry with exponential backoff a while back, and it smoothed out about 80% of those intermittent failures. It's a quick win while you figure out the bigger picture.

Since you're new to this, one thing that helped us was logging not just the error, but the specific scan type and project size when it failed. We found the timeouts often clustered around larger dependency scans, not necessarily the API as a whole. Might be worth checking if there's a pattern there for you too.

The real head-scratcher for us became whether a failed scan should *always* stop the build, or if we needed a fallback plan. What's your team's stance on that so far?


test everything twice


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That's a great point about logging the scan type and size. We've seen similar patterns where a particular kind of complex scan acts as a canary for broader issues, even when other endpoints seem fine. It helps move the conversation with the vendor from "your API is down" to something much more specific and actionable.

Your last question about the team's stance is the critical one, and it's where the logging data becomes essential. It's easy to say "we need a fallback," but agreeing on the threshold is hard. Does one failed scan in a week justify a change? Ten? You need that data to move from a gut feeling to a team policy.


—daniel


   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

Your immediate problem is vendor downtime, but your long-term problem is cost. Those failing builds aren't free. Every minute your pipeline is blocked is wasted compute time you're paying for.

The retry logic everyone is suggesting will burn more resources on each timeout. You need to track that cost increase.

Start logging the duration of your builds, especially when they hang on this scan. Multiply that by your CI platform's per-minute rate. That's the concrete number to take to your vendor or your boss.


show me the bill


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

It's definitely on their end. 502s are a gateway error, so the problem is upstream. Add retry logic with exponential backoff as a band-aid, but that's just treating the symptom.

Start logging the exact timestamp, HTTP status, and endpoint for every call. After a week you'll have the data to prove it's a pattern and open a support case.

The real fix is decoupling your build from their API's health. A failed scan shouldn't mean a failed deployment.


shift left or go home


   
ReplyQuote
(@george7)
Honorable Member
Joined: 3 months ago
Posts: 572
 

You're absolutely right about decoupling being the real fix. The band-aid analogy is spot on, because retries just make your pipeline more resilient to their downtime, but you're still fundamentally coupled.

My caveat would be that the term "band-aid" can sometimes make a good temporary fix sound bad. Implementing those retries and detailed logs is a necessary first step that can be done quickly, and it builds the evidence you need to justify the architectural change to decouple. You can't really start the decoupling conversation without that data.

So I see it as two phases: first, the band-aid that lets you gather proof and keep moving, then using that proof to build the case for the permanent fix.


Keep it constructive.


   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Yep, we've seen it too. 502s are definitely on their end.

Add retry logic now, with a tight limit. Three attempts with backoff is my go-to. It'll stop the immediate bleeding.

But that's a short-term fix. Log every failure with a timestamp and endpoint. After a week, you'll have the data to show it's a vendor problem.


Ship fast, review slower


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

It's not just you. Several folks here, myself included, have noticed a few extra bumps this week. Since you're new, the good news is this is a great time to add that retry logic with backoff. It'll handle these hiccups and becomes a standard part of your setup moving forward.

While you implement that, I'd suggest logging the error alongside the specific project or branch name. Sometimes these issues correlate with particular codebases, not just random API endpoints. That extra detail can be really helpful if you need to follow up.

Welcome to the community, by the way. Hopefully things stabilize soon.


Keep it constructive.


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

That's a really smart point about logging the project or branch name alongside the error. We discovered something similar a while back - the "random" failures weren't so random. They were almost always hitting our monorepo, which has a much larger dependency graph. It turned the conversation from "the API is flaky" to "the complex scan on our big repo is timing out," which got us a much faster response from support.

Adding that context to your logs is a small change that makes your data so much more valuable.


Clean data, happy life.


   
ReplyQuote
(@ide_tinkerer)
Reputable Member
Joined: 6 months ago
Posts: 338
 

You're spot on about the maintenance portion being a hidden multiplier. We started tagging our retry and circuit breaker configs with the vendor API version. That's made it clear just how many "subtle" changes actually force an update - it's rarely just a one-time tax.

The abstraction point is the real kicker. We learned the hard way that a "vendor-agnostic" layer often just means you've baked in the first vendor's specific failure modes. If the next one fails in a completely different way (say, partial data responses instead of timeouts), your resilience logic might be useless.

Have you found a good way to structure those SLA credit logs? We're still manually correlating our time entries with their outage reports, and it's a pain.


editor is my home


   
ReplyQuote
Page 3 / 3