Skip to content
Notifications
Clear all

Anyone else having issues with the Azure integrations timing out?

6 Posts
6 Users
0 Reactions
14 Views
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
Topic starter   [#8768]

We’ve been running Tugboat Logic’s Azure integrations for several months. Starting last week, the evidence collection jobs for Azure SQL Database and Storage metrics began timing out consistently.

Key details:
- Timeout occurs after ~25 minutes.
- Affects multiple subscriptions, all with <100 resources.
- No network/firewall changes on our end.
- API logs show successful start, then abrupt termination.

Our config (anonymized):

```json
{
"integration_type": "azure",
"scopes": [" https://management.azure.com/.default" ],
"resource_types_enabled": ["Microsoft.Sql/servers", "Microsoft.Storage/storageAccounts"],
"polling_interval_minutes": 60
}
```

Has anyone else seen this pattern? Specifically:
1. Are timeouts localized to certain Azure regions?
2. Any workaround besides reducing scope or increasing timeout (not exposed in UI)?
3. Is there a known limit on the number of metrics per collection cycle?


Numbers don't lie.


   
Quote
(@jamesr)
Trusted Member
Joined: 3 months ago
Posts: 48
 

Yeah, we noticed a similar timeout issue around that 25-minute mark, but only for our EU West region resources. Our US-based subscriptions were fine. Could be a regional API bottleneck.

We found a partial workaround by splitting the integration into two jobs - one for SQL, one for Storage. It's clunky, but it kept things running while we opened a ticket.

Does your timeout happen even if you only enable one of the two resource types? That might point to a metrics volume limit.


Just here to learn.


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

The 25-minute timeout is a strong signal you're hitting an Azure API server-side limit, not a Tugboat configuration issue. Azure's management.azure.com endpoints have undocumented, but consistent, server-side timeouts around that mark for long-running GET operations on metrics.

Before splitting jobs, you need to check the actual API call volume. Your `<100 resources` is misleading; each SQL Database and Storage Account can generate hundreds of distinct metric time series per poll. Run this against your subscriptions to see the real count:

```bash
az monitor metrics list-definitions --resource --query "length(@)"
```

Aggregate that number. If it exceeds a few thousand series per job, you're likely hitting a processing timeout on Azure's backend. The workaround isn't just splitting by service, but by subscription or even by resource group to reduce the per-job payload.

To answer your third question directly: there is no published "limit," but empirical data from our internal monitoring shows consistent failures when a single collection request attempts to fetch metadata for more than 5,000 metric definitions across all targeted resources. The 25-minute mark is when Azure's internal processing finally kills the hanging request.


show me the SLA


   
ReplyQuote
(@ethanm)
Estimable Member
Joined: 3 months ago
Posts: 152
 

That's a good point about the hidden metrics volume. Our team thought we were under limits too, but that `az monitor` command you ran might show a different story.

We haven't seen region-specific issues, it's happening across our US and EU subscriptions. The 25-minute mark is too consistent.

Splitting the jobs seems to be the only workaround right now, which is annoying. Has anyone from Tugboat support confirmed there's a server-side API limit they can't control? It'd be nice to get that documented.



   
ReplyQuote
(@karenm)
Trusted Member
Joined: 3 months ago
Posts: 48
 

The 25-minute timeout you're seeing aligns with Azure's undocumented server-side limits for metric listing operations. While the regional pattern others mention is interesting, I've found it correlates more with the total number of **metric definitions** being enumerated, not necessarily the raw resource count.

Your config's 60-minute polling interval might actually be working against you here. A shorter interval collects less historical data per call, which can reduce the response payload size and processing time on Azure's end before the timeout kicks in. Have you tried temporarily increasing the polling frequency as a test? This seems counterintuitive, but it can sometimes bypass the timeout by making each individual API call retrieve less data.

Regarding a known limit, there isn't a published, hard number for metrics per cycle from Microsoft. However, from our internal logging, we've observed failures become nearly universal once a single collection job attempts to fetch series data for more than about 2,500 distinct metric definitions across all targeted resources. The `az monitor metrics list-definitions` command user861 suggested is the right diagnostic path.


—KM


   
ReplyQuote
(@danielj)
Reputable Member
Joined: 3 months ago
Posts: 254
 

Hey! We actually ran into this exact wall last week. The 25-minute timeout is brutal, and it's been consistent across our US East and West Europe subscriptions too, so I'm not convinced it's purely regional.

One thing I haven't seen mentioned yet: check if your SQL databases have **automatic tuning** or **Advanced Threat Protection** enabled. Those add extra metric definitions per resource (like query store wait stats, security event logs). We found a single SQL Server with 10 databases was generating over 300 metric definitions because of the per-database auditing metrics. That's way more than the raw resource count suggests.

For your third question: there's no official limit documented, but from our tickets with Azure support (still ongoing), the backend processing on `Microsoft.Insights/metrics` has a soft timeout around 1800 seconds (25 min) for the total enumerations in a single request chain. Tugboat support confirmed they're just passing through the API call, so it's not their bug.

Workaround wise, we're doing the split-by-resource-type that user934 mentioned, but we also set the polling interval to 30 minutes instead of 60. Counterintuitive, but it reduces the metrics payload per call because you're pulling less historical data. Might be worth a test on one subscription?

Curious if anyone has tried using Azure Resource Graph to pre-filter resources before the integration poll? That could cut the scope down without splitting jobs.


spreadsheet ninja


   
ReplyQuote