You've nailed the core issue that makes SLAs so tricky. That distinction between new and altered resources is exactly where we got burned.
We had a DevOps team spin up a dev Cosmos DB instance with the public endpoint enabled. Because it was a net-new resource during a Friday afternoon deployment spree, it didn't pop in Orca until Monday morning. The incremental scans humming along for our existing databases were completely blind to it.
Your request for a CloudTrail screenshot is fair, but the real proof is in the TimeLag metric between the CloudTrail `CreateTable` event and the scanner's first successful `DescribeTable` call. We built a dashboard tracking that delta, and the spikes during our deployment windows are sobering.
Implementation is 80% process, 20% tool.
The "weird architecture choice" is the real cost. You pay for the scanner, then you pay again in engineering hours to work around its API limits or build a faster event bridge.
Increasing quotas just means hitting the next bottleneck, probably a new line item on your cloud bill for API Gateway or whatever. Building the bridge means you're now maintaining a custom security scanner. Both turn a fixed SaaS cost into a variable resource drain.
always ask for a multi-year discount
Building that event bridge isn't just parallel work; it fundamentally changes your security model. You shift from periodic, comprehensive scanning to reactive, event-driven checks, which misses dormant misconfigurations that aren't triggered by a CloudTrail event.
The correct, albeit more complex, path is a hybrid approach. Use the event bridge for immediate blocking of critical, known-bad patterns on new resources (like a publicly exposed database), while relying on the scanner for completeness and drift detection. This way, you're not building a full scanner, just a fast, narrow gatekeeper for your highest-risk actions. The scanner's latency on new resources becomes less critical because you've already caught the worst offenses in real time.
This hybrid model makes sense, but doesn't it assume you already know your most critical patterns? What happens when a new, critical misconfig type is identified? Your event bridge is blind until you update its rules, so you're back to relying on the scanner's latency for that new threat.
Your 45 minute to 3 hour AWS baseline matches what I've seen. The big catch is new resources, especially during busy deployment periods - that's where it can blow out to hours, or even a full day if quotas get hit.
We use a similar hybrid approach to user777, with a CloudTrail-based Lambda for immediate blocks on high-risk actions like public S3 buckets. It cuts the scanner's latency out for our biggest fears. But it's true, it doesn't help with a newly discovered vulnerability type. For that, you're still waiting on the scan cycle.
That's the trade-off. You're either okay with the scan delay for new threats, or you're building more custom event rules.
dk
You're focusing on the right metric, but the observed latency is almost entirely a function of API quotas, not the tool's inherent scan cycle. The vendor's stated intervals are a best-case scenario in a vacuum.
Your 45 minute to 3 hour AWS range is plausible for incremental changes to inventoried resources during off-peak hours. However, that range expands drastically for net-new resources created during a busy deployment window, as others have noted. We've documented cases where a new, misconfigured Azure Storage Account took over 8 hours to appear because the scanner's service principal was throttled on `StorageAccounts/ListByResourceGroup` calls while other services weren't.
This pushes you towards the hybrid model discussed, but with a specific operational cost: you must now map and monitor the scanner's distinct API quota buckets for each provider to even predict your baseline latency. It becomes a capacity planning exercise.
Support is a product, not a department.
Your 45 minute to 3 hour baseline for AWS is consistent with what I've observed for incremental changes on inventoried resources. However, you're right to focus on the provider variable - the latency profile is different in Azure due to the structure of its resource management APIs.
The biggest gap in your testing is likely storage account encryption. In our environment, toggling `requireInfrastructureEncryption` on an existing Azure storage account often took 5+ hours to alert, far longer than a comparable S3 bucket change. We traced it to the scanner hitting throttling on the `StorageAccounts/GetProperties` call, which is on a stricter default quota than many AWS equivalents.
For real-time threat response, that window is indeed a problem. The hybrid model others mentioned becomes a necessity, not an optimization.
sub-100ms or bust
Exactly, that provider difference is huge. The Azure API quota structure makes latency way less predictable than in AWS. We saw the same thing with `GetProperties` calls.
That 5+ hour window on a storage account toggle is painful. Makes me wonder if we should be pushing vendors for more transparent "effective scan windows" that factor in these API constraints, not just their ideal cycle times. A hybrid model is essential, but you're right, it's really the only way to get close to real-time on those high-risk actions.
cost first, then scale
Your 45 minute to 3 hour baseline for AWS is consistent with what I've seen for incremental changes on inventoried resources. However, you're right to focus on the provider variable, the latency profile is different in Azure due to the structure of its resource management APIs.
The biggest gap in your testing is likely storage account encryption. In our environment, toggling `requireInfrastructureEncryption` on an existing Azure storage account often took 5+ hours to alert, far longer than a comparable S3 bucket change. We traced it to the scanner hitting throttling on the StorageAccounts GetProperties call, which is on a stricter default quota than many AWS equivalents.
For real-time threat response, that window is indeed a problem. The hybrid model others mentioned becomes a necessity.
Integrate or die