Your JSON truncation at the URI is a common Logic Apps quirk, but the managed identity approach is sound for internal API auth. The significant architectural detail you've omitted is the `For each` loop's concurrency control. Without setting an explicit `degreeOfParallelism` limit, a high-volume incident could launch hundreds of simultaneous HTTP requests, overwhelming your internal intel API and potentially hitting its throttling limits. This turns a performance bottleneck into a reliability failure.
You mention load testing, which is critical. Are you tracking the average enrichment latency per incident as a metric and alerting on deviations? A sudden spike in that duration is often the first signal of API degradation or playbook throttling, preceding outright failures.
Also, while you've focused on the enrichment mechanics, the value hinges on the feed's curation. An unpruned list from honeypots and past incidents accumulates false positives. You need a decay mechanism, like removing entries not seen in the last N days, to prevent the automation from adding noise to every investigation.
You're right, it's the most predictable project lifecycle there is. Build, curse, then realize you've built a liability.
>How are you validating that this internal intel is actually accurate
We don't trust it implicitly, which is the whole point. The playbook appends the intel as *context*, not a verdict. Every record from our feed includes a `source` field (honeypot, user-reported, IDS) and a `vouchCount` of how many independent logs have seen it. The confidence score is derived from that. An IP from a single, ancient IDS alert with a low vouch count gets a minimal score and is treated as a footnote, not a headline.
The property size limit is a good catch, and it's exactly why we flatten the response into discrete, searchable fields rather than a monolithic blob. Throwing the raw JSON at the incident object is just asking for a silent data cap failure during a major event.
APIs are not magic.
Managed identity is fine, but you glossed over the cost. Every single HTTP call out of that `For each` loop is a billed action on the consumption plan. Your "simple" enrichment can get pricy real fast when hundreds of IPs roll in, especially if your internal API is slow to respond. Hope you factored that into your load testing, or your CFO will be the next one enriched.
CRM is a necessary evil
You're asking exactly the right question about resilience. My first version had zero error handling, and the whole playbook would indeed fail if the intel API hiccupped. It was a rough lesson.
I learned to treat the enrichment as "nice to have" context, not a required step. So, the playbook now uses a scope action with a `Try` mode around the HTTP call. If it times out or gets a non-200 response, the playbook logs a warning to the incident and proceeds without the enrichment data. The run succeeds, and the analyst still gets the rest of the automated triage.
I'd recommend at least that basic "skip-on-failure" pattern from day one. It's maybe 10 extra minutes of configuration and saves you from a pager alert at 2 a.m. because your internal API had a planned maintenance window you forgot about. You can always add smarter retry logic later if skipping becomes too frequent.
Architect first, buy later
That's a healthy stance, treating intel as context. But your confidence score's foundation sounds suspicious. You're deriving it from vouchCount and source. That's not a confidence score in the statistical sense, it's just a composite heuristic.
So when an IP with a low score gets appended as a "footnote," you're still implicitly trusting the categorization from that single, ancient IDS alert. It's just weighted less. Have you audited whether that footnote context has ever skewed an analyst's judgement? False context can be worse than no context.
Data skeptic, not a data cynic.
Three days for a for-each loop? That's the problem with these no-code tools - you spend all your time working around their limitations instead of actually solving the problem.
Your managed identity auth is fine, but you're ignoring the real bottleneck: that HTTP call per IP. Even with parallelism control, you're still bound by round-trip latency. Should've just dumped the feed into Sentinel as a watchlist and let KQL do the join.
-- old school