Everyone's raving about Panther's out-of-the-box dashboards. So I built a custom one for failed AWS logins across our 50+ accounts. The process was a slog.
The UI for building multi-account queries is clunky. Had to drop to the SQL editor immediately. Even then, aggregating CloudTrail events correctly across accounts and regions is a trap. You'll miss logins from global services unless you know the specific event patterns. My first version only caught 60% of the events.
Here's the core of the query that finally worked. The joins and time filtering are critical.
```sql
SELECT
p_source_label,
awsRegion,
eventSource,
errorCode,
COUNT(*) as failure_count
FROM panther_views.aws_cloudtrail_events
WHERE
eventTime > current_date - interval '7 days'
AND errorCode IN ('AccessDenied', 'InvalidClientTokenId', 'ExpiredToken')
AND eventName IN ('ConsoleLogin', 'AssumeRole')
GROUP BY p_source_label, awsRegion, eventSource, errorCode
```
Now I have to maintain this view and ensure the log sources are mapped correctly every time we onboard a new account. Panther's abstraction leaks badly here. It's a decent alerting engine, but for custom dashboards? You're better off piping the logs to a real observability platform.
Don't panic, have a rollback plan.
The "abstraction leaks badly" line is the real kicker. That SQL you're wrestling with is basically a custom parser for AWS's own inconsistent event schema, which defeats the purpose of a managed service. So much for out-of-the-box.
You mentioned missing global services on your first pass. I'm betting it also doesn't catch failures from STS:AssumeRoleWithSAML or WebIdentity federation, because those have different error patterns. The case study they published last quarter conveniently ignored that part.
Now the real fun begins when you try to correlate these failures with GuardDuty findings or account config changes. Spoiler: you're back to writing more custom SQL and maintaining a dozen brittle views.
cg
That's a very specific set of error codes and event names. Does that query actually capture failures from the Identity Center console? I thought those used different eventSource fields.
Also, if you're maintaining this, how do you handle it when a new account's logs are delayed? Does the dashboard just show zero for that account, or do you have to build in some lag tolerance?
That initial 60% catch rate is exactly what makes this type of dashboard a hidden time sink. You've nailed the pain point with global services.
Even with your refined query, I'd add a check for `eventSource` values like `sso.amazonaws.com`. Identity Center (formerly SSO) console failures can slip through otherwise. You might need to expand your `eventName` list to include `Federate`.
The maintenance burden you mentioned is real. Every new account means verifying the `p_source_label` mapping is consistent, or your aggregate counts are off. For something as critical as authentication failures, that's a fragile setup.
—Anita
You've hit on the universal experience - that initial 60% catch rate is where the real work begins, not ends. The query looks solid for IAM and console, but you'll also want to watch for `errorMessage` patterns, not just `errorCode`. Some failures, especially from federated services, only surface there with a 200 response code, which your filter would miss.
Your point about maintaining the `p_source_label` mapping is the hidden cost. We ended up tagging accounts with a standard naming convention at the source to avoid Panther's label drift. Even then, a dashboard like this needs a weekly validation step against raw logs from a sample account, otherwise you're flying blind 🧑✈️.
Have you looked at the event volume differences between `ConsoleLogin` and `AssumeRole` failures? In our environment, the latter is a much stronger signal for automation issues or attack patterns.
Architect first, buy later
Been there. That query's solid for the basics, but wait until you try to tie it to an actual budget. The real kicker is when finance asks for the cost of those failed logins in terms of wasted IAM analytics capacity or CloudTrail ingest costs for noise. Panther won't surface that.
You're right about the maintenance becoming a hidden SaaS cost. Every new account means another line item on your time sheet to babysit the label mapping. We found the drift got so bad we had to build a separate Lambda just to audit p_source_label versus our actual account tags.
Cloud costs are not destiny.