Solid advice already in the thread, especially checking that custom policy queue health metric.
Since you asked for beginner-friendly steps, here's my quick checklist after an update:
1. Go to **Monitoring > Status** and look for "Custom Policy Scan". If it's not green, that's likely the core issue.
2. If green, pull up one of your broken policies and look at its **Last Evaluated** timestamp and resource count. If it's evaluating zero resources, the resource identifier changed.
3. Run this in the PCQL console to see the *current* resource type names:
```sql
SELECT DISTINCT resourceType
FROM aws.s3.bucket
LIMIT 5
```
Compare the exact string to what's in your policy's `resource` field. Even a capitalization difference can break it.
The 4.0.x updates did tweak some AWS resource naming. I had to change a few from `AWS::S3::Bucket` to `aws.s3.bucket`. 😅
Did the status page show green for the custom scan queue?
Yeah, that's the quickest root cause. The `cloud.service` filter is brittle, though. It's just another alias. Use `provider` and `type` from the raw inventory table instead.
They change those less often. Not never, but less.
If your query's dead after an update, rewriting it against the raw table gives you a few more version cycles before it breaks again.
Just my two cents.
That's a smart defensive strategy. The raw `provider` and `type` fields are indeed more stable because they map to the foundational inventory model, not a curated presentation layer.
However, this approach introduces a new dependency: you now have to join on the asset metadata tables yourself to get attributes like region or tags that `cloud.service` queries often bundle implicitly. You're trading one type of version fragility for the complexity of maintaining those joins, which can also shift subtly between releases if the underlying table relationships change.
For long-term policies, we've started storing these core resource mappings in a small configuration table that we can update centrally post-upgrade, rather than hard-coding identifiers directly into each policy's logic.
Great checklist from user370, but skip step 3 as written. That `DISTINCT` query can lie to you after an update because the console uses cached metadata.
Run this instead to get the raw inventory view the engine actually sees:
```sql
SELECT DISTINCT provider, type, raw->>'cloudType' as cloudType
FROM inventory
WHERE provider = 'aws' AND type LIKE '%s3%'
LIMITๅฎๆ 10;
```
Match *those* values against your policy. If they differ from your policy's `resource` field, that's your break. Updates love to silently rename the presentation layer while the underlying `provider`/`type` stays put. It's a nasty little versioning trick.
The "Last Evaluated" timestamp showing zero resources is the dead giveaway.
- elle
Oh wow, this whole thread is super helpful but also a bit overwhelming 😅 Thanks everyone for the detailed steps. That "Monitoring > Status" tip from user1216 and user370 is probably where I should look first.
Just to confirm, since I'm still new, is that status page under the main admin settings? I don't want to go looking in the wrong place.
And user429, that query to check the raw provider and type sounds like a solid next step if the status is green. I'm hoping it's just a simple queue issue and not a schema thing. Fingers crossed! 🤞
Exactly, that silent zero-resource match is such a frustrating trap. Your point about checking the post-update inventory first is spot on.
I'd just add a caveat to your validation query: sometimes after these updates, even the `cloud.service` alias in the `config` table can lag. I've seen it show old mappings for a few hours while the system reconciles. Running a quick `SELECT * FROM inventory` with a limit, then eyeballing the raw JSON `provider` field, has been a more reliable immediate snapshot for me.
The PCQL doc note is a good find, I'll have to bookmark that section. Thanks for sharing!
cost first, then scale
That lag in the `cloud.service` alias is such a subtle gotcha, and it's why I always treat the first few hours post-update as a "read-only" period for policy validation. The inventory reconciliation job just hasn't caught up yet.
The raw JSON snapshot is definitely the way to go for an immediate check. I'll even take it a step further and query the `inventory_audit` table for the timestamp of the last schema sync, just to confirm the system state.
If that timestamp is very recent, I know any weird mapping I see is likely transient and I should wait before making changes. Saved me from a few unnecessary policy rewrites last quarter.
Prod is the only environment that matters.
That `inventory_audit` check is a solid move I hadn't considered. Makes perfect sense to timestamp the system state before making any policy changes.
My only add is that the reconciliation lag seems worst for global services like IAM versus regional ones. I can see S3 buckets populating correctly within minutes post-update, while IAM roles and policies stay stale in the alias tables for an hour or more. Makes the 'read-only' window you mentioned a moving target depending on what resources your policies actually target.
Oh, that's a really good observation about the lag differing by service scope! I'd bet it's because the global services have a much more complex dependency graph to reconcile after a schema change. S3 is just buckets and objects in a region, but IAM has policies, roles, users, groups, and all their cross-account links. No wonder it takes longer.
Makes me think we should build a simple watchlist of "slow reconciling services" to check first after an update. AWS IAM, Azure AD, GCP IAM... basically anything identity-related is probably going to be in that slow lane. Saves you from checking your fast S3 policies and assuming the whole system is ready.
I've seen the same behavior with clause order. It's less about total failure and more about the query planner picking a non-optimal access path after a metadata refresh. Putting your most selective filter first often forces a more efficient table scan.
On your Azure question: re-saving was never sufficient in our environment if the policy reference was truly orphaned. The policy engine's cache for those `resource` identifiers seems to persist until the rule is cloned and the old one is deactivated. We had to create a new rule, verify it fired, then delete the old one. Just editing left the old cached mapping active, resulting in a permanent zero-resource match.
p-value < 0.05 or bust
That first step in the status page under Monitoring is a good start, but if the standard checks are fine, the queue is probably okay.
You said you checked they're enabled, but did you verify the policy rule itself is still active? I've seen cases where the rule is on but its parent policy got deactivated during an update. It's a permissions sync thing that sometimes happens.
Also, check the rule's "Last Evaluated" time. If it's recent but shows zero resources, then it's definitely the schema mismatch the others mentioned. Your custom check is looking for a resource name that changed.
If the time is old, the rule might just be stuck. Try toggling it off and on again. Sometimes that kicks the scheduler.
That parent policy deactivation point is a great catch, I wouldn't have thought to look there. So you'd check the main policy list, not just the rule itself?
And if the "Last Evaluated" time is old, you mentioned toggling the rule. Does that usually trigger a new scan right away, or does it just get back in the queue for the next scheduled run?
Still learning
The parent policy check and rule toggle are the right first steps. If those don't work, the schema changed.
Run this query against the config table to see what your check is actually looking for now:
```sql
SELECT policy_id, target_resource, cloud_service FROM config WHERE policy_label LIKE '%your_custom_check_name%';
```
Compare that `target_resource` and `cloud_service` to what's in your current inventory. If they don't match, you need to update the rule's resource identifier. This is common after a 4.x update; they normalized several AWS service names.
No logs to check. The rule either evaluates (and shows in the dashboard) or it doesn't.
Numbers don't lie.