Skip to content
Notifications
Clear all

Help: Custom compliance checks aren't firing after the latest platform update.

43 Posts
40 Users
0 Reactions
156 Views
(@infra_ops_learner)
Reputable Member
Joined: 6 months ago
Posts: 297
 

Oh, that silent failure with no syntax error sounds tricky! I've only just started writing custom checks, so I hadn't considered that the `resourceType` alias could change and just fail quietly. Thanks for pointing out the specific PCQL version note.

Is the "fully-qualified internal identifier" you mentioned something you'd find in that inventory table query, or is it separate documentation?


CloudNewbie


   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

Yeah, that's the frustrating part. The inventory query should show the current format. The internal identifier is usually in the `resourceType` column there, but it can look like a long service path, not the friendly alias you used in your policy.

So if you see `aws.s3.bucket` in the query but your check uses just `S3`, that's the mismatch. Always check there first after an update.



   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

The real issue isn't just finding the mismatch, it's *why* the alias becomes a dead pointer overnight. They've moved to fully-qualified namespaces in the inventory schema but kept the old aliases in the policy editor's autocomplete, which is pure misdirection.

So you see `aws.s3.bucket` in the query, you change your check to match, and then next month it's `amazon.aws.s3.bucket.v2`. Chasing the internal identifier is a temporary fix for a process that's fundamentally broken. The platform should version these mappings and fail the policy parse on upgrade, not just return zero resources.


Trust but verify.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Agreed, the root cause is the broken mapping system. But it's not just aliases. The underlying inventory schema version is also decoupled from the policy engine version. I've had checks that used the *correct* fully-qualified name, but the field paths inside the JSON evaluation changed.

You need to correlate three things after every update: the PCQL version, the inventory schema hash (from the system tables), and the policy engine build. If they drift, your query is a dead pointer no matter what you alias it to.

The audit log's `evaluated: 0` is a symptom of that larger drift.


Numbers don't lie.


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

That three-way drift is exactly why my team scripts the version check. Our pre-scan hook pulls the PCQL version, schema hash, and engine build, then compares them against a known-good baseline. If any mismatch, it fails the pipeline and we don't run the checks.

Otherwise you're just wasting cycles on `evaluated: 0`.



   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

Scripting the version check is a logical escalation, but I'm curious about your baseline management. Do you store the known-good triple as a static artifact, or do you dynamically derive it from a previous successful scan? The latter seems more resilient if the platform updates gradually across regions, but it introduces its own race condition when the baseline itself shifts mid-cycle.

Our team attempted something similar, but we found the schema hash wasn't always accessible via a public API after the 8.2 update, forcing us to infer it from the inventory query's column set. That inference added enough latency to offset the cycle savings in some pipelines.


Data over dogma


   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

Ah, the scripting approach. The problem with a "known-good baseline" is that it's a mirage. The vendor doesn't guarantee *any* version stability between those three components. Your baseline is just a snapshot of a transient, undocumented compatibility.

> we found the schema hash wasn't always accessible via a public API

Of course it wasn't. That's the point. They provide just enough rope to hang yourself with automation, then change the knot. Inferring the schema hash from column sets is a performance hit today, and tomorrow it'll be wrong because they'll start returning columns dynamically based on your service tier.

You're not avoiding the drift; you're just building a more complex canary to tell you the mine is collapsing.


Buyer beware.


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

> find the new canonical service name from the updated resource schema documentation

Sure, if that documentation existed. Half the time the "updated resource schema" is just inferred from the API responses. The canonical name you find in the portal today might be gone tomorrow, replaced by a new internal GUID.

Your real fix is to stop using their aliases entirely. Query the raw inventory table and match on the `provider` and `type` fields directly. It's more verbose but at least it breaks loudly when the schema shifts, instead of silently returning zero.


Don't panic, have a rollback plan.


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

You're absolutely right about the raw inventory table being more stable than the aliases. That approach saved us during the 7.x to 8.0 transition.

The caveat I'd add is that even `provider` and `type` aren't completely immune. Last quarter, they split the monolithic `azure` provider into `azure.compute` and `azure.storage` in the inventory schema. Our checks on the raw `provider='azure'` field suddenly missed half the resources. So while it's less volatile, you still need to watch for those namespace refactorings, not just field renaming.

It's a better failure mode, though. Instead of `evaluated: 0`, we got a partial match that flagged in our reports immediately.


customer first


   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

That's a solid checklist to start with. Since you're on 4.0.2, the first thing I'd do is open the audit log for one of your custom policies and look for the `evaluated` count. If it's zero, the problem is almost always a resource identifier mismatch post-update.

The quickest fix: run a basic inventory query in the PCQL console for, say, S3 buckets. Check the exact `resourceType` value it returns now. Compare that string directly to the `resource` field in your custom check's config - even a small difference like `aws.s3.bucket` vs `AWS::S3::Bucket` will break it.

Once you update the resource type, give it a few hours for the next scan cycle. That usually gets them firing again.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@george7)
Honorable Member
Joined: 3 months ago
Posts: 572
 

That's a good, practical first step to recommend. It often is the simple mismatch.

The one caution I'd add is that the `resourceType` column in that inventory query can itself be version-specific. I've seen it show `aws.s3.bucket` in one UI console but a different internal path in the API response for the same scan, depending on which endpoint you hit.

So while checking there is correct, you also have to be sure you're looking at the *same* output the policy engine uses.


Keep it constructive.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Exactly. The policy engine often consumes a denormalized projection of the inventory table, not the raw query output you see. This is a classic impedance mismatch between the user-facing query interface and the internal evaluation pipeline.

You can see this by comparing the query's `EXPLAIN ANALYTIC` output against the engine's debug logs. The column aliasing and path resolution happen at different stages. A mismatch means your check is querying one virtual table while the engine evaluates another.

The vendor's documentation rarely surfaces these internal projections, which is why static version checks and schema inferences fail.


--perf


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Your starting checklist is solid, but I'll add a crucial step zero specific to platform updates.

Before checking the policy config, verify the evaluation engine itself is receiving data. The standard checks often use a different, more stable ingestion pipeline. After our last major update, the custom policy queue was silently paused due to a schema validation mismatch that didn't affect the built-in rules.

Go to **Monitoring > Status** and look for the "Custom Policy Scan" health metric. If it's red or yellow, the update likely broke the ingestion compatibility layer. You'll need to open a support ticket to have them reset the policy evaluation queue - a simple redeploy of the check won't fix it.

For the resource type mismatch, don't just rely on the PCQL console. Pull the exact resource identifier from the internal audit log of a *standard* policy that is firing correctly on the same asset type. That shows you the canonical string the current engine build is using.


Less spend, more headroom.


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

Good advice already, but definitely start with that Monitoring > Status page to see if the custom policy queue is healthy. That's saved me a few times after updates.

If the queue is green, then yeah, compare your policy's resource string to a fresh PCQL query *and* the raw API if you can. I've seen them differ, which is frustrating.

If you're new to this, Palo Alto's own docs for the 4.0.x update notes might list specific resource type changes. They sometimes bury it there. Good luck!


Automate everything.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

The correlation you're describing is real, and it's why our team now tracks those three version numbers as a single composite key in our deployment pipeline. We fail the build if they don't match a known-good tuple from the vendor's release notes.

But your schema hash point is key - it's often only available post-failure. We've had to write a watchdog that runs a diagnostic query before enabling custom policies after an update. It fetches the current schema hash from `sys.table_versions` and compares it against the hash embedded in our last successful policy run's metadata. A mismatch triggers a full policy revalidation before any scans are queued.

Even then, as you note, it's a dead pointer if the tuple drifts. The only reliable mitigation we've found is to version-lock the entire platform deployment, which of course defeats the point of automatic updates.


--perf


   
ReplyQuote
Page 2 / 3