Skip to content
Notifications
Clear all

How do I handle nulls and defaults in Arize without breaking things?

2 Posts
2 Users
0 Reactions
29 Views
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
Topic starter   [#10205]

Dealing with nulls in Arize can be a real tripwire for your pipelines. I've seen teams send data that breaks monitors or skews dashboards because they didn't handle missing values before the payload. Here’s my practical approach.

First, decide on defaults *before* the data hits Arize. Your goal is to send clean, intentional data.
* For **numerical features/preds**, impute with a clear marker (like `-999` or the column mean) in your inference pipeline. Document this in the feature schema.
* For **categoricals**, create a `"NULL"` or `"UNKNOWN"` category. This keeps your data consistent and your drift calculations meaningful.
* For **timestamps**, never send a null. Use a default far past/future date if needed, but filtering it out later is often better.

Then, configure Arize to play nice. In your monitors, use the "Ignore Nulls" setting where it makes sense. For your defaults (like `-999`), set up specific monitor thresholds or exclusion filters so those placeholder values don't trigger false alerts. The key is aligning your upstream defaults with your Arize monitor logic.


Automate the boring stuff.


   
Quote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Exactly. The pre-processing step is non-negotiable. I'd add that the choice of your numerical marker is critical, and `-999` can be dangerous if your feature space could realistically hit negative values with high magnitude. I've had better luck using `np.nan` and then leaning on Arize's "Ignore NaNs" monitor setting, but only if the SDK/client you're using serializes `nan` properly on the wire. It keeps the data clearly invalid for any post-hoc analysis.

Also, for categoricals, your `"UNKNOWN"` category should be tracked for drift itself. If the proportion of `"UNKNOWN"` starts creeping up, that's a signal of upstream pipeline issues you might otherwise miss if you just filter it out. Make it a feature, not just a bucket.


Prod is the only environment that matters.


   
ReplyQuote