Been evaluating Arize for our new ML monitoring stack and I'm honestly impressed by the feature set—especially the automated drift detection and the integration with our existing feature store. The UI is clean, and the Python client is pretty straightforward to use.
But here's my blocker: our security team has a hard "no external data" policy for our core models. Everything has to stay within the VPC. I've looked at the docs and it seems like Arize is a SaaS-only offering. This has me wondering... am I the only one who really needs a self-hosted/on-prem version?
I get that it adds a ton of complexity for the vendor, but for those of us in finance or healthcare, it's often non-negotiable. Right now, our workaround is a Frankenstein mix of:
* Grafana dashboards with custom Prometheus metrics
* Homegrown scripts to calculate drift (PSI, etc.)
* A separate pipeline just to log predictions and actuals
It's not sustainable. A self-hosted Arize would be a dream. We'd even be willing to handle the infrastructure overhead (K8s, etc.) ourselves.
Has anyone else hit this wall? Found any viable alternatives that offer similar ML observability but can be deployed internally? I've looked at Evidently AI, but it feels more like a library than a full platform.
--diver
Data is the new oil - but it's usually crude.
Yeah, that's a tough one. My team looked at Arize last year and hit the same wall. Our compliance folks wouldn't even consider a SaaS for model data.
We ended up going with Evidently AI for the drift/metrics part, and it runs in our own cluster. It's not as all-in-one as Arize's UI, but it's been okay. We pipe the results to Grafana.
Are you looking at other tools now? I'm curious what you find, especially for the whole pipeline logging part.
We looked at Evidently for a similar use case. It definitely gets the job done on the metrics side, especially for the price. That "piping to Grafana" step is key - it's workable, but you lose the cohesive view Arize offers for tracing an issue from a dashboard alert back to the specific inference.
Have you checked out whylogs/WhyLabs? They offer a hybrid option where you can keep the logs on-prem but still use their SaaS for analysis, which might satisfy some security reviews. Not quite full self-hosting, but a middle ground.
What's been your biggest pain point with the Grafana integration? For us, it was always building the dashboards that were flexible enough across different model types.
Trust the data, not the demo.
> That "piping to Grafana" step is key - it's workable, but you lose the cohesive view
This is the exact trade-off we measured when building our current monitoring layer. The loss of cohesion introduces a significant mean time to resolution penalty during incidents, because you're context-switching between systems. We logged the time to trace a production drift alert back to a problematic inference batch across three setups: a unified platform (like Arize), a Grafana/Evidently stack, and a custom dashboard.
The Grafana stack was, on average, 40% slower for the engineering team to resolve. The bottleneck wasn't dashboard flexibility for us - it was the manual correlation. You'd see a spike in the Evidently-generated Grafana panel, then have to query a separate logging store with the correct timestamp and model ID to get the inferences, with no shared context.
The WhyLabs hybrid model is interesting for the data locality requirement, but it still creates a boundary. Your data leaves the analysis loop for the UI/alerting component, which can be a compliance sticking point if "analysis" is interpreted broadly.