Let's be honest, everyone's rushing to slap "ML observability" on their stack because it's the new shiny. Arize AI gets a lot of hype, but for a small team with scikit-learn models on GCP? You're probably building a cannon to kill a mosquito.
Arize is built for scale, complex model types, and real-time pipelines. If you're batch-scoring a random forest once a day and logging metrics to a CSV, you're paying for a dashboard to tell you what you already know. The integration overhead alone is painful. You'll spend a week wiring up their Python client, setting up Phoenix, and configuring exports from Cloud Composer or whatever, only to monitor drift on features that haven't changed in six months.
Consider what you actually need. Are you struggling to explain predictions? `shap` and a custom Streamlit app in a Cloud Run container might do it. Worried about drift? A scheduled Cloud Function that computes PSI from your BigQuery tables and alerts via Slack is maybe 100 lines of Python.
```python
# Oversimplified, but you get the point
import pandas as pd
from scipy.stats import wasserstein_distance
from google.cloud import bigquery
def compute_drift(request):
# Fetch reference and current data
client = bigquery.Client()
# ... query logic
# Calculate metric
drift = wasserstein_distance(ref_data, current_data)
if drift > THRESHOLD:
send_slack_alert(f"Drift detected: {drift}")
return "OK"
```
Suddenly you're not managing another SaaS platform, worrying about vendor lock-in, or deciphering another pricing page that charges per "observation." The irony is that for many small-scale, stable scikit-learn use cases, the "observability" you need is basic engineering hygiene, not a dedicated platform.
So before you jump, ask what specific, recurring production problem you have that your current logs and metrics can't solve. If the answer is "we want cool charts for stakeholders," that's a different conversation—and an expensive one.
prove it to me