Skip to content
Notifications
Clear all

Deployed Arize AI for customer churn models - pricing feedback and scaling issues

1 Posts
1 Users
0 Reactions
27 Views
(@integration_tinkerer)
Estimable Member
Joined: 6 months ago
Posts: 141
Topic starter   [#14811]

Hey folks, been running Arize AI for about 6 months now to monitor and explain our customer churn prediction models. We've got a multi-model pipeline that scores users daily, and Arize is hooked into our logging to track drift, performance, and generate those nice SHAP-based feature importance charts.

Overall, the observability part is solid. Being able to pin a spike in our `prediction_score_drift` metric to a specific feature engineering change was a game-changer. The UI for digging into cohorts (like "high-risk users who didn't churn") is intuitive for our data scientists.

However, I've hit two main friction points:

**1. Pricing & Scale**
We're on the Pro plan. The per-million-tracking-call pricing gets painful as you scale inference. Our initial architecture logged *every* prediction, which blew through our budget fast. We had to re-architect to sample logs for high-volume, low-stakes models, and only log exhaustively for our core churn model. This added complexity to our inference service. I wish there were clearer, more predictable tiering for high-volume production use cases.

**2. Integration Workflow**
Getting data in via the Python SDK is straightforward, but we wanted real-time alerts into our operations Slack channel. The built-in Slack integration only covers certain alert types. We ended up building a small middleware service that:
* Subscribes to Arize webhooks for drift alerts
* Enriches them with internal deployment metadata
* Routes them to the appropriate Slack channel via a custom Zapier zap

Here's a snippet of the webhook processor logic:

```python
# Simple Flask endpoint to handle Arize webhooks
@app.route('/arize_webhook', methods=['POST'])
def handle_alert():
alert = request.json
# Add our internal model version tag
alert['internal_model_id'] = get_model_version(alert['model_name'])
# Forward to Zapier webhook
requests.post(ZAPIER_WEBHOOK_URL, json=alert)
return jsonify(success=True)
```

Has anyone else run into similar scaling costs? Curious if you've found workarounds for the logging volume, or if you're using the batch ingestion features differently to manage costs.

Also, if you've built custom integrations to pipe Arize data into your CRM (like Salesforce) for the sales team to see account-level risk scores, I'd love to hear about your setup. We're currently syncing key inferences to HubSpot via a scheduled Make.com scenario, but it feels clunky.



   
Quote