Skip to content
Notifications
Clear all

wandb alternative for a strict on-premise deployment without cloud dependency

1 Posts
1 Users
0 Reactions
21 Views
(@james_k_consultant)
Estimable Member
Joined: 4 months ago
Posts: 121
Topic starter   [#11573]

Having observed the near-universal recommendation of Weights and Biases for experiment tracking, I find myself compelled to challenge a critical, yet often overlooked, assumption in these discussions: the inherent desirability or even feasibility of a SaaS model for all organizations. While wandb's on-premise offering is frequently cited as the solution for data sovereignty, a deeper examination reveals it to be a complex, dependency-laden deployment that remains tethered to their cloud for critical functions like user authentication and centralized dashboarding in many configurations. This isn't true on-premise independence; it's a leased concession.

For those of us operating under genuine air-gap, strict regulatory (e.g., IRAP, ITAR), or architectural mandates requiring zero external network egress, the search must pivot to solutions designed from the ground up for disconnection. The core requirements are: experiment parameter logging, metric tracking, artifact storage, and visualization—all served from within a private network.

After evaluating several contenders, two primary pragmatic paths emerge, each with distinct trade-offs:

**Path 1: Self-Contained Open Source Platforms**
* **MLflow** is the most direct analogue. Its Tracking Server can be deployed on-premise with a compatible file store (e.g., NFS, S3-compatible on-prem object storage) and backend database (PostgreSQL). Crucially, its artifact storage can be pointed to internal blob storage. The UI is served from the same tracking server.
```yaml
# Example docker-compose for MLflow (simplified)
version: '3'
services:
mlflow-tracking:
image: ghcr.io/mlflow/mlflow
command: >
mlflow server
--host 0.0.0.0
--backend-store-uri postgresql://user:pass@postgres/mlflow
--default-artifact-root s3://mlflow-artifacts
--serve-artifacts
environment:
AWS_ACCESS_KEY_ID: ${MINIO_ACCESS_KEY}
AWS_SECRET_ACCESS_KEY: ${MINIO_SECRET_KEY}
AWS_ENDPOINT_URL: ${MINIO_ENDPOINT}
```
* **Kedro-Viz**, when paired with Kedro, offers a static, pipeline-centric visualization that can be served from any web server without a live backend, post-run.

**Path 2: Purpose-Built, Deployable Observability Tools**
* **TensorBoard** can be run in standalone mode, reading from event files on a shared filesystem. While less feature-rich for comparison, its simplicity for on-prem is a virtue.
* **Grafana + Prometheus** can be co-opted for ML tracking. Log metrics via Prometheus client libraries, store in a long-term TSDB, and build dashboards in Grafana. This leverages existing infra but requires more engineering.

The critical analysis often misses the **hidden costs**: maintaining the database, object storage, and associated CI/CD for updates. MLflow's model registry also introduces additional complexity. The true alternative isn't merely a software swap; it's an acceptance of infrastructure ownership. For small teams, a simple, script-based logging to a shared SQLite DB and a plotting script might be the most pragmatic "migration" of all.

Plan for failure.


James K.


   
Quote