I've been evaluating experiment trackers for a multi-team MLOps pipeline, with a specific requirement for high-cardinality metadata and robust traceability. Weights & Biases is a frequent contender, but its pricing model becomes prohibitive at our scale of experimentation. The common suggestions—MLflow and Neptune—are already ruled out due to internal constraints. MLflow's artifact storage model didn't meet our governance requirements, and Neptune's query performance degraded with our metadata complexity.
I am seeking alternatives that provide similar functionality: hyperparameter and metric logging, artifact versioning, and a strong UI for comparison. However, my observability background places a premium on two aspects often overlooked: the ability to trace a model prediction back to the exact experiment, configuration, and dataset version that produced it (distributed tracing concepts applied to ML), and handling extremely high cardinality in experiment metadata without cost explosions.
I've briefly examined tools like ClearML and Comet.ml. ClearML appears to have a more open-source on-premise stance, which is attractive. Comet's focus seems aligned, but I lack deep hands-on experience. My primary concern is whether these systems treat experiment metadata as true structured events, allowing for the slicing and dicing I need for incident analysis when a deployed model degrades.
Has anyone implemented a tracking system that bridges the gap towards full observability? I'm less interested in feature checklists and more in real-world experiences regarding the cardinality of logged parameters, the ease of exporting all data for custom analysis, and the ability to integrate the experiment lineage into a broader tracing context (e.g., linking to OpenTelemetry traces). Concrete examples of integration pain points or successful scaling would be most valuable.
you can't fix what you don't measure
Your focus on traceability and high-cardinality metadata is the right filter. ClearML's open-source lineage tracking is indeed strong for tracing predictions back to exact dataset versions, as it treats pipelines, data, and models as first-class, versioned objects. However, its metadata system uses a relational backend, so the cardinality issue you mentioned with Neptune might resurface depending on your indexing strategy and scale.
Comet.ml handles high-cardinality tags well in my experience, using a schemaless design that avoids the cost explosion you're wary of, but its on-premise offering is less emphasized than ClearML's. For your traceability requirement, examine whether Comet's artifact lineage provides a true distributed trace graph or just a flat parent reference.
A third option to add to your shortlist is Kubeflow Pipelines, but specifically its metadata component. When decoupled from the full orchestration engine, it offers a pure, graph-based lineage model purpose-built for your exact need of tracing a prediction back through the experiment and data version. The UI for comparison is weaker, but it might serve as a core tracing layer you could front with another tool.
Start with the question.