Hey everyone! I've been diving deep into Freeplay for the last few weeks to track my LLM experiments, and it's been a game-changer for prompt management and evaluation. But it got me thinking—my team also has a bunch of traditional ML models in production (think churn prediction, classification stuff). We're currently using a mix of spreadsheets and custom dashboards, which is... not ideal.
So, can Freeplay handle the traditional ML use case, or is it strictly an LLM-focused tool? I'm specifically wondering about:
* **Model versioning:** Can I track different versions of a scikit-learn or XGBoost model the same way I track LLM prompts?
* **Input/Output logging:** For a model that takes tabular data, can I log the feature set used for a prediction and the resulting score/classification?
* **Evaluation & Testing:** Does it support running offline evaluations on batches of inference data against a new model version before deploying?
I love the workflow for LLMs, and if it can extend to our older models, it would be amazing to have one unified platform. Has anyone tried this or seen it in action? I'm curious about the practical limits.
Beta tester at heart
Great question! I'm also new to Freeplay and using it mostly for LLM stuff, so I'm following this. I did see a mention in their docs about "model sessions" that aren't LLM-specific, but I haven't tried it myself yet.
For your point about input/output logging, I think it can log generic JSON, so your tabular feature set should fit. But I'm not sure if there's a built-in way to link that data directly to a model artifact like a .pkl file. Would love to hear if someone has a working setup for that.
Have you reached out to their support about traditional ML model versioning?
You're correct that the SDK's session logging accepts generic JSON, so tabular features can be serialized in. The linking to a model artifact, like a .pkl file, isn't a direct built-in feature, but you can achieve it through the metadata field. We've set this up by storing a URI pointing to the model artifact in our object storage (S3, GCS) within the session metadata.
The model versioning aspect is more about convention. Freeplay doesn't manage the model binaries themselves. Instead, you create a "model" object in the platform that represents, for example, "churn-predictor-v1.2", and then all sessions logged specify that model name. The version control is effectively delegated to your model registry or artifact repository, with the link maintained in metadata.
It works, but it requires some discipline in your deployment pipeline to keep the metadata consistent.
infra nerd, cost hawk
That's a really practical setup you've described. I think the key insight here is that Freeplay's "model" object becomes a logical wrapper, not a storage system. We use a similar pattern, and the discipline you mention is crucial.
One caveat we found is that if you're using the metadata field for the artifact URI, it can get crowded if you also want to log things like inference latency or the git commit hash of your feature pipeline. We started using a consistent naming prefix, like "artifact_uri:" and "pipeline_commit:", to keep it parseable later.
How do you handle the case where someone deploys a model but forgets to update the metadata link? We built a small validation step into our CI/CD to check the URI exists, but I'm curious if others have a smoother guardrail.
Trust the data, not the demo.
That's exactly what I've been wondering about too! I love Freeplay for LLMs, but our team has so many old-school classifiers running.
The part about offline evaluations before deploying a new model version is super important to me. Did you figure out if you can set up test suites for traditional models, like you do for LLM prompts? Or is the batch testing more of a manual process?