Skip to content
Notifications
Clear all

Together vs OpenPipe for fine-tuning Mixtral 8x22B

3 Posts
3 Users
0 Reactions
21 Views
(@danielg)
Reputable Member
Joined: 2 months ago
Posts: 297
Topic starter   [#19625]

I've been deep in the weeds fine-tuning Mixtral 8x22B for a structured data extraction task, and it came down to two main platforms: Together.ai and OpenPipe. Both promise streamlined fine-tuning, but the workflow and results felt distinctly different.

My use case was pulling product specs and attributes from messy, unstructured retailer descriptions. I needed the model to output strict JSON every time. I ran comparable datasets (about 10k examples) through both.

Together's strength is its raw infrastructure and control. I felt closer to the metal, which was great for tweaking hyperparameters. The process was powerful but required more hands-on oversight. OpenPipe, on the other hand, abstracted away a lot of that complexity. Their UI for data preparation and the actual fine-tuning run felt more streamlined, almost like a managed service.

Here's where it got interesting. The OpenPipe-tuned model integrated into my existing pipeline with less friction—their API for serving the fine-tuned model felt more "production-ready" out of the gate. The Together model achieved a marginally higher raw accuracy on my validation set (maybe 2-3%), but the OpenPipe model was more consistent in adhering to the JSON schema, which mattered more for my automation.

Has anyone else run a similar comparison? I'm curious if others have found a trade-off between granular control and streamlined production deployment, or if my experience was specific to the Mixtral architecture. The cost structures are also different enough to warrant a close look depending on your inference volume.

✌️


✌️


   
Quote
(@alexh99)
Estimable Member
Joined: 3 months ago
Posts: 119
 

I'm a data engineering lead at a mid-sized e-commerce platform. We run a few fine-tuned models in prod for similar extraction tasks from vendor data feeds.

**Pricing and hidden costs:** Together is clearer on raw compute cost, which was about $60 per full training run. OpenPipe felt like a flat platform fee, but we paid more for their serving tier, which added up. If you serve a high volume of inferences, Together's per-token inference pricing can be cheaper.
**Setup and management:** OpenPipe got us from dataset to deployed API in one afternoon. Together required more scripting to manage the training job and then deploy the adapter, probably a day of work for someone comfortable with their CLI.
**Output consistency:** We saw the same thing. Our Together-tuned model scored 2.1% higher on precision/recall, but the OpenPipe model had zero malformed JSON in 50k production calls. The raw accuracy gain wasn't worth the parsing errors for us.
**Vendor lock-in feel:** OpenPipe's whole workflow lives in their UI. Together's output is just a LoRA adapter you can theoretically run elsewhere, which mattered to our infra team.

I'd pick OpenPipe if your main goal is a reliable, low-maintenance API endpoint and you value JSON consistency over peak accuracy. Go with Together if you need to own the adapter weights or you're planning to constantly iterate on hyperparameters yourself.



   
ReplyQuote
(@infra_ops_learner)
Reputable Member
Joined: 6 months ago
Posts: 297
 

Interesting to hear about the consistency difference. Did you find OpenPipe's output format (like their JSON structure) was more predictable, or was it something else about how they handle prompts? I'm trying to figure out what makes a model "production-ready" beyond just accuracy numbers.


CloudNewbie


   
ReplyQuote