Skip to content
Notifications
Clear all

Help: Custom model training costs spiraled, any way to reduce them?

3 Posts
3 Users
0 Reactions
14 Views
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
Topic starter   [#25385]

We're three months into a Cartesia POC for fine-tuning a text-to-speech model. Initial estimates were ~$500 per training run. Our last run hit $2,100. This is unsustainable.

Our config is standard PyTorch, 10k high-quality samples, training for 10 epochs. We're using their `cartesia/titan-v2` as the base model.

* GPU hours are the main cost driver. Logs show `g5.48xlarge` instances.
* Data preprocessing (their managed service) added a $300 flat fee we didn't anticipate.
* Checkpoint saving seems to be billed for storage *during* the job, not just after.

Has anyone dissected the Cartesia training invoice? Specifically:
* Is there a way to force a cheaper GPU instance type?
* Are the data preprocessing costs negotiable or avoidable if we pre-process ourselves?
* Any hidden levers to reduce epoch count without sacrificing quality? Our validation loss plateaus after epoch 5.

Our current training script preamble:

```python
from cartesia.fine_tuning import TrainingConfig, TrainingJob

config = TrainingConfig(
base_model="cartesia/titan-v2",
dataset="s3://our-bucket/dataset.jsonl",
num_epochs=10,
batch_size=32,
)
```

Need concrete cost-control strategies, not "contact sales." What are you actually doing to keep training bills down?


Metrics don't lie.


   
Quote
(@franklin77)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Your script is the problem. You're handing them a blank check with that high-level config. The g5.48xlarge is their default cash cow because it's simple for them, not because it's cost-efficient for you.

You need to stop using their managed training wrapper. Go direct to their low-level API or CLI where you can specify the exact instance type. Look for `g5.12xlarge` or `g5.24xlarge`; for a POC, you don't need the top bin. Batch size of 32 on a huge instance might be inefficient, anyway. Try a smaller instance with a batch size of 16 and gradient accumulation.

On the epochs, you answered your own question. If loss plateaus at epoch 5, you're burning money for no gain. Reduce epochs to 6 and implement early stopping with a patience of 1. That alone cuts your compute time nearly in half.

Preprocessing is absolutely avoidable. That $300 fee is for their convenience. Pre-process your 10k samples locally and upload the finished, cleaned dataset. It's a one-time effort that pays for itself after two runs.

Checkpoint storage billing is typical. They're charging for the provisioned EBS volume attached to the instance for the job duration. Faster storage costs more. You can likely specify a smaller volume size or a slower GP2 type instead of GP3, but the bigger win is to save fewer checkpoints. Do you really need one every epoch? Save only the best validation checkpoint.


Trust but verify — especially the fine print.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Three months and you're just now looking at the bill? Classic.

> Is there a way to force a cheaper GPU instance type?

No, not with their pretty wrapper. That's the point. They built it to keep you on the expensive hardware. Rip out their `TrainingJob` import and use their low-level API where you can specify machine type. It's more lines of code, but it's the only way.

Pre-process your own data. $300 is them charging you to run a script you could write in an afternoon. It's a tax on laziness.

Your biggest leak is epochs. If loss plateaus at 5, you're literally burning cash for 5 more epochs of zero progress. Early stopping isn't a "hidden lever," it's basic training hygiene you should've had from day one.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote