Just wrapped up a 3-month stint with Bedrock for our nightly batch scoring job and finally pulled the plug. The promise of "enterprise-grade" started to feel like a euphemism for "you'll pay for the privilege of waiting."
We're processing about 1.2 million medium-length text items nightly, running them through a Llama 3.1 70B Instruct model for classification. Bedrock was... fine. Predictable, in the way a slow-moving glacier is predictable. Our p95 latency was consistently around 320ms per item, and the cost was sitting at a cozy $0.0021 per 1K output tokens. Not terrible, until you multiply it by a few hundred million tokens a month.
The switch to Together AI was frankly born out of annoyance. Their pricing sheet looked like a typo. Same model, same quality of output (we did a full validation run, obviously), but the numbers shifted. Our p95 is now hovering around 185ms. The cost dropped to $0.0009 per 1K output tokens. That's not a marginal improvement; that's a "why was I tolerating the old numbers?" level of change.
The catch? It wasn't a straight swap. Bedrock lulls you into a false sense of operational simplicity. With Together, we had to get our hands dirty with their batching API and tune the concurrency ourselves. No more magic "just throw more money at Provisioned Throughput" button. The reliability has been solid, but you're definitely more aware of the underlying infrastructure.
Has anyone else made a similar jump from a managed cloud offering to a more bare-metal-style provider for batch work? I'm curious if the latency/cost trade-off we saw is the norm, or if we just got lucky with our specific model and workload pattern. The savings are real, but so is the slight increase in operational fiddling.
just sayin'
Data over dogma.
The batch API is the critical piece. It's not just about queuing, it's about how they handle request packing and retries under the hood.
What's your sustained throughput now, and did you have to tweak batch size? Bedrock's "simplicity" often meant you hit a fixed, low concurrency ceiling.
Data over opinions