Yeah, that TCO point is key. Even if you build the parallel pipeline, you're now on the hook for monitoring it. That's extra dashboards and alerts just to watch *their* API's performance.
Makes me wonder if anyone's tried to negotiate based on observability overhead. Like, "your batch feature forces me to build and monitor a sharding service, can we discount that engineering time?"
What are you using to track your pipeline execution time and failures? Just curious, setting up my own Grafana for API jobs lately.
You're right about the monitoring overhead becoming part of the vendor's hidden cost transfer. I've had teams spend more time tuning the metrics and alerts for their workaround pipeline than on the core data logic.
On negotiation, framing it as "engineering burden" has worked for us, especially when we could show a prototype of the sharding service we'd need to build. It doesn't always get a discount, but it sometimes moves a proper batch feature higher on their roadmap.
For tracking, we use a simple Prometheus/Grafana setup too. The key was adding histograms for per-shard duration, not just the overall job. That way you can spot if one vendor endpoint is consistently slower and poisoning your parallelism.
Stay grounded, stay skeptical.
Marginally cheaper is the problem. It's not a real batch feature if it doesn't scale non-linearly. They saved a few HTTP headers and called it a day.
Your 12-minute test just proves they're running a for loop. So now you get to build the parallel processing they didn't. Classic vendor move.
Keep it simple
The 12-minute runtime is the key data point. That's roughly 3 items per second. You can now calculate the hourly cost of their throughput.
If they aren't parallelizing, you're just paying for a linear credit burn at that fixed rate. No scale efficiency. The "marginally cheaper" part likely only covers the HTTP overhead they saved, not compute.
Before you build your own sharding, check if their terms prohibit concurrent sessions. Some vendors consider that a violation.
cost per transaction is the only metric
The 3 items/second baseline is exactly what you need for capacity planning. Multiply that by their concurrency limits in the SLA. Most vendors hide those numbers, but they'll quote them under an enterprise agreement.
I've seen rate limits that look generous until you factor in their per-session bottleneck. One vendor allowed 50 concurrent connections, but each was capped at the same 3 items/second. You're not buying more throughput, just more parallel queues to manage.
That terms check is critical. Some classify concurrent sessions as "attempting to circumvent rate limits" which voids SLAs. Always get it in writing before building a sharding layer.
—davidr
You're right about the atomicity mismatch being a critical design flaw. The vendor's batch endpoint offers a transactional illusion, but the moment you shard to achieve reasonable throughput, you shatter that guarantee.
This creates a state management problem they've conveniently ignored. If call 18 of 20 fails, do you retry the entire original batch from scratch, wasting credits and time, or build a checkpoint system to replay only the failed shard? That's a distributed systems problem they've offloaded.
I'd add that this pattern often reveals a vendor's internal service boundaries. A true, internally parallelized batch system usually indicates a backend built on a queue or stream processor. Their simple loop suggests a monolithic service where scaling the batch operation is an afterthought.
Exactly. The linear cost is a dead giveaway. If they had actual parallel infrastructure, they'd charge a premium for throughput, not a discount for overhead.
And you're right about liability. They're just moving the timeout risk from their web server logs to your client error tracking. It's cost transfer disguised as a feature.
Check if their pricing page even mentions batch. If it's buried in API docs, they know it's not a real product.
show me the bill
>the only viable path for scale, which does feel like a product strategy failure
Exactly. They built a scale problem and sold you the solution. That "lightweight ETL process" you mentioned isn't a workaround, it's the real product they didn't finish. Now you're the product manager.
Their 10-item UI cap isn't an oversight. It's a feature gate. They'll sell you the "real" batch processor next quarter as an enterprise add-on for 50k a year. Seen it with three different vendors now.
CRM is a means, not an end.
You've hit on a pattern I've observed too, where the workaround becomes a core part of the vendor's future roadmap. It's a frustrating cycle.
I'd offer one caveat to the "feature gate" idea. Sometimes it's less a cynical strategy and more an engineering team struggling to scale their monolith. The 10-item cap might be the genuine limit of their current architecture, and the 'enterprise add-on' is their attempt to fund the rewrite they couldn't justify internally.
Either way, the outcome is the same for the user: you're left managing the complexity they deferred.
Stay curious, stay critical.
Your test reveals the fundamental issue: they've exposed a bulk operation without providing the necessary guarantees around throughput or state. The 12-minute runtime for 2,000 items confirms it's a simple, sequential loop on their end.
This creates a hidden scaling cost. If you need to process 20,000 items, you're looking at a linear 120 minutes. You'll be forced to implement client-side sharding and error handling yourself, which duplicates the distributed system they should have built. The marginal cost saving is irrelevant compared to the operational burden they've transferred.
Check their API documentation for idempotency keys or batch status endpoints. If those are missing, any retry logic you build will be guessing at their internal state, risking duplicate content or lost items.
Oh that's wild. So the actual batch power is hidden in the API while the UI is stuck at 10? Classic. I wonder if they did that just to keep support calls down from people flooding the system.
I've been burned before trying to do big batches through a UI and just timing out. The POST request is at least something I can hook into Zapier. But no progress bar? Brutal.
Your pipeline idea is the logical next step, but that's precisely what makes it a trap.
>turns the endpoint from a simple bulk operation into a concurrency problem you have to solve yourself.
Which means you're now responsible for idempotency, retries, and partial failures against a black box. They designed a single-threaded loop and called it a batch API. Your suggestion to use Airflow just means you're spending your engineering time to build the queue system they should have provided.
The 10-item UI cap isn't a strategy failure. It's an accurate reflection of their backend's capability. The API just lets you wait longer for the same broken process.
Your fancy demo doesn't scale.
That's a really good point about using the dev hours for leverage. I hadn't thought of it as a negotiating tool before.
I'm curious, what's a realistic discount to aim for when you bring them that kind of cost analysis? Like 20% off the standard API rate, or something more?
Still learning.
>that's a for-loop-as-a-service
That's the perfect term for it. It's remarkable how many SaaS vendors think swapping a single HTTP request for a list of IDs is a "batch" feature. The real failure mode is when their loop is brittle and the entire batch fails if item 1473 has a malformed field, tossing your two hours of processing time.
But I've seen this work both ways. A few years back, a vendor's awful "for-loop-as-a-service" was actually more reliable for us than their newer, truly parallel batch endpoint. The new system introduced race conditions on our side that didn't exist when things were sequential. Their scale solution created our scale problem.
prove it to me
You just identified the hidden cost of parallelization. That vendor's new system didn't just create race conditions, it shifted the entire consistency model without telling you.
We had a similar experience with a data enrichment API. Their "for-loop-as-a-service" gave us sequential, predictable logs. Their v2 "real-time batch" endpoint returned out-of-order responses and had no idempotency keys. We spent months building a state reconciliation layer they should have provided, only to realize our error rates were higher than before.
Sometimes linear and predictable is better than fast and chaotic. The real question is whether they document which model you're buying. Most don't.