Yeah, that precision@10 benchmark is always the reality check. I've had nearly identical results tagging support tickets for routing.
One caveat from our own tests: when you're dealing with industry-specific jargon, the free extractors sometimes split compound terms that an API might correctly keep together. We had issues with "L2 cache miss" becoming two separate keywords, "L2" and "cache miss," which hurt routing accuracy. But honestly, a small custom dictionary fixed that.
The cost per 1,000 documents is the real kicker. Once you're past prototyping, that operational overhead adds up fast. It's not just the direct cost, but the system complexity of managing another external dependency. For pure extraction, it's hard to justify.
✌️
Your benchmark aligns with our internal tests on support log streams. We found the latency variance of external APIs to be the bigger issue than cost. A local RAKE-NLTK call completes in 5ms ± 2ms, while an API call is 120ms ± 45ms. That jitter makes it impossible to maintain consistent throughput in a streaming pipeline without significant over-provisioning.
The nuance in API keywords can even be counterproductive for automated routing. We need keywords that match our internal taxonomy for ticket categories. If the API returns "authentication failure" but our routing rule is defined on "login error," we've added a mapping problem.
Exactly, you've hit on the real math. The compute cost for a local lib is negligible, especially if you're already running your own containers. It's just another layer in the Dockerfile.
I've seen teams forget about the integration tax, too. Every external API call is a new failure domain in your error budget. A network blip or a provider outage shouldn't stop your keyword pipeline.
git push and pray