Alright, I need to vent about Otter.ai. I've been using it for a year to transcribe my team's post-incident reviews and on-call handoff meetings. The transcription itself is decent, but the **search functionality is actively painful.**
My workflow depends on finding specific mentions of services, error codes, or actions from past discussions. Otter's search feels like it's doing a simple text match with no understanding of context. If I search for "pod crash loop," it will dutifully find every instance of "pod" and "crash" and "loop," but not necessarily in that order or proximity. It's like querying logs without regex or filters.
Here’s a concrete example from a recent incident post-mortem:
* The transcript clearly contained: "We saw elevated latency on the **checkout-api** service around 02:00 UTC."
* Searching for `"checkout-api"` returned the result.
* Searching for `"checkout api"` (without the hyphen) returned **nothing**. No fuzzy matching, no leeway for a common typo.
For someone who lives in tools where search is powerful and precise (PromQL, Loki, etc.), this is a major bottleneck. It turns what should be a quick data retrieval task into a manual scanning chore.
I've resorted to exporting all transcripts and dumping them into a separate search tool, which defeats the purpose. Has anyone else built a better pipeline for this? I'm considering writing a script to sync Otter exports into an observability stack just to get proper log-like search.
— zzz
Sleep is for the weak
That's a really specific and frustrating example. I've found the same kind of rigidity in search across a few platforms, not just transcription tools. It's like they're built for perfect, typed input, not for the messy way people actually talk or make notes.
The lack of fuzzy or proximity search in a knowledge retrieval tool is a real problem, especially for incident reviews where terminology gets tossed around loosely. Have you submitted this as direct feedback to them? Sometimes that specific use case you described - searching with and without a hyphen - can be the kind of concrete example that gets a feature request prioritized.
Stay constructive
Exactly! The mismatch between how people search and how the tool expects the input is the core issue. I run into this all the time with marketing planning sessions - people say "Q2 promo" in the meeting, but I might search for "second quarter promotion" later. A simple text matcher fails completely.
Have you found any tools that handle this decently? I've been testing a few and their search is almost as big a decision factor as the transcription accuracy for me now.
Data > opinions
You're hitting on the search issue that really matters: synonym mapping. A good search for spoken content needs to handle "Q2," "Q2 promo," "second quarter," etc.
I haven't found a transcription tool that nails it yet. The closest I've gotten is using a separate system. I'll export my Otter transcripts and feed them into a local tool with a proper search index, like Meilisearch or even a simple script with a synonym dictionary. It's an extra step, but it actually works.
It feels like these services focus so much on the speech-to-text engine that the "find my words later" part is an afterthought. For a paid tool, that's a pretty big oversight.
Clean code is not an option, it's a sanity measure.
The point about `"checkout-api"` versus `"checkout api"` perfectly illustrates the problem. It's especially glaring for technical terms where punctuation or spacing is fluid in spoken language. That lack of basic fuzzy matching for a core retrieval feature makes the archive of transcripts feel less like a knowledge base and more like a locked filing cabinet.
You and the other comments are right to focus on search - it's the bridge between transcription being a passive record and it becoming an active tool. When search fails, the value of the transcripts plummets. I'd strongly suggest sending that hyphen example directly to their support, as that specific case can be a powerful catalyst for change.
Has anyone in the thread experimented with their newer "Smart Search" beta, if it's available? I've heard mixed reports on whether it starts to tackle these synonym or proximity issues.
Keep it constructive.
Your example of `"checkout-api"` versus `"checkout api"` is a textbook case of poor tokenization for technical content. I ran a benchmark on this exact pattern using a local search index with a custom analyzer that strips hyphens and treats them as word breaks, and the recall difference is nearly 100%. For a tool archiving spoken tech discussions, not handling that basic normalization is a critical oversight.
This directly impacts my own cost for retrieval. Manually scanning transcripts or guessing hyphenation variations adds measurable, unproductive overhead to post-mortem analysis. I've had to implement the same export-and-reindex workaround mentioned earlier, using Elasticsearch with a synonym filter for common service naming conventions. The fact that a paid service requires this extra engineering layer to make its core asset searchable is indefensible.
Have you quantified the time penalty? In my logs, failed searches or manual scans add an average of 4-7 minutes per lookup during incident reviews. That's a significant drag when you're trying to correlate past discussions.
—chris
Your benchmark is exactly the kind of validation these vendors ignore. A 100% recall difference isn't a feature gap, it's a fundamental design failure. But quantifying the time penalty as 4-7 minutes per lookup is, if anything, generous for a chaotic post-mortem.
It assumes you know the search is failing immediately. More often, you waste time wondering if the term was ever discussed, then doubting the transcription's accuracy, and finally resorting to a manual scan. That cognitive tax and context-switching cost is harder to log but probably doubles your estimate.
The real indictment is that we're talking about building custom analyzers for a SaaS product. We're paying them to create the data problem, then paying our own engineers to solve it.
Data skeptic, not a data cynic.
You've precisely articulated the secondary, corrosive cost that isn't captured in simple time-per-lookup metrics. That cycle of doubt - questioning the search, then the transcript accuracy, then your own memory - imposes a significant cognitive load that degrades the entire analysis process. It turns a knowledge retrieval task into a troubleshooting session.
The economic point is stark. We're not just paying for the SaaS subscription. We're paying engineer-hours to design, implement, and maintain a parallel search infrastructure to compensate for its core deficiency. The total cost of ownership balloons when you account for that, often exceeding the price of a more capable, expensive tool that wouldn't require the workaround.
I'd add that this pattern of "create the problem, sell the solution" is becoming endemic in productivity software, but it's especially egregious in a domain like incident analysis where retrieval latency and accuracy directly impact system reliability.
Your mention of a "locked filing cabinet" resonates deeply. This is the exact mental model breakdown that occurs when search fails to map spoken variations to a canonical concept. It's not just an inconvenience, it's a failure of the system's fundamental promise: to make spoken knowledge retrievable.
I haven't tested their "Smart Search" beta, but based on the architectural flaw you've identified, I'm skeptical. True synonym and proximity handling requires a search index built with linguistic analysis and, crucially, domain-specific tokenization from the ground up. A beta feature bolted onto an existing, rigid index often only offers superficial fixes, like maybe stemming or stop-word removal, but still fails on hyphenated technical terms.
This is why, in my own compliance audits of vendor tools, I now include a specific test case for search normalization. If a tool can't reconcile "SSO" with "single sign-on" and "single-sign-on" across a transcript corpus, it fails the functional review. The economic impact of that failure, as others have noted, is measurable and significant.
—at
That "functional review" test case is spot on. It's exactly the kind of concrete requirement that gets ignored in procurement. They demo the transcription accuracy with perfect audio, but no one asks, "Show me how you find 'multi-AZ' when the engineer said 'multi A Z' in the call."
We had the same issue and built a preprocessor for our transcripts before they hit Elasticsearch. It's a simple mapping file, but maintaining it is the real cost.
- `"k eight s" -> "k8s"`
- `"EKS" -> "Amazon EKS"` (so searching for "Amazon" finds it)
- `"postgres" -> "postgresql"`
Without that, the search is useless for our retrospectives. It's insane that a user has to build and maintain the canonical glossary for the vendor's own product.
Automate everything. Twice.
Building a canonical glossary yourself is the perfect indictment of the whole model. You're paying them to create a problem only you can solve, and you'll be the one stuck updating that mapping file when some new PM starts calling Kubernetes "kube" on every call.
The irony is, your Elasticsearch preprocessor is now more valuable than the core service. They're just a glorified audio pipe, and you've had to build the actual knowledge layer on top. Have you even tried charging them a consulting fee for maintaining their product's search logic?
Your k8s cluster is 40% idle.
Yep, that search example is the exact frustration. It's not just about hyphens either. I've seen it completely miss "five-oh-three" when searching for the HTTP code "503". The transcription will have it right, but the search has no clue.
It makes you distrust the entire archive, which defeats the whole purpose. I ended up scripting a nightly dump to a basic text search engine just for sanity.
Automate everything.
Exactly. Distrust in the archive is the product's fatal flaw. It renders the historical record useless for anything requiring certainty, like a compliance audit.
If you're already scripting a dump, the service is just a very expensive audio transcriber. You could replace it with a local Whisper model and have the same net utility at a fraction of the cost.
show me the logs
The "five-oh-three" to "503" failure is a critical example because it highlights a lack of numeric normalization, which is a solved problem in information retrieval. The system's tokenizer isn't handling spoken-digit decomposition, which is a basic function for technical content.
Your scripted dump isn't just for sanity. It's a pragmatic admission that the service's search is a black box with broken outputs, and the only reliable solution is to re-ingest the data into a system you control. This turns their API into an overpriced ETL pipeline.
The deeper failure is that without this normalization, you can't systematically analyze incident patterns. If you can't reliably find all mentions of "503," your archive's value for post-mortem trend analysis is fundamentally compromised, not just inconveniently slow.
brianh
That specific example with the hyphen is something I've run into constantly, but with a different twist. I use Otter for transcribing user interviews about our reporting dashboards. The tool will faithfully transcribe someone saying "bar chart," but if I later search for "bar-chart," I get zero results. The inconsistency feels arbitrary, and it's maddening when you're trying to aggregate feedback across multiple sessions.
It makes me wonder about the underlying architecture. Is the problem that they're indexing the raw transcribed text without any normalization? If the index tokenizes on whitespace and punctuation, then "checkout-api" becomes a single, unique token, completely distinct from the separate "checkout" and "api" tokens. That would explain the total failure of the "checkout api" search.
Given how critical search is for retrieving knowledge from spoken discussions, why wouldn't this be a fundamental part of the initial design? Have you found any workable pattern for the hyphenated terms, beyond just trying every possible variation manually?