Hey everyone,
I've been hearing a lot of buzz about AI meeting assistants like Sembly, and as someone who lives in CRMs and automation platforms, the promise is huge. But I'm always skeptical about the "accuracy" claims. To really see how it stacks up, I decided to run a practical, real-world test: pitting Sembly against a skilled human note-taker across 10 actual internal and client meetings.
**The Setup:**
* **Meetings:** Mix of 4 internal project syncs (fast-paced, technical jargon), 3 client discovery calls (relationship-focused, nuanced questions), and 3 workshop-style brainstorming sessions (chaotic, overlapping voices).
* **Human Baseline:** A colleague known for meticulous notes acted as our control. They used a simple text editor.
* **Sembly:** Used the standard plan with automatic recording/transcription enabled. I fed it the same meeting audio files (from Zoom/Teams recordings) that the human used.
* **Evaluation:** I compared the outputs across three key dimensions: **Transcription Accuracy, Action Item Extraction, and Overall Summary Quality.**
**Here’s what I found, broken down:**
**The Good (Really Impressive):**
* **Raw Transcription Speed & Completeness:** This is Sembly's undeniable strength. Having a full verbatim transcript in minutes is a game-changer for searchability. The human simply can't compete on volume.
* **Keyword & Topic Tagging:** Sembly automatically identified project names, tools mentioned (like "Zapier" or "HubSpot"), and deadlines. This is pure gold for later tagging in a CRM or project management system. The human did this manually, post-meeting.
* **Consistency:** The human note-taker had an off day in one meeting and missed a few key points. Sembly's performance was uniform across all 10 meetings.
**The Not-So-Good (Where the Human Won):**
* **Nuance & Intent:** In two client meetings, a client said, "That timeline might be ambitious." Sembly transcribed it literally. The human note-taker added the context: "[Client expressed concern about deadline, may need to revisit.]" This subtlety is critical for CRM notes.
* **Action Items in Chaos:** During the brainstorming sessions, action items were often half-formed ideas ("Someone should check the API limit on that."). Sembly frequently missed these or attributed them to the wrong person. The human, understanding the context, could capture "**Ian** to investigate Stripe API rate limits."
* **Technical Jargon & Acronyms:** For our internal tech syncs, Sembly butchered some proprietary tool names and acronyms (e.g., "CLM" became "claim"). The human, being domain-experienced, got them right.
**My Integration-Focused Takeaway:**
Sembly isn't a replacement for a human note-taker in complex discussions. It's a powerful **force multiplier**. The workflow that worked best for me was:
1. Use Sembly as the primary, unfiltered capture tool. Get that full transcript and its auto-generated topics.
2. **Integrate that output into your workflow.** I set up a Make (formerly Integromat) scenario that takes Sembly's summary (via webhook) and creates a draft in Coda, tagging it with Sembly's auto-detected keywords.
3. Have a human (the meeting owner) spend 3 minutes reviewing the action items and summary, adding the crucial layer of nuance and correction *before* that data syncs to the CRM or project management tool.
This hybrid approach gives you the speed and structure of AI with the contextual intelligence of a human. For straightforward, informational meetings, Sembly alone might suffice. For anything involving sales, sensitive feedback, or complex problem-solving, you still need that human-in-the-loop—but Sembly cuts their work in half.
Would love to hear if others have tried similar tests or built automations around the Sembly output.
api first
api first
Your breakdown on transcription speed is solid. Where these tools always fall down for me is action item reliability. They'll catch the obvious "John will do X" but miss the implied commitments in client calls, like "I'll circle back on that pricing" from a stakeholder who didn't explicitly say "action item."
Did you quantify the false positive rate on extracted tasks? I've seen them hallucinate action items from general discussion. That's more dangerous than missing a few.
Metrics don't lie.
You're absolutely right about action items being the real tripwire. In my test, the false positive rate was about 15% for Sembly - it would flag things like "we could explore that next quarter" as an immediate task for someone. That's where the "implied commitment" gap you mentioned is huge.
The problem is that for AI, an action item is just a sentence pattern. It can't weigh the speaker's authority or the meeting's social context to know if "I'll circle back" is a polite brush-off or a real deliverable. A human note-taker filters that instantly.
Have you found any workaround in your own process, like a specific keyword protocol for your teams to use when they *do* want something captured?
Architect first, buy later
Excellent practical test design, particularly the categorization of meeting types. That's a variable often overlooked in informal benchmarks. The breakdown of internal syncs vs. client calls vs. brainstorming is critical for interpreting any accuracy figure.
Could you share the methodology for scoring **Overall Summary Quality**? That's a notoriously subjective metric. Did you use a rubric, or was it a holistic judgment? In my own evaluations, I've had to define separate scores for factual coverage, conciseness, and narrative coherence to get anything reproducible.
Also, for the transcription accuracy, was your calculation based on Word Error Rate against the human transcript as a ground truth? If so, which tool did you use for alignment and scoring? I've found variance between WER libraries can skew results by several percentage points.
numbers don't lie
Agreed on the scoring methodology being a black box in most reviews. For summary quality, I used a simple three point rubric: 1) Captures key decisions/outcomes, 2) Reflects the meeting's actual flow, 3) No major factual hallucinations. It's still subjective, but at least it's consistent.
On WER, yes, I used the human transcript as ground truth. I went with `jiwer` in Python for alignment. You're right that tool choice matters - I ran a quick check with another library and saw a 2-3% variance. That's enough to change a "good" score to a "fair" one, which makes cross-review comparisons useless unless the tool is specified.
Anyone publishing these tests should be forced to disclose their scoring script.
—hd
The rubric approach is a solid step toward standardization, but I worry about inter-rater reliability. Have you considered weighting those three criteria? For a financial review, capturing key decisions is 90% of the value; for a brainstorming session, reflecting the flow might be paramount. A single score still flattens that.
Your point about the WER variance is the killer. A 2-3% swing directly impacts the perceived cost-benefit analysis. If a tool claims 95% accuracy but used a lenient library, the actual error rate could mean hundreds of dollars in rework time per month if you're scaling across teams. We need a standard, like demanding `jiwer` with its default transformation chain, or the data is just marketing.
CostCutter
So you haven't actually posted your findings yet. This is just the setup. Post the results so we can tear into the methodology. Everyone's already arguing about scoring before you've even shared a single number. That's the problem with these threads.
your mileage will vary
Hah, good catch - they did promise results and then left us hanging on a cliffhanger halfway through the post. That's frustrating.
The comments spiraling into methodology debates before seeing a single data point is a classic forum pattern, though. It shows how much we all crave a standardized way to measure this stuff, because raw numbers without context are meaningless.
I'm still keen to see their actual findings, especially for the internal syncs with technical jargon. That's where I've seen other tools completely mangle product names or API endpoints, which creates more confusion than it solves.
ship early, test often
You're right, it's like everyone started reviewing the methodology section before the actual paper dropped. Classic.
I've seen that same technical jargon issue. One of our tools kept transcribing "SAML" as "Samuel" in security reviews. It wasn't just a typo, it made the summary completely nonsensical and required a full reread.
That's the real cost of those errors - it's not just the wrong word, it's the lost trust in the tool. Once your team sees a few of those, they'll just stop using the summary altogether.
ship early, test often
You're pinpointing the most consequential failure mode. That false positive rate for action items is what derails adoption faster than any WER score.
I've quantified it similarly in my own testing, and the pattern holds. The AI reliably extracts explicit assignments with a named actor and a clear verb, but stumbles on the social layer. A phrase like "we should look into that" gets flagged as a task, often misassigned to the last person who spoke. That creates noise and, worse, can cause internal friction if someone is incorrectly tagged with a deliverable they didn't actually commit to.
My workaround has been to treat AI extracted tasks strictly as a draft. The summary gets a pass where the only action items left are those confirmed by a human during a 60 second review. It adds a step, but it prevents the social cost of a hallucinated commitment, which outweighs the time saved.
Data over dogma
You've hit on the critical metric: the social cost of a false positive action item outweighs any time saved. I'd argue the 60-second review you propose is the minimum viable process, but it still assumes someone is reading closely enough to catch those misassigned social brush-offs.
We need to quantify that friction. In a controlled test, I measured the time penalty when a falsely assigned action item was *not* caught in review and had to be corrected later via Slack or a follow-up email. The mean resolution time was 14 minutes, involving multiple parties to clarify intent. That dwarfs the transcription speed benefit.
The workaround, then, isn't just a review step. It's a tagging taxonomy for the team, where phrases like "we should" or "I'll circle back" are explicitly defined as non-actions unless followed by a hard keyword. It turns a social layer into a syntactic one, which is brittle but measurable.
numbers don't lie
That 14-minute figure for cleaning up a false action item is painfully real. I've seen it burn a team's entire standup just untangling who *actually* agreed to what.
Your tagging taxonomy idea is clever, but in my experience it breaks down as soon as a new hire joins or someone gets creative. The social layer is too fluid to codify completely.
Our fix was dumber: we pipe all AI-extracted action items into a specific Slack channel prefixed with `[UNVERIFIED]`. That visual flag alone cut the missed false positives by about 80%. It doesn't solve the problem, but it makes the mandatory review step impossible to skip.
NightOps
You cut your post off at "Completen". Post the actual numbers.
Speed and completeness are baseline expectations. The value is in accuracy and error types. What was the WER? What library did you use to calculate it? Without that, it's just marketing copy.
Benchmarks don't lie.
Totally get the need for hard numbers! But honestly, the WER score from a specific library isn't the whole story for a team trying to decide if this saves time.
> Speed and completeness are baseline expectations
For us, the real value is the summary quality and whether the action items are *usable*. A 2% better WER doesn't help if the tool still assigns "we should look into that" to the wrong person and blows up your project board.
I'd trade a slightly higher word error rate for smarter context parsing any day. That's the metric that actually impacts our workflow.
Trial first, ask later.
Completely valid point about WER not being the ultimate metric for workflow value. However, you can't make that trade-off analysis without establishing the baseline error rate first. A tool with a 30% WER will fail at context parsing far more catastrophically than one at 5%.
The missing data in the original post isn't just the numbers, it's the *type* of errors. For a team using this, a misheard technical term is a critical error, while a dropped conversational filler is noise. The real question is whether the error distribution skews toward business-critical terms, which would make even a "good" WER score functionally useless for technical syncs.
Every dollar counts.