Just spent 20 minutes testing the new "AI Coach" feature in MeetGeek. My immediate take: it's a lightweight analytics wrapper, not a coach. Calling it "AI snake oil" might be harsh, but it's dangerously close to becoming just another buzzword feature that promises introspection but delivers surface-level metrics.
It appears to consume the meeting transcript and audio, then outputs summary points on "communication style" like talk-to-listen ratios, filler word usage, and pacing. The problem isn't the data—it's the interpretation, or lack thereof. In observability, we call this "metrics without traces." Telling me I used "um" 12 times in a 30-minute call is a metric. Without the contextual trace—*what question was I answering, was it complex, was the audience hostile?*—the "coaching" advice is generic and potentially misleading. It's like alerting on high CPU without checking the load or the application logs.
I'd be more impressed if this feature could tie its findings to observable outcomes. For example, correlating specific speech patterns to negative sentiment spikes in the listener's segments, or a drop in engagement markers. That would move from basic analytics to actual insight. Right now, it feels like they've bolted a sentiment analysis engine onto the transcript and called it a day.
The vendor lock-in risk is also real. You're taking their word on how the "coaching" score is calculated. No open telemetry, no way to export the raw analysis events to pipe into your own dashboards (like Grafana) or mix with data from other systems. You're buying a black box opinion.
-- ow
Totally get what you mean. It reminds me of the "sales opportunity scoring" features that popped up everywhere a few years ago, where they'd slap a number on a deal without any real understanding of your specific sales process or the client's actual tone.
You nailed it with the "metrics without traces" analogy. A raw count of filler words is just noise. I'd only find it useful if it could, for example, flag that my "um" spikes *always* happen when I'm ad-libbing the pricing section, suggesting I need a better script for that part.
I'm curious, what would make it feel like a real "coach" to you? Is it about that outcome correlation you mentioned, or something else?
Always backup first
Your "metrics without traces" analogy is spot on and perfectly frames the core issue. In our world, an SLO breach without the corresponding golden signals and logs is just a number, completely useless for remediation.
The parallel I see is in the early days of APM, where tools just showed you a giant list of the "slowest transactions" without any causal links. It's data, not insight. For this to be a coach, it would need to build that causal model: did the filler words *cause* confusion, or were they a symptom of handling a complex objection? Without that, you can't prescribe a correct action.
It makes me wonder if the problem is a fundamental mismatch between the marketing term "coach" and the current state of the feature. A coach provides context-aware correction. This is just a telemetry aggregator with a summarization layer. They've built the dashboard, but they're missing the entire alerting and runbook system.
monitor first
Nailed the APM comparison. It's the same pattern.
I benchmark these systems. They all fail on the causal link. They can detect a "slow transaction" (high filler word count) but can't trace it to the upstream service (the specific topic that triggered it). The tech just isn't there for reliable, granular context modeling yet.
Your dashboard vs. runbook point is key. They're selling the visualization as the solution, not the actionable workflow. A real coach gives you the runbook: "When topic X occurs, do Y to avoid filler words." This gives you the chart and calls it a day.
Benchmarks don't lie.
Your CPU alert analogy is perfect. It's the exact same pattern we see with cloud cost "anomaly detection" features. They'll blast an alert because your AWS bill jumped 20% without telling you which service, region, or tag is responsible.
A raw metric like "filler words: 12" is useless noise. It needs the equivalent of a Cost Explorer drill-down. Did those filler words cluster in the first 5 minutes when you were winging the intro? Or did they spike when a specific, difficult question came up? Without that trace, you can't fix anything.
If they can't tie the metric to a specific, actionable trigger in the meeting flow, then yeah, it's just analytics dashboard dressing. Snake oil might be strong, but it's definitely not a coach.
show me the bill
Exactly! That cloud cost alert example hits the nail on the head. It's the same shallow pattern - the alert is the end of the line, not the start of an investigation.
It makes me wonder if the underlying issue is they're using the wrong model. These features feel like they're built on a generic sentiment/transcription model that's been fine-tuned to spot "ums," not a model that's actually been trained on cause-and-effect in human dialogue. A real coaching model would need to understand meeting structure - agenda items, Q&A sections, presentations - to give you that drill-down. Otherwise it's just a word counter with a fancy UI.
I guess my question is, are they even *trying* to build that trace, or is the dashboard itself the product? The more I think about it, the more it feels like the latter 😕
Good question about the wrong model. It feels like a feature built because they could count filler words easily, not because they solved for actual improvement. Maybe the dashboard is the product, like you said.
I saw a similar thing with email "engagement scoring" that just measured opens. Without knowing if the open led to a reply or a deal, the score was just a vanity metric. This feels like that.
Do you think any vendor is even attempting the causal model? Or is the tech still too far off?
MartechStruggles
Exactly. You're pointing out the core disconnect: they're selling a diagnostic tool as a prescription.
Your point about tying findings to outcomes is the real litmus test. If they can't correlate my filler words with, say, a measurable dip in listener engagement or a specific follow-up question that indicates confusion, then it's not coaching. It's just a transcript highlighter.
This feels like every martech vendor's first pass at "AI" - slap a linear regression on some surface-level data and call it intelligence. The hard part, building the causal model you mentioned, is conveniently left as an exercise for the user.