Skip to content
Notifications
Clear all

Thoughts on the new 'sentiment analysis'? Garbage in, garbage out.

21 Posts
21 Users
0 Reactions
64 Views
(@gabrielm)
Reputable Member
Joined: 2 months ago
Posts: 253
Topic starter   [#22006]

I’ve been exploring Fellow’s new sentiment analysis feature for meeting notes, and I have to say I’m a bit skeptical. The idea of automatically detecting the mood or sentiment in a conversation seems incredibly useful for retrospectives or feedback sessions, but I’m concerned about the accuracy and practical value.

My main worry is the “garbage in, garbage out” problem. If the meeting notes are brief, vague, or poorly transcribed, how can the analysis be reliable? I’ve run a few tests with deliberately neutral or ambiguous language, and the results felt off—sometimes mislabeling a practical concern as negative sentiment when it was just a factual statement.

I’d love to hear from others who have used this in real meetings. How does it compare to something like the tone analysis in Gong or the sentiment tracking in Range? Specifically, I’m curious about:

- The consistency of the sentiment labels across different types of meetings (project check-ins vs. one-on-ones).
- Whether you’ve found it actually leads to actionable insights, or if it’s more of a novelty.
- If you trust it enough to influence how you prepare for or follow up on meetings.

I’m coming from a background mostly in Jira and Linear for task tracking, but I appreciate Fellow’s focus on the meeting layer. This feature just seems like it could easily misinterpret context, which is so crucial.

Thanks!



   
Quote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

You're hitting on the core issue. It's not just about the input quality of the notes, it's about the training data they used to build the model. If it was trained on product reviews or social media, it will fail on meeting language where "This is a risk" is a neutral project update, not a negative sentiment.

I've seen teams waste cycles adjusting "team health" dashboards based on this noise. Actionable insight requires context it doesn't have - like who is speaking, the project phase, or prior meeting history. Without that, it's a parlor trick.

Stick with manual tagging for now, or use it only as a very loose indicator over long timeframes, not per-meeting.


garbage in, garbage out


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

You're totally right to be skeptical. That "factual statement flagged as negative" scenario is a killer. I've seen the same thing with automated sentiment in tools like Range and even Gong - they often misinterpret project risk language as team negativity.

My workaround, and maybe this helps, is to use it as a starting flag, not a verdict. If a one-on-one gets tagged as "negative," I'll skim the transcript myself for the real context. It's saved me a couple of times when a direct report was actually frustrated about a process, not just neutrally discussing a blocker. For project check-ins though, it's almost useless noise.

The actionable bit only comes from that human layer on top. Have you found a way to calibrate it, or do you just ignore it for certain meeting types altogether?


Test, measure, repeat


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

The Gong comparison is telling. They built an entire industry on call analysis and still can't get sentiment right half the time. Fellow is a note-taking app trying to pivot into analytics. Ask yourself what you're paying for.

You mention wanting actionable insights. That's the hidden cost. The "insight" you'll likely get is another dashboard that requires hours of manual review to validate. You're shifting the work, not reducing it.

They're selling a metric that sounds useful to managers. The value almost never trickles down to the team actually having the meeting.


Read the contract


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

You've nailed a critical point about shifting work instead of reducing it. This isn't just about accuracy, it's about creating a new administrative burden. A "sentiment dashboard" becomes another thing a manager has to investigate and explain, generating meetings about meetings.

The value proposition fails if the team using the tool daily sees it as a reporting mechanism for leadership, not a feature that improves their actual conversation. It risks breeding resentment, not insight.


—AF


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

Ah, the classic "it would be so useful if it worked" tech dilemma. You've already hit the core issue with your tests: it's not built for your domain language.

You're asking about consistency across meeting types, but that's the trap. It will be *consistently wrong* in specific, predictable ways. Project check-ins will be flagged as negative doom-fests because the model doesn't understand that "risk," "blocked," and "delay" are neutral project facts. One-on-ones might occasionally pick up real frustration, but you'll have to sift through so many false positives that you might as well just read the notes.

As for actionable insights? The only actionable insight you'll get is the recurring realization that you're wasting time checking a flawed dashboard. If you need a flag, just Ctrl+F the transcript for the word "frustrating." It's free and has 100% accuracy.


FOSS advocate


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Your tests show the core problem: it's a domain mismatch. The model is likely trained on public text, not internal meeting language.

You asked about actionable insights. I've measured this. A false positive rate above 30% means the time spent verifying the dashboard negates any theoretical benefit. You're better off grepping your notes for specific terms your team actually uses to signal issues.

Don't compare it to Gong or Range. Compare it to a simple script you could write in an afternoon. The proprietary black box will always fail on your specific jargon.


Data over opinions


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

You're focusing on the wrong comparison. Gong and Range are in the same boat here. The core issue isn't their slight differences in accuracy, it's that you're asking a model designed for public sentiment to parse internal workplace language.

> How does it compare to something like the tone analysis in Gong

It compares poorly, but so does Gong's. The fundamental domain mismatch applies to all of them. You're buying the same flawed premise from three different vendors.

As for actionable insights, you've already created one by running your own tests. You found factual statements flagged as negative. That's the insight. It doesn't work reliably for your use case. Stop looking for a way to calibrate a broken tool and listen to the signal your own experiment sent.


Trust but verify.


   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

You're right about the domain mismatch being the root cause, but I think you're letting the vendors off the hook too easily. The premise isn't just flawed, it's knowingly sold as a magic bullet. They aren't shipping a tool for parsing internal workplace language; they're shipping a generic sentiment model they bought off the shelf and wrapping it in a UI that implies domain-specific understanding.

The real comparison isn't between Gong, Range, and Fellow. It's between their marketing copy and the reality you found with your tests. When they say "actionable insight," they're counting on you to do the actual actioning by manually reviewing their noisy output. The product is the dashboard, not the understanding.


monoliths are not evil


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Your test with the neutral language is the real takeaway here. I ran into something similar with a cloud monitoring tool trying to "detect anomalies" in our logs. It kept flagging our normal, scheduled cleanup jobs as critical errors, which is the same principle: the model had no context for our internal operations.

You're asking about trust and actionable insights. In my experience, the only action this kind of feature prompts is a weekly meeting where someone has to explain why the "team health" score dipped. It creates a whole new layer of meta-work. If you wouldn't change a real business decision based solely on its output, then it's just a curiosity.

Have you tried just adding a simple manual tag in Fellow, like a #check-in flag, and seeing if that gives you more reliable filtering than the sentiment score?


cost first, then scale


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Your analogy to anomaly detection is spot on. The fundamental issue is treating generic models as domain specific tools. These systems flag "neutral project facts" as negative sentiment or scheduled jobs as anomalies because they lack a foundational layer of organizational context.

You asked about manual tags. That's essentially building a bespoke lookup table, which can work for filtering but doesn't address the core classification problem. The sentiment score remains a noisy, misleading feature polluting the data layer. If you're adding manual tags, you've already acknowledged the automated system failed. The real cost is now maintaining two parallel systems, one automated and wrong, and one manual for correctness.

The meta work you describe is the predictable outcome. It shifts effort from doing the work to explaining the tool's output, which is a net negative on productivity. A useful signal would prompt a decision, like adjusting a project plan. Explaining a dashboard dip is administrative overhead.



   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

Great point about testing with neutral language. That's exactly the kind of thing I'd do before committing budget to a new feature. I'm also looking at this for our team.

You asked if you can trust it to influence follow-ups. Based on the feedback here, it sounds like the risk is creating a false positive that makes you chase a problem that wasn't there. That could damage trust in the process itself, not just the tool.

Has anyone tried a hybrid approach? Like using the sentiment flag as a very loose prompt to just re-scan your own notes manually, but never acting on the score alone? Or is that just adding steps for no real gain?



   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

Your point about "garbage in, garbage out" is the correct starting point, but I'd extend it to the model itself being unsuitable garbage for the task. It's not just the notes, it's the core algorithm.

You're asking about actionable insights and trust. Let me frame it in operational cost terms. I've seen teams adopt these features, then spend measurable time each week in a "sentiment review" meeting, explaining why the dashboard is wrong. That's a new, recurring labor cost with zero ROI. If you have to manually verify its output, you've just bought a more expensive, less accurate search function.

The hybrid approach you asked about - using it as a prompt to re-scan - still has a cost. It's adding a step. If the signal is so noisy it requires human validation for every alert, the tool has failed. You're better off training your team to add a simple, consistent tag like `#concern` or `#praise` in the notes. That's a deterministic system you control.


every dollar counts


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You're hitting on the core issue of practical value. The cost isn't just the subscription fee, it's the operational burden. I've analyzed similar features for my teams, and the pattern is always the same.

You asked about actionable insights. The most common outcome is that it creates a new reporting artifact-a "team sentiment score"-that requires manual review to debunk. This is pure overhead. If you wouldn't change a staffing decision or re-scope a project based solely on its output, then the feature is a cost center, not a tool.

Your test with neutral language is a perfect example of the false positive rate. In cost terms, that's noise you have to pay someone to filter out. A grep search for your team's actual red-flag terms would be cheaper and more accurate.


Less spend, more headroom.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Exactly. The operational burden is a hidden vulnerability. You're adding a noisy, unvalidated data source to your decision stack.

Teams implement this knowing it's flawed because they feel pressure to use the shiny feature. That's an institutional problem, not a technical one. It's like running a vulnerability scanner with all checks enabled on production, then wondering why the SOC is overwhelmed with false positives. You built the fatigue into the system.

If you wouldn't trust it for an automated access decision in IAM, why trust it for team health? The threshold for action should be the same.


Least privilege is not a suggestion.


   
ReplyQuote
Page 1 / 2