Skip to content
Notifications
Clear all

Thoughts on the new 'sentiment analysis'? Garbage in, garbage out.

8 Posts
8 Users
0 Reactions
0 Views
(@gabrielm)
Estimable Member
Joined: 2 weeks ago
Posts: 62
Topic starter   [#22006]

I’ve been exploring Fellow’s new sentiment analysis feature for meeting notes, and I have to say I’m a bit skeptical. The idea of automatically detecting the mood or sentiment in a conversation seems incredibly useful for retrospectives or feedback sessions, but I’m concerned about the accuracy and practical value.

My main worry is the “garbage in, garbage out” problem. If the meeting notes are brief, vague, or poorly transcribed, how can the analysis be reliable? I’ve run a few tests with deliberately neutral or ambiguous language, and the results felt off—sometimes mislabeling a practical concern as negative sentiment when it was just a factual statement.

I’d love to hear from others who have used this in real meetings. How does it compare to something like the tone analysis in Gong or the sentiment tracking in Range? Specifically, I’m curious about:

- The consistency of the sentiment labels across different types of meetings (project check-ins vs. one-on-ones).
- Whether you’ve found it actually leads to actionable insights, or if it’s more of a novelty.
- If you trust it enough to influence how you prepare for or follow up on meetings.

I’m coming from a background mostly in Jira and Linear for task tracking, but I appreciate Fellow’s focus on the meeting layer. This feature just seems like it could easily misinterpret context, which is so crucial.

Thanks!



   
Quote
(@data_pipeline_guy_42)
Estimable Member
Joined: 1 month ago
Posts: 82
 

You're hitting on the core issue. It's not just about the input quality of the notes, it's about the training data they used to build the model. If it was trained on product reviews or social media, it will fail on meeting language where "This is a risk" is a neutral project update, not a negative sentiment.

I've seen teams waste cycles adjusting "team health" dashboards based on this noise. Actionable insight requires context it doesn't have - like who is speaking, the project phase, or prior meeting history. Without that, it's a parlor trick.

Stick with manual tagging for now, or use it only as a very loose indicator over long timeframes, not per-meeting.


garbage in, garbage out


   
ReplyQuote
(@ethanc)
Trusted Member
Joined: 2 weeks ago
Posts: 38
 

You're totally right to be skeptical. That "factual statement flagged as negative" scenario is a killer. I've seen the same thing with automated sentiment in tools like Range and even Gong - they often misinterpret project risk language as team negativity.

My workaround, and maybe this helps, is to use it as a starting flag, not a verdict. If a one-on-one gets tagged as "negative," I'll skim the transcript myself for the real context. It's saved me a couple of times when a direct report was actually frustrated about a process, not just neutrally discussing a blocker. For project check-ins though, it's almost useless noise.

The actionable bit only comes from that human layer on top. Have you found a way to calibrate it, or do you just ignore it for certain meeting types altogether?


Test, measure, repeat


   
ReplyQuote
(@craigs)
Estimable Member
Joined: 2 weeks ago
Posts: 110
 

The Gong comparison is telling. They built an entire industry on call analysis and still can't get sentiment right half the time. Fellow is a note-taking app trying to pivot into analytics. Ask yourself what you're paying for.

You mention wanting actionable insights. That's the hidden cost. The "insight" you'll likely get is another dashboard that requires hours of manual review to validate. You're shifting the work, not reducing it.

They're selling a metric that sounds useful to managers. The value almost never trickles down to the team actually having the meeting.


Read the contract


   
ReplyQuote
(@amandaf)
Estimable Member
Joined: 2 weeks ago
Posts: 103
 

You've nailed a critical point about shifting work instead of reducing it. This isn't just about accuracy, it's about creating a new administrative burden. A "sentiment dashboard" becomes another thing a manager has to investigate and explain, generating meetings about meetings.

The value proposition fails if the team using the tool daily sees it as a reporting mechanism for leadership, not a feature that improves their actual conversation. It risks breeding resentment, not insight.


—AF


   
ReplyQuote
(@finnj)
Estimable Member
Joined: 2 weeks ago
Posts: 72
 

Ah, the classic "it would be so useful if it worked" tech dilemma. You've already hit the core issue with your tests: it's not built for your domain language.

You're asking about consistency across meeting types, but that's the trap. It will be *consistently wrong* in specific, predictable ways. Project check-ins will be flagged as negative doom-fests because the model doesn't understand that "risk," "blocked," and "delay" are neutral project facts. One-on-ones might occasionally pick up real frustration, but you'll have to sift through so many false positives that you might as well just read the notes.

As for actionable insights? The only actionable insight you'll get is the recurring realization that you're wasting time checking a flawed dashboard. If you need a flag, just Ctrl+F the transcript for the word "frustrating." It's free and has 100% accuracy.


FOSS advocate


   
ReplyQuote
(@gracep)
Trusted Member
Joined: 2 weeks ago
Posts: 52
 

Your tests show the core problem: it's a domain mismatch. The model is likely trained on public text, not internal meeting language.

You asked about actionable insights. I've measured this. A false positive rate above 30% means the time spent verifying the dashboard negates any theoretical benefit. You're better off grepping your notes for specific terms your team actually uses to signal issues.

Don't compare it to Gong or Range. Compare it to a simple script you could write in an afternoon. The proprietary black box will always fail on your specific jargon.


Data over opinions


   
ReplyQuote
(@danielr)
Estimable Member
Joined: 2 weeks ago
Posts: 87
 

You're focusing on the wrong comparison. Gong and Range are in the same boat here. The core issue isn't their slight differences in accuracy, it's that you're asking a model designed for public sentiment to parse internal workplace language.

> How does it compare to something like the tone analysis in Gong

It compares poorly, but so does Gong's. The fundamental domain mismatch applies to all of them. You're buying the same flawed premise from three different vendors.

As for actionable insights, you've already created one by running your own tests. You found factual statements flagged as negative. That's the insight. It doesn't work reliably for your use case. Stop looking for a way to calibrate a broken tool and listen to the signal your own experiment sent.


Trust but verify.


   
ReplyQuote