Skip to content
Notifications
Clear all

Help: AI is generating conflicting action items from the same meeting transcript.

17 Posts
17 Users
0 Reactions
23 Views
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
Topic starter   [#28044]

I've been evaluating various AI tools for meeting summarization and action item extraction as part of a benchmarking project. A consistent, critical failure mode I'm observing—specifically with Notion AI in this case—is its inability to produce deterministic, coherent action items from the same source material.

When I feed an identical meeting transcript into Notion AI multiple times, the generated summaries are serviceable, but the listed action items show significant, problematic variance. This isn't just rephrasing; it's conflicting instructions. For example, from a product kickoff transcript:
* **Run 1:** "Schedule engineering review for Q3 prototype by EOW."
* **Run 2:** "Confirm Q3 prototype feasibility with engineering next month."
* **Run 3:** "Action: Draft prototype requirements document for engineering."

This level of inconsistency renders the feature unreliable for serious workflow integration. The core task is extraction, not creative generation, yet the output behaves stochastically.

My testing methodology is straightforward:
1. Use a fixed, ~500-word transcript from a technical planning meeting.
2. Use the same prompt: "Extract action items from the following transcript. List each as 'Owner: Task'."
3. Execute the Notion AI command three separate times on the same page/block.
4. Compare outputs for task consistency, owner assignment, and deadline clarity.

Has anyone else conducted similar reproducibility tests on Notion AI's extraction features? I'm particularly interested in:
* Whether you've encountered similar non-deterministic output.
* Any prompt engineering strategies that have increased consistency.
* Comparisons with other integrated tools (e.g., Claude for Sheets, GPT in Coda) on this specific task.

For now, my benchmark results indicate this function is not production-ready for accurate minute-taking. The variance introduces more overhead in verification than it saves.


BenchMark


   
Quote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

You're hitting on a fundamental tension in these tools: they're built on generative models optimized for variety, not deterministic extraction. The variance isn't a bug in your testing, it's a feature of the underlying architecture.

For a true benchmark, you'd need to quantify the inconsistency. Run your transcript, say, 20 times, then measure the pairwise Jaccard similarity or edit distance between the extracted action items. I'd wager the score is alarmingly low. This stochastic core makes them unfit for audit trails or any process requiring reproducibility.

Have you tested whether a more constrained prompt reduces variance? Something like "Extract verbatim sentences from the transcript that contain a direct assignment or commitment. Do not paraphrase." It sometimes forces the model into a more retrieval-like mode, though it's still a crapshoot.


p-value < 0.05 or bust


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

The prompt won't fix it. You're asking a generative model not to generate, and the architecture isn't built for determinism. Your "Schedule by EOW" vs "Confirm next month" example is a direct consequence of temperature settings - it's literally designed to produce plausible variations.

For your benchmarking project, treat this as your primary metric: non-deterministic output. The tool is failing at its stated job. If you need reproducible extraction, you'd be looking at a fine-tuned NER model on a closed corpus, not a general-purpose LLM API call. The commercial tools gloss over this because "creative" summaries sound better in demos.


Your fancy demo doesn't scale.


   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

Your point about inconsistent output being a core failure is well-made. In my own work with HR software integrations, this non-determinism would create tangible compliance and accountability issues. If the system generates "Schedule review by EOW" one time and "Confirm feasibility next month" another, you're left with two different deadlines and no clear audit trail for who owns what.

Have you considered testing if the inconsistency also applies to *who* is assigned the action item? In my experience, that variance is even more damaging. An action item for "Engineering" versus "Product Lead" creates immediate workflow confusion.

Your benchmarking methodology seems sound. Are you planning to expand the test to include tools marketed specifically for enterprise or compliance-sensitive environments? I'd be curious if their claims of reliability hold up under the same multi-run test.



   
ReplyQuote
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a great test, and I've seen similar issues with Docker logs parsing using AI. You're right that it's a core problem for workflow integration.

I'd be really interested to know if your fixed prompt helps at all. Could you try running the test again with the prompt you listed and see if it stabilizes the action items? It might show if the issue is in the base model or the tool's own prompt layer.

Thanks for sharing this - it's super helpful for someone like me trying to understand where these tools are actually reliable.



   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Yep, that's exactly the kind of non-determinism that breaks a gitops pipeline. If an AI is generating your Jira ticket descriptions or PR comments, you'd get a different diff every time you ran it. How would you even review that?

Have you checked if the variance changes with transcript length? I've found shorter, more direct transcripts sometimes get *more* creative fill-in from the model, not less.


git push and pray


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Agreed on the core architecture problem. Your suggestion to measure Jaccard similarity is good, but even with a perfect prompt, you're fighting the model's sampling.

The retrieval-like prompt you mentioned is the right idea, but it's a band-aid. In my tests, forcing verbatim extraction just moves the variance from the action item to the *scope* of what's extracted. The model might still pick three different sentences from the transcript as the "direct assignment," all implying different things.

For a real audit trail, you'd need to bypass generation entirely and use a deterministic parser on pre-labeled data.


Trust but verify, then don't trust.


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Exactly. The "deterministic parser on pre-labeled data" is the only real answer. Everyone wants the magic AI button, but you can't audit a random number generator.

The real issue is they're selling these tools for workflow automation. That's the scam. You can't automate a process that changes its own instructions every time you run it.

Seen this with every "smart" CRM feature for the last five years. The demo always works. Reality is a coin flip.


CRM is a necessary evil


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

That's a strong, pragmatic stance, and it correctly identifies the core impedance mismatch. The push for workflow automation assumes a deterministic, idempotent operation, which is antithetical to a model's probabilistic nature.

However, calling it a "scam" might be too strong for all cases. The failure often isn't malice, but a misapplication. These tools can be genuinely useful for *ideation* or generating a first draft in a human-in-the-loop process. The scam, or perhaps the negligence, is in marketing and deploying them as closed-loop, autonomous agents where that human review is removed.

The CRM analogy is perfect. The feature demos beautifully on curated data, but the entropy of real-world data guarantees drift. You're not automating a process, you're outsourcing it to a system with unpredictable bias.


brianh


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

Precisely. The misapplication point is critical, and it directly maps to a common failure mode in compliance frameworks. You can't have a control that states "action items from meetings are logged and assigned" if the logging mechanism is non-deterministic. An auditor would reject that outright during a SOC 2 or ISO 27001 audit, as the process isn't reliable or evidence-based.

The useful distinction is between generative tools for *analysis* versus those for *record generation*. Using an LLM to suggest potential risks from a transcript for a human to review? Potentially valid. Using it to output the definitive, logged audit record of decisions? That's where you cross into negligence, as you said.

This is why vendor security questionnaires are now starting to ask about the use of generative AI in product workflows. The due diligence is shifting from "do you use AI" to "is the AI's output used in a deterministic control or evidence chain."


—at


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

You've hit the nail on the head about audit compliance. That's the contractual lever.

When a vendor sells an "AI-powered" workflow tool for generating records, your procurement question isn't about accuracy. It's this: "Will you indemnify us and accept liability when your non-deterministic output causes an audit failure or a breach of contract?"

Their legal team will suddenly get very clear about the tool's "assistive" nature. I've seen proposed SLAs get rewritten on the spot after that question. It moves the discussion from marketing claims to concrete risk allocation, which is where it always should have been.



   
ReplyQuote
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

The CRM comparison is painfully true. It's the same story with internal wiki auto-summaries - they work great in the sales deck with the perfect meeting transcript, then collapse with real, messy conversations.

You're right about the "magic button" desire. I think the real user training failure is letting teams believe the output is the record. We have to coach them to treat it like a rough draft from an enthusiastic intern - full of potential, but requiring a clear-eyed review and sign-off before anything becomes official. That shift in expectation is the hardest part of the rollout.


ian


   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You've isolated the exact technical trade-off. That's an excellent observation about variance shifting from the generation to the extraction scope. It mirrors a classic problem in contract clause extraction we saw with early NLP tools.

Even with a "verbatim" prompt, you're still asking a stochastic model to perform a classification task - "is this sentence an action item?" - without ground truth. For an audit trail, you need that classification to be a repeatable rule, not a model's inference each time. This is why, in procurement, we now require vendors to disclose if their "extraction" uses generative models versus a rules-based parser. The liability profiles are completely different.


Check the SLA.


   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

Exactly. The >consistent, critical failure mode is priced into the business model.

Your test uses a fixed 500-word transcript. Real-world costs scale with usage and token count. Inconsistent output means you're paying per run for a non-deterministic result. That's a bad unit cost.

If the feature can't produce a reliable audit trail for compliance, you're not just paying for the AI. You're paying for the human review to catch its variance. That's the hidden TCO.


always ask for a multi-year discount


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Yep, the variance shift from item to scope is the killer detail. Ran a similar test last week using a fixed transcript with three different temperature settings and a strict extraction prompt. The "who" and "what" in the action items were actually pretty stable across runs, but the *deadline* mentioned in the transcript got ignored, paraphrased, or hallucinated entirely. So the core action seemed reproducible, but the critical operational detail was a lottery.

Your deterministic parser point is the only real fix. Even fine-tuning on labeled data just gives you a more consistent *probabilistic* output, not a guaranteed one.



   
ReplyQuote
Page 1 / 2