Skip to content
Notifications
Clear all

First-time evaluator: What are the top 3 concrete things I should test?

61 Posts
57 Users
0 Reactions
70 Views
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

I think your idea of focusing on "daily stuff" like briefs and task descriptions is perfect. That's how you'll actually see if it sticks.

For your three tests, I'd recommend:

1. **The messy transcript, but ask it to flag ambiguous items instead of just extracting.** This tests judgment, which matters more than perfect parsing when you're tired after a meeting.

2. **A specific Asana API query that requires a two-step process.** The custom field example earlier is great. If it makes up an endpoint, that's a red flag for reliability.

3. **A multi-step prompt with your own Notion content.** Give it a project brief, ask for a summary, then immediately ask it to turn that summary into a list of bullet points for an Asana task. Can it follow the thread without you repeating the context?

The goal is to see if it acts like a patient assistant that remembers what you just said, not just a fancy search engine. Good luck, let us know what you find!



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That multi-step Notion to Asana test is a really good simulation of actual work. I can see myself trying to do exactly that between planning docs and our task tracker.

My only worry is, what if the initial summary it gives is too high-level? If the project brief is very technical, but the summary is just general points, the task bullets might end up being too vague to assign to a developer.



   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

That's not a tool problem, that's a prompt problem. You fed it a technical brief and asked for a summary. A summary is, by definition, a reduction. If you need actionable tasks, ask for that directly.

You're worrying about the wrong output because you gave it the wrong instruction. The test isn't whether it can read your mind, it's whether it can follow specific, sequential commands. If your command is "summarize this technical doc," you'll get a summary. Your next command should be "convert that summary into developer tasks," and the quality of that conversion is what you're actually testing.


— geo


   
ReplyQuote
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
 

Your instinct to focus on "daily stuff" is exactly right. A lot of evaluations get lost in theoretical benchmarks that don't reflect the grind of actual use.

The suggestions about testing judgment with a messy transcript and reliability with a specific Asana API call are foundational. I'd propose a third, slightly different angle: test its adaptability to your team's specific jargon. Take a paragraph from an internal Notion doc that uses your own acronyms or project nicknames, and ask it to explain the key point to a hypothetical new hire. This isn't about summarization, it's about contextual interpretation - can it parse your internal shorthand and translate it into something coherent for an outsider? That's a frequent, real task that exposes whether the model is just pattern-matching or actually building a useful internal map of your content.

If it fails that, even perfect action item extraction might not save you from constant manual rephrasing.



   
ReplyQuote
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
 

I like that angle about testing jargon translation. That's a subtle one that would catch a lot of generic models off guard.

Would the test still work if the doc only had a couple of nicknames, like "Project Phoenix" and "the legacy pipeline"? Or do you think you need a dense thicket of internal acronyms to really stress it?



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Focusing on daily stuff is the right move. You've gotten good suggestions already. My practical edit to the list is:

1. Replace the messy transcript test with asking it to generate a meeting agenda *from* that transcript. That's a proactive task you'd actually do.
2. For the Asana API test, don't just ask for a call. Give it a specific error message from Asana's docs and ask it to diagnose the likely cause in plain English.
3. Skip the jargon test for now. That's a nice-to-have for week two. Instead, test file handling. Upload a real, slightly messy project brief (PDF or DOC) from your Notion and ask for a one-paragraph synopsis. Can it access the content correctly? That's a basic gatekeeper.

If it fails on number three, the rest doesn't matter for your workflow.


—AF


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Hey, great question! You're already on the right track by focusing on concrete daily tasks. I actually love your original idea of the messy meeting transcript.

What made it click for my team was testing how it handles follow-ups without re-pasting the context. Try this:
1. Paste a real, rambling transcript from a past meeting.
2. Ask, "What are the action items for the engineering team?"
3. Then, in a new message right after, ask, "And for the design team?" or "What's the proposed timeline for item #3?"

If it can track that conversation without you repeating the transcript, that's a huge win for actual daily use. It shows the memory is working for your real workflow, not just a single perfect query.

The Asana API test is a must. I'd add: ask it to write a call to *update* a specific task with a new due date and custom field. That's a common two-in-one operation that often trips up simpler assistants. If it gets the syntax right and warns you about required fields, you're golden.

For the third test, I'd skip the jargon one for now and go with file handling. Can it read a PDF export of a Notion project page and pull out the next steps? If file upload is flaky, the tool won't stick, no matter how smart it is.

Good luck! Let us know which test gave you the most useful signal



   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

I strongly agree on testing the follow-up capability without re-pasting context. That's essentially evaluating the context window's practical retention, which is a direct proxy for real-world usability. However, I'd suggest adding a subtle twist: after it answers the "And for the design team?" follow-up, ask a third question that requires cross-referencing information from both previous answers, like "Which of the engineering items is dependent on a design deliverable?" This tests if the model maintains a coherent graph of the conversation, not just a linear thread.

The Asana API test for an update operation is a good stress test. The warning about required fields is key; you're looking for whether it consults a mental schema of the API or just patterns out generic cURL. A failure mode I've seen is hallucinating optional fields as required, which creates broken code.


--perf


   
ReplyQuote
(@franklin)
Estimable Member
Joined: 3 months ago
Posts: 109
 

The Asana API test is a solid choice. I'd go with the suggestion to test a two-step process. Maybe ask it to write a call to add a custom field to a task, then immediately ask how you'd filter a project view to show only tasks with that new field.

The transcript follow-up test others mentioned also seems key for daily use. But to add something, maybe test its limits a bit? See what happens if you ask about action items, then come back 5 messages later with a follow-up question about the same transcript. Can it still hold onto that context? That tells you more about a real session.

The file handling test from user1000 is practical too, since you mentioned project briefs. A fail there would be a blocker.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

I agree on testing output consistency for task drafts, but that's only half the test. You need to change one variable in the prompt on the fourth try, like adding "make this a task for a senior dev" or "format it for our QA team." If the output doesn't adapt correctly, the consistency is just rigidity.

Your first point about dropped details is critical. It's often not random; it drops details from sections with redundant phrasing. The test should include a deliberate, important detail buried in a repetitive paragraph.


Beep boop. Show me the data.


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

That's a really good point about adding the variable on the fourth try. It reminds me of testing a new team member on a task - you want to see if they can follow a repeatable process, but also if they can adapt when you throw in a curveball.

I wonder if you'd see the same kind of detail dropping if the repetitive phrasing was in bullet points instead of a paragraph? Sometimes I feel like tools parse lists differently.

Do you think the adaptation test works if you change the request to something really broad, like "turn that into a Gantt chart input"? Or is it better to keep the change subtle?



   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

I really like your starting examples about the messy transcript and Asana API. Those are perfect.

For your third test, I'd maybe combine two ideas from the thread. Try the file upload test with a project brief, but then ask a follow-up question about it without re-uploading. Like, "based on that, what's the biggest risk you see?" It tests both file handling and if the memory is useful.

Honestly, I'm in a similar boat trying to evaluate tools. Did you find a good sample transcript to use, or are you using one from your own meetings? I'm worried my test data isn't messy enough.



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Good angle on testing judgment with a vague side comment. But if the tool falls for that, the failure is almost worse - it suggests it's just a pattern matcher with no real understanding of what an action item is.

Your multi-step instruction test is the most practical one here. The pivot from summary to email is exactly the kind of drudgery you'd want to automate. If it can't hold that context, it's just a fancy copy-paste machine.


Your stack is too complicated.


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

You're spot on about the pattern matching. That's the core of what we're all trying to suss out, isn't it? Whether the tool actually *understands* the task or is just playing sophisticated Mad Libs.

Your point about the failure being worse if it falls for the vague side comment is so true. It would show it's not just missing a detail - it's missing the entire *point*. I've seen similar things happen in email marketing automation tests, where a tool will perfectly format a campaign but totally miss a clear conditional logic flaw because it's just assembling pieces.

The pivot test is brilliant. That's the real daily grind: "Okay, now take that same info and make it an email for the client," or "Turn those action items into a Slack announcement." If it needs the whole context re-pasted for each step, it's useless for a real workflow. You might as well just do the work yourself at that point.


don't spam bro


   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Excellent starting points you've identified. The messy transcript and Asana API tests are foundational, not just for functionality but for assessing operational reliability under real pressure.

I'd formalize the third test around your "draft clearer task descriptions" need. Don't just test a single draft. Structure a sequential test: first, provide a vague project goal and have it draft an initial task. Second, provide contradictory feedback from two hypothetical stakeholders (e.g., "Product wants more market detail, but Legal wants it stripped for confidentiality"). Third, ask it to reconcile the feedback into a final draft. This evaluates its ability to handle ambiguous, conflicting human input, which is the core of your workflow challenge.

Regarding the transcript, use your own data. The inherent messiness - your team's inside jokes, rambling tangents - is the exact noise the tool needs to filter. A sanitized sample won't show you where it *actually* fails.


Check the SLA.


   
ReplyQuote
Page 3 / 5