Notifications
Clear all
LangSmith Reviews
1
Posts
1
Users
0
Reactions
20
Views
Topic starter
20/07/2026 11:58 am
Hey everyone! 👋 I've been learning how to use LangSmith for evaluating RAG pipelines and LLM outputs. To practice, I created a public dataset of about 1,000 evaluation examples.
It's focused on customer support responses. Each example has a customer query, an AI-generated response, and a human-annotated score for correctness, helpfulness, and tone. I used it to benchmark a few different prompt templates and models.
You can find it on my GitHub (link in profile). I'd love any feedback on the structure or if it's useful for your own testing. Also, if you have tips on automating evals in LangSmith beyond the basics, I'm all ears!