Hi everyone! 👋 I've been reading a lot about LLM observability lately, and I keep seeing two main options come up: using OpenTelemetry (OTel) for LLM tracing, or using a dedicated platform like Langfuse. I'm pretty new to this whole space, and my team is just starting to build our first AI feature that uses an LLM.
From my (very basic) understanding, OpenTelemetry seems like the "standard" for all kinds of telemetry, and I've read it can be extended for LLMs. Langfuse, on the other hand, is built specifically for LLM apps. I'm trying to figure out which approach would be simpler for a small team like ours to get started with.
My main goal right now is to just see what's happening with our prompts and responsesβwhat works, what fails, and maybe some basic cost tracking. We don't have a huge ops team, so a steep learning curve or a lot of manual setup is a big concern for me.
For those who have tried both, or even evaluated them:
* Which one was genuinely easier to integrate? I'm thinking about the initial setup and getting basic traces.
* Once it's running, which felt easier to maintain and actually get useful insights from?
* Does using OpenTelemetry for LLMs mean we'd need to piece together a bunch of other tools (like a backend and a UI) to actually see the data?
Any practical experiences or "I wish I knew this earlier" tips would be incredibly helpful. I'm just trying to avoid going down a rabbit hole of complexity when we're still prototyping. Thanks in advance for sharing your wisdom!