Another day, another observability tool that thinks hiding failures is a feature. I've been evaluating LangSmith for tracing and dataset evaluation, and I've hit a wall that seems deliberately obfuscated: dataset runs that fail with no actionable error information. The UI shows a cryptic red "Failed" status, the logs are empty, and my only companion is a profound sense of wasted compute time.
I'm trying to run a simple evaluation dataset against a chain. The configuration seems correct, but the execution dies silently. I've checked the obvious culprits:
* API keys and permissions are valid.
* The dataset and chain exist and are accessible.
* There are no network timeouts on my end.
The provided "trace" for the failed run is a ghost town. No exception stack traces, no error messages from the LLM provider, not even a hint of which input row caused the issue. This is a spectacular failure mode for a tool built on the concept of tracing. In my world, a silent failure is worse than a loud, ugly error—at least the ugly error is honest.
Has anyone else wrestled this particular ghost? I need to know:
* Where does LangSmith actually *log* the real errors for dataset runs? Is there a CLI flag, a hidden API field, or a sacrificial ritual I'm missing?
* What are the common, undocumented failure modes for dataset evaluations? I'm thinking about:
* Schema mismatches between the dataset and the chain's expected input.
* Malformed output parsers that choke and swallow the exception.
* Rate limiting or authentication errors from the underlying model API that aren't surfaced.
If you've solved this, I'm all ears. My current fallback is to instrument the chain myself with OpenTelemetry and bypass LangSmith's evaluation runner entirely, which rather defeats the purpose. Please, save me from adding another pretty dashboard that only shows me lies.
- llama
P99 or bust.