Looking for eval frameworks that can run on-prem, integrated with a local LLM stack on Kubernetes. LangSmith is cloud/SaaS. Need open-source or self-hosted.
Evaluating:
* **AIConfig** - YAML-based. Could run the evaluator pods in-cluster.
* **Trulens** - Tracks chain steps. Their `Feedback` functions could be containerized.
* **Langfuse** - Open-source. Can deploy via their Helm chart. Direct alternative.
* **Phoenix by Arize** - OSS. Focus on tracing and evals. Good K8s fit.
Key requirements:
* Deployable via Helm/Kustomize.
* Storage (Postgres, etc.) can use in-cluster PVs or external DB.
* Minimal external API calls; everything local.
Example `values.yaml` for a Langfuse Helm install targeting an internal LLM:
```yaml
langfuse:
server:
environment: "production"
database:
internal:
enabled: true
openai:
# Point to local LLM gateway service
baseUrl: "http://llm-gateway.llm-stack.svc.cluster.local/v1"
apiKey: "dummy"
```
Biggest challenge is routing all tracing/eval calls internally. Need to patch SDKs to use internal service URLs, not cloud endpoints.
Who's running this setup? What's your stack (vLLM, TGI, etc.) and how are you managing eval datasets?
null