Just spun up an OpenPipe inference endpoint backed by vLLM on AWS EKS, and overall, it's a huge win for cost and latency. But I hit a few configuration snags that weren't immediately obvious from the docs, so I'm sharing here to see if anyone else ran into these.
My main gotchas were around the IAM permissions for the S3 cache and getting the node autoscaling to play nicely with the vLLM backend.
* **S3 Cache Permissions:** The OpenPipe pod needs *both* `s3:GetObject` and `s3:PutObject` on the cache bucket, obviously. But it also needs `s3:ListBucket`. The vLLM backend does a check on startup that fails without it. Took me a bit to trace that 403.
* **Resource Requests/Limits:** If you're using the provided Helm chart, double-check the CPU/memory requests for the `openpipe` container. I found the defaults a bit too lean for stable batching, leading to some OOM kills during longer inference jobs. Bumping them up smoothed things out.
* **Node Selection:** You'll want to ensure your nodes have enough GPU memory (obviously) but also that the node labels/taints match your pod tolerations. My cluster is mixed, and the first deployment landed on a CPU-only node because I missed a nodeSelector.
Has anyone else deployed this stack? Specifically:
- Did you manage to get the cluster autoscaler to work well with the vLLM worker pattern?
- Any secrets around optimizing the `vllm` config in the `values.yaml` for lower latency on LLaMA 70B?
- Are you using the S3 cache or found it better to use a different backend for the model cache?
Ship fast, measure faster.