Skip to content
Notifications
Clear all

Help: Can't get Freeplay to work with our on-prem Llama2 instance

2 Posts
2 Users
0 Reactions
24 Views
(@sre_tales)
Eminent Member
Joined: 6 months ago
Posts: 15
Topic starter   [#2772]

Alright, gather 'round the virtual campfire, folks. I’ve got another tale from the trenches of trying to make shiny new SaaS tools play nice with our fortress-of-solitude, on-prem infrastructure. This one’s about Freeplay, which we were genuinely excited to try for streamlining our LLMOps and prompt management. The sales pitch was all rainbows and unicorns about connecting to any model endpoint. Our reality? A bricked integration and a support ticket thread longer than the RCA doc for that time we took out the entire payment service with a malformed config map.

So here’s the deal: we’re running Llama2 (the 70B variant, if it matters) on our own Kubernetes cluster, behind our standard API gateway with its own authentication. It’s not rocket surgery; it’s a `/v1/completions` endpoint that speaks mostly OpenAI-compatible JSON. Freeplay’s documentation makes it seem like you just plug in a “Custom OpenAI-Compatible” endpoint URL, add your API key, and you’re off to the races.

Spoiler alert: we are not off to the races. We are, in fact, stuck at the starting gate watching the error logs pile up.

The main pain points, in glorious bullet-point form:
* **Authentication Headers From Hell:** Freeplay seems to assume a very specific auth header format. Our on-prem setup uses `Authorization: Bearer ` but also requires a secondary `X-API-Key` for internal service routing. Freeplay’s custom endpoint config only allows for a single “API Key” field, which it transforms… somewhere… into a header. No option for custom header shapes or multiple headers. The result is a relentless 401 from our gateway.
* **The Health Check Black Box:** The integration has a “Test Connection” button. It’s about as useful as a pager that only goes off when the sun is shining. It’ll report “Connection Successful” even when our logs show nothing but auth failures, which implies it’s not actually testing the model call, just maybe a TCP handshake? Fantastic.
* **Prompt Templates Meet Silent Death:** Even when we jury-rigged a proxy to strip and reshape the auth (don’t @ me, it was for testing), we could *sometimes* get a connection. But then, using a Prompt Template in a Playground just spins for 30 seconds and fails with a useless “Model invocation error” in the UI. Our model logs show the request came through, but the response format—while valid JSON—wasn’t what Freeplay expected, I guess? Zero logging on their side to show what it *did* expect.

I’ve been through similar integration wars with incident management tools (looking at you, PagerDuty and your cryptic “Integration Key” dramas), so I know the drill. But this feels particularly opaque.

Has anyone else in this community tried to connect Freeplay to a truly custom, on-prem LLM? Not Hugging Face Inference Endpoints, not Azure OpenAI, but something you host yourself, behind your own walls. Did you have to build a custom adapter service? Are we missing some hidden config panel? Or is the cold, hard truth that “any model” really means “any *cloud* model they’ve already tested”?

Our team’s enthusiasm is currently waning faster than the battery on my laptop during a major incident bridge. Any war stories or tactical advice would be appreciated.


Postmortems are not blame sessions.


   
Quote
(@observability_watcher_2025)
Eminent Member
Joined: 7 months ago
Posts: 24
 

That authentication header issue is a classic. We hit something similar with a different monitoring tool. The gateway's custom header format just wasn't being passed through.

What's the actual error in the Freeplay logs? Sometimes the SaaS tool strips or reformats headers in a way you wouldn't expect. Could be a case mismatch.

Also, have you checked the network egress from Freeplay's IP range? Our on-prem setup initially blocked what looked like "non-standard" cloud provider IPs trying to hit internal endpoints.



   
ReplyQuote