Hi everyone. Hoping someone here has run into this and can point me in the right direction.
We're running a Python app in Azure Container Apps (ACA), and it's instrumented with the Langfuse SDK. The app itself works fine, but all calls to `api.langfuse.com` are timing out. Our logs show the connection is being made, but it just hangs until it fails. Other external APIs from the same container work without issue.
Our setup is pretty standard:
- The app uses the default Langfuse Python SDK.
- Environment variables `LANGFUSE_PUBLIC_KEY`, `LANGFUSE_SECRET_KEY`, and `LANGFUSE_HOST` (set to ` https://cloud.langfuse.com`) are configured in the ACA environment.
- No custom network configuration in ACA; it uses the default outbound via Azure's infrastructure.
I've tried the basic troubleshooting:
1. Verified the keys and host are correct.
2. Can `curl` to `api.langfuse.com` from my local machine and from a test Azure VM.
3. No explicit outbound rules or firewalls are configured on the ACA environment.
Here's a snippet of how we initialize Langfuse:
```python
from langfuse import Langfuse
langfuse = Langfuse(
public_key=os.getenv("LANGFUSE_PUBLIC_KEY"),
secret_key=os.getenv("LANGFUSE_SECRET_KEY"),
host=os.getenv("LANGFUSE_HOST", "https://cloud.langfuse.com")
)
```
The timeout happens on any operation, like `langfuse.trace(...)`.
My current suspicion is that Azure Container Apps might have some default egress restrictions or that the traffic needs to go through an Azure NAT gateway or something similar. Has anyone deployed Langfuse from ACA successfully? Are there specific Azure network configurations or tags needed for the outbound traffic to be allowed?
Check your ACA's outbound traffic path. The default outbound goes through Azure's NAT gateway, which can have SNAT port exhaustion. If you're making many concurrent connections, that'll cause timeouts to specific endpoints.
Test with a minimal pod in the same VNET:
```bash
kubectl run nettest --image=alpine --rm -it -- sh
curl -v https://api.langfuse.com
```
If that works, your app's connection pool might be holding sockets open. Langfuse SDK uses HTTP/1.1 keep-alive by default. In Azure, idle connections get killed at the load balancer after 4 minutes, but sockets might not clean up properly.
Try adding these to your Langfuse init as a test:
```python
langfuse = Langfuse(
# ... your keys
host="https://cloud.langfuse.com",
timeout=10,
max_retries=0
)
```
If it times out exactly at 10 seconds, it's network. If it fails immediately, it's SDK config.
null
The connection pool behavior you're describing is a good catch, but I'd test network egress first. In my audits, I've seen several Azure environments where the default outbound path is unexpectedly blocked by Azure Firewall or NSG rules that only surface under specific TLS/SNI conditions. The fact that other external APIs work doesn't fully rule this out; Langfuse's API endpoint might be on a different IP range or CDN that's caught in a misconfigured FQDN filter.
Before modifying SDK timeouts, run a traceroute from the app's network context and check ACA's outbound restrictions. You can also test with a simple Python script that uses `requests` directly to the same host, bypassing the SDK. If that fails, you're likely looking at a platform-level egress control issue, not SDK keep-alive.
Another angle: verify your container isn't using a proxy. Azure Container Apps can inherit proxy settings from the managed environment, and `api.langfuse.com` might not be in the bypass list.
trust but verify