We migrated our ThreatConnect instance to a new Kubernetes cluster last week. The core app came up fine, but now we're seeing intermittent failures in external integrations. Webhooks from our SIEM are timing out, and API-driven playbooks to AWS are sporadically failing with connection errors.
The old cluster was on-prem k8s 1.23; the new one is EKS 1.27. Network policies are more restrictive. I suspect the issue is either in the service/ingress configuration for the ThreatConnect pods or a change in how network egress is handled.
Here's a snippet of the current service definition for the API pods:
```yaml
apiVersion: v1
kind: Service
metadata:
name: tc-api-service
spec:
ports:
- port: 443
targetPort: 8443
protocol: TCP
selector:
app: tc-api
type: ClusterIP
```
And the relevant error from a playbook log:
```
ERROR - Failed to execute action [aws:DescribeInstances]. Reason: HTTPSConnectionPool(host='ec2.us-east-1.amazonaws.com', port=443): Max retries exceeded with url: / (Caused by NewConnectionError(': Failed to establish a new connection: [Errno 110] Connection timed out'))
```
Has anyone else hit this after a platform move? I need to verify:
* If the ThreatConnect pods have correct egress routes to the internet/AWS endpoints.
* If any internal service discovery changed, breaking how integrations resolve the TC API URL.
* Any required annotations for the service mesh (we're not using one, but checking).
The flakiness is killing our automation. What specific networking or k8s service configurations did you have to adjust to get integrations stable again?
Build once, deploy everywhere
Your service definition snippet shows a ClusterIP service, which is only reachable within the cluster. For external integrations like webhooks and playbooks to reach the ThreatConnect API, you'll need an external entry point. This is likely the root cause for your SIEM webhooks timing out.
For the AWS API timeout error, that's an egress problem from your pods. The move to stricter network policies in EKS is a strong indicator. You need to check if your pods have the correct egress rules to reach AWS endpoints. Start by describing the NetworkPolicy applied to the namespace and run a quick test with a curl pod to verify outbound connectivity.
If you're using the VPC CNI, also verify that the subnets have routes to the internet or VPC endpoints for AWS services. Security groups attached to the worker nodes could be blocking outbound traffic on port 443.
infra nerd, cost hawk
The ClusterIP diagnosis is correct, but incomplete. The SIEM webhook timeout could also be due to the load balancer health check configuration, not just its existence. If the LB health check is hitting the wrong path or port, it'll route traffic to a failing pod, causing intermittent timeouts.
On the egress point, you're right about checking NetworkPolicy, but describing it won't help if there isn't one. EKS doesn't enforce policies by default. The more common culprit is the node security group. It needs outbound rules for 443 to AWS endpoints, and the VPC likely needs a NAT gateway route. Don't forget about potential DNS resolution issues in the new cluster's VPC.
Run a quick test from a pod: `curl -m 5 https://sts.amazonaws.com`. If that passes, your egress is fine and the problem is in the ThreatConnect app's config or IAM role.
Good point about the health check config. We had a similar issue where the LB was healthy but routing to pods that were failing readiness probes due to a long startup time. The LB health check path needs to match an endpoint that actually reflects app readiness, not just a generic `/`.
For the egress test, `sts.amazonaws.com` is a solid choice. If that passes, I'd also check DNS resolution from inside a pod with `nslookup` or `dig`. We once chased our tails on "connection errors" that were actually CoreDNS config issues in the new VPC.
Clean code, happy life