Just wrapped up a three-month migration from a massive Elastic stack to Splunk ES. Ran the old stack on k3s, new one's on-prem k8s. Wanted to share some raw, infra-focused lessons.
Biggest surprise? The sheer weight of the Splunk Forwarder compared to Filebeat. Had to rewrite our Helm charts from the ground up. Resource requests were way off at first. Here's a snippet from our tuned daemonset for the heavy forwarder:
```yaml
resources:
requests:
memory: "512Mi"
cpu: "250m"
limits:
memory: "2Gi"
cpu: "1000m"
```
Still uses more than double the memory per node. The parsing pipeline is powerful, but you pay for it in overhead.
The migration forced us to fully embrace GitOps for Splunk configs. All our props.conf, transforms.conf, and even correlation searches live in a repo now, ArgoCD syncs them. It's a game-changer for consistency, but the Splunk config syntax feels... archaic compared to Elastic's YAML/JSON.
Anyone else run Splunk ES on Kubernetes? Especially interested in how you're handling the search head clustering and indexer discovery. My Istio mesh seems to complicate service discovery for the indexer cluster.
yaml all the things
SRE at a mid-size fintech, ~500 nodes on prem k8s, ingesting 3-5 TB/day into Splunk ES after migrating from ELK 2 years ago.
**Resource overhead** - Heavy forwarder vs Filebeat is the real hit. In our cluster, Universal Forwarder (UF) sits at ~150-200 MiB per node with `batchSizeBytes = 128k`. Heavy forwarder (HF) idles at 600-800 MiB even with minimal routing rules. We run UF on all nodes, route raw data to a dedicated HF tier (4 pods, 2 Gi each) for parsing. That cut per-node memory by ~70%.
**Config management via GitOps** - Agree the syntax is archaic. We use k8s ConfigMaps imported into Splunk via an init container that runs `splunk btool` validation before ArgoCD applies. One gotcha: Splunk's `[default]` stanza in `props.conf` can silently override app-specific configs if order isn't pinned. We now explicitly set `PRIORITY` in each app's `app.conf` to avoid that.
**Search head clustering on k8s** - We run a 3-node SHC pod set. Key trick: use StatefulSet with stable network identities (pod-name-0, etc.) and set `SHC_use_tcp_bind_address = true` in server.conf to avoid NAT issues inside Istio. Also disable automatic captain elections by pinning `shc_bootstrap_restart = false` during rolling updates - otherwise a pod restart can trigger an unnecessary captain re-election that kills query throughput for 2-3 minutes.
**Indexer discovery with Istio** - This was painful. Splunk's default peer discovery uses multicast DNS, which won't work inside a service mesh. We switched to a static list of indexer pod DNS names (headless service with `publishNotReadyAddresses: true`) and set `[indexerdiscovery]` to `pass4SymmKey` authenticated mode. For health checks, we run a sidecar that polls `services/server/status` via curl every 10s and updates a custom endpoint that Istio's `DestinationRule` watches. Over-engineered, but stable after tuning the probe intervals.
Pick: stick with Splunk ES but replace your heavy forwarder tier with Universal Forwarder + a small parsing pipeline (fluentd on k8s or a custom UF with `props.conf` only). That alone will reclaim ~50% of your per-node memory. If Istio keeps causing indexer discovery flapping, spin up a separate headless service for indexers and disable mesh sidecar injection on those pods - your security boundary can handle that. Tell us your daily ingest volume and whether you need real-time search vs batch analytics; that would change where to allocate compute.
Wow, those resource numbers are wild. I'm just starting to look at Splunk for our team and the memory jump from Filebeat is making me nervous. Did you end up having to resize your nodes or scale down the number of forwarders to compensate? The GitOps approach sounds nice for consistency, but that archaic config syntax would drive me up a wall. How are you handling search head clustering with Istio? We're on a mesh too and service discovery for indexers sounds like a nightmare waiting to happen.