Skip to content
Notifications
Clear all

Complete newbie here - how do I know if it's installed right?

6 Posts
6 Users
0 Reactions
35 Views
(@joshuaa)
Trusted Member
Joined: 3 months ago
Posts: 45
Topic starter   [#10628]

Hey everyone. I see this question pop up a lot from folks just starting their service mesh journey, and it's a great one. That initial "is this thing even working?" feeling is totally normal. Since Fathom operates at the infrastructure layer, it doesn't come with a flashy UI that screams "I'm on!", so we have to rely on some good, old-fashioned diagnostics.

Let's walk through a few concrete checks. I'll assume you've followed the basic installation guide for your platform (Kubernetes, VMs, etc.). The goal here is to validate the data plane (the proxies alongside your services) and the control plane (the management components).

**First, check your pods or workloads.** If you're on Kubernetes, the most straightforward sign is seeing the Fathom proxy containers running alongside your application containers. A simple kubectl command can show this. Look for containers with names related to the proxy (often `fathom-proxy` or similar).

```bash
kubectl get pods -n -o jsonpath='{.items[*].spec.containers[*].name}'
```
You should see your app container *and* the Fathom proxy container listed for each pod where injection was enabled.

**Second, verify the control plane is healthy.** The Fathom control plane components (like the discovery service, configuration server, etc.) should be running without constant crash loops. Check their logs for any obvious errors.

```bash
kubectl get pods -n fathom-system
kubectl logs -n fathom-system deployment/fathom-discovery --tail=50
```
Healthy logs should show startup messages and periodic sync activities, not a barrage of connection errors or panics.

**Third, generate some traffic and look for the mesh's fingerprints.** This is the most telling test. The core function of the mesh is to add observability and control. So:
* Do your application logs still work? (They should.)
* Can you see new, mesh-generated metrics (like in Prometheus) for the traffic between services? Look for metrics with labels indicating source and destination workloads.
* Are distributed traces being generated for requests that flow through multiple services? Check your tracing backend (Jaeger, Tempo, etc.).

If those three pillars are in place—proxies running, control plane stable, and telemetry flowing—you can be very confident the installation is correct. The next step is exploring specific features like traffic shifting or fault injection.

It's a complex architecture, so take it one step at a time. Which installation method are you using, and what are you seeing so far?

—Josh


Design for failure.


   
Quote
(@jimmyb)
Trusted Member
Joined: 3 months ago
Posts: 37
 

Thanks, this helps! But I'm still a bit lost on the control plane part. What does "healthy" actually mean? Like, are there any specific logs or status pages to look for?


Learning the ropes


   
ReplyQuote
(@latency_king_2)
Estimable Member
Joined: 5 months ago
Posts: 78
 

You're right to focus on the control plane health definition, as it's often poorly documented. A healthy control plane means it's actively managing the data plane without internal errors.

For logs, you need to check the Fathom controller component specifically. Look for periodic heartbeat logs or status sync messages. The absence of repeated connection errors to the data plane proxies is a positive signal, but a truly healthy state shows regular configuration broadcasts. Some deployments include a basic /health endpoint on the controller's admin port (often 9901 by default, but check your config). A simple curl can verify liveness, though that endpoint typically doesn't validate all internal subsystems.

I'd also recommend setting up a minimal canary service with Fathom injection. If the control plane is functioning, that service's proxy should receive its configuration from the controller within seconds of startup. You can verify this by checking the proxy's logs for the initial config fetch.



   
ReplyQuote
(@juliap)
Estimable Member
Joined: 3 months ago
Posts: 100
 

Oh, the classic "look for the proxy container" advice. That's like saying you've installed a security system because you see a box on the wall. It tells you it's there, not that it's wired correctly or even powered on.

Seeing the container is step zero. I've seen deployments where the proxy container is just sitting there idly, not actually intercepting any traffic, because the init-container or sidecar injection config was borked. The real question is, are your application's connections actually getting routed through it? A quick check is to look at the proxy logs for *your app's* traffic, not just the container's existence. No inbound requests logged? Then it's just a fancy, resource-consuming placeholder.


Your free trial ends today.


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Exactly. Seeing the container is necessary but not sufficient. It's the difference between "it's installed" and "it's operational".

I'd add a step zero before your checks: verify the injection actually happened. On Kubernetes, look for the `sidecar.fathom.io/injected: "true"` annotation on the pod, not just the container list. That annotation is the clearest sign the admission controller did its job. Without it, even a running proxy container might be ignored.

Your point about checking for *app traffic* in the proxy logs is key. That's the only real proof it's working. No logs means it's just a decorative sidecar.


Sleep is for the weak


   
ReplyQuote
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
 

Spot on about the annotation. That's the smoking gun for injection.

But I've been burned by that "decorative sidecar" scenario too many times. Even with the annotation and traffic in the logs, you can still have a broken config where the proxy *sees* the traffic but doesn't actually apply any routing rules or policies. It's just a very expensive TCP logger.

The final check for me is always to apply a dead-simple, observable policy. Like a header injection or a canary route split. If that works, the whole pipeline is alive. If not, you're just admiring a fancy paperweight.



   
ReplyQuote