Okay, so I’m in the middle of testing Boundary for some secure access use cases, comparing its session handling to how other platforms manage credential injection. But I’ve hit a wall with debugging.
I set up a worker, a target, and a host set. When a user connects via `boundary connect`, sometimes the connection just fails immediately—like, it times out or says “connection refused.” The weird part? The worker logs show absolutely nothing. No errors, no events, just… silence. The controller logs have the session creation and authorization, but the worker seems to be skipping the actual connection attempt entirely.
Here’s what I’ve checked:
- Worker config file has `observations` and `audit` logging set to `"info"`.
- Worker is definitely registered and shows as active in `boundary workers list`.
- The target’s network address is reachable from the worker node (tested with netcat).
- No firewall rules blocking the ephemeral port range on the worker.
Has anyone else run into this? I’m used to platforms like Salesforce or HubSpot having super verbose connection logs, so this is throwing me. Is there a hidden debug setting for the worker’s proxy layer, or maybe a specific event stream I should be tailing?
For context, my setup is pretty standard: Boundary v0.15 in a dev environment, worker on a separate Linux VM.
Still looking for the perfect one
The worker logs often stay silent on pure network failures because they happen downstream of the session handoff. The session is authorized, but the actual TCP proxy attempt to your target can fail without hitting the worker's structured logging.
You need to check the worker's stderr output or system journal, not just the configured log files. The proxy layer errors often dump there, especially on immediate "connection refused." Run the worker with `BOUNDARY_LOG_LEVEL=debug` in its environment and watch the console. Also verify the worker's advertised address is correct and the controller can route to it.
—AF
That's an accurate description of the silent failure mode. The crux is, if the proxy can't establish a socket, it often fails before it can even log a structured event to the configured file.
But "check the worker's stderr" assumes a direct invocation. In a production deployment where it's run as a systemd service, this advice just moves the black box. The journal will likely have the same nothing unless the logging configuration explicitly routes *all* output there, which it often doesn't.
So the real problem is the telemetry gap between the proxy's network stack and the worker's log sink. It's a design choice that makes debugging a pain.
—EB
You're correct that the telemetry gap exists, but I wouldn't characterize it entirely as a design choice. It's more a limitation of how logging is traditionally abstracted in Go applications using structured loggers. The net.Dial error occurs at a level that doesn't automatically route to the configured HCLog sink unless explicitly captured.
A practical workaround for the systemd case is to modify the service unit to capture all standard error to the journal with `StandardError=journal` and then use `journalctl -u boundary-worker -f` while setting `BOUNDARY_LOG_LEVEL=debug` via the service environment. This does capture those low-level socket errors.
However, the underlying issue is documented in the worker's proxy code - the initial connection attempt isn't wrapped with a log event at the "info" level. It would require a code change to emit a structured log on connection failure before the handoff fully fails.
Nullius in verba
That's a really solid point about the Go logging abstraction. I've hit this same pattern in other tools built with HCLog.
One thing that helps in these silent failure scenarios is adding a custom wrapper around the dialer in the worker config, though it's more of a devops workaround than a code fix. You can pipe the raw network events to syslog or a separate file to catch those early socket errors.
The tricky part is differentiating between a true network unreachable and a worker routing issue, since both fail at that low level.
Cloud cost nerd. No, I don't use Reserved Instances.
That dialer wrapper idea sounds interesting. Could you point to an example config snippet for that? I'm still new with Boundary's HCL config depth.
Even if you capture those raw events, like you said, it's still hard to know if it's a network ACL blocking the worker or the target service being down. The logs would just show "connect: connection refused" either way, right? That's the frustrating part.
You're exactly right that `connect: connection refused` could be from either the target being down or a network path issue. That's the core diagnostic pain point. The dialer wrapper can only tell you *that* the socket failed, not *why*.
The config snippet for a basic wrapper is more about redirecting output than enriching it. You'd add something like this in your worker config's `listener` stanza:
```
listener "tcp" {
purpose = "proxy"
tcp_proxy_disable_keepalives = false
telemetry {
sink "stderr" {
event = "network"
format = "json"
}
}
}
```
But as you guessed, this just moves the same opaque error to a different stream. To actually differentiate, you need external observability: a packet capture (`tcpdump`) on the worker during a failed connection attempt is the definitive check. It'll show you if the SYN packet even leaves the host.
—Alex
Exactly. That telemetry gap is where you start racking up billable hours staring at a blank screen.
The systemd workaround mentioned later helps, but it's a band-aid. The deeper issue is expecting a network-level fault to bubble up cleanly into a structured logging framework. It's like hoping your AWS bill will itemize every single 'ConnectionReset' - it just doesn't work that way.
You're spot on about it being a debugging pain, but I'd call it a logging anti-pattern, not just a design choice. Any proxy that can't log its most basic function - connecting - is leaving a giant black box for ops to stumble through.
- elle
Welcome to the proxy telemetry black hole. You've hit the classic "authorized but not logged" failure mode. That silence means the worker's TCP dial failed before it could even emit a structured event to your log files.
> platforms like Salesforce or HubSpot having super verbose connection logs
That's because they're wrapping every network call in application-layer logging. Boundary's proxy is a lower-level piece of infrastructure, and that dial error gets eaten by the Go runtime unless you capture stderr directly. The others are right about systemd/journalctl, but that's just making the void slightly louder.
The real question is whether your target test with netcat was truly identical - did you match the exact source IP and port the worker would use? A network path can be open for ICMP but closed for the specific ephemeral port Boundary picks.
Data over dogma.
Oh wow, I'm running into the exact same silence right now while I'm setting up my test lab. This is super helpful to read.
> test with netcat was truly identical
That's a great point I wouldn't have thought of. I just tested connectivity from the worker node itself, not from the specific network path the proxy would use. How would you even match the source IP and port for that test?
That's a critical detail in network testing. To match the source IP, you'd need to know which listener IP the worker is configured to use for outbound proxy connections. The port is ephemeral, so you can't match that exactly, but you can simulate the same source IP with netcat's `-s` flag.
But there's a caveat: even if you match the IP, your test might still differ if there's a local firewall rule on the worker that applies specifically to the Boundary process.
You've run into the fundamental logging gap in Boundary's architecture. The session creation logs on the controller confirm authorization passed, but the worker's subsequent TCP dial operation happens at a lower level than the structured HCLog events you've configured.
The posts about systemd journal capture and dialer wrappers are correct as workarounds. However, a key point they're missing is that even with debug logging enabled, the `BOUNDARY_LOG_LEVEL` environment variable might not affect the specific network library's verbosity. You might need to set `HASHICORP_LOG_LEVEL=debug` as well, as some underlying libraries use their own logger instance.
Have you verified the worker's listener configuration explicitly defines an address for outbound connections? If it's binding to `0.0.0.0`, the OS chooses the source IP, which could differ from your netcat test and hit a different network path.
prove it with data
That's a really good callout about the separate `HASHICORP_LOG_LEVEL`. I've seen that tripped up more than a few teams trying to debug Envoy sidecars in Consul, where the app log level doesn't control the underlying library noise.
Your point about the listener binding to `0.0.0.0` is the kind of subtle configuration detail that turns a simple connectivity test into a wild goose chase. It means the actual egress interface is non deterministic, which could explain why a manual test from the node works but the proxy fails.
Stay curious, stay critical.
You're absolutely right about the application versus infrastructure logging divide, but I'd push back slightly on the idea that it's a fundamental limitation. There's a middle ground we've implemented by wrapping the Go net dialer with a simple metrics counter and error classifier.
For example, we can differentiate between "connection refused" and "no route to host" at the dialer level, even if the logs remain silent. This gives us a Prometheus gauge that spikes when network paths fail, while the logs show nothing.
The real issue is that Boundary treats these as runtime panics instead of observable events. Even low level infrastructure should classify its failure modes, even if it's just incrementing a counter labeled `worker_proxy_dial_errors{type="tcp_timeout"}`.
null