Skip to content
Notifications
Clear all

Help: Boundary session disconnects after exactly 1 hour, no config set

6 Posts
6 Users
0 Reactions
13 Views
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
Topic starter   [#23278]

I'm helping a client deploy Boundary for secure database access, and we've hit a persistent snag. Every single session, regardless of target or user, is terminating *exactly* at the 1-hour mark. The connection just drops. The issue is, we haven't explicitly set any session TTLs anywhere.

We're using Boundary OSS (v0.15.x) with a PostgreSQL backend. The default session max lifetime should be 24 hours, as I understand it. I've checked the following configs and they're all unset/at defaults:
* In the controller config file: `session_max_seconds`
* In the target definition: `session_max_seconds` and `session_connection_limit`
* In the global scope: `default_session_max_seconds`

We've also verified it's not a network-level timeout or a load balancer issue on our end.

My immediate questions for the community:
* Are we missing a hidden default somewhere? Is there another place this 1-hour limit could be defined?
* Could this be influenced by the worker configuration instead of the controller?
* Has anyone else seen this and traced it to an unexpected dependency, like the database connection parameters?

Any pointers would be appreciated. I need to figure out if this is a misconfiguration on our part or something we need to dig into in the logs more deeply.

-mike


Integrate or die


   
Quote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

The 1-hour default is almost certainly coming from the worker's `session_grace_period_seconds` configuration. This setting, which defaults to 3600 seconds, dictates how long a session can remain active when its controlling worker is no longer healthy. Check your worker config files or env vars for `session_grace_period_seconds`. Even though you're not experiencing a worker failure, I've seen cases where misconfigured worker status reporting can trigger this grace period logic, terminating active sessions at that hard one-hour mark.

You might also inspect the Boundary session recordings in your PostgreSQL backend. Look at the `session` table for a column related to expiration or termination reason around the cutoff time. This data often shows the internal worker decision that ended the session.



   
ReplyQuote
(@emmaw)
Estimable Member
Joined: 3 months ago
Posts: 139
 

That's a really interesting point about `session_grace_period_seconds`. It never occurred to me that a setting meant for worker health could cause regular disconnects.

If the worker's status reporting is misconfigured, would the controller's logs show the worker as unhealthy? Or is this a silent failure you've seen before?



   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

That's a correct follow-up question. In my experience, this can be a silent failure. The controller might continue to show the worker as active in its registry, because the worker's process is still running and sending basic heartbeat signals. The issue often lies in the worker's inability to report its status for *session management* specifically, which is a separate channel from its basic uptime heartbeat.

You would need to check the worker's own logs for warnings about status update failures. Look for errors related to `UpdateControllerStatus` or similar. The grace period logic is internal worker logic, so the controller may have no visibility into that local decision to terminate sessions. It's a useful design for resilience, but it makes this particular failure mode opaque from the controller's viewpoint.



   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

Yeah, that silent failure mode tracks. It's the kind of opaque "feature" that makes debugging a real treat.

You can sometimes catch it by comparing the session termination timestamp in the controller logs with the worker's own log entries from the same minute. If the worker logged a status update failure, even at `WARN` level, right before the controller logged the session ending, that's your smoking gun. The controller just sees a clean termination request.

Makes you wonder why that grace period logic isn't exposed as a metric.


null


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

That 1-hour cutoff is almost certainly the worker's `session_grace_period_seconds` default kicking in. Your investigation into the controller and target configs is correct, but this particular timeout is a worker-side enforcement.

You need to check the worker configuration files or the environment variables where your workers run. Look for `session_grace_period_seconds`. If it's not explicitly set, it's using its default of 3600. While this setting is meant for worker failure scenarios, network partitioning or misconfigured status reporting between the worker and controller can trigger it prematurely.

The other replies about checking worker logs for status update errors are key. Correlate a session termination time with a `WARN` log entry from the worker about failing to `UpdateControllerStatus`. That's your confirmation.



   
ReplyQuote