Skip to content
Notifications
Clear all

My experience deploying to 200 Mac users: The client broke FileVault for three.

18 Posts
18 Users
0 Reactions
22 Views
(@charlotte1)
Estimable Member
Joined: 3 months ago
Posts: 94
Topic starter   [#25192]

Hi everyone. I’ve been quietly reading this forum for a while, trying to learn about all sorts of tools for my own small business, but today I need to share something from a different angle. I recently helped a client—a creative agency with around 200 employees all on Macs—deploy FortiSASE. My own expertise is more in bookkeeping and invoicing software, so this was a bit outside my comfort zone, but I was brought in to help manage the transition from a project perspective.

The deployment itself, the agent install and policy setup, seemed to go smoothly over a phased rollout. The users connected, and everything appeared fine for the first week. Then, out of the blue, three of the designers reported they couldn’t unlock their Macs after a restart. Their FileVault encryption was… broken. The recovery key didn’t work, and the normal password just spun forever before failing. It was honestly terrifying—these people had critical client work on those machines.

After a lot of panic and working with Fortinet support, we traced it back to a specific security policy profile pushed via FortiSASE that, for some reason, interfered with the FileVault decryption process on these particular Macs running a certain OS point-release. It wasn’t all 200, just these three, which made it so strange and hard to catch beforehand. We had to boot into Recovery Mode and use a combination of terminal commands and disabling the FortiSASE service to finally get them back in. We lost a full day of work for each of those users.

I guess I’m writing this partly to vent, but also to ask: has anyone else run into anything like this with SASE agents and Mac FileVault? We’re now hyper-cautious about any policy change, testing on a dozen different Mac configurations first. The experience made me realize how a tool meant to secure everything can sometimes have the opposite effect in a really dramatic way. For those of you managing larger Mac fleets, what’s your rollout testing strategy like now? I feel like I learned a very hard lesson about assuming uniformity, even in what seems like a homogeneous environment.



   
Quote
(@helenb)
Estimable Member
Joined: 3 months ago
Posts: 128
 

That's a startling outcome to manage. Even from an accounting perspective, a lost machine with unrecoverable client data would be a compliance and insurance nightmare. Did Fortinet support clarify if the policy was modifying something like the Secure Token, or was the interference more indirect?



   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

That's exactly the kind of vendor black box I'm always on about. The policy "for some reason" interfered, and you had to wait for their support to maybe figure it out. What was the actual system call or permission change? They'll never tell you.

This is why I push for transparent, scriptable tooling. You'd see the `diskutil` command that borked the token before you sent it to 200 machines.

So what's the fix? A support engineer gives you a magic config toggle and swears it won't happen again? Hope the next update doesn't break something else.


Your vendor is not your friend.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Support didn't clarify the mechanism, they just provided a remediation KB. The three affected machines had a specific firmware version. It's a known but poorly documented conflict that silently invalidates the Secure Token.

From a security perspective, that's unacceptable. The agent shouldn't be able to touch that layer without explicit, logged authorization and a pre-flight check. The "compliance nightmare" you mentioned is exactly why this isn't just a support ticket. It's a fundamental design flaw in their endpoint privilege model.



   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

That "known but poorly documented conflict" is the whole issue. It means they shipped a release without proper firmware version detection in their compatibility matrix. That's a basic QA failure, not a design flaw. A design flaw would be intentionally touching the token. This is negligence in their release process.

You can't have a zero trust model if the agent itself is a trusted component that bricks devices based on an undocumented condition. The pre-flight check you mentioned wouldn't have existed because they didn't know to look for the conflict. Their testing coverage was insufficient, full stop.

This is how you fail a control audit. You can't demonstrate due diligence when your security tool is the cause of the data loss event.


— geo


   
ReplyQuote
(@cloud_cost_owen)
Reputable Member
Joined: 5 months ago
Posts: 181
 

Oh wow, that's a nightmare scenario. It highlights why phased rollouts are so crucial, even when things seem smooth. That moment where the recovery key fails must have been pure panic.

So glad you caught it at just three machines. Could've easily been dozens before anyone restarted. Makes me think about adding a forced reboot step to our own deployment checklists.



   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 5 months ago
Posts: 403
 

Yeah, phased rollouts saved them. Forced reboots in the checklist are smart, but they'd have just found the bricked state faster, not prevented it. Still better than letting it simmer.

The real gap is a post-install validation step. That agent should've been able to report back "Secure Token Status: Compromised" before anyone ever logged out. The deployment was "done" when the agent said hello, not when it proved it didn't break the OS.



   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

You got lucky with the phased rollout. Most clients push for big bang deployments.

These all-in-one endpoint agents are a mess. They promise to manage everything but really just add a new layer of opaque failure points. You can't audit what you can't see. Should've stuck with something simpler, maybe just a VPN config profile.


Keep it simple


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 2 months ago
Posts: 377
 

Totally feel you on clients pushing for big bang. The pressure to call it "done" is huge.

But ditching the all-in-one for just a VPN profile isn't always practical. The promise is threat prevention and data filtering, not just connectivity. The real problem is the black box. I need logs I can query, not a "trust me" from the vendor.

A simpler tool that does less but shows you every step would be my pick too. But sometimes the requirement list from the client's security team forces your hand.


data over opinions


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

The requirement list is the problem. Security teams read a Gartner slide and check boxes without understanding the operational risk.

Your job is to push back. Show them this thread. "Threat prevention" is useless if the tool itself becomes the threat because you can't audit it. A simpler tool that passes an audit beats a complex one that fails it.


-- old school


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

Phased rollouts saved you from a much larger disaster, but they're a tactical mitigation, not a strategy. The core failure here is the lack of a validation gate in your deployment pipeline.

Your pipeline considered the job "successful" when the agent phoned home. It should have required a successful post-install health check that included verifying critical system functions like Secure Token state. I'd add a simple scripted check that runs `fdesetup status` and `diskutil apfs list` post-install, failing the deployment stage if anything looks off.

Without that, you're just crossing your fingers.


shift left or go home


   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 2 months ago
Posts: 227
 

That's a critical data point. The fact it was a specific policy profile, and not the core agent install, changes the failure model. It suggests the problem isn't just in their compatibility matrix but in their policy translation layer.

When a policy is compiled and pushed to a Mac via MDM, it becomes a configuration profile. A flaw in how Fortinet's backend generates that profile for certain firmware versions could absolutely corrupt system state. This moves it from a QA gap to an engineering flaw in their profile payload logic.

You'd need to isolate the exact policy setting. Was it a restriction on removable media, a specific kernel extension setting, or something else? The audit trail should be in the FortiSASE logs for what was pushed to those three machines versus the 197 others. That correlation is your only hope of preventing a repeat.


Data over dogma


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

That distinction between a QA gap and a payload logic flaw is interesting, but it feels like splitting hairs. The outcome for the user is the same: a broken device from an opaque process.

You're right that the FortiSASE logs are the only source of truth, but good luck getting a useful audit trail out of them. Their logging is notoriously geared toward compliance checkbox reports, not actual forensic detail. Even if you could correlate, you're now reverse-engineering a black box to understand its failure modes.

It shifts the blame from "they didn't test" to "their engine generates bad code", but it doesn't make the tool any more trustworthy.



   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

That "seemed to go smoothly" phase is the most dangerous part of any deployment. The real test only happens at the next reboot or login.

Three machines is a lucky break. The real question is whether that specific policy profile is the only trigger, or if you just got unlucky with the timing of those three restarts. Have you checked if any of the other 197 have restarted since the policy push?


YMMV


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Exactly. The forced reboot in your checklist is a smart mitigation. In my experience, that kind of post-deployment check is essential, but it needs a defined scope.

Forcing a reboot will surface a problem like a broken Secure Token immediately. But you have to be ready to act on what you find. It's not just about checking a box; it's about having the rollback procedure documented and tested before you trigger that first forced restart. Otherwise, you're just staging your own crisis.

What would your rollback plan be if that reboot test failed for 20 machines?


—Anita


   
ReplyQuote
Page 1 / 2