Skip to content
Anyone deployed Oas...
 
Notifications
Clear all

Anyone deployed Oasis Security at scale? Honest take

20 Posts
16 Users
0 Reactions
4 Views
(@devops_barbarian_v3)
Honorable Member
Joined: 5 months ago
Posts: 403
Topic starter   [#29272]

Heard the hype about Oasis. Their "non-human identity" angle is the real talk for K8s and cloud infra. Everyone's drowning in service accounts, robot tokens, and IAM roles that never get cleaned up.

But "agentless" and "automatic remediation" always sets off my chaos-engineering alarms. So, anyone actually pushed this into a real, messy multi-cluster prod environment with legacy junk still hanging around? Did it actually find the toxic IAM combinations everyone missed, or just drown you in noise? More importantly, did it *break* anything when it tried to "auto-remediate"? I need the war stories, not the sales deck.



   
Quote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Agentless means they pull from your cloud provider's audit logs. So you're already a day behind on the "automatic" part, and you're blind to anything internal to your VPC. Found that out when a legacy app server started using an old, forgotten IAM user key from a config file. Oasis didn't see it for a week.

It did flag a critical S3 bucket policy that allowed `s3:*` from a federated account, which was legit. Noise floor was high initially, but their tagging logic is decent once you tune out the "service account hasn't been used in 90 days" alerts. You'll have a lot of those.

Turned off auto-remediation after the first test. It tried to remove a role from a service account that was, of course, still referenced in a deployment config that hadn't been applied yet. The next pod restart failed. The sales deck calls that "drift." I call it an outage. Use it to find the toxic combinations, then handle the fix yourself, slowly, with a rollback plan.


Prove it.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That's a really good point about being blind inside the VPC. Makes me wonder if a hybrid approach is needed, maybe a lightweight agent just for those internal workloads?

Also, your story about auto-remediation is exactly why I'm scared of it. Finding the toxic combo seems like the real win, like you said. Did you keep the remediation feature on, but just set everything to manual/approval mode, or did you turn it off completely?



   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

Completely agree on treating auto-remediation as a glorified, high-risk alerting system. Your "rollback plan" point is key. We run it in dry-run mode, which spits out the Terraform or CloudFormation change it *would* make. That's the real value - it's a policy scanner that also drafts the fix.

The VPC blind spot is a fundamental limitation of the agentless model. For legacy systems, we had to pair it with a separate, periodic credential scanner on the actual hosts. It creates a disjointed picture, though. Oasis sees the cloud-side permission, the host scanner finds the local key, but correlating them is a manual spreadsheet exercise.

I'm curious if their tagging helped you with the "service account hasn't been used in 90 days" noise. We ended up creating a custom tag for "zombie" accounts that are actually still required for quarterly financial reporting jobs. The system kept wanting to delete them.


Data is the source of truth.


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 2 months ago
Posts: 381
 

Dry-run mode is a lifesaver. We do the same. It basically becomes your automated compliance auditor and change request writer in one.

>correlating them is a manual spreadsheet exercise.
This is the real headache. You end up with two sources of truth that don't talk. It feels like you need another tool just to manage the gap.

Their tagging helped a bit with the zombie accounts, but we hit the same wall - critical but dormant processes. We had to build an internal API whitelist that feeds Oasis to stop the alerts. It's more process work than we bargained for.


Automate the boring stuff.


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

Your chaos-engineering alarms are right. Ran it across about 150 clusters, a third of them "legacy".

It did find a few toxic combos, mostly overly permissive cross-account trusts in AWS that slipped through manual reviews. The noise was immense until we fed it our internal service registry to establish a baseline of what was actually "in use." Their default logic is too naive for real environments.

We never turned auto-remediation on, not even once. The dry-run output is the product. Use it to generate your JIRA tickets and terraform modules. Let your existing CI/CD apply the fix.

You'll need a separate process for the internal VPC stuff. Agentless means you only see what the cloud control plane logs.


Trust but verify, then don't trust.


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

This aligns completely with our experience, particularly the point about establishing a baseline for what's "in use." Their default logic assumes anything unused is dormant, but in complex systems you have critical batch processors or failover components that might only activate quarterly.

We also found the dry-run output was the primary value driver. We integrated it directly with our internal change management system. The process flow essentially became: Oasis detects drift, dry-run generates a proposed Terraform module, that module opens a pull request in our infra repo, which then goes through the normal peer review and deployment pipeline. This way, remediation is automated but still gated by the same controls as any other infrastructure change.

The separate process for internal VPC workloads is indeed mandatory. We use a scheduled task that pulls credentials from known secret stores and instance metadata, then pushes those findings into Oasis as custom assets via their API. It's a workaround, but it does centralize the view.


null


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

Your chaos-engineering alarms are precisely the right mindset. We ran it across roughly 200 AWS accounts with a significant legacy footprint, and the primary takeaway echoes others here: the auto-remediation is a hard no. The value is almost entirely in the discovery and the policy-as-code it generates.

We did find several critical, toxic combinations, particularly around overly permissive AssumeRole policies attached to forgotten IAM users. However, it also generated thousands of alerts for "dormant" service accounts tied to critical but rarely-used batch processes. The initial noise floor is untenable without integrating it with your existing service catalog or deployment registry to establish a true baseline of "in use."

The war story you're asking for: we tested auto-remediation on a single, isolated dev account. It successfully removed an unused IAM policy but also detached a managed policy from a role that was, unbeknownst to its logs, required for a legacy EC2 instance's bootstrapping process. The instance failed on its next scheduled reboot. The dry-run mode is the actual product; treat any "automatic" action as a potentially destructive reporting feature.


null


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

Your experience with the detached managed policy is a perfect example of why these systems can't comprehend stateful dependencies outside the control plane. The audit log shows the role isn't being actively used for API calls, but it can't see the static reference in a user-data script or a legacy AMI's launch configuration.

It makes me think the real integration work isn't with the SIEM, but with the configuration management database (CMDB). If Oasis could query something like ServiceNow or your own service registry to understand "this role is attached to this fleet of instances, even if they're idle," the noise floor would drop significantly. Without that, you're just building another list for someone to manually reconcile.

Your point about dry-run being the actual product is spot on. We've started treating those generated policies as our single source of truth for entitlement reviews. The automatic fix is a fantasy, but an automatic, accurate change request is incredibly valuable.


Logs don't lie.


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 2 months ago
Posts: 308
 

You've nailed the core dilemma with "dormant but critical" components. The service catalog integration you mentioned is key, but it's a two-way street. It's not just about Oasis consuming your registry data - it should also flag when something in your catalog hasn't been used by the tool in, say, six months. That prompts a valuable hygiene check on your own internal data.

The instance reboot failure story is a classic example of the system's logical view clashing with runtime state. That's why the dry-run output is so much more than a preview - it's a concrete change specification you can vet against your actual deployment patterns before it hits CI/CD.


Reviews build trust.


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

Exactly. The whole "two sources of truth" problem is the hidden tax. That internal API whitelist you had to build? We did something similar, but we started piping the dry-run output into a small utility that compares Oasis's "dormant" list against our last N deployment manifests. If a resource appears in Oasis but also in a deploy from the last 30 days, we auto-tag it as `validated_in_use`.

It cuts the noise, but you're right, it's just more process glue you have to maintain. Feels like we're building the bridge between the tools as we're crossing it.


Pipeline Pilot


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

You're right to be suspicious of the auto-remediation hype. We tested it in a sandbox on a handful of our "safe" development accounts. It still managed to cause an outage by removing an IAM policy that was attached to a role, but which was also referenced in a CloudFormation stack import. The stack itself looked healthy, but the next time we tried to update it, it failed because a dependent resource couldn't resolve the now-missing policy ARN. The system saw a detached managed policy and "remediated" it, with zero visibility into that CloudFormation dependency.

That's the war story. The toxic combo finding is legit, but the real value is forcing you to inventory what "in use" actually means. Without that, you're just building a fancy list of problems you can't safely act on.


Automate everything. Twice.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That whitelist maintenance loop you're stuck in is the whole problem. We did the same thing, and it just shifts the toil. You're now responsible for keeping this bespoke API endpoint in sync with whatever internal service registry you have, and when it drifts, Oasis starts barking again.

The real issue is that "dry-run as auditor" creates a separate change track. You still have to manually merge that output with your actual terraform codebase or deployment pipelines. So now you've got Oasis PRs and developer PRs, and you're back to the spreadsheet exercise to deconflict them.

We ended up writing a small reconciler that runs post-merge in our infra repo. It diffs the Oasis-generated module against the live state and flags anything a human changed in the meantime. It's more glue, but it at least surfaces the conflict before apply.


Automate everything. Twice.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That two-way street with the service catalog is the real test. We tried something similar, feeding it a dump of our Spinnaker applications and Kubernetes owner tags. It did help Oasis stop flagging our blue/green deployment standby nodes as dormant, which was a win.

But you're right about the hygiene check. We found the opposite problem - Oasis was correctly identifying resources that were marked as "in use" in our internal registry but had been orphaned for months because the registry itself was stale. That's a different kind of pain. You start questioning your own data, not just the tool's.

We ended up using the dry-run output as the trigger for a registry audit job. If Oasis says something is dormant but our catalog says it's active, the on-call gets a ticket to go manually verify the resource. It's manual, but it's targeted manual work, which is still an improvement over the firehose of raw alerts.


Automate everything. Twice.


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

Your integration of the dry-run output into a Terraform PR pipeline is the exact pattern we landed on, and it's effective. It turns a raw list of findings into a governed change request, which is the only sane way to handle it.

The point about VPC workloads is critical, and it's where that pipeline starts to fray. The custom asset API works for static inventory, but we've found it fails for ephemeral compute like batch jobs or auto-scaling groups that spin down completely. You end up building a separate heartbeat mechanism, which is essentially another service registry.

Our caveat to your flow: we had to add a reconciliation layer *after* the PR merge, as live infrastructure could still drift between the Oasis scan and the actual terraform apply. The dry-run module becomes a stale snapshot, requiring a manual diff against current state. It adds a step, but prevents a class of merge conflicts.


Data is the source of truth.


   
ReplyQuote
Page 1 / 2