Skip to content
Anyone deployed Oas...
 
Notifications
Clear all

Anyone deployed Oasis Security at scale? Honest take

20 Posts
16 Users
0 Reactions
5 Views
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Absolutely, that two-way street with the service catalog ends up being the most valuable, but also the most painful, part of the exercise. We used it to surface stale entries in our catalog that pointed to decommissioned services, which was a nice cleanup win.

But it gets tricky with things like DNS records or VPC endpoints that are absolutely critical, but show zero API traffic. Oasis would flag them as dormant, and our catalog said they were active, but proving that to the system meant building a whole separate "business critical" tagging schema just to keep the peace. It starts to feel like you're documenting your infrastructure twice.


cost first, then scale


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

We tried the lightweight agent route for VPC-internal workloads, using a sidecar pattern on our Kubernetes pods and a custom daemon on EC2 instances. It created as many problems as it solved. The agent itself became an attack surface, its telemetry was noisy, and we now had to manage its lifecycle and scaling. You're essentially building a distributed monitoring system, which is a massive undertaking.

On auto-remediation, we turned it off completely after the CloudFormation incident another poster described. The approval mode still carries the risk of a human rubber-stamping a change they can't fully comprehend due to hidden dependencies. The "toxic combo" findings are the product's core value, but acting on them automatically assumes a perfect, complete model of your infrastructure's state - which never exists.

Our compromise was to treat Oasis purely as a discovery engine. The dry-run output feeds a separate, manual cleanup pipeline owned by each service team. It shifts the burden of proof, which is the only scalable approach.


infrastructure is code


   
ReplyQuote
(@henryg78)
Estimable Member
Joined: 3 months ago
Posts: 165
 

Your point about policy-as-code being the primary value matches our analysis. The generated Terraform modules from dry-run scans have become a key input for our quarterly IAM reviews.

However, we quantified the "noise floor" problem. In our environment, 87% of initial alerts were false positives due to batch or seasonal workloads. The service catalog integration only reduced that by 60%, leaving a significant manual triage burden.

The reboot failure scenario you described is a dependency graph problem. We found Oasis's model doesn't account for indirect or temporal dependencies, like a role used quarterly for financial reporting. This makes even the approval mode risky without a separate dependency audit, which negates the automation benefit.


EXPLAIN ANALYZE


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Your false positive numbers are brutal. We saw similar noise with seasonal data pipelines, but our biggest headache was the hidden API gateways. A staging VPC gateway with zero monthly traffic might only get hit during a biannual pen test. Oasis would mark it as dormant every single time.

We also use the Terraform output for IAM reviews, but found we had to run it through a separate "last used date" check from the cloud provider's own audit logs. Oasis's snapshot model misses those long-tail dependencies entirely, so the generated module might be suggesting you delete something that's technically still required. It's useful for hygiene, but not a safe source of truth on its own.



   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Exactly. The snapshot model is the fundamental flaw for any environment with cyclical or bursty traffic. We found the same with pen test assets and quarterly financial reporting jobs.

The last used date check is mandatory, but even that's incomplete. We had to supplement it with a custom Prometheus metric for "critical but idle" resources. If something is tagged as business-essential, it's scraped for a heartbeat signal (TCP port check, DNS resolution) regardless of API calls. Without that, you're just building a list of things you can't delete anyway.


Metrics don't lie.


   
ReplyQuote
Page 2 / 2