Skip to content
Rolled out Wiz to 5...
 
Notifications
Clear all

Rolled out Wiz to 500 users - what broke and what we fixed

1 Posts
1 Users
0 Reactions
20 Views
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
Topic starter   [#14026]

Our platform engineering team recently completed a full-scale deployment of Wiz across our entire cloud estate, encompassing approximately 500 developers and SREs. The primary drivers were unifying our cloud security posture management (CSPM) and cloud workload protection platform (CWPP) under a single agentless CNAPP, and improving our ability to shift-left security findings into developer workflows. The implementation was successful, but not without significant, unanticipated friction points that required substantial architectural and process adjustments.

The most critical breakage occurred not in our production environments, but within our CI/CD pipelines and developer experience. We had underestimated the volume and velocity of findings that would be generated from Wiz's deep integration with our cloud provider accounts and Kubernetes clusters. This led to three primary failure modes:

1. **Alert Fatigue and Notification Storm:** Our existing PagerDuty integration was immediately saturated. Wiz's default alerting rules for "High" and "Critical" severity issues, particularly those related to publicly exposed storage buckets and overly permissive IAM roles, generated over 2,000 alerts in the first 24 hours. This drowned out our existing operational alerts and caused the security team to be functionally ignored.
2. **Pipeline Performance Degradation:** We integrated Wiz's IaC scanning into our Terraform and Helm deployment pipelines. The default scanning configuration added a 90-120 second delay to every pipeline run, which for our microservices architecture meant a significant aggregate slowdown in developer velocity. Developers began seeking ways to bypass the scanning step.
3. **Context Overload for Developers:** A Wiz finding for a misconfigured `SecurityGroup` is richly detailed, linking to the cloud resource, the relevant code repository, and the owning team. However, the initial presentation to developers via Slack bots or Jira tickets was overwhelming. The links to the offending Terraform code were several layers deep (e.g., linking to a module in a private registry), and the remediation steps were generic CSPM advice, not actionable for our specific platform abstractions.

Our remediation strategy focused on triage, integration, and contextualization.

**Phase 1: Triage and Signal-to-Noise Ratio**
We immediately moved to implement a robust exclusion and suppression framework. This was a multi-layered approach:
* **Global Exclusions:** Established via Wiz's `Project` and `Cloud Account` tagging to exclude non-production sandbox accounts from critical severity alerts.
* **Targeted Suppressions:** Used Wiz's API to create automated suppressions for known, accepted risks that were already part of our risk register (e.g., specific S3 buckets requiring public read for legacy compliance).
* **Alert Pipeline Refactoring:** We replaced the direct PagerDuty webhook with an internal mediation service. This service consumes all Wiz findings, applies our business logic for prioritization, and only forwards alerts that meet the following criteria:
```yaml
# Example mediation service rule (pseudo-logic)
if (finding.severity == 'CRITICAL' &&
finding.status != 'RESOLVED' &&
resource.environment == 'production' &&
!finding.isSuppressed() &&
finding.age > '24h') {
forward_to_pagerduty(severity='critical', team=resource.owning_team);
} else {
create_jira_ticket(auto_assigned_to=resource.owning_team);
}
```

**Phase 2: CI/CD Optimization**
We addressed pipeline delays by implementing a two-tier scanning strategy:
* **Rapid Pre-Submit Scan:** A lightweight check in the pull request, using Wiz's CLI with a restricted rule set focused only on critical security errors (e.g., hardcoded secrets, public facing load balancers). This scan completes in under 15 seconds.
* **Full Post-Submit Scan:** A comprehensive scan runs after merge to the main branch, utilizing the full rule set. Its findings are aggregated and reported as a daily digest to the responsible team, not as a blocking gate.

**Phase 3: Contextualizing Findings for Platform Users**
We built internal tooling to bridge the gap between Wiz's generic recommendations and our platform's conventions. For example, a finding about a Kubernetes Pod running as root no longer just linked to the raw YAML. Our tooling uses Wiz's API to fetch the finding, then correlates it with our internal deployment metadata to generate a pull request against the correct Helm chart in our GitOps repository, with the specific `securityContext` change required.

**Key Metrics and Outcomes**
* **Alert Volume:** Reduced from 2,000+ to an average of 5-10 actionable PagerDuty alerts per week.
* **MTTR:** Mean Time to Remediation for critical misconfigurations decreased from 14 days to 48 hours, primarily due to clear, contextualized ticketing.
* **Developer Sentiment:** Tracked via internal surveys; net promoter score for security tooling improved from -45 to +12 after the optimizations were deployed.

The core lesson was that deploying a powerful CNAPP like Wiz is not a simple install-and-forget operation. Its agentless model provides immense visibility, but that visibility is a double-edged sword. The real work shifted from *implementation* to *orchestration*—building the glue code and business logic to filter, prioritize, and contextualize the firehose of data into actionable, role-specific insights for our 500 users. The tool's effectiveness is now directly proportional to the quality of our internal integration layer.


No free lunch in cloud.


   
Quote