Skip to content
Notifications
Clear all

Anyone using Wiz to scan Hugging Face models in CI/CD pipelines?

2 Posts
2 Users
0 Reactions
14 Views
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
Topic starter   [#25615]

I've been evaluating the integration of Wiz into our MLOps pipeline for several weeks, specifically targeting the security scanning of externally sourced Hugging Face models before they are promoted to our internal registry. The primary goal is to intercept potentially malicious or vulnerable models at the CI/CD gate, rather than during runtime or in a production staging environment. While the concept is sound, the implementation presents several nuanced trade-offs between security coverage, pipeline latency, and operational complexity.

Our current workflow, prior to Wiz integration, looked roughly like this:

1. A data scientist submits a PR that introduces a new dependency on a Hugging Face model (e.g., `model = AutoModelForCausalLM.from_pretrained("author/model-name")`).
2. The CI pipeline downloads the model artifacts, runs unit tests, and performs a lightweight vulnerability scan using a static analysis tool.
3. Upon success, the model artifacts are bundled and pushed to a private artifact repository.

The integration point we are testing with Wiz is between steps 2 and 3. We've configured a pipeline job that, upon downloading the model files (safetensors, config.json, tokenizer files), uses the Wiz CLI to initiate a scan. The core configuration challenge revolves around the scan policy and the handling of the extensive dependency tree often found in PyTorch/TensorFlow-based models.

```yaml
# Simplified GitLab CI job example
scan_hf_model:
stage: security
image: python:3.10-slim
script:
- pip install wizcli
- wiz auth login --client-id "$WIZ_CLIENT_ID" --client-secret "$WIZ_CLIENT_SECRET"
- |
wiz scan create
--project "$WIZ_PROJECT_ID"
--path ./downloaded_model_dir
--scan-policy "$WIZ_HF_MODEL_POLICY_ID"
--json > scan_report.json
- python ./scripts/evaluate_wiz_results.py scan_report.json
```

The key questions and observations from our testing thus far are:

* **Scanning Depth vs. Time:** A full recursive scan of a large model (e.g., 7B parameters) can take 8-12 minutes. This is a significant addition to our CI feedback loop. We are experimenting with policy exclusions for certain file types (e.g., `.bin` data files) to reduce time, but this obviously introduces a coverage gap.
* **Policy Configuration:** Defining what constitutes a "critical" issue for a model is non-trivial. Issues like "Model Artifact Contains Embedded Script" are high-priority, but findings related to "Dependency with Known Vulnerability" often cascade into the entire underlying framework (PyTorch), which is a separate, managed infrastructure component for us.
* **False Positives and Baseline Management:** The initial scans flagged numerous "secrets" (e.g., Hugging Face API tokens in configuration templates) that were false positives. Establishing a clean baseline for our model repository and tuning the policy to ignore certain paths has been an ongoing effort.

I am particularly interested in how others are navigating the latency versus comprehensiveness trade-off. Are you performing a deep scan on every PR, or are you using a two-tiered approach (lightweight scan on PR, deep scan on merge to main)? Furthermore, how are you handling the orchestration of scan results—failing the build on any critical finding, or using a more nuanced risk-acceptance workflow?

From a systems design perspective, the ideal would be to have the scan be asynchronous and non-blocking for the developer, but the need to prevent the merge of a hazardous model requires a gate. This inherent tension between speed and safety is the central problem we're trying to optimize.


brianh


   
Quote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

Good analysis of the trade-offs. The latency hit is real, especially with large models. I'd add that you need to factor in the cost of the Wiz compute for those scans over hundreds of pipelines. It's not just time, it's actual spend.

Also, scanning just the artifacts you download might miss poisoned training data or malicious code in the model card. If your threat model includes supply chain attacks, you'll need a broader scope.



   
ReplyQuote