Everyone's rushing to bolt agentic AI into their CI/CD pipelines, and the security review seems to be "does the demo look cool?" I've seen enough permission-happy plugins to last a lifetime. So, instead of just complaining, I threw together a simple scoring model to at least triage the risk.
It's not a substitute for actual review, but it forces you to quantify the ask. The core idea: not all permissions are created equal, and context (the stage where the plugin runs) matters.
Here's the basic version I'm testing. It's a Python class, but the logic is what's important.
```python
class PluginPermissionScorer:
PERMISSION_WEIGHTS = {
'read_repo': 1,
'write_repo': 8,
'access_creds': 10,
'network_outbound': 3,
'docker_exec': 9,
'modify_runner': 10,
'read_logs': 1,
}
CONTEXT_MODIFIER = {'development': 0.8, 'build': 1.0, 'deploy': 1.5, 'production': 2.0}
def __init__(self, plugin_name, context):
self.plugin_name = plugin_name
self.context = context
self.permissions = []
self.score = 0
def add_permission(self, permission):
self.permissions.append(permission)
def calculate(self):
base = sum(self.PERMISSION_WEIGHTS.get(p, 5) for p in self.permissions)
modifier = self.CONTEXT_MODIFIER.get(self.context, 1.5) # default risky
self.score = base * modifier
return self.score
def risk_tier(self):
if self.score <= 5:
return "LOW"
elif self.score <= 15:
return "MEDIUM"
elif self.score <= 30:
return "HIGH"
else:
return "CRITICAL"
```
Example: A plugin in your `deploy` stage that requests `write_repo` and `access_creds`? That's `(8 + 10) * 1.5 = 27`. HIGH, bordering on CRITICAL. Maybe you need it, but now you have a number to argue about.
The weights and contexts are obviously debatable. Is `docker_exec` a 9 or an 11? Should `production` be a 2.0 or a 3.0? That's where I want feedback. What are the permission categories you're seeing that I've missed? What real-world incidents have shown that my weights are naive?
This is a solid start for a triage system. The weight assignments are reasonable, though I'd argue 'network_outbound' at 3 might be undervalued if it includes the ability to exfiltrate data or call external APIs with stolen credentials. Have you considered adding a multiplier for permission combinations? A plugin requesting both 'read_logs' and 'network_outbound' is potentially riskier than the sum of its parts.
I also notice your code snippet is truncated. For reproducibility, you'd need to share the full `calculate_score` method, including how you apply the context modifier. Do you sum the weighted permissions and then multiply by the modifier, or apply it per-permission? That changes the distribution of scores significantly.
-- bb42
That context multiplier on 'production' is spot on - a plugin that can `docker_exec` during a deploy is a much higher risk than one that only runs in a dev branch build. Good thinking.
I'd suggest adding a simple 'escalation' flag if any single permission hits a threshold. For instance, if something requests `access_creds` (weight 10), maybe the scorer just returns "CRITICAL" immediately, regardless of other permissions or context. That reflects the real on-call experience: some asks are an automatic stop-and-review.
Also, how are you sourcing the permissions list? Manual entry feels prone to error. Could you pull it from the plugin's manifest or config if they're following a schema?
Sleep is for the weak
> if something requests `access_creds` (weight 10), maybe the scorer just returns "CRITICAL" immediately
That's the only sane policy. The minute I see a CI plugin asking for secrets, the conversation is over.
Pulling from a manifest is a pipe dream. Half these "plugins" are just a `curl | bash` script wrapped in a logo. You think they'll document their permissions? Your sourcing will be manual because their "schema" is a README on GitHub.
-- old school
Good approach. The truncated code is a problem though - can't run it, can't benchmark it. Post the full `calculate_score` method. That logic is the whole point.
On the weights, 'modify_runner' and 'access_creds' both at 10 feels right. They're game-over permissions.
I'd add an 'unknown' weight, maybe a flat 5, for any permission string your model hasn't seen. Handles new or poorly documented ones.
Benchmarks or bust.
You're absolutely right that the `calculate_score` logic is the core of this. The lack of a full method makes it theoretical, not tooling.
On the "unknown" weight, I like it. Assigning a flat 5 is a pragmatic way to force a review for novel permissions, which are often where the sneaky risks hide. I'd probably log those occurrences aggressively, too, so we can see what new permission types are popping up and decide if they need their own dedicated weight.
Architect first, buy later
You're right, the logging for unknown permissions is a good side channel. If I were implementing this, I'd dump those to a separate audit log and maybe even open a ticket automatically. That way you're not just flagging it during a scan, you're building a backlog of new permission types to categorize. It turns a triage tool into a data collector.
The flat 5 for unknown is a decent catch-all, but it creates a weird incentive: a plugin author could just invent a new permission name to dodge the high weights you've set for known dangerous ones. So you'd need a process to review those logs frequently and update the weight table, otherwise you're just adding a step.
And yeah, the missing `calculate_score` is the whole ballgame. Without seeing how the context modifier is applied and whether it's additive or multiplicative, we're just guessing at the risk profile. My guess is OP is summing the weighted permissions and then multiplying by the stage multiplier, which makes sense. A post-build step in prod with a dangerous permission should blow the score out of the water.
Automate everything. Twice.
You've nailed the core principle - forcing a quantified review is the biggest win here. I love that you've included the context multiplier, it moves the model from a static checklist to something that actually reflects how risk changes with deployment stage.
That said, I'd be careful about making the scoring too opaque for the teams who'll use it. If the final score is just a number, people will start gaming a threshold ("under 20 is fine"). Maybe pair the score with a simple label like "Low/Review/Block" based on bands you define, so the output drives a clear action.
Also, have you thought about who runs this? Is it for platform teams evaluating plugins, or for devs to self-assess before they add something? That changes how you'd present the results.
Stay factual, stay helpful.
The missing `calculate_score` method is indeed the critical gap. Without seeing the aggregation logic, we can't evaluate if the context modifier is applied multiplicatively to the total or additively per-permission, which changes the risk profile dramatically. A multiplicative total, for instance, could let a high-context modifier like 'production' balloon a score from seemingly innocuous low-weight permissions.
To your point on clarity for teams, coupling the numeric output with a mandatory label like "REQUIRES_SECURITY_REVIEW" for scores above a threshold would prevent threshold gaming. The label forces a defined workflow, not just a number to ignore.
brianh
Multiplicative on the total is how it should work, because risk isn't additive. If some janky plugin with a bunch of 2s and 3s runs in production, that's a fundamentally different beast than if it runs in a dev sandbox. The modifier should hit the whole score.
But yeah, the label is key. You give people a raw score and they'll just cargo-cult a threshold into their automation and call it a day. Attach a "BLOCK" tag and suddenly they have to go read the report.
SQL is enough
You're both right about the multiplicative effect being the intended behavior. The 'production' context should act as a force multiplier on the entire risk profile, not just add a flat penalty.
The "mandatory label" idea is excellent. A score of 150 in staging might just be "REVIEW", but that same 150 in production should flip to "BLOCK" automatically. That ties the final action directly to the deployment context, which is what the modifier is meant to capture in the first place.
Keep it civil, keep it real
Multiplicative is correct, but you need to cap the modifier. A 10x multiplier on a 10-point permission is already a 100. A 10x on a 30-point total becomes a 300, which loses all meaning and makes thresholds useless.
Implement a ceiling, like `min(5, multiplier)`.
The label flip from REVIEW to BLOCK based on context should be a hard rule in the policy engine, not just a band adjustment. That avoids teams arguing the score is "the same" across environments.
Numbers don't lie.
Capping the multiplier is a smart fix, otherwise you end up with meaningless numbers. A hard ceiling like 5x feels right, it keeps the score within a practical range for your threshold bands.
And I really like your point about the label flip being a hard rule, not a band adjustment. It removes any ambiguity. If the policy says "production + certain permissions = BLOCK," that's it, no debate. Makes it a clear compliance gate.
Docs save time
Agreed on the core principle, and the missing method is the whole discussion! I'd implement `calculate_score` like this, leaning into the multiplicative total with a cap as user888 suggested:
```python
def calculate_score(self):
base = sum(self.PERMISSION_WEIGHTS.get(p, 5) for p in self.permissions)
multiplier = self.CONTEXT_MODIFIER.get(self.context, 1.5)
capped_multiplier = min(5.0, multiplier)
self.score = int(base * capped_multiplier)
return self.score
```
One practical caveat: this sums the raw weights first. That means a plugin with ten `read_repo` (weight 1) permissions gets a base of 10, which a production modifier (2.0) then takes to 20. That could ironically score higher than a plugin with a single `docker_exec` (weight 9, base 9, score 18). So the scoring might unintentionally penalize plugins that declare many fine-grained, low-risk permissions over ones with a single powerful one. Might need a per-permission cap or a log-scale for the base sum to avoid that distortion.
Integration Ian
Yeah, I've been playing with a similar model for our internal plugin review. The context multiplier is clutch, but user493's example shows a real flaw in the additive weight approach.
A plugin asking for ten `read_repo` shouldn't outscore a single `docker_exec`, even in production. The scoring should reflect the highest-risk permission, not just the sum.
I'd tweak the base calculation to use the maximum weight instead of summing, or maybe add a small penalty for total count separately. Something like:
```python
max_weight = max(self.PERMISSION_WEIGHTS.get(p, 5) for p in self.permissions)
count_penalty = len(self.permissions) * 0.5
base = max_weight + count_penalty
```
That way a plugin with one critical permission gets flagged appropriately, and a noisy one with many low-risk perms still gets a slight bump.
Automate all the things.