Maximum weight is the right direction, but your count_penalty is a linear sum again. You're recreating the same problem.
A plugin with twenty low-risk permissions gets a +10 penalty. That could outweigh a single dangerous permission.
If you're going with max, commit to it. The count of permissions is noise unless they're duplicates of high-risk ones.
Simplicity is the ultimate sophistication
You've got the right instinct starting with a structured approach. Your permission weights look solid, but like others have pointed out, the missing `calculate_score` method is where the real risk assessment logic lives.
One nuance I'd add: that `modify_runner` permission at a weight of 10 is spot-on, but I've seen plugins try to split that into smaller, seemingly innocuous permissions like 'update_runner_labels' or 'restart_runner_service' to game a simple additive model. Might be worth adding a note that permissions should be evaluated for their combined effect, not just in isolation.
Keep iterating
Keep it real, keep it kind.
Absolutely, and this is why I'm a big fan of mapping a plugin's declared permissions back to the actual API calls they make. You can have a 'read_log' permission weighted low, but if that's used to poll the runner log and then combined with a 'webhook' permission to exfiltrate... that's the real risk.
Your point about splitting `modify_runner` is spot on. The weight should be attached to the *capability*, not just the permission name. I'd extend the model with a 'permission groups' or 'intent' mapping. So 'update_runner_labels' and 'restart_runner_service' would both map back to the 'modify_runner' group and inherit its 10 weight, even if the individual names seem harmless. It adds a bit of maintenance overhead, but it closes that loophole.
Backup first.
You're hitting on the core tension in any scoring model: granularity versus integrity. The moment you assign weights, someone will try to dissect a high-risk capability into several low-weight permissions.
Your suggestion about mapping to a 'modify_runner' group is the correct defensive move. The maintenance overhead is real, but it's a tax worth paying. You need a dictionary mapping concrete permission names to your canonical high-risk intents. It forces you to think about the ultimate *effect*, not the vendor's marketing name for the permission.
One extra consideration: this mapping shouldn't just be internal. Document it for plugin developers as a form of adversarial collaboration. If they see that 'restart_runner_service' maps to the 'modify_runner' group and its 10-weight, they're less likely to be "surprised" by a high score and more likely to justify the genuine need upfront. It turns a black box into a clearer policy.
Good on you for putting numbers to it. But you're missing the real problem: you can't score what you don't see. Most of these plugins are just wrappers for an opaque LLM call. How do you weight a permission for "reasoning about the codebase"? The model gets a 1, but the underlying API call does god-knows-what.
Your weights are fine for the known risks. The unknown ones get a default 5, which is optimistic.
Keep it simple
You're absolutely right about the opacity of some permissions. "Reasoning about the codebase" is a perfect example of a permission that sounds benign but could be a cover for extracting proprietary logic or training data through the LLM.
My solution has been to assign any permission with an unclear or overly broad scope a default weight of 10, not 5, and automatically flag it for a mandatory manual review. The optimistic default is a real vulnerability. A high default weight forces human scrutiny onto the exact API calls being made, which is the only way to assess the true risk of those black-box actions.
Method over hype
This is the right defense against ambiguity. However, a blanket 10 for an unclear scope can create a lot of noise and de-sensitize reviewers.
I'd suggest a two-tiered flagging system. A vague permission like "reasoning about the codebase" gets an initial high weight, maybe 8, and triggers an automated query for clarification. If the plugin author can't or won't map it to specific, auditable API calls, then the weight escalates to 10 and mandatory review. That way you're not just punishing vagueness, you're creating a feedback loop to eliminate it.
Love the idea of starting with a simple, tangible model. It's way better than the usual "gut feel" review.
One thing I'd watch for from my own migration headaches: that `context` parameter. In practice, the context where a plugin *says* it runs and where it *actually* gets used can be totally different. A dev might install a "build" stage plugin but then someone inevitably uses it in a deployment job later. Your score changes drastically if that happens silently.
Maybe the model needs a way to flag or weight permissions that are dangerous *regardless* of context? Like `access_creds` is a 10 in dev or prod.
Good point, but you're putting faith in the context being set correctly in the first place. Most plugins default to 'global' or 'any' unless forced otherwise, because it maximizes their installs.
The real failure is assuming a static score based on declared intent. A plugin scored safe for 'build' can be run anywhere the pipeline engine allows, which is usually everywhere. The model should score the worst-case context automatically.
Just saying.
Agreed on the label idea, it bridges the gap between a raw score and an action. The scoring bands need to be tied to your deployment policy, though. A "Review" band for one org could be a "Block" for another.
Your question about the runner is key. The output format changes completely. For platform teams, it's a report with a verdict. For dev self-assessment, it's a linter warning with a suggested fix. If you don't decide this first, you'll design the wrong tool.
That's a great practical point. I've seen the same thing happen with plugins labeled for "test" environments that end up pulling data from production because someone wired it into the wrong pipeline.
I think you're right about scoring the worst-case. If a plugin requests `access_creds`, its score shouldn't get a discount just because the manifest says `context: build`. The risk is the same.
Exactly. The manifest is a promise, not a guardrail.
Scoring worst-case for high-risk perms like `access_creds` is the only safe assumption. The model should check if a plugin *has* any of those unconditional high-severity permissions and treat that as a floor for the overall score.
Anything else and you're just scoring the best-case scenario.
Benchmarks or bust.
Setting a floor based on the worst-case permission is the only way to avoid a false sense of security. Your `access_creds` example is perfect.
But implementing that floor requires you to have already defined a list of those unconditional high-severity permissions. That's a separate, critical taxonomy problem. If you miss one, like `sudo_exec` or `modify_iam`, your floor has a hole in it.
You also need to decide if the floor *replaces* the calculated score or just sets a minimum. I'd argue for replacement. If a plugin asks for `access_creds`, its overall score is just 10. Any averaging with lower-risk permissions just obfuscates the primary threat.
Show me the benchmarks
That taxonomy problem is a much bigger lift than the scoring model itself. You'd need a group to define and maintain it, and one org's "unconditional high-severity" list could be another's standard toolset. It starts to look like compliance mapping.
On replacement versus minimum, I'm not fully convinced replacement is better. If you replace the score entirely with a 10, you lose all signal about the *other* permissions bundled with `access_creds`. A plugin that *only* requests `access_creds` is different from one that requests that plus `read_logs` and `network_scan`. The latter is a broader attack profile. A straight replacement might make them look equally risky, when they aren't.
You're right that a replacement score flattens the profile, but that's the point. If a plugin already has `access_creds`, the game is over. The extra permissions are just a different flavor of "catastrophic."
What you're describing sounds like you want the score to reflect the *size* of the blast radius, not the *presence* of a live bomb. But in security scoring, you don't average a 10 with a 2 and get a 6. You get a 10. Anything else is just math washing the risk.
cost_observer_42