Skip to content
Notifications
Clear all

Just built a Grafana dashboard from Mend's API data.

30 Posts
29 Users
0 Reactions
47 Views
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

The napkin math is everything. I've watched teams build gorgeous Grafana dashboards on top of a data lake that costs them $12k a month in BigQuery scans, when the vendor dashboard license for the whole org was $8k. You have to bill your pipeline's infra to the same cost center that would pay for the seats.

But the seat cost argument only works if the vendor dashboard actually gives you what you need. If you're building this because their portal lacks a specific metric or drill-down, then the cost isn't just the seat. It's the opportunity cost of *not* having that data, which is harder to quantify but often the real driver.

> watch that "average time to fix" like a hawk

Absolutely. We track it, but we also track the *maximum* time to fix for each severity band. That one critical sitting there for six months skews the average less than you'd think, but it'll show up as a glaring red bar in a max-time chart. Makes it politically impossible to ignore.



   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 7 months ago
Posts: 427
 

That max time per severity is such a good idea. I'm stealing that for my next iteration.

The cost thing is real, but you nailed it with the "opportunity cost of not having that data." My little cron job on a cheap VM is probably fine for now, but I hadn't thought about billing it back. If this gets adopted by the team, I'll need to figure that out before it scales.

> that one critical sitting there for six months skews the average less than you'd think
Yeah, that's exactly the kind of stat that gets a manager's attention in a review. Averages can be too safe.



   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

"Billing it back to the same cost center" is the quiet part that every side-project-turned-critical-infrastructure forgets until the AWS bill lands. I've seen that movie.

Your max-time chart point is exactly right. The average is a conversation ender. The max is a conversation starter, especially when you can link the oldest critical to the specific blocked feature or deployment freeze. It turns an abstract risk into a tangible business cost.


YMMV


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

I hadn't thought about billing the pipeline costs back, but that's a crucial step. My current job runs on a shared VM, so the cost is invisible. If this dashboard takes off, who's budget gets charged for the extra resources?

The "opportunity cost of not having that data" is what pushed me to build this in the first place. The vendor portal's drill-down is too limited for our team leads.

Your max-time chart idea sounds useful. How do you present it? Is it just a bar chart with the single oldest ticket per severity, or do you show the top five?



   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

The API-driven approach is fine for now, but you've just created an unsanctioned data pipeline. Before this scales, you need to document its security controls.

Where is your API credential stored? Is it rotated? Are you logging the extract job's access attempts? If that credential leaks, someone has read access to your entire vulnerability inventory.

Also, check your data retention. Grafana's default might keep those license compliance snapshots longer than your agreement with Mend allows. You're responsible for that data once it's in your system, regardless of where it came from.


Where is your SOC 2?


   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Cool project, but you're probably measuring the wrong thing. Average time to fix is a management pacifier. It'll drop when your team fixes a bunch of trivial warnings, making everyone think they're safe. Meanwhile, that one critical license violation in your core library stays open for six months because it's a headache.

Track the oldest issue in each severity band instead. That's what actually blocks releases.


your mileage will vary


   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Nice! I'm looking at building something similar for our team's AWS cost data. How are you pulling from the API? Just a Python script on a cron job?

The max time per severity idea from later comments is really smart. I might try that instead of just average fix time.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 3 months ago
Posts: 542
 

That's exactly how I started a couple years back. The API is functional, but you need to be ready for its rate limiting and pagination quirks. Fetching data for a large project portfolio can turn a simple script into a multi-hour job if you don't handle the offset parameters correctly.

The metrics you've picked are a solid start, but average time to fix will mislead you. I track the 90th percentile fix time per severity band instead. It shows you the tail of your problem, not the center. A flood of trivial fixes will make your average look great while that one critical from last quarter still hasn't been touched.

For license risk, don't just look at compliance status. Add a panel for license types grouped by project. You'll often find a single problematic license like AGPL-3.0 scattered across dozens of projects, which is a much clearer call to action than a simple red/green status.



   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

That sounds like a solid start! I've been thinking about doing something similar to get our Jenkins pipeline data into Grafana, but Mend's API is new to me. Are you using their REST endpoints directly with something like `curl` in a script, or do they have a Python library you'd recommend?

Also, on the API quirks, did you run into any weirdness with pagination or rate limits when pulling historical data? I've had that bite me before with other services.


Learning by breaking


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

The REST endpoints, unfortunately. Their Python library exists, but it's a thin wrapper that doesn't help much with the pagination.

The real kicker is their "offset" pagination combined with a fairly low rate limit. You can't just fetch all vulnerabilities in one go for a large project. You need to loop per severity, per project, and respect the `x-ratelimit-remaining` header. A naive script will get throttled and your Grafana panel will be empty half the time.

For a quick start, you can use something like this in a cron job, but add sleep and error handling:

```python
def fetch_with_pagination(url, headers):
all_data = []
offset = 0
limit = 100
while True:
# ... fetch with offset
# check ratelimit header, sleep if needed
# break if data is empty
```

What API are you using for your Jenkins pipeline data? I found the Metrics plugin output a bit clunky to parse.


YMMV


   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

You're right to bring up the cost question, I haven't done a formal calculation. Right now the script runs on an existing team VM, and we use a shared Grafana instance, so the incremental cost feels near-zero. But I see the trap you're describing - if this becomes "critical infrastructure," that VM's resource allocation won't stay invisible, and the shared cost will get scrutinized.

I don't have a break-even point because I'm treating the cost as sunk. But that's probably a mistake. The value for us is in the customized metrics, which I couldn't get from an extra portal seat. So maybe the comparison isn't 'DIY vs. extra seat' but 'DIY vs. no actionable data at all.' Is that a valid justification, or just self-deception?



   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

It was pretty straightforward for the basic metrics like total issues by severity. Their API docs list the endpoints and query parameters clearly.

The tricky part was figuring out the exact JSON path for some nested fields in the response, like finding the specific license name inside the `library` object. I ended up adding a debug step to log a raw response snippet to my script first.

If you're starting out, I'd recommend using Postman or `curl` to manually call the endpoint once and see the structure before you write any parsing logic. Saved me a lot of time.


Infrastructure as code is the only way


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Totally agree on the debug step. That's saved me from so many headaches.

I'd add that if you're working in Python, the `json` module's `indent` parameter when dumping to a file makes those nested structures instantly readable. A quick `print(json.dumps(response_snippet, indent=2))` before you write any parsing logic is my go-to.

Did you find that the API's JSON structure was consistent across different endpoint categories, like vulnerabilities vs. licenses? I've seen some services where the nesting patterns change wildly.



   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Spot on about json.dumps, it's a lifesaver. I use that or `pprint` all the time.

>Did you find that the API's JSON structure was consistent

Unfortunately, no. The vulnerability and license endpoints have totally different nesting. The license data is buried deeper in my experience. Always best to test each endpoint you need separately.


Automate the boring stuff.


   
ReplyQuote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Average fix time. Right.

Ever notice how averages vanish problems? One ancient critical bug among a hundred trivial fixes pulls the number down nicely. You're measuring comfort, not risk.

Track the max age per severity instead. Then you'll see the rot.


-- old school


   
ReplyQuote
Page 2 / 2