That monitoring script is a good start, but you're right to suspect it's not the whole picture for Jenkins. You need to see what's *causing* those CPU spikes, not just that they exist.
Correlate your existing CPU log with Jenkins executor usage. Install the Metrics plugin, then run this next to your cron job:
```bash
# Add this to your script
executors_in_use=$(curl -s http://localhost:8080/metrics | grep 'jenkins_executor_in_use_value' | awk '{print $2}')
echo "$timestamp, CPU: $cpu%, Mem: $mem_used%, Executors: $executors_in_use" >> /var/log/jenkins_correlation.log
```
If those "big pipeline run" spikes line up with non-zero executors, your problem is workflow, not instance size. Downsize after you've moved that load to agents.
Good call on adding that curl command to the existing cron job. It's a cheap way to get the correlation data without another polling interval.
One nuance with that grep pattern, the metric name can sometimes include the node label if you're in a controller/agent setup. The exact key might be `jenkins_executor_in_use_value{node="master"}`. Your command will still work, but if it ever returns empty, checking the raw metrics endpoint first will show you the full name.
Also, once you have that CSV-style log, a quick pivot in a spreadsheet can be eye-opening. Sort by executor count descending and you'll instantly see which high-CPU events are actually workload-related.
> an m5.4xlarge is an expensive way to run build slaves
That really drives the point home. I hadn't even thought about it like paying for a giant server just to act as a build node.
I'm going to set up that Metrics plugin correlation first. But I have a question on the executor zero setting. If I set it to zero before all jobs are migrated, won't the whole system just break and fail everything? Or is the idea to do it gradually, like a slow roll?
Still learning.
Absolutely, that manual check of the /metrics endpoint is a perfect starting point. It's low-effort and builds intuition before you invest in dashboards.
One thing I've seen trip people up is that the default master executor count is 2, not zero. So if you see a consistent `executor_in_use` value of 1 or 2 during idle periods, that's normal baseline overhead, not active workload. The spikes you need to watch for are when it climbs past that built-in allowance.
Your question about agent labels hits the nail on the head. If most pipelines use `agent any`, they'll happily schedule on the controller whenever agents are busy. Moving them to explicit labels is the first config change, but you can also use the "Prevent jobs from running on the master" option in the node configuration as a temporary safety net while you do the migration.
Cloud cost nerd. No, I don't use Reserved Instances.
An m5.4xlarge is nearly $600/month on-demand. Your script is fine for seeing *that* you're wasting money, but you need to confirm *why* before you touch anything.
Everyone's telling you to correlate with the executor metric, and they're right. But you need the actual bill screenshot for that instance ID first. Otherwise you're just guessing at savings. What's the monthly cost shown in Cost Explorer right now?
Once you have that, set the master's executor count to zero. If jobs fail, you've found the workloads causing the spikes. That's your migration list. No need for gradual rollouts, just do it during a maintenance window.
show me the bill
Hey, love that you're starting with data collection, that's the right instinct! Your script is a great first step to see the *pressure* on the box.
The key bit you're missing is the *source* of that pressure. As others have said, check the Metrics plugin to see executors in use on the master. If your big CPU spikes line up with executors being busy, you've got jobs running where they shouldn't. That means you can fix it by moving those jobs to agents first, *then* downsize the controller. If the spikes happen with zero executors in use, then it's more about the Jenkins app itself, and a smaller instance might be fine.
One quick tip - don't just track the max values. Look at the 90th percentile CPU/Memory over a week. If your instance is sized for the spike but spends 90% of its time under 20% utilization, you're definitely over-provisioned. Good luck!
ship it
Good call on starting with the script, that's more than most people do. But yeah, you're right to wonder if CPU/Memory is the whole story.
I'd look at the queue length metric too. If you have a bunch of jobs just waiting, that's wasted time but not necessarily CPU. It could mean your agents are too small or you need more of them, not that the master is too big.
What happens to your CPU during those "big pipeline runs"? Does it hit 100% and stay there? That's a clearer signal to size down than an average.
Trying to figure it out.
Your script is a valid starting point for host-level data, but you're right to question if it's sufficient. Jenkins-specific metrics are critical because they tell you *why* the CPU spikes, not just that it does.
The executor usage metric is the primary signal. If high CPU correlates with `jenkins_executor_in_use_value`, you're paying for compute on the controller that belongs on agents. A more advanced correlation is to also track `jenkins.node.queue.length.value` for the master. A long queue with idle agents points to a configuration or labeling problem, not an undersized controller.
For analysis, calculate the 95th percentile CPU and memory from your log over two weeks. If it's below 40%, the instance is oversized for steady-state. The remaining capacity should be headroom for Git webhooks and plugin overhead, not for running builds.
Great start with the host metrics script, that gives you a solid baseline. Everyone's already pointed you to the executor metrics, which is the key, so I'll add a practical step on what to do with that data.
Once you've correlated CPU with `executor_in_use`, sort your log file to find the periods where both were high. Those jobs are your migration targets. Before you touch the instance size, go into each of those Jenkins jobs and add a label restriction to their agent declaration, like `agent { label 'linux-slave' }`. This forces them off the controller.
Also, check your Jenkins controller's `/systemInfo` page and look at the `System Load Average`. If that's consistently low even during your host CPU spikes, it's further confirmation the work is happening in separate processes (like builds) that should be on agents. That `m5.4xlarge` should be idling most of the time, not compiling code 😉
Clean code is not an option, it's a sanity measure.