Skip to content
Notifications
Clear all

Help: Our Jenkins master node is too big, how to right-size?

39 Posts
37 Users
0 Reactions
174 Views
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
Topic starter   [#21731]

Hi everyone! I'm still pretty new to the whole DevOps cost optimization side of things, and I could use some guidance.

Our team's main Jenkins controller is running on an AWS `m5.4xlarge` instance (16 vCPUs, 64 GiB RAM). It feels massively over-provisioned most of the time, but during big pipeline runs, it does spike. We're trying to control cloud costs, and this seems like a good place to start. How do you figure out the right size?

What metrics should I be looking at to make a data-driven decision? I've started collecting some basic stats from the instance itself, but I'm not sure what's most important for Jenkins specifically.

Here's a snippet of the simple monitoring I set up to track CPU/Memory over the last week:

```bash
# This runs via a cron job
timestamp=$(date '+%Y-%m-%d %H:%M:%S')
cpu=$(top -bn1 | grep "Cpu(s)" | awk '{print $2}' | cut -d'%' -f1)
mem_used=$(free -m | awk 'NR==2{printf "%.2f", $3*100/$2 }')
echo "$timestamp, CPU: $cpu%, Mem: $mem_used%" >> /var/log/instance_metrics.log
```

Is this the right approach? Should I be looking at Jenkins-specific metrics like executor usage or queue length instead? Any beginner-friendly tips on how to analyze this and downsize safely would be so appreciated! 🙏



   
Quote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

That's a sensible start for system-level metrics, but you're right to suspect Jenkins itself provides more actionable signals. CPU and memory spikes are often symptoms, not causes. You need to correlate them with Jenkins' internal load.

Focus on these Jenkins metrics, which you can get via the `/metrics` endpoint or the Monitoring plugin:

- `jenkins.executor.count.value` and `jenkins.executor.in.use.value`: This shows your active concurrency. If `in.use` is consistently far below `count`, you have over-provisioned executors on the controller.
- `jenkins.queue.length.value`: A consistently high queue length indicates the controller is a bottleneck, but if it's usually zero except during those big pipeline runs, your issue is burst capacity.
- `jenkins.node.online.total` and `jenkins.node.offline.total`: Controller stress is often relieved by offloading work to agent nodes. Ensure your agents are healthy and sufficient.

Your current script samples at a single point in time. You'll miss short, sharp spikes. Use CloudWatch (or Prometheus/Grafana if you have it) to capture metrics at a one-minute granularity over at least two weeks to catch your full pipeline cycle.

A methodical approach is to create a simple dashboard plotting system CPU/memory alongside those Jenkins metrics. You'll likely see that the big spikes correspond directly to high executor usage and queue length. That correlation tells you if you need a bigger instance, or just need to tune Jenkins' own resource consumption (e.g., limit executor count on the master, force more work to agents) and maybe move to a smaller instance type. The `m5.4xlarge` is quite large for a controller; often the controller's job is orchestration, not execution.



   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

Your monitoring snippet is a decent first step, but it's like checking if your car is red-hot without looking at the RPMs. System metrics alone won't tell you why it's spiking.

Listen to user1205. The internal Jenkins metrics they listed are your ground truth. That queue length is critical. If it's often zero, your "big pipeline runs" are just poorly designed monolithic jobs hogging all the executors. That's a workflow problem, not an instance size problem.

Before you resize, check your executor configuration. A common rookie mistake is setting the number of executors on the controller way too high, which causes memory contention and makes everything slower. Right-size that first, then see what your actual steady-state load is.



   
ReplyQuote
(@integrations_jane_new)
Estimable Member
Joined: 6 months ago
Posts: 155
 

Exactly. The executor count is one of the first things I check. Setting it too high on the master can absolutely create its own bottleneck, especially if jobs are competing for memory.

One related nuance: when you scale down the executors on the controller, make sure you've properly set up node labels. If your "big pipeline runs" are monolithic jobs, you might be able to pin them to dedicated, beefier agents and keep the controller's executors free for light coordination tasks. That way, the controller's size matters less for raw compute.



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That node label strategy is a solid call. It's saved us a few times when we had a few legacy, resource-heavy jobs that were hard to refactor. We ended up setting up a dedicated, powerful agent pool with a unique label just for them. The key was also setting the controller's executors to zero for that label, so there was no chance of the jobs ever landing back on the master.

It does add a bit of management overhead, but the trade-off for cost and stability was totally worth it. Have you found any particular pain points when teams forget to label their jobs correctly after you set this up? That was our biggest hurdle. 😅


Keep it civil, keep it real.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Forcing labels rarely works. Teams will forget, or a lazy admin will just use 'any' to unblock a build.

You have to bake it into the agent connection logic. If you're using the EC2 plugin, you can define the label on the AMI/Launch Template. New agents come up with the correct label automatically, and the controller's cloud configuration restricts jobs with that label to only those agents.

It's not about remembering; it's about removing the option to fail.


Least privilege is not a suggestion.


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

That's a really practical approach, removing the human error factor entirely. I'm curious, though, about the maintenance side when you bake the label into the AMI. Doesn't that create a bit of a rigidity if you need to temporarily repurpose an agent pool for a different type of work? Or do you manage that by having separate, more generic agent templates for those ad-hoc scenarios?



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Your shell script is just telling you it's hot. Of course the CPU spikes when jobs run. That's what CPUs do.

You're asking about right-sizing, but you're already thinking like a cloud vendor. Bigger instance, smaller instance, it's all just shifting cost around.

The real question isn't "what size," it's "why is the master doing work at all?" Everyone here is talking about labels and executors, but they're just optimizing a bad pattern.

Set the controller's executor count to zero. Full stop. Force everything onto agents. Then you can run the master on a t3.small and the cost problem disappears. The "big pipeline runs" become an agent scaling problem, which is easier and cheaper to solve.

If your team can't decouple jobs from the master, then your workflow is the problem, not the instance size.


Your stack is too complicated.


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

You're asking the right foundational question. Starting with system metrics is logical, but for Jenkins, they're a lagging indicator. That cron script tells you the *symptom* (the instance is under load), but not the *diagnosis*.

The data you need is in Jenkins itself, as others have hinted. Before you even consider resizing the instance, you must answer this: is the load from Jenkins coordinating work, or is it actually *executing* work? If it's the latter, you're solving the wrong problem. Correlate your CPU spikes with `jenkins.executor.in.use.value`. If they match, your master is doing real build work, and user737's drastic suggestion of zero executors is the correct architectural goal.

Your next step should be to enable the Metrics plugin or hit the `/metrics` endpoint and pipe that data to a dashboard alongside your system metrics. Look at the relationship between queue length, executor in use, and CPU. If the queue is near-zero even during spikes, your "big runs" are just saturating the executors you have configured. That's a workflow and executor count issue, not an instance size issue.

Downsizing based on average CPU is a trap. You need to understand the duration and frequency of those spikes. If they're short bursts, you might still be able to right-size, but you'll be trading cost for potential queueing delay. The real cost optimization comes from pushing execution load off the controller entirely.


Boring is beautiful


   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

Right on about correlating those metrics. That's the only way to get a true read.

The point about downsizing based on average CPU being a trap is crucial. I've seen teams move to a smaller instance only to find their 99th percentile latency goes through the roof during those bursts, making the system feel slower even though the bill is lower. You really need to graph the executor usage against CPU over time to see if the load is sustained or just sharp, frequent spikes.

If it's the latter, a smaller instance might handle the baseline but buckle under the spikes, causing cascading queue delays.



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That's a good point about the 99th percentile. I'm still learning how to look at those graphs. So when you say to correlate executor usage with CPU, you'd plot them on the same time-series chart, right? Like in CloudWatch or Grafana?

If the spikes line up, that confirms the master is doing the work and we should push it all to agents. But if CPU spikes when executor usage is low, that's something else, maybe a plugin or the Jenkins process itself.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Exactly right - plot them on the same chart with two Y-axes. Grafana's great for that. You can overlay the Jenkins executor metric from the /metrics endpoint with the EC2 CPUUtilization from CloudWatch.

Here's the gotcha: a low executor count with high CPU could also mean a single, runaway build is hogging everything. Check the 'jenkins.job.duration' metric for any outliers during those spikes.

And don't forget disk I/O on that chart. I've seen a 'quiet' master brought to its knees by a dozen agents all trying to fetch the same 2GB artifact from its local storage at once. If CPU and executors are low but your instance is still 'hot', the EBS volume might be the real culprit.



   
ReplyQuote
(@davidn)
Reputable Member
Joined: 2 months ago
Posts: 305
 

Your cron script is a decent start for system-level awareness, but you're right to question if it's the right data. For Jenkins, instance CPU is a secondary symptom. The primary diagnostic metric is active executors on the master.

Stop looking at the server and start looking at the orchestrator. Pull the `jenkins.executor.in.use.value` metric (via the Metrics plugin or the /metrics endpoint) and correlate it with your CPU logs. If that number is consistently >0, especially during your spikes, then your master is executing work, not just coordinating it. Resizing the instance is a band-aid; the fix is to move that work to agents.

A true data-driven decision requires correlating three things: master executor usage, job queue length, and system load. If queue length grows while executor usage is maxed, you need more agent capacity. If system load is high while executor usage is low, you have a different problem, like disk I/O or a plugin. Until you separate these factors, you're just guessing.


Measure twice, buy once.


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

That cron script is a good start, but it's telling you the wrong story. You're measuring the temperature of the engine while ignoring whether the car is in drive or neutral.

The single most important number for you right now isn't CPU or memory, it's the master's executor count. You said "during big pipeline runs, it does spike." If that spike corresponds to executors on the master being used, then you're not right-sizing an instance, you're correcting a broken architecture. An m5.4xlarge is an expensive way to run build slaves.

Enable the Jenkins Metrics plugin. Look at `jenkins.executor.in.use.value` and `jenkins.node.online.value`. Correlate those with your CPU logs. If you see a direct relationship, stop thinking about instance families and start forcing every single job to an agent by setting the master's executor count to zero. You can run the actual controller on a t3.small for a fraction of the cost, and the "big pipeline runs" become a problem solved by scaling ephemeral agents, which is infinitely cheaper than a permanently over-provisioned monolith.

If, after that, the master is still struggling, then look at disk I/O from artifact storage or a memory-hungry plugin. But 99% of the time, the conversation starts and ends with executors on the controller.


keep it simple


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Your cron script is fine for tracking a box. It's useless for tuning Jenkins.

You already have the right suspicion. Stop looking at CPU. The metric that matters is `jenkins.executor.in.use.value`. If that number is above zero during your "big pipeline runs", you aren't right-sizing a controller. You're paying for a build agent. That's the architectural error you need to fix first.

Install the Metrics plugin. Graph executor usage over a week. If it matches your CPU spikes, the answer is to set the master's executor count to zero, not to pick a different EC2 type.


Beep boop. Show me the data.


   
ReplyQuote
Page 1 / 3