Hi everyone! I'm still pretty new to Jenkins and I've hit a really confusing wall. My declarative pipeline runs fine most of the time, but about 20% of runs just... stop. The console output ends abruptly with no error message, and the build is marked as FAILURE. It's usually at different stages too.
Here's a simplified version of my `Jenkinsfile`:
```groovy
pipeline {
agent any
stages {
stage('Build') {
steps {
sh 'docker build -t my-app .'
}
}
stage('Test') {
steps {
sh 'docker run my-app npm test'
}
}
// Sometimes it fails here, sometimes later
stage('Push') {
steps {
sh 'docker tag my-app my-registry/app'
sh 'docker push my-registry/app'
}
}
}
}
```
The console log just cuts off, sometimes after `docker push`, sometimes after `npm test`. I've checked disk space and memory on the agentβseems okay. Is this a common thing for beginners to run into? Any idea where I should even start looking? Thanks in advance for any guidance!
Ugh, random failures with no logs are the worst. Since you've already checked disk and memory, I'd look at network timeouts next. Docker push/pull commands hanging can silently kill a pipeline if the registry is flaky.
You could try adding explicit `timeout` blocks around your `sh` steps to force a clearer error. Something like wrapping your `docker push` step in a `timeout(time: 5, unit: 'MINUTES')`. That might at least convert a silent hang into a timeout failure you can see.
Also, have you checked the Jenkins controller logs? Sometimes the agent connection gets dropped, and the error only shows up there, not in your pipeline console.
Ship fast. Learn faster.
Yeah, those silent failures can really eat up your day. Since it's failing at different stages, I'd suspect something external to your actual commands. Two things I always check in these scenarios:
First, watch your agent's connection to the controller. If the agent VM or container is under heavy load (even from other jobs), the heartbeat can drop and Jenkins just abandons the build. The controller logs will show a "sending interrupt signal" or similar.
Second, try adding `-tt` to your `docker run` command for the test stage. If `npm test` spawns a subprocess that doesn't get proper signals, it can appear to hang forever. Jenkins might be killing the whole pipeline after a hidden timeout. Wrapping steps in explicit `timeout` blocks like user1473 said will at least give you a clearer failure point. Good luck!
Data doesn't lie, but dashboards sometimes do.
Great point about the agent connection. The controller logs are key, but the agent logs often hold more specific clues about why the connection severed. Look for JNLP agent launch failures or sudden spikes in system load averages right before the drop.
Regarding `-tt` for Docker, that's smart for interactive signal forwarding. However, be aware it allocates a pseudo-TTY, which can sometimes cause its own issues with log aggregation or buffering in Jenkins. A more deterministic approach is using a process supervisor like `dumb-init` inside your test container to handle signal propagation correctly.
connected
Absolutely right about the controller and agent logs being complementary. I've found that searching the controller log for the specific build number right after a failure often points you straight to the agent logs, which saves a ton of time.
Your note on `dumb-init` is spot on for a more reliable fix, especially in production. The `-tt` flag can be a decent quick diagnostic to see if signal handling is the root cause, but moving to a supervisor is the way to go for a stable pipeline. Have you found one you prefer over another?
~Harry
Yeah, dumb-init works until you need to debug pid 1. Then you're in a world of hurt.
For pipelines, I just use tini as the Docker entrypoint. It's in the official repo, zero config, and actually made for containers.
Prove it.
That's a really good point about external load on the agent. I hadn't thought about other jobs causing the heartbeat to drop. Would this still happen if the agent has plenty of free CPU and memory? Or is it more about network connectivity under load? 🤔
Also, thanks for explaining *why* the -tt flag might help with npm test. I'll try that as a test. If it works, maybe I can look into tini for a permanent fix like others mentioned.
Still learning.
> yeah, dumb-init works until you need to debug pid 1
True, that's the trade-off, isn't it? Tini is definitely the more container-native choice and I use it too, but it's not a silver bullet. If your base image already has a minimal init system (like buildpack-deps), adding tini can sometimes cause signal conflicts or double-init weirdness.
For the OPs Jenkins issue, the real win is just having *any* proper init to catch the SIGTERM from Jenkins when a step times out. Makes me wonder if their Dockerfile's `ENTRYPOINT` is just a raw `npm start` command. That'd explain the random hangs perfectly.
pipeline all the things