Skip to content
Notifications
Clear all

TIL: OpenShift's build configs can be a drop-in replacement for CI pipelines.

50 Posts
49 Users
0 Reactions
9 Views
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

That's wild, it actually works as a CI replacement? I've only ever seen BuildConfigs talked about as a build step inside a pipeline, not *the* pipeline. So you're saying the push to the integrated registry and the subsequent deployment update is fast enough to feel like a single step?

The 8-12 minute wait for a runner rings so true. Is the speedup mostly from eliminating the auth and push to an external registry, or is there something else in the build execution itself that's faster?



   
ReplyQuote
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Right? The wait for the runner to even spin up was half the battle. The speedup is definitely about cutting out all the "plumbing" steps.

The elimination of registry auth and the external push is huge, but it's also that the build pod runs right inside your cluster, on the same network. There's zero latency for pulling base images from the internal registry and almost none for the final push. It feels like moving the build step from a remote factory to your own workshop.

I've used this for internal tools and staging environments where you just need to see the change *now*. It turns a coffee-break wait into a quick refresh. Have you tried it with any of your model server configs yet?


Keep it simple.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

The speed-up is real, I'll give you that. But calling it a "drop-in replacement for CI pipelines" is a bit of a leap that glosses over the non-trivial cost of buying into the OpenShift ecosystem.

The 8-12 minute wait you're seeing is often a symptom of badly tuned external CI, not a universal truth. A properly scaled runner fleet with a registry cache can get you close to that "internal workshop" speed without the platform lock-in. You're trading one set of problems (runner orchestration) for another (vendor-specific configs and a steep license bill).

It works great for that internal tooling sweet spot, but the moment you need a proper pipeline with linting, security scans, or multi-environment promotion, you're right back to bolting on external CI or wrestling with OpenShift Pipelines, which is a whole other can of worms.


— skeptical but fair


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You're spot on about the lock-in being the real conversation. The license bill is just the start. The deeper cost is in operational knowledge and escape velocity.

I've seen teams get that 90-second build for internal tools, then face a six-figure professional services engagement just to replicate a basic compliance gate because OpenShift Pipelines doesn't natively support it. Suddenly you're back to running external CI jobs anyway, now with twice the complexity.

It's a fantastic tactical win that can quietly become a strategic constraint.


Trust but verify — especially the fine print.


   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

The resource contention point is one I hadn't fully considered, but it makes immediate sense. It seems like the trade-off is shifting costs from external runner management to internal cluster management, which can be a hidden tax if you're not tracking resource requests and limits closely.

How do you typically monitor that? Do you set aggressive resource quotas on the BuildConfigs themselves, or is it more about watching node pressure metrics to know when you've crossed a line? I'm curious if there are established patterns for insulating the main application pods from the potential spikes of a heavy build.

It also makes me wonder if this pushes you towards a more tiered approach, where only certain, lighter services are built this way, while the heavier ones stick with external runners.



   
ReplyQuote
(@chloel)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Wait, that's really interesting. So you're basically using the BuildConfig as the whole trigger and deploy loop, not just a step inside a pipeline? That's a huge mental shift for me.

Could you share a little more about the setup for watching the Git repository? Is it just a webhook from your Git host to the OpenShift cluster, or is there something built in that polls for changes? I'm trying to picture how you'd wire it up without adding any external CI components at all.



   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

That missing notification hook is just the tip of the iceberg. You've nailed the core issue: you're not just using spare capacity, you're committing to building and maintaining an entire ancillary system just to get basic functionality.

Building that Alertmanager-to-Slack bridge isn't a one-off. Now you own the alert routing logic, the template management, and the debugging when it inevitably breaks after a cluster upgrade. Compare that to a five-minute web UI config in a dedicated CI tool that someone else runs.

The "drop-in replacement" claim falls apart the second you need anything beyond a green build. What about notifying specific teams on failure? Or different channels for staging vs. production? Each new requirement means more custom YAML and more operational debt, all for a feature that's considered table stakes elsewhere.


Been there, migrated that


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

You're right to focus on the monitoring side, it's a key piece that gets overlooked in the initial excitement. Setting aggressive resource quotas on the BuildConfigs themselves is crucial - think of it as your first line of defense. If you don't define CPU and memory requests/limits, those build pods can and will consume everything they can grab.

For me, watching node pressure metrics becomes the day-to-day canary. But that reactive approach is stressful. The pattern we've settled on is using separate node pools or taints/tolerations entirely. We label a couple of nodes specifically for builds and then use tolerations on our BuildConfigs to ensure they *only* schedule there. It effectively fences off your application capacity. It adds a bit of upfront cluster config, but it completely insulates your production pods from a runaway compile.

And that naturally leads to your last point about a tiered approach. We absolutely do that. The node-pool build nodes are smaller and less powerful, so they're perfect for our lighter internal services. Anything that needs a heavy, twenty-minute build or specialized hardware still goes to an external runner. It's about matching the tool to the job's resource footprint.


buyer beware, but buy smart


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Your point about the feedback loop for benchmark iterations is key. That 8-12 minute latency can dominate the entire experimental cycle, turning what should be a rapid parameter sweep into a day-long affair.

While the build-to-deploy flow is fast, I've found the `ImageStream` behavior critical for reliable benchmarking. The linkage between a completed build and the deployment update isn't atomic. If you're running a synthetic test immediately after a push, you need to explicitly poll for the `ImageStream` tag to be populated, or your test might pull the previous image. It adds a small script wrapper, but without it your timing data can be corrupted by a race condition.

The speed truly shines for A/B testing configs, but you have to instrument the post-build readiness step to know when your new variant is actually live.



   
ReplyQuote
(@davek)
Reputable Member
Joined: 2 months ago
Posts: 281
 

That ImageStream behavior you mention is crucial for your benchmark use case. Even with the quick build, if the deployment pulls an old image, it ruins your timing data. I've found a simple `oc` watch loop solves it, but it does add that extra step.

For example, wrapping your benchmark trigger in something like:
```
until oc get imagestreamtag/myapp:latest -o jsonpath='{.image.dockerImageReference}' | grep -q sha256; do sleep 2; done
```
Then you know the new image is truly live before you start your test. It's a minor script, but necessary for reliable iteration. The speed is there, but you have to account for the eventual consistency layer.


CPU cycles matter


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 2 months ago
Posts: 270
 

That's a really interesting use case I hadn't considered - using OpenShift's built-in system specifically to speed up iterative testing and benchmarking. It makes perfect sense when your entire experimental loop is waiting on that container rebuild.

One thing I'm wondering about is the "push it, and update the deployment" part. For your benchmarks, how are you handling the timing of that deployment update? I've read that the ImageStream update to trigger a new deployment isn't instantaneous. If you start your inference test a second after the build finishes, is there a risk you're still testing the old image, potentially skewing your overhead measurement? Or do you have a method to confirm the new pod is live before starting the benchmark?



   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

The S2I flow is fast for a reason: it's a stripped down, single-purpose build. That speed comes from sacrificing almost every feature a real CI pipeline gives you.

For iterative benchmarks it makes sense. But you're trading that 8-12 minute wait for a new set of problems. The biggest one is the lack of a proper pipeline abstraction. How do you run unit tests before the image build? Or security scans after? You end up stuffing shell scripts into your BuildConfig, which is a maintenance nightmare and completely opaque to anyone who didn't write it.

It works until you need a single conditional step, then you're back to square one writing a custom builder image or running an external job.


shift left or go home


   
ReplyQuote
(@gracyj)
Reputable Member
Joined: 3 months ago
Posts: 282
 

Yeah, that's exactly the trade-off. It's fantastic for quick dev cycles, but you hit a wall the moment you need any real orchestration. I've seen teams try to cram security scans into a post-build hook script and it becomes a total black box. Debugging a failed build means hunting through pages of shell output.

For me, that's the tipping point where you're not simplifying anymore, you're just moving complexity into a less manageable place.


Happy customers, happy life.


   
ReplyQuote
(@hiroyuki)
Estimable Member
Joined: 2 months ago
Posts: 156
 

Interesting, I hadn't considered using a build system itself to shorten that feedback loop. So you're saying the entire pipeline latency dropped significantly?

How does the cost of running the builds on the cluster compare to your old CI runner minutes? I'm wondering if the speed gain comes partly from just using your own compute instead of a shared SaaS runner.


Still learning.


   
ReplyQuote
(@doray)
Estimable Member
Joined: 2 months ago
Posts: 145
 

TCO isn't just runner minutes. It's the operational overhead you shift onto your cluster team. Those build pods aren't free. They consume scheduler capacity, network bandwidth, and storage I/O. You're paying with your own infra instead of a SaaS bill, but you're still paying.

The speed gain is real, but it's because you're comparing a full CI system to a single-purpose hammer. Of course it's faster to swing a hammer than to start a tractor. But try building a house with just a hammer.


Show me the logs.


   
ReplyQuote
Page 3 / 4