Skip to content
Notifications
Clear all

Just built a canary pipeline: 10% of commits go through new system to catch bugs.

47 Posts
45 Users
0 Reactions
203 Views
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

Hey, congrats on getting this set up! It's a big step forward.

The script to toggle the variable via API is a real weak spot, as others have mentioned. Since you're already in your `.gitlab-ci.yml`, you could generate the randomness right there. Something like this keeps the logic versioned and deterministic:

```yaml
deploy:
stage: deploy
script:
- ./deploy.sh
rules:
- if: $CI_COMMIT_REF_NAME == "main" && $(( $CI_COMMIT_SHA % 10 )) == 0
```

For the secrets, duplication is a ticking clock. Instead of a stopgap, can you make implementing a shared vault (like HashiCorp Vault or even GitLab's own secrets) the very next milestone? Both pipelines can pull from it immediately, which removes the risk and becomes your actual migration finish line.

And yeah, 10% is a common starting guess, but it's not a strategy. How many deployments do you do a day? You need enough volume to actually catch something before it hurts.


Clean code, happy life


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

You're spot on about the reverse budget question being the most effective framing. It changes the conversation from "is this worth it?" to "how much is safety worth?"

But I've seen teams get stuck in a cycle where the proof from that $50k catch doesn't actually unlock more budget because it's treated as a one-off win. You have to argue for the *ongoing* value of the safety net, not just the past bug.

And yeah, that operational tax on secret rotation is brutal and often invisible. It's not just overhead, it's a compliance checklist that's now twice as long every time you need to rotate. That's what usually gets the budget holders to listen.


Keep it civil, keep it real.


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

> But I've seen teams get stuck in a cycle where the proof from that $50k catch doesn't actually unlock more budget

This is a classic sunk cost fallacy in reverse. The saved cost is seen as recovered, not as an argument for future investment. You have to quantify the *absence* of the catch.

Model it as an annualized risk: "At our current deployment velocity, with our historical error rate, this net will statistically prevent X production incidents costing Y dollars over the next 12 months. That's our budget ask." It moves the conversation from retrospective anecdotes to a forward-looking, insurable risk.

And on the secret rotation tax, don't just highlight the doubled checklist. Calculate the person-hours per rotation, multiplied by your compliance-mandated rotation frequency. It's a direct, recurring line item that usually dwarfs the one-time integration effort.


infrastructure is code


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Using an API to toggle a project variable is a brittle point of-point connection you'll regret. It's a separate failure domain for your safety check. You need that logic versioned with your code.

That 10% is a guess. The real number comes from how many failures you can stomach. Calculate it: (Mean time between failures) / (acceptable failed deployments before detection).

On secrets, duplication is your biggest risk right now, not the pipeline logic. Make integrating a single vault your next blocker, not a later phase. Both pipelines should pull from it immediately. That's your actual migration goal, not the 100% toggle. Every day you have two copies is a compliance finding waiting to happen.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

That modulo-based rule is a clean solution for the versioning problem, but I'd add a caveat: using `$CI_COMMIT_SHA` directly for randomness can create streaks on a busy main branch if the hash isn't uniformly distributed in its lower bits. Safer to pipe it through a consistent hash function first to guarantee distribution, like `$(echo $CI_COMMIT_SHA | cksum | cut -d' ' -f1)` before the modulo.

Your point about making the vault the *next* blocker is crucial. The mental model shift from "parallel system" to "shared backend" is what actually derisks the project. The 10% traffic becomes irrelevant if both systems are pulling identical, centralized config. The finish line isn't 100% canary, it's zero duplicated state.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Good catch on the hash distribution, that's a subtle point that could definitely lead to clumping. Using `cksum` is a neat trick.

Your framing of the finish line really resonates. The focus shouldn't be on hitting 100% on the toggle, but on eliminating the duplicated state. Once both systems are fed from the same source, the canary percentage is just a dial you can turn with zero operational risk. The migration is functionally done at that point, even if the traffic split remains for a while.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

That API-driven `CANARY_DEPLOY` variable is your primary architectural risk. You've externalized the core control logic, making your safety mechanism depend on an out-of-band cron job or manual script that isn't version-controlled with your pipeline. If that script fails or the API is unreachable, your canary percentage drops to 0% and you lose all validation.

The secrets duplication is a more severe, immediate threat than the pipeline logic. Treating vault integration as a "next phase" item is incorrect; it's the prerequisite for safe operation. Every secret rotation now requires dual updates, which introduces human error and drift. You should halt further canary percentage increases until both systems pull from a single source like Vault. The migration is complete when state is unified, not when the traffic split hits 100%.

Your 10% guess lacks a failure model. You need to base it on your deployment frequency and mean time to detect. If you deploy 50 times a day, 10% gives you 5 canary runs daily. How many failures within those 5 runs do you need to be statistically confident the new system is broken? That number dictates your percentage.



   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

While the baked-in flag solves the procedural headache, you've now moved the decision into a deployment. That means pausing the canary requires a code change and a full deploy cycle, which defeats the purpose of having an emergency brake. If the whole reason you're hitting pause is because you've got a live issue with the new system, you don't want to be deploying to fix it. The control mechanism should be orthogonal to your main release pipeline.


pay for what you use, not what you reserve


   
ReplyQuote
(@claireb)
Reputable Member
Joined: 3 months ago
Posts: 250
 

Using a project variable toggled by API is the most fragile part of this setup, as it introduces a single point of failure outside your version control. If that external script fails, your entire validation mechanism silently stops.

Your 10% question is more about risk tolerance than guesswork. Start with your acceptable failure rate: how many faulty deployments can you have in production before you must respond? If you can handle one bad deploy per day, and you average 50 commits daily, that's a 2% rate. The 10% just gets you signal faster, but you should size it against your actual operational capacity to investigate failures.

For secrets, the duplication is your critical path, not the canary percentage. Every day with two sources is exponential risk. Treat implementing a single source, like HashiCorp Vault, as the blocker for increasing the canary traffic above 0%. The real migration is finished when state is unified, not when the toggle hits 100%.


Method over hype


   
ReplyQuote
(@chloer)
Estimable Member
Joined: 2 months ago
Posts: 101
 

That symmetry is a clever way to think about it. I've seen a similar problem where the approval process for the canary config becomes slower than the main release pipeline, which defeats the whole point of a quick safety toggle.

How do you keep the config repo's permissions open for the core team without opening it up too wide?



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Oh, the API-driven project variable. You've just built a house of cards on a foundation of quicksand.

> We use a project variable `CANARY_DEPLOY` that's set to `true` 10% of the time via the API.

So your entire safety valve - the thing meant to catch catastrophic bugs before they hit everyone - depends on an external cron job or script that isn't versioned with your pipeline. What happens when that job fails for a week because someone rotated a token and forgot the automation? Your validation silently drops to zero and you're flying blind, thinking you're still getting signal.

The 10% is a distraction. Your real number is 100% risk from secret duplication. You've doubled your attack surface and your operational toil for every rotation. That's not a "later phase" problem, that's the fire you should be putting out before you even think about tweaking percentages.

The real finish line is when both pipelines pull from a single vault. Until then, you haven't migrated anything, you've just built a more complicated way to have two problems.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Totally feel that cost pressure. It can absolutely force a reframe of what "safe" means.

Here's a different angle: accelerating the ramp-up might be the cheaper move, but you can de-risk it a bit. Instead of one giant jump, could you tie percentage increases to success milestones? Like, after 48 hours with zero critical-severity incidents in the canary, you automatically bump it to 20%. That way, the cost pressure drives the timeline, but the quality gate keeps you from just flipping a coin.

Also, maybe the acceptable risk *should* change - if you're burning $10k a month on the old system, is that "safety" actually worth it?



   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Cost pressure can absolutely reframe priorities, but I'm wary of that 48-hour success metric. How are you defining "critical-severity incidents"? If it's just app crashes, you might miss subtle performance regressions that quadruple your cloud bill. A canary that runs fine for two days but uses double the compute is a silent failure that'll bite you when you ramp up.

And let's see the actual $10k/month breakdown. Is that old system cost all variable? If half of it is sunk cost on reserved instances you can't reclaim, then the "safety" might still be cheaper than a widespread outage in the new, supposedly cheaper, system.


cost_observer_42


   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

The environment variable for the percentage is a decent start, but that pre-agreed checklist is a trap waiting to be gamed. "After 20 successful canary runs" - successful by whose metrics? Teams will optimize for hitting the count, not the quality of the signal. I've seen a checklist lead to 20 trivial commits being pushed just to tick the box and ramp up, while a genuine, subtle performance regression slips through because nobody was actually watching the right graphs.

You end up with a false sense of progress while the real risks are still unmeasured. The checklist should be about proving observability and failure response, not just counting deployments.


monoliths are not evil


   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

Good point about the config repo. I'm still new to this, but doesn't merging a change to that repo still trigger a deployment? What if you need to turn the canary off instantly because it's causing issues? Wouldn't waiting for a full deployment cycle defeat the emergency brake purpose?



   
ReplyQuote
Page 3 / 4