That's the exact question I've been asking while evaluating our options. Most discussions talk about it breaking at scale, but I'm still fuzzy on where the true cutoff is.
You mentioned tracking on-call and logging post-mortems. Has anyone compared the effort of keeping a Slack-based log versus using a simple template in a shared doc? Does the threading model actually save time, or does it just move the fragmentation problem elsewhere?
Also, for routing alerts, how does this compare to a free tier of something like Opsgenie? Is the Slack workflow genuinely simpler for a 4-person team, or does it just feel simpler until the first schedule conflict?
That export issue hits hard. We had the same problem when our security team needed a timeline for a minor data leak. The flattened thread lost all context of which messages were replies to specific diagnostics, making it look like decisions were made out of order.
I'd even add that the "hidden debt" becomes tangible when you need to prove compliance. Screenshot stitching is one thing, but having an auditor question the integrity of your log because timestamps are in a human-friendly format is a different kind of pain.
Ship fast, measure faster.
You're spot on about the hidden cost, but you're lowballing the rate. A junior engineer at 3am on a critical outage? That's not a junior's rate, it's a senior's panic rate. Call it $250/hr fully loaded.
The real cost is context switching. That 15 minute workflow edit means 45 minutes to get back to deep work. That's the compounding debt, not the tool cost.
Teams ignore this because it's "just a few minutes." It's the definition of penny wise, pound foolish.
If it's not a retention curve, I don't care.
Absolutely, the context switching cost is the silent killer that teams just don't track. Your 45-minute figure might even be optimistic after a 3 a.m. interruption - the brain's fried.
We tried to account for this by measuring "time to resolution" before and after a proper alerting tool. For a 6-person team, we weren't just saving on workflow edits. We cut the "mental reload" period in half because people weren't juggling three places (Slack, dashboard, sheet) to trust the alert state. That's where the real "pound foolish" hits - it's not a line item, so it feels free.
Test, measure, repeat
You're right about the mental reload cost, but the part I'd add is that it's often misattributed to the people, not the system.
We saw the same pattern, but leadership kept saying our team was "bad at context switching." The real issue was the system demanding constant verification. Having to check a sheet, a dashboard, and then Slack to confirm an alert wasn't moving the needle just burned focus. A proper tool gives you a single source of truth for state, which cuts the verification loops.
It's easy to blame the engineer's focus, but the real fix is building a system that doesn't require so much of it.
catdad
That's a critical distinction. The blame shifts from a performance issue to a system design failure. It reminds me of the classic "Swiss cheese" model of accidents - you're layering a series of flawed barriers (checking a sheet, then Slack) and hoping nothing slips through, instead of building one solid wall.
A related problem I've seen is that when teams get used to the verification loops, they start to see them as normal diligence. It becomes part of "the way we do things," which makes justifying a tool that eliminates those steps even harder.
That Swiss cheese model reference is on point. I've seen this exact pattern in infrastructure teams where the verification loops get so ingrained they're codified into runbooks. You end up with a step like "Confirm alert in #alerts channel, then check rotation sheet tab 'Q3-overrides', then validate current assignee in PagerDuty UI" because the trust in any single system is broken.
The justification problem is real, but there's a technical angle. These ad-hoc systems often fail silently. A Slack workflow that's supposed to tag the on-call might just post to the channel if the user group hasn't been updated. The sheet might have a stale link. The engineer, now trained to check all three, catches it and thinks they've prevented a failure. What they've actually done is compensated for a chronic, invisible system fault. You're not proving diligence, you're proving the system is unreliable.
Measuring that silent fault rate is how you break the "normal diligence" cycle. How many times did the primary alerting method actually fail, and the manual verification loop catch it? That's the data point needed to move from blaming context switching to fixing the root cause.
infrastructure is code
Totally agree it's a tempting idea. I've seen teams try it, and the routing alerts part works okay at first. You can set up a webhook to post to a channel and even tag a user group.
The real limit for "effectively routing alerts" shows up when the alert needs logic. What if the primary person is on vacation? A real paging tool can escalate. A Slack workflow just pings a user or channel and then... stops. You're basically building a dead-end.
The post-mortem notes in a thread is the other big one. It feels organized until you need to reference something from a prior incident. Then you're scrolling through a channel forever. It breaks the moment you need any kind of historical analysis or reporting.
Still looking for the perfect one
Yes, I've seen this attempted, and the answer depends entirely on how you define "basic." For a true two-person team with maybe a single weekly alert, it's a clever hack. But the moment you need even one conditional - like skipping a vacationing engineer - you're manually rebuilding a pager in a chat client, which is a trap.
The first thing that breaks isn't scaling to ten people; it's the loss of statefulness. Slack is a stream of events. An incident is a state machine (triggered -> acknowledged -> investigating -> resolving -> resolved). Without a dedicated system to manage that state, you're relying on humans to correctly interpret and update a static post, which falls apart under stress or when someone forgets to update the thread title. You'll have two engineers diving in because the "status" is ambiguous.
Tracking on-call rotation in a Google Sheet referenced by a workflow seems fine until the sheet link breaks or a permissions hiccup means the alert fires into a channel instead of to a person. Then you've built a silent failure mode that looks like it's working.
infrastructure is code
That state machine point is key. We tried using the thread title as the state, like "[INVESTIGATING] API Latency". It worked until someone joined the channel late and just replied to the main message, creating a parallel thread. Now you have two states.
How do dedicated tools solve this? Is it just a dedicated UI that prevents branching?
We tried it for about six months on a team of four. The alert routing via user groups worked fine until someone left the company and we forgot to update the group - that was a fun Friday night 😅. Logging post-mortems in threads got messy fast because you can't easily pull data out for a report later. It's okay if your only goal is to "notify someone," but it doesn't help you *manage* the incident.
That's the patchwork approach, and it often creates the same verification loops people are trying to avoid. If the shared doc isn't in the initial automated workflow, someone has to create and link it manually. That's a step that gets missed at 3 a.m.
You're better off using a tool that creates the incident log as the single source of truth from the start, then posts a link to Slack. Not the other way around.
Least privilege is not a suggestion.