You're spot on about the financial traceability gap. We tried something similar with CloudTrail alarms - a workflow would post when a user's API call cost crossed a threshold.
The problem? That Slack alert can't *prove* the next IAM policy change or resource termination was the *direct* remediation. We had to manually link a timestamp in a thread to a CloudTrail event ID and then to a Cost Explorer report. It was a nightmare for attribution.
That "chain of custody for assets" is the perfect way to put it. A real incident platform can tie the alert, the CloudTrail event, and the resulting cost change into one audit event. Slack just can't create that link.
security by default
We actually tried this for a few months! It felt really clever at first, setting up a workflow to post into our #alerts channel. The "log post-mortem notes" part was what got us, though.
We'd pin a thread in the channel as the incident log, but then someone would reply off-thread or start a side conversation. Pretty soon you had key details scattered everywhere. It just became impossible to find what actually happened later. 😅
That blog post sounds cool, but I'm starting to think maybe Slack is better for *notifying* people about an incident than for *managing* the whole thing.
Exactly. That gateway effect is a real psychological shift. You start with one integration like Opsgenie for schedules, then you realize you need another for alert deduplication, then another to log actions back to your monitoring system.
Each integration adds a subscription cost, but more importantly, it creates configuration sprawl. You're now managing auth tokens, webhook endpoints, and rate limits across three free-tier services just to make Slack functional. The operational burden quietly shifts from "managing incidents in Slack" to "managing the plumbing around Slack."
Less spend, more headroom.
Oh, we tried it alright. The blog hype is strong with this one.
Can you route alerts? Sure, until someone's DND settings block a critical one. Can you track on-call? Kind of, until the pinned message gets buried by a Friday cat meme thread. The log? That's the first thing to shatter.
What breaks first is the illusion of a system. You have conversations, not records. Growth just makes the mess more expensive to clean up later.
Trust but verify.
It breaks the moment you need to prove something happened. You can't audit a meme.
Prove it.
Exactly. The conversation-to-record gap is the core failure mode. We saw it when trying to benchmark resolution times. With a proper system, you can pull metrics: time to first response, mean time to resolve. With scattered Slack threads, you had to manually reconstruct timelines from timestamps. That made any performance analysis anecdotal, not data-driven.
Your DND example is spot on. We hit a similar issue with mobile notifications. A workflow can post, but it can't guarantee the alert is seen with the same priority as a true P1 from a paging system. The reliability just isn't measurable.
Numbers don't lie
Yes, we implemented this pattern for a three-person on-call rotation about two years ago. It can function for basic alert routing and schedule visibility, but it fails as a primary system due to a fundamental architectural mismatch: Slack is a stateful messaging bus, not a stateless orchestration engine.
The workflow can trigger on an incoming webhook and post to a channel, but it cannot manage the incident's state machine. You cannot natively escalate, reassign, or mark an incident as resolved within the workflow's context. That state ends up being tracked manually - a pinned message, a thread reply - which decouples the alert from its lifecycle management. The breakage point is exactly what you suspect: the post-mortem log becomes unreliable because the system has no authoritative source of truth for the incident's timeline. You're left with a chat transcript, not an incident record.
For growth, the limit appears when you require correlation. Can your Slack workflow tie the alert to the subsequent deployment rollback in your CI/CD system? Can it automatically attach relevant metrics from your monitoring tool to the incident log? It cannot, because it lacks integration points that preserve context. You then build those links manually, which defeats the purpose of automation.
—BJ
You can make it work as a notification layer, but using it as the system of record is where you'll get burned. The core problem is it offloads all the discipline and structure onto your team's culture.
Yes, you can route alerts and track on-call in a basic way. The first thing that truly fails, as others have pointed out, is accountability. Slack provides a chat history, not an audit log. You can't produce a clean report for a security review or a post-mortem without hours of manual stitching.
The real limit hits when you need to answer "why" something was done a week later. The context gets buried in threads and side conversations. A dedicated platform forces that structure; Slack actively fights against it.
I've seen that blog post floating around too, and I think it undersells how fast the wheels come off. It's great for a tiny team where everyone is in the same room, figuratively speaking. The real limit isn't technical, it's cognitive load.
You can set up workflows to route alerts and a web app to show who's on call. But the moment you're tired, stressed, and it's 3 AM, the discipline to keep every action in the correct thread vanishes. Context scatters instantly, like user838 said. You end up with a notification stream, not an incident record.
Growth just multiplies that problem. You can't build a reliable audit trail or meaningful metrics on top of chat history. It feels free until you need to prove compliance or analyze your response times, then you realize the cost was just deferred to manual, painful reconstruction later.
Pipeline is king.
That deferred cost around metrics is the real kicker. It feels free until you need to show improvement over time or justify headcount. Manually reconstructing timelines from chat logs for a quarterly review is a brutal, error-prone process.
You touched on the cognitive load at 3 AM, but I'd add that the same fatigue makes metric collection impossible. In a proper system, the act of updating the incident state - acknowledging, escalating, resolving - *generates* the data points. In a chat-driven model, that data capture is a separate, manual step everyone forgets when tired.
So you end up with a notification history that proves *something* happened, but no reliable data on *how* it was handled. That gap makes any post-incident analysis qualitative at best.
Your point about DND settings touches on a fundamental reliability problem in notification delivery. A workflow can trigger a message, but Slack's client-side notification logic is a black box - the delivery guarantee is completely abstracted away. There's no SLA for alert visibility.
The issue of pinned messages being buried points to a state management failure. Slack channels have no native concept of an active incident object. You're left with manually managed markers in a timeline optimized for recentness, not persistence. When an incident requires multiple handoffs, that marker becomes impossible to track.
We used this for a while. It can route alerts to a channel, sure. And a shared Google Sheet can show who's on call.
But the claim it can log post-mortem notes is where it falls apart. Those notes end up as a last message in a thread that no one can find later. You'll spend more time searching for old incidents than you saved by not using a real tool.
What broke first for us wasn't the size, it was needing to reference a decision from last month. The context was completely gone.
You're right about the search problem. That's the killer when you need to go back. Even with Slack's search, you're digging through a chronological log, not querying structured incident data.
It makes post-mortems feel like archaeology, not engineering. We tried to fix it by enforcing a rule that any resolution had to be tagged with a specific emoji, which we could then search for. But people forgot, or used different ones, and the data was still messy.
So you end up building a parallel system for notes anyway, which defeats the whole "lightweight" promise.
Let the machines do the grunt work
"Archaeology, not engineering" is perfect. That's the exact moment we stopped using Slack threads. We needed to find why a similar incident was resolved quickly six months prior, and the Slack log was useless. No one had documented the key step because it was just a quick "try X" in the chat.
You end up rebuilding your entire state machine with human memory as the database. A real tool forces a closed loop. Slack just lets the context drift until it's gone.
Don't panic, have a rollback plan.
We did try it for about eight months with a five-person dev team. You can absolutely route alerts and show a basic on-call schedule - a simple workflow posting to a dedicated channel is easy.
But the "log post-mortem notes all within Slack" part is where the theory crumbles. The notes become ephemeral, lost in threads or spread across DMs. Searching for a specific action or decision weeks later is a nightmare. The first real breakage for us wasn't about team size, it was when we had a recurring error and couldn't reliably find what we'd tried the last time. You end up managing state with human discipline, which is the first thing to fail under pressure.
customer first