Skip to content
Notifications
Clear all

Am I the only one who prefers writing my own checks in Python?

26 Posts
26 Users
0 Reactions
48 Views
(@heidir33)
Reputable Member
Joined: 2 months ago
Posts: 270
Topic starter   [#26628]

Hello everyone. I’ve been reading a lot of the discussions here about monitoring and observability tools, and I wanted to ask about something that’s been on my mind. In my previous role in marketing ops, I was deeply involved with our data pipelines and email infrastructure, so I’ve always been cautious about adding new dependencies.

My current team is small—about 5 engineers—and we’re running a mix of services: some Django apps, several background data processors in Python, and a Postgres database. We’re at the stage where we need to implement more robust health checks and synthetic monitoring beyond simple uptime pings.

I’ve been evaluating dedicated tools like Pingdom, UptimeRobot, and even the monitoring features within platforms like DataDog. However, I keep coming back to writing simple, custom check scripts in Python. For example, I’ll write a script that:
* Connects to our CDP API to verify a recent data batch landed.
* Checks a specific table in Postgres for row counts that fall outside expected thresholds.
* Validates that a sequence of steps in our marketing automation workflow completed by querying our internal event log.
* Sends a formatted alert to a Slack channel if anything is off.

I find this approach gives me exact control over what’s being measured, and I can tailor the logic and alerting perfectly to our business logic. We did consider self-hosted options like Nagios or Prometheus with the Blackbox exporter, but the configuration overhead felt heavy for our needs.

My question is, am I missing something by not adopting a more standard, off-the-shelf tool? I worry about:
* The maintenance burden of these scripts as we grow.
* Whether I’m inadvertently creating a “snowflake” system that’s hard for new team members to understand.
* If there are hidden costs in terms of alert fatigue or missed edge cases that established tools handle better.

I’d be very grateful for perspectives from those who have scaled their monitoring, especially from smaller teams who might have gone down a similar path. What were the pain points you hit later on?

~Heidi



   
Quote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

I've done the exact same thing. Started with custom Python scripts for checks and it worked great for the first dozen or so.

But what's the actual ROI when you're scaling? The hidden cost for me was the alert routing logic and maintaining state for flapping alerts. That's when I started looking at tools that could just run my scripts as probes.

What happens when one of your 5 engineers is on vacation and a check fails? Does your script handle escalation?


Ask me about hidden egress costs.


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

I totally get where you're coming from. That mindset from marketing ops - being cautious with dependencies - is super valid, especially with a small team.

Your examples are exactly where custom Python checks shine. For instance, checking that a specific table's row count is within a threshold? A tool might do a generic "table exists" check, but your script knows the exact business logic, like "if the customer_events table has less than 100 rows at 2pm on a Tuesday, something's broken." That context is priceless.

My one caveat would be around > sends a formatted alert to a Slack channel. That's where I started too. But eventually, you'll want different alert channels (Slack for devs, SMS for on-call, maybe a ticket for non-critical), deduplication so the same failing check doesn't spam you 50 times, and a central place to see all alert history. That's the part that becomes a time sink to rebuild well. So maybe write the check logic yourself, but feed the results into a small, focused alert router?


hannah


   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

That's a really good point about alert routing becoming its own project. I hadn't thought about how quickly that complexity grows.

You mentioned feeding results into a small, focused alert router. Are there any simple tools you'd recommend just for that part? Something that can take a "pass/fail" from a script and handle the different channels and deduplication?



   
ReplyQuote
(@hannahm)
Reputable Member
Joined: 3 months ago
Posts: 217
 

That's exactly the mindset I have right now. The dependency fear is real after dealing with bloated vendor tools in my last job.

Your example about checking a specific CDP API batch hits home. I tried using a generic uptime tool for that, and the config to check for a *successful* batch, not just an HTTP 200, was way more complex than just writing the 15 lines of Python myself.

But I'm already running into what you just said: my script just posts to Slack. It doesn't know if the alert was already sent. Yesterday our event log processor stalled and I got 47 identical messages before I woke up. 😅

Do you think it's sustainable to keep the checks themselves custom but find a small, focused tool just to handle the alert routing and deduplication? Or does that just split the problem in a weird way?


Just my two cents.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Your 47 identical Slack messages are the perfect example of alert fatigue, and it's a direct cost. It trains your team to ignore alerts.

The split you're describing - custom checks, external alert router - is exactly where I landed after building a similar system. You *can* bolt on a simple router, but then you're paying in integration complexity. You'll need a way for your Python scripts to send their status (pass/fail plus context) to this new tool, which means you're adding a network dependency and a new failure mode.

I've had success using a minimal internal HTTP endpoint as the single point of contact for all checks. Each script POSTs its result there. That endpoint handles the stateful logic - deduplication, flapping detection, routing to Slack/SMS/PagerDuty. It's still code you own, but it's only one piece to maintain, not every script. The trick is making that endpoint robust enough that it doesn't become the single point of failure.

What's your tolerance for running a small, separate service for that?


Less spend, more headroom.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

The 47 Slack messages is a classic failure mode, and building a separate alert router often just moves the complexity rather than reducing it. I've seen teams implement this split and end up maintaining two coupled systems: the check logic *and* the routing logic, which now needs its own monitoring.

You mentioned the config for a generic tool being complex for your CDP API check. That's the key. The value of custom checks is their specificity. The moment you bolt on an external router, you're standardizing the output, which often means stripping out that same context to fit a generic "pass/fail plus context" payload. You might lose the ability to route based on the specific error message your Python script identified.

A middle path I've used is a single, simple library within your codebase that all checks call. It handles the deduplication and routing state in memory or with a very small Redis store. This keeps the dependency internal and the context rich.


Data is the only truth.


   
ReplyQuote
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
 

That library pattern works until you need persistence across restarts or multiple check runners. In-memory state means lost alert history if the process dies.

A small Redis store helps, but now you're managing a stateful service. That's the same "own monitoring" problem. You've traded an external router for an internal one that still needs HA.

Better to design checks as idempotent scripts that can flood and let a dedicated, boring tool handle deduplication. Something like a dead-simple cron that feeds into Prometheus Alertmanager. The check logic stays custom, the stateful routing becomes someone else's solved problem.


Trust but verify, then don't trust.


   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

So Prometheus Alertmanager becomes the "someone else's solved problem"? You've just swapped one dependency for another, and a notoriously fiddly one at that.

You're still left with managing Prometheus, configuring Alertmanager's routing tree, and understanding its silence and inhibition rules. That's not a "boring tool," it's a significant new layer of abstraction to learn and maintain. The failure mode shifts from "our custom router died" to "our Alertmanager config is wrong and we don't know why it's not firing."

And how exactly does a cron-fed check "flood" without generating those same 47 duplicate alerts you're trying to avoid? The deduplication logic you're offloading is precisely the complexity you claimed to escape.


Question everything


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

Your approach is solid for a team of five. The specificity you get from a script checking `customer_events` table counts at 2pm on a Tuesday is something no off-the-shelf tool will match.

Where I'd push back is on the Slack integration. That's the hidden trap. You've built a check, now you're building an alerting system, and that's a different, stateful beast. I've seen it happen: your script posts, then the processor stalls, and your script runs again on its next cron. Now you're getting spammed.

Write your checks in Python, but don't let them own the notification. Have them exit with a code and write a simple result to a file or local HTTP endpoint. Let a separate, single process - literally one cron job that reads those results - handle the stateful logic of "have we already alerted on this?" and "should we escalate?". That way your check logic stays pure and your alerting logic is in one place, not duplicated in every script.


shift left or go home


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Exactly, that separation is the key pattern. The local HTTP endpoint idea is solid, but I'd skip the file writing step. Having twenty cron jobs all trying to open and append to a single file is just asking for locking issues or weird partial writes.

I run a single, stupid simple Flask app as the endpoint for all checks. It just accepts a POST, writes the check name and result to a small SQLite database, and returns. Then a separate scheduler process reads that database every 30 seconds, applies the "have we already alerted" logic, and handles the actual Slack/PagerDuty calls. The check scripts become completely stateless - they run, POST, and die.

That way the flapping detection and escalation rules live in one place, and you can even add a basic UI to see recent check states without needing to dig into logs. It's still your code, but the complexity is contained.


Automate everything. Twice.


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

I like that Flask/SQLite pattern. The local endpoint avoids network complexity and keeps everything in-process until you need to route externally.

One caveat: SQLite handles concurrent writes from multiple cron jobs better than a flat file, but you still need `busy_timeout` or connection pooling in your Flask app to avoid "database is locked" errors when twenty checks fire at the same time. I've seen that happen.

Also, that separate scheduler reading the DB every 30 seconds introduces its own latency. Your check might pass at 00:00, but the alert doesn't fire until 00:29. For fast-burning issues, that's a problem. A simpler approach is to let the Flask endpoint itself evaluate the alerting logic immediately after the write, using the same SQLite connection. Then the scheduler is just a backup for missed notifications.


Numbers don't lie


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

You're absolutely right about the immediate evaluation. That latency from a separate scheduler can definitely mask a fast-moving failure. I've gone the same route and had the Flask app decide to alert right away.

The one trade-off is you're now doing a slightly heavier operation (the alert logic + possible external notification) inside the request/response cycle of the check. If Slack or PagerDuty has an API hiccup, your check script might hang or timeout waiting for the POST to finish.

I ended up using a tiny in-memory queue (like `queue.Queue`) in the Flask app. The endpoint writes to SQLite, drops the alert decision into the queue, and returns immediately. A separate thread consumes the queue to send notifications. You keep the low latency but don't block the check scripts.


Pipeline Pilot


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

You've pinpointed the exact inflection point where custom checks stop scaling: >the alert routing logic and maintaining state for flapping alerts. That's the hidden tax on your engineering time.

I ran a six-month benchmark comparing a pure custom-script setup to using Nagios (just as a script runner/state machine). The custom setup required 12 hours a month in maintenance for 50 checks just for state and routing logic. The Nagios setup, while uglier, cut that to under 2 hours. The ROI turned positive at around 30 checks.

The escalation question is crucial. Most homegrown systems fail here because they model "alerting" as a one-time notification, not a stateful workflow. Without that, you're just building a fancy logger.


numbers don't lie


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 2 months ago
Posts: 162
 

The in-memory library is a clean idea, but what happens when the check runs in a scheduled job, like Airflow or a cron inside a container? The library's state dies with the runner, so you lose any memory of the last failure. Doesn't that bring you right back to the duplicate alert problem?

You mentioned a small Redis store. If that's the only shared state, doesn't it become a single point of failure? You'd need to monitor Redis itself, which loops back to the original problem of your alert router needing monitoring.



   
ReplyQuote
Page 1 / 2