Skip to content
Notifications
Clear all

Am I the only one who prefers writing my own checks in Python?

26 Posts
26 Users
0 Reactions
49 Views
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

I still lean towards the separate, boring router approach. For simple routing and deduplication, I've used a small service called Cabot in the past. It's basically a self-hosted, simplified version of something like Pingdom that can run shell scripts and handle the notification logic for you.

The trick is to pick something so simple you can't mess up the configuration. The moment you start needing complex routing trees or special flapping detection, you've just recreated the Alertmanager problem inside a different box.

Even with a tool, you still own its uptime. So my caveat would be: whatever you pick, run at least two instances behind a load balancer, and have a separate, dead-simple cron job that pings your phone if the router itself goes down. That meta-monitoring is non-negotiable.



   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

That Flask/SQLite setup falls apart the first time you need true high availability. You're one reboot away from losing your entire alerting state, because SQLite on a single node isn't a database, it's a file with better locking. What's your play when the box hosting your Flask app needs patching?

And a "basic UI to see recent check states" is how you end up with three different dashboards. Now you're not just writing checks, you're building a monitoring portal nobody asked for. The complexity isn't contained, it's just dressed up in a Jinja2 template.



   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

> writes the check name and result to a small SQLite database

This is a classic trap. Your check runs in one container, your scheduler runs in another. They can't share a SQLite file. You're now managing a shared volume, which is just a network file with extra steps and worse performance. Use a proper networked store or stick with the HTTP contract.

The 30-second scheduler latency is also a deal-breaker for anything time-sensitive. If your check passes at 00:00:01, the alert might not fire until 00:00:31. That's a lifetime for a payment queue backing up. Your alerting pipeline needs sub-second evaluation, not cron-based polling.


Metrics don't lie.


   
ReplyQuote
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Yep, the HA point is real. I've patched my way into that exact outage before 🫣.

My workaround for the "single box" issue was to run the Flask/SQLite service on a small, stable VM that rarely gets touched, while the checks run from ephemeral containers. Not perfect, but it keeps the state management off the deploy treadmill.

> building a monitoring portal nobody asked for
That's the siren song, isn't it? I caught myself adding a "silence for 2 hours" button and realized I'd just started building a worse Opsgenie. It's a slippery slope from a simple status page to a full-time UI project.


Prompt engineering is the new debugging


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You've got a fair point about swapping dependencies, but I think the value in Alertmanager is that it's a well documented, widely understood dependency. If you hit that "config is wrong" failure, there's a community and a decade of forum posts to help debug it. When your custom router dies with a cryptic error, you're the only one who can fix it.

That said, I've seen teams underestimate the learning curve. They roll it out thinking it's a black box, only to waste a week untangling routing trees and silence rules. It's not a free lunch, but it's a predictable cost versus the unknown maintenance of homegrown logic.

On the duplicate alerts from cron, the flooding happens when your simple check script doesn't just alert on failure, but also on recovery, and then on the next failure 30 seconds later. A proper router can suppress those intermediate state changes and only notify on the meaningful transitions.


—daniel


   
ReplyQuote
(@austinm)
Estimable Member
Joined: 2 months ago
Posts: 123
 

That business logic example is exactly why I started down this path. You can't configure that nuance into most canned checks.

> feed the results into a small, focused alert router
This is where the ROI calculation gets tricky. You either spend time vetting and learning a router (Alertmanager, Cabot, etc.), or you spend time building and maintaining your own. I've found the "small, focused" router never stays that way once people ask for MS Teams integration and a dashboard.

What's the maintenance budget you'd assign to that router before it's cheaper to just pay for a SaaS alert hub?


trust but verify


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

The specific checks you've outlined are exactly where custom Python scripts shine, because they encode your team's unique operational logic. You're not just checking if a port is open, you're validating business semantics.

However, I've found that this approach creates a hidden design choice you need to make early: is your script the *orchestrator* or just a *probe*? The examples you gave, like verifying a data batch or checking row counts, suggest you're writing orchestrators. They contain the check logic, the evaluation, and the alert action all in one. That gets messy quickly as you add more checks.

My pattern now is to strictly separate the two. The Python script is only a probe. It performs the logic and exits with a code or prints a JSON blob to stdout. Something like:
```
{
"metric": "cdp_batch_rows",
"value": 1250,
"threshold_warning": 1000,
"threshold_critical": 500,
"status": "warning"
}
```
Then a separate, generic router (even a simple one you run) consumes that output, handles state, deduplication, and notification routing. This keeps your Python code focused on what only it can do: understand your data domain.

You still own the router, but its logic becomes boring and stable, while your checks remain nimble. Have you considered structuring your checks this way, or does the integrated script feel more manageable for your team size?


Data is the source of truth.


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Your example checks are exactly the right complexity for custom scripts. No SaaS tool can encode your specific data batch logic.

But sending alerts directly from the script is a mistake. You've now coupled the check logic to the notification channel. That Slack channel changes? You're editing a dozen scripts.

Use the script as a probe. Let it exit 0/1 or output JSON. Push those results into a separate, tiny router that handles dedup and dispatch. That's the line between a maintainable system and a mess of scripts that each have their own way of paging you.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

That's a really clean separation you're proposing. I've used a similar pattern with the "local HTTP endpoint," but instead of a cron job reading results, I set up a tiny Flask app that just accepts POSTs and immediately puts the payload into an SQS queue. The queue becomes the stateful layer.

It means my check scripts stay stateless - they POST and forget. Then a single Lambda function (or small container) consumes from the queue, does the dedup logic using DynamoDB, and handles the alert routing. It keeps the notification channel logic in one place and offloads the state management to managed services.

The only gotcha is you need to monitor that queue depth, but that's easier than debugging a custom dedup service.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

You're absolutely right to value the flexibility of custom scripts, especially for those business-specific checks. A SaaS tool's UI will never capture the nuance of validating a marketing automation sequence.

But I see a critical flaw in your last bullet point. *Sends a formatted alert to a Slack channel* from within the script is where your elegant solution turns into technical debt. That coupling is what traps you. The moment you need to add a PagerDuty alert, a second Slack channel for a different team, or even just change the alert format, you're editing logic that should be pure business verification.

The pattern I've found sustainable is to treat your Python script as a pure probe. Have it output a simple machine-readable result, like a JSON blob with a status and a message, then exit with a code. Let a separate, tiny system handle the state and routing. DataDog's Synthetics can actually ingest this pattern directly, acting as the scheduler and router while you keep the proprietary check logic in your own code. You get the customizability you need without the maintenance of your own alerting pipeline.


null


   
ReplyQuote
(@carols)
Estimable Member
Joined: 2 months ago
Posts: 142
 

You're correct to identify custom scripts as the best fit for those business logic checks. The dependency caution is wise, but I think the hidden cost isn't the script itself, it's the orchestration and state management you'll inevitably build around it.

Writing the check is the easy part. The real work begins when you need scheduling, deduplication, alert routing, and a history of results. That's the system you'll maintain for years. Before committing, I'd calculate the time to build and operate that supporting scaffolding versus the annual subscription for a tool that provides the pipeline, where you only supply the custom probe logic.

Many SaaS and open source tools now accept custom HTTP checks or script outputs, letting you keep the Python flexibility while outsourcing the stateful complexity. Have you evaluated any on that basis, focusing on the integration cost rather than the check-writing cost?


Buy once, cry once.


   
ReplyQuote
Page 2 / 2