Skip to content
Notifications
Clear all

Check out what I made: A Prometheus rule generator that writes back to Git

3 Posts
3 Users
0 Reactions
10 Views
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
Topic starter   [#27925]

Hey folks! I've been wrestling with Prometheus alerting rules in our GitOps flow for months. We'd update a rule in a PR, merge it, but then the actual alert would fire hours later when ArgoCD finally synced. 😩 I wanted something that could react *immediately* to metric changes.

So I built a little operator that bridges the gap! It watches a specific Prometheus `rules.yml` file in a Git repo, and when a new alert condition is detected (like a service's error rate spiking), it **generates a new, temporary alert rule on the fly and writes it back to a branch** for review. This gives us instant visibility while still keeping the Git record.

Here's the core idea in a nutshell:

1. **Watches** Prometheus for a set of "template" alert conditions (e.g., `job:error_rate > 0.05`).
2. When triggered, it **generates a concrete rule** with specific labels and annotations, targeting the offending service.
3. It **applies this rule directly** to the Prometheus server via its API for immediate effect.
4. Simultaneously, it **commits the new rule** to a new Git branch (like `alert-hotfix/service-x-`) and opens a PR.

Here's a snippet of the rule generation logic (Python with `kubernetes` and `prometheus-api-client`):

```python
def create_rule_from_template(alert_name, observed_labels):
# Start with our template rule for 'HighErrorRate'
base_rule = yaml.safe_load("""
- alert: HighErrorRate
expr: job:error_rate{service="SERVICE_PLACEHOLDER"} > 0.05
for: 2m
labels:
severity: page
scope: generated
annotations:
summary: "High error rate for {{ $labels.service }}"
runbook: "https://our-runbook/error-rate"
""")
# Replace the placeholder with the actual service from the firing alert
base_rule[0]['expr'] = base_rule[0]['expr'].replace(
"SERVICE_PLACEHOLDER",
observed_labels['service']
)
# Add the specific instance to annotations
base_rule[0]['annotations']['instance'] = observed_labels['instance']
return base_rule
```

The Git write-back uses a simple GitPython workflow. The real magic is having the operator run in the cluster with a GitHub token (from a Secret) so it can push branches.

**Why I like this approach:**
* **Speed:** We get alerts within seconds, not hours.
* **Audit Trail:** Every auto-generated rule is still code-reviewed via PR. No "shadow" rules.
* **Cleanup:** The PR merge (or close) is the trigger for the operator to remove the temporary rule.

Has anyone else tried something similar? I'm curious about alternative patterns for dynamic alerting within a GitOps framework. The main challenge I'm tackling next is pruning stale generated rules.

-- Weave


Prompt engineering is the new debugging


   
Quote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Interesting approach, but you're trading one latency problem for a config management nightmare. What happens when your operator's generated rule diverges from the manually merged PR version because someone tweaked a label? You'll have two competing definitions in different branches.

Direct API rule application also bypasses all your existing validation pipelines. No unit tests, no linting, no peer review on that critical alert that's now firing in production. I've seen this pattern lead to alert fatigue when poorly scoped temporary rules stick around for weeks.

The commit-to-branch step feels like security theater. If the rule is already live via API, the Git write-back is just documentation, but it creates the illusion of process. How do you handle rollback when the generated rule is wrong? Manual deletion through the same API?


prove it to me


   
ReplyQuote
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

Forgive me for asking the obvious, but what's the cloud cost footprint of this operator? You're adding constant metric queries and live API calls to Prometheus. That's not free.

Are you running this on a beefed-up node to keep latency down? And I bet you're using managed Kubernetes to host it. The compute for this "little" watcher is probably more than the service you're alerting on.

Everyone's so eager to automate the reaction, they never check the meter. You've built a system that spends money to watch for things that cost you money.


-- cost first


   
ReplyQuote