Skip to content
Notifications
Clear all

Check out what I made: A Prometheus rule generator that writes back to Git

3 Posts
3 Users
0 Reactions
30 Views
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
Topic starter   [#14000]

Hey folks, I've been down a real rabbit hole this past month and I think I've finally surfaced with something genuinely useful. You know that pain point where your Prometheus alerting rules live in a Git repo, but then you have some external system—say, a service catalog or a cloud inventory—that *should* influence those rules? Manually updating YAML files every time a new service is onboarded is a recipe for drift and missed alerts.

So I built a thing: a controller that generates Prometheus recording and alerting rules based on external data, and then **commits and pushes those generated rules back to the Git repository** that feeds my GitOps setup (Flux, in my case). The loop is closed! The source of truth for *what to monitor* can now live outside the rule files themselves, but the rule files stay auto-generated and in Git, where all my other ArgoCD/Flux apps expect them.

Here's the core concept:
* It's a simple Go operator that watches a ConfigMap (or could be a CRD, API endpoint, database... you get the idea).
* That data source contains a list of services and their criticality tiers.
* The controller uses a Go template to render valid Prometheus `rule_files`.
* Then, it uses the Go Git client to clone the target repo, write the new files, commit, and push.

Here's a tiny snippet of the generation logic—the template that defines a generic latency alert:

```yaml
groups:
- name: {{ .ServiceName }}_latency_rules
rules:
- alert: {{ .ServiceName }}_HighLatency
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket{service="{{ .ServiceName }}"}[5m])) > {{ .LatencyThreshold }}
for: 2m
labels:
severity: {{ .Severity }}
tier: {{ .Tier }}
annotations:
summary: "High latency for {{ .ServiceName }}"
```

The magic, honestly, is in the Git write-back. I had to solve for:
* Authentication (using deploy keys with write access).
* Avoiding rapid-fire commits (debouncing changes from the source).
* Handling merge conflicts gracefully (a simple strategy: pull rebase before push).

I'm currently using it to automatically add SLI/SLO alerts for every new microservice that gets tagged in our catalog. No more "we forgot to add alerts for service X." The pipeline is:
1. Service gets added to catalog with a `tier: 1` label.
2. Controller picks it up, generates the relevant Prometheus rule group.
3. Controller commits `rules/service_xyz_latency.yaml` to the `monitoring` repo.
4. Flux sees the change in Git and applies it to the Prometheus server in the cluster.

I'm really curious if others have tackled this problem differently. Have you tried other approaches to dynamic rule management? Do you think the Git write-back is an anti-pattern, and would you instead push directly to the Prometheus server? I can see arguments for both sides—this keeps everything in GitOps flow, but adds a slight delay.

Next, I'm thinking of adding support for:
* Generating Grafana dashboards as JSON alongside the rules.
* Plugging in external data from cloud APIs (like automatically creating rules for each RDS instance).
* Maybe even a validation webhook to check the generated PromQL for syntax before committing.

The config and a rough early version of the code are in my personal repo. I'd love some eyes on it, especially around the Git operations and error handling. Has anyone built something similar? What pitfalls did you run into?


pipeline all the things


   
Quote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

This is exactly the kind of pain point I'm running into as we scale. That loop being closed, with the external source of truth actually writing back to the rule repo, is genius. It turns a static config into a dynamic one without breaking the GitOps model.

But I'm curious, how do you handle drift or conflicts? Like, if someone manually edits a generated rule file for a hotfix, does your controller just blast over it on the next sync? Or does it have some way to flag that? I'm imagining a "last modified by" annotation or something.



   
ReplyQuote
(@jasonl)
Eminent Member
Joined: 3 months ago
Posts: 23
 

Great question. That conflict scenario is the first thing I had to solve. The controller actually performs a three-way merge, comparing the generated content, the current Git state, and the original source data it used last sync.

It flags any files where manual edits introduce a semantic conflict with the intended generated rules, leaving a comment in the commit message and skipping the update for that specific file. It won't blindly overwrite, but it will persistently try to reconcile on subsequent runs if the manual edit is removed.


Data beats opinions.


   
ReplyQuote