Skip to content
Notifications
Clear all

Check out what I made: An agent that monitors error logs and creates Jira tickets.

22 Posts
21 Users
0 Reactions
102 Views
(@integration_maven_jane)
Reputable Member
Joined: 5 months ago
Posts: 156
Topic starter   [#22151]

Hi everyone. I've been deep in the weeds with AutoGen for a few months now, primarily exploring how to bridge monitoring systems with developer workflows. I wanted to share a project that’s been saving my team a significant amount of time: an agentic system that automatically monitors application error logs and creates detailed, actionable Jira tickets.

The core problem was our alert fatigue. We’d get a flood of error notifications, but triaging and translating them into properly scoped development tickets was still a manual, context-switching heavy process. This setup aims to emulate that initial triage step.

Here’s a high-level view of the agent group I configured:

* **Monitor Agent:** This is the "listener." It's configured to consume filtered error streams from our application monitoring tool (in our case, Sentry). It doesn't act on every single error, but uses logic to identify new, escalating, or critical error patterns.
* **Analyst Agent:** This is the "brain." When the Monitor Agent signals a ticket-worthy event, it passes all the context (error message, stack trace, frequency, user impact) to the Analyst Agent. This agent's job is to summarize the issue, identify the likely service/component, and suggest a priority based on predefined rules.
* **Jira Creator Agent:** This agent takes the structured output from the Analyst Agent. It formats a Jira ticket with a clear summary, description, labels, priority, and assigns it to the correct team's backlog based on the component mapping. It interacts directly with the Jira Cloud API.

The real magic is in the conversation workflow and the custom prompts that guide each agent. For instance, the Analyst Agent is prompted to always ask itself:
* Is this a new error signature or a recurrence?
* What part of the codebase is implicated in the stack trace?
* Based on keywords and our severity matrix, what's the suggested priority (P0-P3)?

Some pitfalls and learnings from getting this to run smoothly:

* **Information Overload:** The first version dumped the entire error payload into the conversation. We had to refine the Monitor Agent to extract only the most salient 10-12 lines of the stack trace and the key error message to keep the context window manageable.
* **Idempotency is Key:** You must build in checks to avoid creating duplicate tickets for the same recurring error. We implemented a simple cache of recent error signatures that the Monitor Agent checks against before initiating a new conversation thread.
* **Human-in-the-Loop Gate:** We initially had it create tickets automatically, but we’ve since added a final step where the proposed ticket is posted to a dedicated Slack channel for a team lead to approve with a simple emoji reaction before the Jira Creator Agent executes. This safety net has been invaluable.

The result isn't fully autonomous, and it shouldn't be. It's a force multiplier that handles the tedious parts of synthesis and data entry, freeing up the engineering team to focus on actual problem-solving. It’s been running for about six weeks and has created over 60 well-structured tickets we might have otherwise missed or delayed.

I'm happy to dive deeper into the specific agent configurations or prompt structures if anyone is interested. Has anyone else built similar bridges between monitoring/observability tools and task management systems? I'm particularly curious about alternative approaches to the deduplication problem.

~Jane


Stay connected


   
Quote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

This sounds really cool! I'm new to this kind of automation. How do you handle authentication between your agent and Jira? Do you store API keys in the agent config, or use something like secrets manager? Also, does the Analyst Agent ever create duplicate tickets for the same error? That's my biggest worry setting something like this up.


Still learning


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Secrets in the config is a major red flag in any audit. Use a vault or secrets manager, period.

For duplicates, you need a deduplication hash. Without one, yes, it will create endless tickets. The core logic is as basic as checking if the error signature already exists in a lookup table.


Trust, but audit.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

This sounds like a system that exists to apologize for another overly noisy system. You're basically automating the creation of busywork tickets from your Sentry alert spam.

Do your devs actually find these auto-generated Jira tickets useful? In my experience, they end up being another notification stream to ignore, filled with generic summaries that miss the actual root cause.

You built a whole "Analyst Agent" just to mimic what a decent monitoring dashboard with proper alert grouping and suppression rules already does. Why not fix the *first* problem instead of adding a second, more complex layer on top of it?


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

Hey, this is genuinely cool. We're actually wrestling with a similar issue, but it's on the marketing automation side - alert fatigue from email/webhook event streams that *should* trigger a customer journey step but fail due to weird formatting or API hiccups.

Your Analyst Agent approach feels really familiar. We've been experimenting with a "workflow triage" agent that sits between our send failure logs and our internal task board (we use Linear). It doesn't just summarize, it tries to categorize the failure bucket: is this a Mailgun parsing issue, a malformed merge tag from the CRM, or a SPF/DKIM config hiccup that needs ops to look at? That categorization step alone has saved so much "wait, whose problem is this?" pinging.

One thing I'm curious about, because it's our constant battle: how do you handle the quality of the ticket description? We found the first drafts from the agent were too technical. We had to coach it to include the *business* impact, like "This is blocking the welcome sequence for 2% of new sign-ups," not just "HTTP 422 error on endpoint X." Did you run into that?


don't spam bro


   
ReplyQuote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
 

>how do you handle the quality of the ticket description?

You're right, it's the whole game. The business impact prompt is the *only* thing that makes these tickets not garbage.

But you're solving the symptom again. If you have to prompt-engineer the agent to ask "how many sign-ups are blocked?" then your monitoring stack isn't surfacing that metric in the first place. You're using the LLM to glue over an observability gap.

We hardcoded impact tiers based on log volume and user cohort tags that already existed. No "coaching" needed if your logs are structured right. Your Mailgun parsing issue should already be tagged with the campaign ID and audience size.

Otherwise you're just building a fancy, expensive guesser.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You've posted a code snippet in your thread starter. Please edit the post to use a proper code block so the formatting is preserved for others. Use the code button in the toolbar.


Beep boop. Show me the data.


   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

So you're essentially automating the creation of Jira noise because your monitoring is too noisy to begin with. You've admitted the Analyst Agent needs "logic to identify new, escalating, or critical error patterns" - that logic should be in your monitoring tool's alert rules, not a downstream bot trying to clean up the mess.

What's your false positive rate on tickets? If this thing fires off a ticket that turns out to be a non-issue, you've just created more work, not less. The support and dev time wasted on bad tickets will quickly outweigh the triage time you saved.


Show me the data


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

That's a fantastic breakdown, and it directly hits the nail on the head for where the real work starts. The distinction between the Monitor and Analyst agents is the key piece.

I've found that the "logic to identify new, escalating, or critical error patterns" is the make-or-break layer. If you get that wrong, everything downstream is garbage. We ended up implementing a two-phase check: the first pass is a simple rule-based filter in the Monitor Agent (like, ignore anything below a certain error rate for a specific endpoint), and then the *second* pass uses the Analyst Agent to look for correlations the rules might miss.

It's that handoff where the Analyst tries to guess "is this a one-off blip or the start of a cascade?" that's both the most useful and the most headache-inducing. It sounds like you've got that same core concept nailed. How are you defining "critical" for your Monitor Agent's filtering? Is it purely volume-based, or are you feeding it any user or business metric thresholds?


Test, measure, repeat


   
ReplyQuote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

Absolutely agree on the vault point. It's non-negotiable for production.

On the deduplication hash, your basic logic is correct, but the trick is in defining that "error signature." A naive hash of the full error message can still create duplicates if timestamps or incidental IDs change. You need to strip out the variables first. We use a simple regex to pull out the static message template before hashing.

That lookup table also needs a TTL or a cleanup process, otherwise it grows indefinitely from unique, one-off errors.



   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

That two-phase approach makes a lot of sense. The handoff from simple rules to the correlation check is exactly the gap I'm trying to understand for our own alert streams.

When you define "critical" for that first filter, do you factor in the type of user affected? Like, if an error only hits free-tier users versus accounts on an enterprise plan, would that change the threshold?



   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

Yeah, that's a great question. We actually ran into that problem - our first filter only looked at error volume. But a glitch affecting a single enterprise customer's data sync is way more urgent than a thousand free-tier users seeing a UI flicker, even if the volume is lower.

We ended up tagging logs with a user segment field. The first filter rule checks if the error occurred in a segment tagged "enterprise" or "high-value." If yes, it uses a much lower threshold to pass it to the Analyst Agent for review.

Does your logging pipeline have a reliable way to attach user plan info to errors?



   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

True, but the bigger red flag is treating the vault as a magic bullet. I've seen teams vault their credentials and then leave the vault token in an environment variable, or bake it into a deployment script. You just moved the secret one layer down without solving the access control problem.

Your deduplication hash is the easy part. The hard part is maintaining that lookup table across deployments, team handoffs, and database purges. If your hash store gets wiped, you're back to ticket spam. And if you never clean it, you're stuck querying a million-row table for every single error.


Test the migration.


   
ReplyQuote
(@catherine)
Reputable Member
Joined: 3 months ago
Posts: 195
 

The vault token problem is precisely why we require ephemeral credentials for production service accounts. The vault isn't the endpoint, it's the broker for a JIT-issued credential with a 5-minute TTL. If your deployment process can't support that, you shouldn't be using an autonomous agent.

On the hash table, your point about maintenance cost is valid but incomplete. A million-row table is trivial for a key-value lookup if it's indexed correctly; the operational burden is in the cleanup logic, not the query performance. We pair the hash with an "error state" field and a last-seen timestamp. A monthly cron job expires any entry older than 90 days that hasn't transitioned from "active" to "resolved." The real failure mode isn't a wiped table, it's a schema change or data migration that invalidates the hash algorithm, causing a silent flood of duplicates.


Trust but verify.


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Totally get the "solving the symptom" angle, and you're right about needing proper tagging from the start. But in messy legacy systems, you don't always have clean logs with campaign IDs pre-attached.

Our workaround was a small enrichment step *before* the agent. We run the log through a quick lookup to attach user segment data, then feed that enriched context into the prompt. It's still a band-aid, but it gets us 80% there while we fight the bigger observability fight.

It's a fancy, expensive guesser for now, but it's better than manual triage while we fix the plumbing.


Trial first, ask later.


   
ReplyQuote
Page 1 / 2