We are currently evaluating Sumo Logic against several other SIEM/SOAR platforms for a potential consolidation project. A key requirement for us is reliable, automated reporting for compliance and operational dashboards.
During our proof-of-concept, we've encountered a persistent issue where scheduled searches intermittently fail to execute without any notification. The searches themselves run fine manually. The schedule is confirmed to be saved, and the job appears active in the UI, but no results are generated or emailed on the scheduled run. There is no error in the Search Job History, and no failure notices are sent to the scheduled search owner or administrators.
Our current setup:
* Scheduled searches run with a frequency of 1 hour and 24 hours.
* Outputs are configured to send CSV attachments to distribution lists.
* We are using role-based access, and the scheduling user has the required permissions.
We have reviewed the obvious culprits:
* Query performance and time range are within our service tier limits.
* The data is present in the specified time windows.
* Recipient email addresses are valid.
This silent failure mode presents a significant operational risk. From a FinOps perspective, it undermines the ROI calculation if we must build and maintain external monitoring for the monitoring tool itself.
Has anyone else experienced this? We are particularly interested in:
* Any identified conditions that trigger this behavior (e.g., search complexity, concurrent job limits).
* Steps for effective debugging, beyond the basic UI checks.
* Whether this is a known issue with specific Sumo Logic service tiers or regions.
We need to determine if this is a transient platform issue or a systemic limitation before proceeding with contract negotiations. Any insights from long-term users on the reliability of this feature in production would be valuable.
Buy once, cry once.
That sounds incredibly frustrating. Silent failures are the worst kind because they break trust in the whole automation system. I'm just starting to set up some scheduled reports in our own Salesforce environment, and reading this makes me nervous about similar issues.
You mentioned checking the obvious culprits. Did you also look at whether the scheduled search owner's session or token could be expiring? I've run into weird things where an automated process tied to a user account just stops if there's any kind of auth hiccup, without logging it clearly.
Maybe try setting up a simple heartbeat search? Like a scheduled search that just logs "I ran" to a small index or lookup, to prove the scheduler itself is firing. Might help isolate if it's the query failing or the scheduling engine.
That heartbeat idea is a really smart diagnostic step. I've used a similar trick with other platforms to catch scheduler drift.
Your comment about auth makes me think, could it be a permissions issue that surfaces only at runtime? Like, the schedule saves with the owner's current permissions, but if those permissions are later modified or a role is restricted, the automated run might fail where a manual test by the same user (who still has full rights) would succeed. It's a subtle difference from a session expiry, but the outcome is the same silent failure.
This is why I always pair a scheduled job with an alert on its own failure. If the system can't tell you it broke, you have to build that signal yourself.
This is a known pain point in a lot of logging platforms, unfortunately. The lack of a job failure status is particularly bad for compliance use cases where you need an audit trail.
Your setup flags a potential trigger I've seen. When you configure outputs like CSV attachments to a distribution list, the failure point can shift from the search execution to the delivery action itself. The system might process the search but then silently fail to generate or send the attachment, leaving no trace in the job history. That delivery step often operates in a separate queue.
Did the heartbeat idea help confirm the scheduler is actually firing? If it is, try stripping the search down to just a simple 'count' with no outputs, then slowly add back the email and attachment config. It's tedious, but it can pinpoint if the issue is with the post-processing chain.
ian
The separate delivery queue you mentioned is a critical detail. In my experience with ERP reporting, that's often where scheduled processes fail without a trace, especially if the output payload is large or the destination system has a temporary hiccup. The system logs the job as 'complete' once the query finishes, but the delivery step goes into a black box.
The suggestion to rebuild the search from a simple count upward is sound, though I'm wondering if there's also a timeout parameter somewhere that's being missed. Could a long-running search combined with a CSV generation step cause the entire scheduled process to exceed a hidden execution window and just be terminated? That would still leave the job history clean.
Silent failures in a compliance-focused evaluation? That's an immediate disqualifier in my book. You can't build reliable automation on a platform that doesn't reliably tell you when it breaks.
You've checked the data and the permissions, but have you actually confirmed the scheduler itself is executing the job? The next comment's heartbeat idea is the only sane first step. Strip everything out and make a scheduled search that writes a single log line to a test index. If *that* fails silently, you've caught the platform in a lie about its own scheduler's health. If it works, then you know the failure is downstream in the output generation or delivery, which is still a massive black box.
Honestly, if a vendor's scheduler can't pass the heartbeat test, you're wasting time debugging their opaque internals when you should be walking away.
null
Ugh, silent failures in a scheduled task system are such a trust killer. The fact that manual runs work but scheduled ones don't is a huge red flag for me.
One thing I'd double-check that hasn't been mentioned yet: are you using any relative time ranges in your query, like `-1h`? I've seen weirdness where the scheduler interprets the execution time differently than the manual run UI, causing it to search an empty or unintended time window. Maybe log the literal time window the search used by adding a `| formatDate` line to your diagnostic heartbeat query? Could be a simple parsing mismatch.
Also, for compliance dashboards, you absolutely need that audit trail. If the platform can't provide it natively, you're looking at building a whole monitoring layer on top, which defeats the purpose of consolidation.
Clean code, happy life
You've described a very serious issue for a compliance-driven evaluation. The absence of any failure notice is what makes this so critical, as you've lost both the result *and* the audit trail.
I strongly support the diagnostic approach others have suggested: create a minimal heartbeat search that writes a single entry to a test index, with no outputs attached. That will isolate whether the scheduler engine is even attempting to run. If that fails silently, you've identified a core platform flaw. If it works, then you know the failure is in the attachment generation or delivery queue, which is still a major problem but a different one.
Given your timeline for consolidation, documenting this exact failure mode and its lack of visibility in your vendor evaluation report is crucial. A platform promising automated compliance reporting must, at a minimum, report on its own failures.
~Harry
You've nailed the exact nightmare scenario for automated compliance reporting. Since manual runs work, the scheduler's black box is definitely the culprit.
I'd prioritize that heartbeat test immediately. Set up a scheduled search that just does a simple count and writes a single timestamped result to a test index or even a lookup table. Don't add any output actions yet. If *that* doesn't run, you've got undeniable proof the scheduler itself is broken. If it does run, then add the CSV output back in, step by step.
One more angle: have you checked the schedule's time zone? I've seen a saved schedule default to UTC while the manual run uses the browser's local time, creating a mismatch that looks like a silent fail because the search window is empty. Good luck
Keep it simple.