Running a dev environment with Sumo Logic can get stupid expensive fast, especially if you're ingesting full application logs, debug-level stuff, or telemetry from non-production services that only run 9-to-5. I got tired of manually toggling expensive log sources or begging devs to remember to shut things off.
So I built a script. It uses the Sumo Logic API to find collectors and sources based on a tag (like `env=dev`), and enables/disables them on a schedule. This cut our dev ingest costs by about 65% last month. No magic, just automation.
Here’s the core approach:
* Tag your dev collectors and sources consistently in Sumo Logic. This is the key. Without a reliable filter, you're sunk.
* The script calls the `sources` endpoint, filters by tag, and then toggles the `enabled` state.
* I run it via cron on a lightweight container. A disable job runs at 8 PM, an enable job runs at 7 AM. Weekends are handled separately (disabled Friday night, enabled Monday morning).
You'll need an API access key with permissions to manage sources. The main loop logic is straightforward:
1. Get all collectors.
2. For each collector, get its sources.
3. If a source has the target tag (e.g., `env:dev`), set `enabled: false` (or `true`).
4. PUT the change back to the API.
Biggest pitfalls:
- API rate limiting if you have a huge number of sources. Add a small delay between calls.
- Not all source types support being disabled via API. Mostly works for local file and script sources.
- If your tags are a mess, this won't help. Clean those up first.
This isn't a silver bullet, but if you have predictable dev hours, it's an easy win. I can share the basic script structure if anyone's interested. It's just bash and curl.
—JW
—JW
That's a solid approach for Sumo, though it hinges heavily on source-level API control. In Datadog, you'd typically manage this at the ingestion pipeline with exclusion filters based on tags, which is more centralized and doesn't require toggling individual sources on and off. The cost control is baked into the pipeline rules.
You could achieve a similar scheduled effect by automating those filter rules via Terraform or the Datadog API, updating the `query` to exclude your `env:dev` resources during off-hours. It's a different layer of control, but it prevents the data from ever entering the platform, which can be cleaner than disabling collection at the source.
null
This is brilliant. The weekend handling is such a smart addition. I'm going to pitch this exact approach at work, though I'm a bit nervous about our tagging being consistent enough.
Do you ever have issues with the script if a source is in use or locked when it tries to toggle? Or does the API handle that gracefully?
Learning every day
Yeah, the pipeline filtering approach in Datadog is a great point. It does feel cleaner to stop data at the gate rather than turn off the tap after it's already flowing.
I've found that while central pipeline rules are simpler to manage, they can sometimes make devs feel less accountable for their own spend, since the cost control is invisible to them. A source toggle, even if clunkier, keeps that cause-and-effect connection visible.
Have you ever run into pushback from dev teams when implementing blanket ingestion filters? Like, if they suddenly need to debug something off-hours, is there a quick override process, or does it require a pipeline change?
sales with substance
Your point about dev accountability is crucial. We had a similar debate when centralizing our Datadog exclusion filters. The compromise was implementing a two-tier tagging system: every service requires an `env=dev` tag, but we also introduced a discretionary `debug_enabled` tag. The pipeline filter excludes data where `env=dev` AND `debug_enabled != true`. This gives teams a self-service, temporary override by simply adding that second tag to their service, which takes effect in the next pipeline scan.
It requires discipline, but we built a small CLI tool to manage the `debug_enabled` tag and log its usage for review. This keeps the primary control centralized and automated while providing that visible, opt-in mechanism for off-hours debugging. The friction of having to actively enable it does seem to correlate with more responsible use.
Data over dogma