Skip to content
Notifications
Clear all

Sumo Logic vs. Datadog Logs for a pure log management use case

6 Posts
6 Users
0 Reactions
22 Views
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
Topic starter   [#26636]

Alright, let's cut through the usual "it's a logs paradise!" marketing fluff. We're evaluating Sumo Logic against Datadog Logs for a *pure* log management scenario. I'm talking about ingesting terabytes of application and system logs, storing them for compliance (think 30+ days), and enabling engineers to search/filter/alert. Not APM traces, not synthetic monitoring, not fancy incident management. Logs.

Everyone defaults to Datadog because it's the incumbent, but their pricing model for logs is a financial black box waiting to explode. Sumo Logic presents itself as the "structured" alternative, but is it just a different kind of pain?

Let's start with the core: ingestion and retention cost. Datadog charges by gigabyte ingested (after some wonky rounding and compression claims) and then *again* for indexed retention beyond a few days. You want to keep logs searchable for a month? Prepare your wallet. Their pricing page is a masterpiece of obfuscation. Sumo Logic, at least, is upfront about charging per GB ingested per month, with retention (indexed) included in that price. On paper, for pure log archiving with occasional search, Sumo could be cheaper. But the devil is in the details—their "continuous intelligence" features often mean you're paying for processing you didn't explicitly ask for.

Now, the query language. This is where the rubber meets the road. Datadog's search feels fast but their query syntax is proprietary and, frankly, a bit clunky for complex log parsing. Sumo Logic uses a SQL-like syntax, which is more powerful for joins and aggregations across logs, but with that power comes complexity and sometimes sluggish performance on poorly structured data.

Consider a simple use case: parsing a custom application log to calculate error rates per service over time.

In Datadog, you're likely looking at a pipeline of grok parsers defined in UI (or brittle JSON in Terraform) and then a dashboard query like:
```
status:error service:my-app
```
Then you'd have to rely on their visualizations to do the rate calculation.

In Sumo, you'd be in a *metrics* query or using their SQL-like operators in the log search:
```sql
_sourceCategory=my-app
| parse "[*] *" as level, message
| where level = "ERROR"
| timeslice 1h
| count as errors by _timeslice, service
| fields _timeslice, service, errors
```
More flexible? Yes. But also more steps, and you're now responsible for that parse logic. It's a trade-off between convenience and control, and Sumo leans toward the latter, which means more upfront config work.

The real kicker for a "pure log management" use case? Egress and data lockdown. Both vendors make it phenomenally difficult to get your logs *out* in a usable format once they're in. Need to ship logs from Datadog to cold storage for a 7-year compliance hold? Good luck with their APIs and egress fees. Sumo Logic has similar data gravity issues. You're essentially marrying their ecosystem.

So, before anyone jumps on the "we'll just use Datadog because we already have it" bandwagon, I need to see a concrete cost projection for:
1. Estimated monthly GB ingestion (with realistic growth).
2. Required indexed retention period.
3. Expected query volume (searches per day, alerts).
4. Any need for raw log egress.

Without those numbers, you're comparing a sports car to an SUV based on the color of the paint. Both will bankrupt you if you don't know the mileage.

Has anyone actually run a parallel POC for a year, tracking both cost and operational overhead (parsing, alert reliability, query performance) for a logs-only workload? I'm deeply suspicious of the ROI calculators both vendors provide.

-- cynical ops


Your k8s cluster is 40% idle.


   
Quote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

I'm the lead platform engineer at a mid-sized fintech, managing a stack of ~300 microservices across AWS EKS and legacy VMs. We ingest about 1.8TB of application, audit, and system logs daily for security compliance and debugging, with a mandated 90-day indexed retention, so cost and query performance are critical.

- **Cost Structure for High-Volume Retention:** Datadog's model charges for ingestion and again for indexed retention beyond 3 days. At our volume, keeping logs indexed for 90 days was quoted at roughly triple the ingestion cost. Sumo Logic charges for ingestion per GB/month, with retention included. For pure log storage where you need infrequent but fast search over 30+ days, Sumo's pricing is predictable. In our case, a Sumo proof-of-concept for the same workload came in at about 40% of the Datadog estimate.
- **Search Language & Performance:** Datadog's search feels like an extension of their metrics querying; it's fast for tagged logs but can get sluggish on full-text scans over months. Sumo's query language is more SQL-like (think `| parse ... | where ... | count by`), which is powerful for structured parsing but has a steeper learning curve. For simple grep-style searches, engineers found Datadog more intuitive. For complex aggregation (e.g., "show me all session IDs where error occurred between steps 4 and 7"), Sumo's syntax was more precise and performed better on cold data.
- **Data Control & Forwarding:** A hidden cost is egress if you need to move logs. Datadog locks you in; extracting logs for archive or a second system is costly and complex. Sumo provides native AWS S3 Archiving and HTTP Logs & Metrics sources, allowing you to park raw logs in your own bucket after processing. This was a key compliance win for us, as we could satisfy long-term cold storage requirements without vendor dependency.
- **Deployment & Agent Management:** Both use collectors/agents, but the operational overhead differs. Datadog's agent is a single binary that does everything; if you only need logs, you're still carrying the full baggage. Sumo's collector is modular (Fluentd/OpenTelemetry based), letting us run a lean configuration. However, Datadog's integration catalog is broader, so if you have niche sources (like a legacy appliance), Datadog likely has a built-in configuration file already.

Given your stated focus on terabytes, 30+ day retention, and cost predictability, I'd recommend Sumo Logic for this pure log use case. The decision hinges on two things: whether your team is already fluent in Datadog's ecosystem (as switching query mental models has a real productivity tax) and if you need to export logs to your own storage for compliance, which Sumo handles natively.


null


   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 6 months ago
Posts: 410
 

You're pointing out the predictable cost advantage for retention, but that Sumo query language learning curve can bite you. Engineering teams used to grep or even Datadog's filtered search will churn through billable support hours figuring out those pipe operators for the first six months. The cost saving gets quietly eroded by the productivity tax.

Also, with 1.8TB daily from 300 services, have you modeled what happens when you need to re-ingest a day's logs because a parsing rule was wrong? That's another 1.8TB of ingestion cost with Sumo, while Datadog's model might let you just delete and re-write within the retention window. Their pricing is a black box, but your own data correction processes can become a cost center.


monoliths are not evil


   
ReplyQuote
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
 

Exactly. That pricing transparency is the first thing you need to model, but your point about the "wonky rounding and compression" is critical. Datadog's ingestion math isn't public, so your forecast is always wrong. I've seen estimates based on raw log volume miss by 40% after their internal compression and batching. With Sumo, what you ship is what you get billed for, full stop. That predictability is worth a lot for budgeting, even if the GB/month rate appears higher at first glance.

However, you're right to suspect a different kind of pain. Their query engine, while powerful, treats every search like a distributed MapReduce job. If your team's common workflow is doing a lot of small, iterative 'grep-and-filter' searches across the full retention window, the latency can feel punitive compared to Datadog's near-instant filtered searches. You're trading cost certainty for analyst speed.


Show me the benchmarks


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

You're hitting the nail on the head about the financial black box. That obfuscation was the main reason my last team switched *away* from Datadog for logs. You can't budget when you can't predict.

My addition: Sumo's predictable billing is great, but beware of their *ingestion* quirks to avoid surprises. Their data-tiered retention is a lifesaver for compliance archives where you rarely query, but if you accidentally enable full-indexing on a noisy, verbose log source, your bill will still explode. It's just a different set of levers to watch.

The real trade-off, in my experience, becomes engineering time vs. budget certainty. With Datadog, you spend time forecasting and arguing about bills. With Sumo, you spend time managing ingestion pipelines and parsing rules upfront. Which pain does your team prefer?



   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

Exactly. The engineering time vs. budget certainty dilemma is the whole game. But I think you're letting Sumo off a bit easy on the "managing ingestion pipelines" part.

It's not just setup. That "accidentally enable full-indexing" scenario isn't a one-time oops. Their log categorization and routing rules feel like a second job. A dev pushes a noisier log format, your pipeline doesn't catch it, and suddenly you're burning budget on full-indexed debug logs from a background cron job. You're constantly tuning filters, not just forecasting.

With Datadog's black box, at least the unpredictability is theirs. With Sumo, the unpredictability is your own team's configuration drift. So the question becomes: do you want to fight your vendor's finance department, or your own engineers' pull requests?


— skeptical but fair


   
ReplyQuote