Having participated in numerous vendor proofs-of-concept, I have found that the most common point of failure is an inadequate assessment of log ingestion throughput. Teams often rely on vendor-provided synthetic tests or extrapolate from minuscule samples, leading to catastrophic miscalculations in projected costs and performance once the platform is under actual production load. The core issue is that ingestion pipelines are sensitive to log volume, structure, and cardinality in ways that are not always linear.
To design a rigorous throughput test, you must replicate your production environment as closely as possible within the constraints of a POC. This involves several critical components:
**1. Test Data Fidelity**
* **Source:** Do not use generic Apache logs. Use a representative sample of your actual application logs, ideally anonymized if necessary. The log format, average line length, and field diversity directly impact parsing overhead.
* **Volume & Burst:** Calculate your average logs-per-second (LPS) but, more importantly, identify your 95th and 99th percentile burst rates. Your test must sustain the average and withstand the bursts.
* **Cardinality:** Ensure your test data includes the high-cardinality fields you expect in production (e.g., `trace_id`, `user_id`, `request_id`). High cardinality can dramatically affect ingestion performance and costs on many platforms.
**2. Test Infrastructure & Tooling**
* Your test client should be isolated from the system under test to avoid resource contention. Use a dedicated machine or instance.
* For controlled, repeatable tests, a tool like `logcli` (for Loki), `fluent-bit` in benchmark mode, or a custom script using the vendor's SDK is necessary. The goal is to push logs at a precise, measurable rate.
* Monitor the client's resource usage (CPU, network) to ensure it is not the bottleneck.
**3. Measurement Methodology**
* Instrument the test to collect the following metrics:
* **Client-side:** Logs sent per second, error rates (backpressure/retries), and network egress.
* **Platform-side (via vendor API/dashboard):** Logs ingested per second, throttling events, ingestion latency (acceptance to queryability).
* **Cost:** If the platform provides per-ingestion pricing, calculate the effective cost per GB based on your test payload.
* Run tests in phases:
* **Ramp-up:** Gradually increase LPS to find the point where errors or latency increase non-linearly.
* **Sustained:** Run at your target average LPS for a minimum of 30-60 minutes to identify stability issues.
* **Burst:** Spike to your 99th percentile LPS for 5-10 minutes.
Here is a conceptual example of a simple test harness using a script and `curl` to post logs, though you would ideally use the vendor's official ingest client.
```bash
#!/bin/bash
LOG_FILE="representative_logs.ndjson"
TARGET_URL="https://ingest.vendor.com/api/logs"
AUTH_HEADER="Authorization: Bearer $API_KEY"
RATE_LPS=10000
DURATION_SECONDS=3600
# Read logs in a loop, controlling the rate
log_generator /dev/null &
done
```
**4. Key Success Criteria**
The test is not merely about whether the platform accepts the logs. You must verify:
* **No data loss:** All sent logs are eventually queryable.
* **Predictable latency:** The time from ingestion to availability remains within your SLO (e.g., under 30 seconds) under sustained and burst load.
* **Cost alignment:** The observed ingestion volume matches the vendor's metering, and the projected cost aligns with your FinOps model. Pay particular attention to how compression, batching, and indexing affect the billed volume, as it often differs from raw byte count.
Ultimately, a well-structured throughput test will expose weaknesses in the vendor's ingestion pipeline, clarify true costs, and provide the data necessary to negotiate committed-use discounts or Savings Plans with confidence, as you transition from a POC to a production deployment.
Spreadsheets or it didn't happen.
We run a microservices platform for a mid-market fintech, handling about 20 TB of logs daily from ~400 production pods across multiple AWS regions, using a combination of Fluent Bit, OpenSearch, and Loki.
- **Test Data Generation**: You must use real log shape. We wrote a Go producer that reads anonymized production log samples and replays them at a target LPS, varying the cardinality of key labels (like `user_id` or `transaction_id`). This showed us that a vendor's agent choked on high-cardinality JSON fields, cutting throughput by 60% compared to their synthetic test.
- **Pressure Test the Full Pipeline**: Isolate the ingestion endpoint and hammer it. In our last POC, we used a dedicated Kubernetes namespace to deploy multiple instances of the log agent, each fed by a custom generator. We found the vendor's managed intake service started discarding logs at about 12k EPS, while their docs claimed 20k EPS per node. The specific detail was network egress costs from our cloud spiked, which wasn't in their calculator.
- **Measure the Actual Cost Driver**: For SaaS, don't just track EPS; track ingested GB/day. In one test, our gzipped JSON logs expanded 3.2x on ingestion due to indexing, turning a projected $3k/month bill into nearly $10k. We now always run a 48-hour sustained test matching our peak day profile and request the detailed usage report from the vendor.
- **Validate the Control Plane**: Test scaling actions. With a major vendor, triggering an index rotation during a sustained 8k EPS load added 2 seconds of latency across the board for 5 minutes, which would violate our SLOs. The fix required a custom buffer configuration they didn't mention in the POC guide.
I'd recommend building a custom test harness with `docker-compose` or a local K8s cluster first to baseline your log characteristics. Then, for the vendor POC, you need to test their agent and their intake. Tell us your average log size and whether your bursts are regional or global to get a clearer recommendation.
Absolutely correct on cardinality. It's the silent killer in these tests.
The biggest miss I see is teams not testing for cardinality explosions during incidents. Your normal production sample might have modest cardinality, but a cascading failure can spike unique error IDs or request IDs by orders of magnitude. This can flatline an ingestion pipeline that was fine with "representative" data.
Your POC test must include a cardinality stress phase. Generate logs with a ramp of unique field values that mimics your worst-case incident scenario. If the vendor can't show you the ingestion rate curve as cardinality climbs, walk away.
Five nines? Prove it.
Spot on about the burst rates. Too many POCs just aim for the average and call it a day, but that 99th percentile spike is exactly when you need your observability platform to be rock solid, not falling over.
One caveat from moderating these discussions: be crystal clear with the vendor on what "sustaining" a burst means. Is it just ingesting without loss, or does it also include the logs being queryable within an acceptable SLA? We've seen setups that buffer the burst beautifully but then take hours to index it, which defeats the purpose during an incident.
Stay constructive
Precisely. The buffer-and-delay scenario you've described is a common failure mode that often gets obscured in vendor metrics. I'd insist on defining a clear "time to query" SLA as part of the burst test. For instance, during a test, you might see ingestion rate hold steady at 100k logs/sec, but if the 95th percentile log takes 15 minutes to become available in a simple `error` filter, the system is functionally degraded.
To test this, your load generator needs correlatable timestamps and a separate, low-volume query client that polls for known log entries post-ingestion. The gap between the ingestion timestamp and the successful query timestamp is your observable latency. You can't manage what you don't measure, and ingestion rate alone is an incomplete metric.
every dollar counts
Your point about measuring the "gap between the ingestion timestamp and the successful query timestamp" is the core operational metric most POCs miss. I call this "Indexing Latency at Ingest," and it's crucial to measure its distribution, not just an average.
In a recent benchmark, we saw a system hold a steady 80k logs/sec, but the 99th percentile log took over 300 seconds to become queryable due to a hidden internal queue for field extraction. The mean was a deceptive 2 seconds. This is why you must capture a histogram of that queryable latency during the test.
A practical method is to embed a unique, random UUID in each log payload from your generator and have a separate, time-synchronized process perform idempotent GETs for a sample of those UUIDs. Plot the latency percentiles against the concurrent ingestion rate. If the curve doesn't stay flat, the system is trading ingestion stability for queryability, which is a critical architectural flaw.
numbers don't lie
The histogram approach for "Indexing Latency at Ingest" is the correct diagnostic tool. However, the UUID polling method you describe can inadvertently create a "quiet probing" effect that fails to capture the system's behavior under a true concurrent query load, which is the realistic stress condition.
The latency distribution for a handful of background GET requests will often differ significantly from the latency experienced when the platform's query interface is under load from actual users or automated dashboards refreshing during the incident. You need to simulate that concurrent query pressure. A valid test should add a background query workload that mimics your typical investigative patterns - say, a mix of full-text searches and field filters - while the high-cardinality burst ingestion is ongoing.
Without this, you're only measuring the pipeline's best-case query latency, not the latency under operational duress when the data is needed most.
Nullius in verba
Wow, that "hidden internal queue for field extraction" bit hits close to home. We saw something similar where logs would ingest fine but searching for a specific container_id would timeout for ages. The mean was totally fine, just like you said.
Your UUID polling method sounds smart, but how do you keep that query load from affecting the ingestion performance you're trying to measure? Isn't there a risk of the polling itself adding load to the query nodes and making the latency look worse?
You're right about the sensitivity to log shape, but you're skipping the most common trap in "representative sample" selection. Teams grab a quiet Tuesday afternoon snapshot and call it a day. That sample has none of the pathological edge cases that actually break parsers.
You need to deliberately include the malformed, the oversized, and the oddly structured logs that inevitably slip through your logging libraries. I once saw a vendor's grok parser choke for 30 seconds on a single log line with nested JSON inside a quoted string that our real app occasionally spits out. Their synthetic test data was all clean RFC5424 stuff. The throughput numbers from their demo were off by a factor of four once we fed it a week's worth of real, dirty logs.
And don't just anonymize; you have to preserve the byte distribution of the original fields. A simple string replacement that turns a 16-character user ID into a 16-character hash is fine. Turning it into a sequential integer wrecks the compression and indexing behavior you're trying to test.
Ugh, the "real, dirty logs" point is so true. Everyone's test data is pristine, and then the vendor tries to upsell you on their "Advanced Parse Engine" module to handle the mess your actual apps make. I've seen that add-on cost more than the base ingestion tier.
But let's be honest, if your platform can't handle a nested JSON-in-a-string without a 30-second grok choke, it's broken. You shouldn't need a premium SKU for that.
Preserving byte distribution is the sneaky bit most people miss, though. I'd also argue you need to preserve the *ordering* of fields in your JSON logs. Some parsers rely on field order for performance, and reshuffling everything alphabetically during anonymization gives them an unrealistic boost.
—DW
You've nailed two crucial aspects of using "real" data: the vendor upsell trap and the hidden performance cliff of field ordering. It's frustrating when core reliability is gated behind a premium SKU.
Your point about field ordering is subtle but can be a massive differentiator. I've seen vendors whose documentation quietly recommends alphabetical ordering for "optimal performance," which is a red flag that their ingestion pipeline can't handle the entropy of a real environment. If they require sanitized, normalized data to hit their throughput numbers, those numbers are marketing fiction.
This is exactly why the cardinality stress tests mentioned earlier are so important - they combine with your "dirty logs" point to create a truly representative worst-case scenario. A high-cardinality incident rarely produces clean, well-ordered logs.
Stay curious, stay critical.
Oh wow, the point about alphabetical ordering being a "red flag" is something I'd never have thought to check. That's scary.
I'm pretty new to this, but that makes me wonder: how do you even test for that? Do you just take a random sample of your production logs and see if the field order is all over the place? I guess you'd have to capture that "entropy" first.
And then you'd have to make sure your test suite actually sends it over in the same messy order... which sounds tricky.
Your point about cardinality is exactly where things get slippery. Everyone talks about volume and burst, but I've heard teams get absolutely wrecked by an explosion in unique values for a single field, like a user ID that gets logged in a ton of places.
If you're going to **ensure your test data includes realistic cardinality**, how do you actually simulate that growth over a test run? You can't just replay a static sample, because cardinality is about the rate of new values appearing. Do you need to script your load generator to inject an increasing percentage of new, unique identifiers as the test progresses to model a real day's accumulation?
Absolutely you need to simulate growth! A static replay just measures how well they handle a snapshot, not the accumulation.
I script my load tests to gradually increase unique values in a key field (like request_id) over the test duration. For a 1-hour test, I might start with a pool of 10k and add 500 new unique values every minute. That mimics a real day's new users or sessions.
Big caveat: watch out for vendor demos that use a "high-cardinality" dataset that's actually just *large*, not *growing*. If their cardinality count is flat from minute 1 to minute 60, they're cheating.
Trial first, ask later.
Great point about scripting the growth. That's the only way to see how indexing strategies handle the accumulation over time.
One thing I'd add: you also need to watch for cardinality "spikes" within your growth curve. Real traffic isn't linear. A batch job might mint 10k new unique IDs in 30 seconds, then nothing for an hour. So I'd modify that script to occasionally inject a burst of new values, not just a steady trickle. That's often where you see memory pressure or indexing lag really jump.
Totally agree on the vendor demo cheat. A flat cardinality chart is a dead giveaway they're just replaying a static file.
ship it