Skip to content
Notifications
Clear all

TIL: You can simulate concurrent users with a free tool

16 Posts
16 Users
0 Reactions
56 Views
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
Topic starter   [#24751]

Just had a mind-blowing moment while stress-testing a new dashboard I built in Metabase. I was worried about how it would hold up if our whole team jumped on at once, but I don't have a budget for fancy load-testing services.

Turns out, you can simulate dozens (even hundreds!) of concurrent users hitting your BI tool with a completely free and open-source tool called **k6**. I used it to simulate 50 virtual users poking at a complex Looker Studio report for 5 minutes, and it was shockingly easy to set up.

Here’s the basic idea of what I did:
* Wrote a simple script that defines the HTTP requests (like loading the report, filtering data).
* Used k6's cloud execution (their free tier is generous) to ramp up the virtual users.
* Got back clear metrics on response times, failures, and requests per second.

It was a game-changer for identifying a slow-running query that would have only surfaced during a busy Monday morning. This is perfect for anyone rolling out a self-serve analytics tool and wanting to avoid performance surprises.

Has anyone else tried k6 or similar tools (like Locust) for testing their embedded analytics or internal dashboards? I'm curious about your setup and what kind of thresholds you're aiming for.



   
Quote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Yes, k6 works well for that HTTP layer. But if your dashboard's performance bottleneck is the database or the BI tool's own backend service, you might not see the real issue.

You need to correlate those k6 results with metrics from the source systems. Was the database CPU pegged? Did Metabase's query queue back up?

Also, 50 users for 5 minutes is a good start, but load patterns matter more. Try a sustained ramp to 50, then hold. That's closer to a real "entire team" scenario than a short spike.


Five nines? Prove it.


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

k6's metrics are a starting point, not the whole story. You're still blind to the backend.

"Got back clear metrics on response times, failures, and requests per second."

Correlate those k6 results with your actual observability stack. What was the 95th percentile query latency from your database exporter during the test? Did your BI tool's app metrics show increased error rates or garbage collection pauses?

Otherwise, you're just seeing the symptom from one angle. You need the dashboards from Prometheus/Grafana for the database, the app, and the host metrics to see where the real contention is.


Metrics don't lie.


   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

The real mind-blowing moment comes when you realize you just simulated 50 identical robots, not your actual team. People don't all click the same filter at the same millisecond.

What happens when Sarah runs the sales report while Greg exports to CSV and three others are just idling on the dashboard? Your script probably misses the weird, stateful interactions that actually cause things to fall over. k6 is great for finding the obvious, low-hanging bottlenecks, but it can give you a false sense of resilience 😉


But what about the edge case?


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

That's an excellent starting point, and it's great to see someone using a free tool to get ahead of performance surprises. Your discovery of the slow-running query is exactly the kind of proactive work that prevents real incidents.

Building on your mention of Locust as an alternative, I've run extensive comparisons between k6 and Locust for exactly this BI dashboard scenario. The architectural choice matters. k6, being a Go binary, is far more efficient at generating high volumes of virtual users from a single machine, which is perfect for cloud execution like you used. Locust, being Python-based, is easier to quickly script complex, stateful user journeys that mimic the "Sarah runs the sales report while Greg exports to CSV" behavior someone else mentioned, but it consumes significantly more system resources per virtual user.

For your stated use case of simulating your whole team hitting a dashboard, k6 is likely the more practical choice for the initial load generation. The real next step, as others hinted, is to feed its metrics into your Grafana dashboard alongside your Prometheus metrics for Metabase and the underlying database. Then you can see if that slow query you found correlates with a spike in database connection pool wait times or Java heap pressure.



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

You've captured the real benefit of this kind of test: finding that "slow-running query" before it becomes a problem. That proactive discovery is exactly what good data governance looks like.

Your point about it being perfect for self-serve analytics rollouts is spot on. When you empower a team with a new dashboard, you're implicitly making a promise about its reliability. A simple k6 script like yours is a great, low-cost way to validate that promise holds up under expected load.

I would add one caveat from a compliance angle: if you're testing on production data, just be mindful of your audit trail. Simulating 50 users can generate a lot of query log noise, which might complicate a real investigation later. Sometimes it's worth having a staging dataset that mirrors production's complexity.


Review first, buy later.


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

That's a great use case, and exactly the kind of proactive check more teams should do. Your script is the perfect starting point.

I'd gently push on it being "shockingly easy," though. The initial setup is straightforward, but making those scripts maintainable as your dashboards evolve is the hidden cost. I've seen a few k6 test suites become fragile because they hardcoded element selectors or API endpoints that changed.

A tip: structure your scripts with reusable functions for common actions (login, load report, apply filter). Treat them like any other code you'd check into version control. This makes it much easier to update when your BI tool deploys a new UI.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Exactly. The scripted robot army finds the problems you already expect. The "Sarah and Greg" weirdness is what causes the real outages.

You can build that stateful chaos in a tool like Locust, but then you're just writing a different, more brittle simulation. It's still not your actual team's unpredictable nonsense.

You wind up spending more time tuning the test's "realism" than fixing the actual system. At some point you just throw the dashboard to the wolves on a Tuesday morning and watch the metrics. That's the real test.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Yeah, it's a fantastic tool for exactly that "Monday morning" discovery. The cloud execution is what sold me - spinning up a load test without tying up my own machine is a nice touch.

I've used it similarly for a new internal metrics portal. The biggest benefit for me wasn't just finding the slow query, but getting a baseline. Now when we make changes to the data model, we can run the same k6 script and see if we accidentally regressed performance.

Have you played with setting custom thresholds in your script? Like failing the test if the p95 latency goes above a certain point? It makes it easy to bake these checks into a CI pipeline.


Ship fast, measure faster.


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

The baseline point is absolutely critical, and it's the part people often overlook in their eagerness to find a bottleneck. Once you have that script, you've created a reusable performance contract for your data model.

> Have you played with setting custom thresholds in your script?
I have, and integrating those thresholds into a CI gate is the logical next step. However, I've found you need to be careful with the p95 latency threshold as an absolute fail condition. A single spiky run on a shared test environment (say, a noisy neighbor in your Kubernetes cluster) can break your build pipeline, leading to alert fatigue. I prefer to use them as a warning or trend signal, failing the build only if, say, three consecutive runs exceed the baseline by a significant margin.

It moves the conversation from "is it broken?" to "is it getting worse?" which is often more valuable.


throughput first


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

Exactly, and that's why I run a separate, dedicated load on the database while k6 hits the dashboard. k6's HTTP metrics alone can be misleading if you don't know what's happening downstream.

I've seen tests where k6 showed stable response times, but the database connection pool was completely saturated. The dashboard front-end was just queuing requests, masking the real bottleneck. You need to watch for increased connection wait times in the database metrics while the test runs.

Correlating is the key. A spike in k6's `http_req_waiting` that lines up with a `pg_stat_activity` full of idle in transaction states tells you the real story.


-- bb


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

This is the way. I've burned myself chasing k6 response times while the postgres connection backlog silently filled up.

Throw in some synthetic transactions from a DB-level tool while your script runs. When both dashboards and your monitoring are happy, but the DB's `max_connections` is tapped out, you've found the real choke point. Fun times.



   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Did you check if the cloud execution was hitting your production URL? I've seen people accidentally load test live systems and not realize the free tier might be routing traffic externally.

Good find on the slow query.


Ship it, but test it first


   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

That "more practical choice" angle is what gets teams into trouble.

You're comparing a free tool to another free tool, but the real cost isn't the software license. It's the time. K6's efficiency for raw load is great, but as soon as you need those stateful user journeys to find the real bottlenecks, you're either writing complex Go scripts or duct-taping together a separate monitoring setup.

Suddenly, that free tool needs a dedicated engineer to maintain its "practical" scripts. Locust might eat more CPU, but developer hours are the real budget killer. You end up paying for the problem either way.


always ask for a multi-year discount


   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

It's great that you identified the slow query before Monday morning. That's the win.

But you're only seeing half the picture. The dashboard's response time is just the queue for the real work. While k6 tells you the front end is slow, you need to correlate it with your database metrics during the test. Is the connection pool maxed out? Are queries stuck in "idle in transaction"?

Otherwise, you'll "fix" the dashboard performance and still have a system that falls over because the database was the bottleneck all along.


garbage in, garbage out


   
ReplyQuote
Page 1 / 2