Skip to content
Notifications
Clear all

Anyone have experience with BI tools that handle >100M rows?

4 Posts
4 Users
0 Reactions
6 Views
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 424
Topic starter   [#26165]

Hey everyone. Just switched to a DevOps role and I’m learning the observability stack. I’ve been setting up Grafana with Prometheus for metrics, but now my team is asking about BI tools for analyzing huge application logs and historical data.

We're talking over 100 million rows from Postgres and some clickstream data. I see tools like Tableau, Power BI, and Looker mentioned, but I’m unsure how they perform at this scale. Does anyone have real-world experience with this?

My main concerns:
- Query performance without pre-aggregating everything.
- Cost for this data volume.
- How they connect to cloud data warehouses (BigQuery, Redshift).

I built a simple Grafana dashboard for real-time metrics, but this feels different. Here's a basic Prometheus query I'm used to:

```promql
rate(http_requests_total[5m])
```

For BI, I guess I need something that can handle complex joins and historical trends on massive datasets. Any recommendations or things to avoid? Thanks



   
Quote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 514
 

Based on your Prometheus query example, I think you're correctly identifying a key difference. BI analysis on 100M+ rows is fundamentally about aggregated historical trends, not real-time rates. Your concerns about query performance and cost are spot on for this scale.

In my experience, the tool itself matters less than the underlying data warehouse architecture at this volume. Tableau or Power BI connected directly to a transactional Postgres instance will struggle. The performance you're asking for without pre-aggregation will almost always require pushing the heavy lifting to a dedicated analytical store. That's where your connections to BigQuery or Redshift become critical - the BI tool then acts as a visualization layer on top of the warehouse's computed results.

Have you considered a two-tiered approach? Use your Grafana setup for the real-time operational metrics you're already comfortable with, and implement a separate cloud data warehouse for the historical log analysis. Then you can evaluate BI tools primarily on their cost and UX for connecting to BigQuery/Redshift, since they won't be executing the raw joins themselves. Looker is built around this model, while Tableau and Power BI can work in a similar fashion but may try to pull too much data locally if not configured properly.


Data > opinions


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 420
 

Your third point is the only one that matters. The BI tool doesn't handle the rows, the warehouse does. You're asking the wrong question.

Forget "performance without pre-aggregation" for 100M rows and joins. That's a fast path to a dead dashboard. The real cost is the compute you'll burn in BigQuery or Redshift every time someone changes a filter. The tool's job is to let you build queries that don't bankrupt you.

Looker forces you into a semantic layer, which can prevent analysts from writing stupid-expensive queries. Power BI pushes more down to your local machine, which falls over. Tableau sits in the middle. Pick based on how much you trust your team's SQL and your company's cloud budget.


Your CRM is lying to you.


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Absolutely agree that the warehouse does the heavy lifting. The point about the semantic layer preventing expensive queries is so key, especially when you start adding sales and marketing teams who just want to slice data every which way.

I'd add one caveat from the CRM world: sometimes that "forced" semantic layer in Looker can become a bottleneck for urgent, one-off investigations if your data team is swamped. You trade cost control for some agility. Tableau's "middle" approach has burned us with a surprise BigQuery bill when someone built a viz on a massive, joined dataset without understanding the scan sizes.

Your final line about trust is the real answer. It's less about the tool and more about governance.


Pipeline is king.


   
ReplyQuote