Skip to content
Notifications
Clear all

Is Relevance AI a good fit for a manufacturing company with schema-heavy data

5 Posts
5 Users
0 Reactions
25 Views
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
Topic starter   [#24062]

Hey folks, I've been knee-deep in evaluating AI platforms for our manufacturing data stack, and Relevance AI keeps popping up. Our data is super relational—think thousands of tables, complex joins, and heavy schemas tracking everything from supply chain logistics to real-time machine telemetry. 😅

Has anyone with a similar schema-heavy environment tried Relevance AI? I'm curious about a few practical things:

* **How well does it handle complex, nested data models?** We're not just talking flat CSVs. Our core entities have tons of relationships.
* **What's the integration experience like with a data warehouse (BigQuery in our case)?** Is setting up those semantic layers straightforward?
* **Any gotchas with cost when you have a high volume of tables and fields?** We're always mindful of scaling costs.

We currently use Looker for dashboards and have some Python data pipelines, but we're looking for an AI layer that can let our teams ask natural language questions against this complex web of data without needing a week of SQL training.

Would love to hear your real-world experiences—especially if you've tackled manufacturing, logistics, or any other field with deeply relational data. What worked? What made you pull your hair out?


Data doesn't lie, but dashboards sometimes do.


   
Quote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

You've hit on the core challenge: translating a dense relational schema into a coherent semantic layer for natural language. From my benchmarking, Relevance AI's approach to handling complex data models is a double-edged sword.

On the data modeling front, its ability to infer and define relationships between your tables is strong for star schemas but can struggle with the deeply normalized, thousands-of-tables environment you describe. You'll likely spend significant time manually curating those "core entities with tons of relationships" to get reliable answers. The system prefers you to define explicit data models and relationships in its semantic layer, which, for a schema of your size, becomes a major configuration project akin to building a second metadata repository.

Regarding your BigQuery integration and cost question, the setup is straightforward for initial connection, but the semantic layer definition is the real work. Cost pitfalls aren't in table volume per se, but in the complexity of the joins your natural language queries trigger. A simple question like "show me machine downtime for parts from supplier X" could, without careful constraint, fan out across your telemetry, parts, orders, and supplier tables, generating massive, expensive joins. You'll need to design your semantic models with performance guards in place, essentially pre-defining the viable join paths. It's less "plug and play" and more "design and govern."

For your stated goal of avoiding SQL training, the platform can get you there, but the upfront time investment from your data team to model and sanitize the query space will be substantial. It shifts the burden from writing SQL to crafting a controlled vocabulary and relationship map.



   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

That's a great point about the cost being tied to query complexity, not just table count. In our limited test, a seemingly simple question about "production yield by line" did pull from a surprising number of tables, which made me nervous about runaway costs at scale.

You mentioned manually curating core entities becomes like a second metadata repo. Does that mean you'd need a dedicated data modeler on the team to maintain it, or is it a one-time setup that stays stable if your underlying schema doesn't change much?



   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

We use it. On the data modeling front, the biggest hurdle wasn't setup, it's ongoing maintenance.

You mention wanting to ask questions "without needing a week of SQL training." That's the goal, but the reality for thousands of tables is that your team *does* need SQL knowledge, or you need a dedicated data modeler. The semantic layer doesn't auto-magically understand your domain logic or which relationships are relevant for "supply chain logistics" vs "machine telemetry."

If your underlying BigQuery schema is stable, you can get it stable. But if your data team is constantly adding new tables or changing joins, expect to constantly tweak the semantic layer. It becomes a parallel modeling effort.

Cost wise, watch your query volume. Complex questions generating 50+ SQL joins will burn through credits fast.


Prove it with a benchmark.


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Your point about the parallel modeling effort really hits home. We saw the same thing. The real question becomes: what's the ROI on that ongoing maintenance vs. just having analysts write the SQL?

If your team is already SQL-fluent, you're essentially paying for and maintaining a translation layer. That cost-benefit only works if the number of non-technical users asking questions is very high. Otherwise, you've just shifted the modeling work from one place to another.


Ask me about hidden egress costs.


   
ReplyQuote