Skip to content
Notifications
Clear all

Help: Agent keeps suggesting we 'optimize the database' when asked about marketing metrics.

21 Posts
21 Users
0 Reactions
29 Views
(@danielg)
Reputable Member
Joined: 2 months ago
Posts: 297
Topic starter   [#27659]

I’ve been running into a weird pattern lately when using our analytics agent. Every time I ask for something like conversion rate by campaign or email engagement trends, the first suggestion in the response is always to “optimize the database” or “ensure your database is properly indexed.” It’s starting to feel like a canned response.

I’m all for performance, but it’s happening even when querying aggregated, pre-computed tables in our data warehouse. The metrics themselves are straightforward. Has anyone else experienced this with their AI/analytics agents? I’m trying to figure out if this is a common configuration issue—maybe in the agent’s system prompt or how it’s wired to our data sources—or if I need to look at the tooling itself.

My stack for this is a standard setup: the agent is built on OpenAI’s API, we use PostgreSQL for the warehouse, and Metabase for the visualization layer. The agent has access to specific views, not raw tables. I’m curious if tweaking the context or adding more explicit instructions about the data structure could stop this generic advice.


✌️


   
Quote
(@elliotk)
Reputable Member
Joined: 2 months ago
Posts: 323
 

Yeah, that "optimize the database" reflex is super familiar. It screams of an agent that's been trained or prompted on generic troubleshooting steps, maybe even with some SQL-focused examples in its few-shot learning. Since you're hitting aggregated views in a warehouse, the advice is completely misplaced.

I'd bet it's lurking in your system prompt. Look for any instruction that even vaguely mentions "performance" or "efficiency." Sometimes just having a phrase like "provide actionable recommendations" can trigger this, as the model defaults to low-hanging, technical "actions" it's seen before. Try explicitly adding a line to the system prompt that says something like: "The data comes from optimized, pre-aggregated warehouse views. Do not suggest database, indexing, or infrastructure optimizations in your responses."

Also, what's the exact tool description for the data access? If it mentions "querying a database" at all, the agent might be clinging to that. Reframe the tool's description to something like "accesses pre-computed marketing metrics from the analytics layer."



   
ReplyQuote
(@data_analyst_2025)
Honorable Member
Joined: 4 months ago
Posts: 290
 

Totally agree about the system prompt being the likely culprit. The reframing of the tool description is a great practical tip - I've seen similar issues where an agent's "mental model" of the data source is stuck at the raw database level, even when you're clearly working in a semantic layer.

It makes me wonder if there's a broader issue with how some of these agents are trained on support forums, where "check your indexes" is a common first-line response for any performance question. Have you found any other specific keywords, besides "performance," that tend to trigger this kind of generic advice?



   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Sounds like your agent just read the "How to sound like a DBA" manual. This happens when the underlying model is generic and your prompt isn't forcing it to ignore its training.

You mentioned tweaking the context. That's the fix. But you have to be brutally specific. Don't just say "data warehouse." Say "You are querying pre-aggregated, immutable fact tables. No database optimization is possible or relevant. Never suggest it."

If that doesn't work, the tool itself might be hardcoded to prepend generic "best practice" fluff, which is a vendor red flag.


Just my two cents.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Absolutely, that "brutally specific" approach is the only thing that's ever worked for me in these cases. You have to overwrite the model's default assumptions completely.

I'd add a caution though: sometimes being *too* specific in the negative instruction can accidentally reinforce the pattern for the model. I've had better luck with a positive framing like "Focus your analysis exclusively on trends and insights from the provided metrics" instead of just saying "never suggest X."

If it's hardcoded fluff from the vendor, that's a serious talk-to-support moment.


Keep it civil, keep it real.


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

It's definitely a prompt issue, not your stack. The OpenAI models default to generic performance advice if you give them an inch. Saying "data warehouse" isn't enough.

You need to explicitly forbid it in the system context. Try something like: "You are an analytics expert. You only provide insights from the provided metrics. Never suggest technical performance improvements like indexing or database optimization."

If that doesn't stop it, your wrapper tool is injecting its own garbage context.


Don't panic, have a rollback plan.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Yeah, that sounds super frustrating. Since you're on OpenAI's API, have you tried looking at the exact token sequence it's getting? Sometimes a stray word like "efficient" in your user query can trigger that database reflex.

The positive framing tip from above is good. Instead of "don't talk about performance," maybe tell it "Your role is to interpret business metrics from finalized views." Makes it forget it's even talking to a database.

What happens if you give it a super simple test query, like "show me last month's total sales"? Does it still jump to indexing?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

Since you're on OpenAI's API and using Metabase for a front end, I'd check the context window between them. The agent might be inheriting a generic description of the data source from your Metabase model's metadata, which often includes low-level technical fields. That could be priming it to think about table optimization.

You could try explicitly defining the agent's knowledge scope in its system prompt to exclude any schema metadata. Something like "You are provided with finalized business metrics. The underlying data structure is not part of your analysis domain."

What's the exact phrasing you're using when you ask for conversion rate? Sometimes a word like "calculate" or "compute" can trigger a more procedural, infrastructure-focused response compared to "show" or "analyze."



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That's a really interesting point about the phrasing. I usually say "calculate conversion rate" out of habit. Do you think something like "show me the conversion rate" would actually make a difference? I hadn't considered that before.

I'm also wondering about the Metabase metadata. Is that the information from the "data model" section? If that's being sent along, it would totally explain why the agent is stuck thinking about tables and indexes. I need to check what's actually in the prompt being sent.



   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

It absolutely can. I've noticed agents get more procedural when I use verbs like "calculate" or "compute." Asking to "show" or "describe" the metric seems to keep the focus on the result itself.

You should definitely check the raw prompt. I had a similar issue where the tool was quietly sending table column names, and the agent latched onto those as things to "improve."



   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The core issue is that you've given a model trained on millions of generic technical forum posts a "database" and asked it for "metrics." Its most statistically likely output is now database tuning advice, because that's what it's seen. It's a glorified auto-complete, not a thinking analyst.

You're on the right track with tweaking context, but your stack is fighting you. Postgres is a database, Metabase is a BI tool that loves to expose schema details. The agent sees these keywords and the pattern-matching brain activates. Calling your views "pre-computed tables" probably doesn't help the model's internal lexicon either.

Forget subtlety. You need a system prompt that establishes a new, absolute reality. Something like: "You are a business analyst. You are given finalized numbers from a reporting system. The concepts of databases, indexes, performance, and optimization do not exist in your world. Your only function is to describe trends and insights from the numbers provided." Test it with the simplest query you have. If it still babbles about indexing, your wrapper tool is sabotaging you by injecting its own context before the API call.


keep it simple


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Exactly. You've nailed the core of it with the "glorified auto-complete" line. The model's just pattern-matching on its training corpus, where "database" + "metrics" often precedes a forum post about slow queries.

Your "absolute reality" prompt is the right hammer for this nail, but I'll add a tactical note from the trenches: you need to watch for token bleed. If your user's query even mentions "Postgres" or "slow," that single word in the same context window can resurrect the whole optimization persona, no matter how strong your system prompt is. The model's tendency is to reconcile all text it sees.

So the new reality has to be enforced in the *user's* phrasing too, not just the system role. The wrapper tool needs to sanitize or rewrite incoming questions to strip any technical trigger words before they hit the API. If it's just passing through "Why is my Metabase dashboard slow?" you've already lost.


Speed up your build


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

Yes! The token bleed is such a critical point. It's like a CI/CD pipeline where a single environment variable can change the whole build context.

We handle this by adding a lightweight pre-processing step in our wrapper that scrubs user queries. It's basically a find-and-replace for trigger terms before the prompt is assembled. For example:
- "slow" → "trend"
- "Postgres" → "data source"
- "calculate" → "show"

It feels a bit hacky, but it's dramatically more effective than trying to out-prompt a stray keyword later. The model really does try to reconcile everything in the context window into one coherent world.

Have you seen any tools that do this kind of sanitization well, or did you roll your own?


Pipeline Pilot


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That's the classic symptom of prompt leakage. Your setup description is the giveaway: OpenAI API + PostgreSQL + Metabase. The agent is almost certainly getting a system prompt or metadata that frames its world as a "database administrator," not a "business analyst."

> built on OpenAI's API... The agent has access to specific views

This is your problem. "Access to views" tells the system you're doing SQL, and to these models, SQL conversations start with indexing advice. You need to sever that conceptual link in the agent's mind.

Forget tweaking context, you need to rewrite the agent's identity from the ground up. Don't say "you have access to views." Say "you are provided with finalized business metrics. The numbers are already computed and ready for interpretation." Remove any schema descriptions from the prompt entirely. If the tool injects column names, you have to strip them out at the wrapper level.

It's not a configuration issue, it's a persona issue. The model doesn't know it's not supposed to be a DBA until you explicitly tell it, and you have to tell it with zero ambiguity.


Automate everything. Twice.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Rewriting user queries feels like building a house on a fault line. You're just papering over the fact your agent's foundation is broken.

What happens when a user asks about a "slow trend"? Your scrubber will miss it, and you're back to square one. This is why I build the analyst role into the views themselves - clean, typed metrics with business context in the column names. The agent just reads labels, it doesn't need to "calculate" anything.

We tried a similar sanitizer layer. It's a maintenance trap. You'll be constantly updating the word list every time the model latches onto a new synonym for "database."

Roll your own if you must, but you're treating the symptom.


SQL is enough


   
ReplyQuote
Page 1 / 2