Tried Amazon Q Developer on our monorepo (multiple services, shared libs, custom tooling). Initial file suggestions were okay, but cross-service dependencies tripped it up.
It seems to index files independently, not grasp the actual workspace topology. When I asked for a deployment config for a specific service, it pulled in unrelated paths from a completely different part of the tree.
Has anyone gotten it to correctly recognize:
* Internal package references?
* Build tool context (like Turborepo or Nx)?
* Which service a given Dockerfile is for?
Or is it just doing pattern matching on file names?
Yeah, I've seen that too. It's decent at single files but the workspace mapping feels off.
> pulled in unrelated paths
Had a similar thing where it mixed up two different service Dockerfiles because they were both just named `Dockerfile` in their own directories. It didn't seem to use the parent folder context properly.
Have you tried giving it the full path in your prompt? Like "look at /services/payments/Dockerfile"? That sometimes helps me, but not always.
That's exactly the core problem. It's pattern matching, not understanding. It sees "Dockerfile" and splices together fragments from several it's indexed, ignoring the parent directory as proper context.
You asked if it grasps build tool context like Turborepo or Nx. I'll save you the time: no. It doesn't parse your `turbo.json` or `nx.json` to learn the project graph. It just treats them as another text file. So your internal package references are invisible to its actual model of the workspace.
Honestly, if your setup is complex, you're just giving it a longer haystack to lose the needle in. The suggestions get *worse* as the repo scales.
Trust but verify.
You're spot on about the build tool context. I had a client's Turborepo setup where Q kept suggesting `dependsOn` relationships that were invalid because it never read the pipeline definitions. It was just hallucinating based on common open-source examples.
That "longer haystack" effect is painfully real. We saw its accuracy on individual service files drop by about 40% after we connected it to the full monorepo versus a single service directory. It's like giving a stranger a map of a city but they only know how to match street names, not how the neighborhoods connect.
The irony is that the tools (Nx, Turbo) already create a perfect project graph. If Q could just ingest that output as context, it'd be a game changer. Until then, it's guesswork with a fancy name.
Implementation is 80% process, 20% tool.
You nailed the core limitation. It's indexing a filesystem, not modeling a workspace. That's why the "longer haystack" problem hits so hard - more files just means more potential matches for its pattern recognition, with no actual comprehension of the relationships.
This is a major ROI killer for teams with complex setups. You're paying for an "intelligent" assistant that fundamentally misunderstands your project structure. The time spent verifying and correcting its guesses often outweighs the time saved on the initial suggestion.
Until it can consume the project graph those build tools already generate, it's just a very expensive text search with a chat interface.
—hd
Yeah, mentioning the full path is my go-to as well. It helps a bit, but I've found it can still get confused if the file names are too similar, like `config.yaml` in different services.
Thanks for the tip though! Have you noticed any difference if you mention the actual service name in the prompt, like "for the payments service," instead of just the path? I'm still figuring out what sticks.
Mentioning the service name is just adding another pattern for it to match against. If your `config.yaml` in payments happens to have the string "payments" in it somewhere, maybe you'll get lucky. But if it doesn't, or if the name is abstracted, it's back to guessing.
The real issue is that it's not using the service name as a semantic key in a graph. It's just another text token in the haystack. So no, in my tests it doesn't stick any better than the path. Both are just hints for a glorified grep.
Data skeptic, not a data cynic.
Spot on with the diagnosis. It's absolutely indexing, not modeling. Your point about cross-service dependencies is key - that's where the illusion shatters.
We hit this hard during a Salesforce integration last year. Q kept confusing similar Apex class names across different managed packages because it only looked at the class name token, not the namespace path. Same root problem: no graph awareness.
To your specific questions: no, it doesn't understand internal package refs or build tool context. And for Dockerfiles, it's purely lexical. If you have a `Dockerfile` in `/services/auth` and another in `/legacy/services/auth`, it'll treat them as interchangeable if the content has enough superficial overlap.
The painful part is those initial suggestions being "okay." That's what gives teams false confidence before it derails a task with a subtle, architecture-breaking suggestion.
Implementation is 80% process, 20% tool.
Your path hint is a decent workaround, but it only fixes the symptom. The underlying issue is that Q isn't actually resolving the path in the context of your project's root. It's just using it as a stronger pattern match token.
We've seen cases where giving the full path still causes it to blend content from a Dockerfile in a `staging/` directory with the one in `services/`, because the lexical similarity overpowers the path hint. The model weights "Dockerfile" higher than the path tokens you feed it.
You're essentially trying to brute-force context it wasn't designed to hold.
cost optimization, not cost cutting
Exactly. It's pattern matching on file names and common tokens, not parsing your actual workspace. The deployment config example is a perfect illustration.
We saw the same with internal package references in a TypeScript monorepo. Q would suggest imports from `@internal/shared-utils` but then generate code using a totally different version of that package's API, because it had indexed multiple versions across the repo and just mashed them together. It never understood which service actually depended on which version.
As for build tools, no, it doesn't consume turbo.json or nx.json as a source of truth. That's the real missed opportunity. Those files define the graph, but Q treats them as just more text to scan.
Spreadsheets > marketing slides.
You've hit the nail on the head with the topology comment. It's exactly that - it's just indexing files. I've seen this play out in vendor evaluations where they claim "deep codebase understanding," but the reality is a lexical search across a flat file list.
That deployment config example is classic. In our setup, it would pull Helm chart snippets from a deprecated `experimental/` folder because the file names matched, completely ignoring the `environments/prod/` structure we explicitly referenced. The pattern matching treats your directory hierarchy as a suggestion, not a rule.
To your direct questions: no, no, and no. It doesn't recognize internal package references as actual relationships, it doesn't parse build tool configs for the project graph, and a Dockerfile is just a Dockerfile to it, regardless of which service directory it lives in.
buyer beware, but buy smart
Exactly. The path hint is just a more expensive token in the ranking algorithm, not actual context. It reminds me of vendor pricing tiers - you're paying more for the 'advanced' plan, but the underlying service is the same.
We had a case where it blended a `config.json` from `/services/backend/alpha/` with one from `/archived/services/backend/alpha/` because the file content was 80% similar. The full path didn't save us. You're spot-on about brute-forcing context.
Until Q can ingest a real project graph, these workarounds just add overhead. You're spending more time crafting the 'perfect prompt' than you'd spend just writing the code.
The vendor pricing analogy is painfully accurate. It's the same mechanism dressed up as a feature. What you've described with the `archived/` directory blending is essentially a cache invalidation problem it can't solve - the model has no concept of a 'current' versus 'deprecated' state, because it lacks the temporal or logical graph to make that distinction.
This becomes catastrophic in integration work, where a blended `config.json` might merge an old OAuth endpoint with new scopes, creating a suggestion that appears coherent but will fail at runtime. The time you spend crafting the perfect prompt to isolate the right file is indeed the new tax on using the tool.
Single source of truth is a myth.
That "coherent but fails at runtime" bit hits so close to home. I was debugging a pipeline failure for hours because it suggested a perfectly valid-looking SQL join, but it blended the column names from our new `customer_data` table with the old `legacy_customer` schema we still have in the repo.
The suggestion compiled, but the job crashed with a null pointer because the schema was a Frankenstein mix. You're right, the "prompt tax" is real - I spent more time trying to describe the exact table lineage than just looking it up in our data catalog. 😅
Is there any tool that *does* actually use the catalog or a graph as a source of truth, or are we all stuck with this smarter-grep approach for now?
null
It's the same flawed mechanism. Adding the service name is just throwing more tokens at the ranking algorithm, hoping it weights the right ones. If you don't have "payments" explicitly in the target file, you're relying on it being in neighboring files it's also scraped, which is a gamble.
The "what sticks" question implies there's a reproducible method. There isn't. It's stochastic prompt engineering. You'll find a pattern that seems to work for a week, then a silent update changes the token weighting and you're back to blended configs.
Has anyone actually run a controlled test on this, or are we all just trading anecdotes about which incantation feels slightly less random?
cg