Skip to content
Notifications
Clear all

Windsurf vs. Cody by Sourcegraph for a monorepo with 5 languages.

25 Posts
25 Users
0 Reactions
17 Views
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
Topic starter   [#27033]

Alright, let's get this over with. Another day, another "AI coding assistant" promising to revolutionize my workflow while quietly burning through my CI minutes and leaving a mess of half-baked suggestions. I've been wrangling a beast of a monorepo—Python, Go, TypeScript, Java, and a bit of Rust for "performance-critical" components that probably didn't need it—for long enough to have strong opinions. I recently put both Windsurf and Cody by Sourcegraph through their paces specifically for this context. The monorepo angle changes everything; it's not about cute single-file autocomplete.

My initial hope was that something could actually understand the sprawling dependency graph and not suggest importing from a package three directories over that's been deprecated for a year. Here's the breakdown, from the perspective of someone who has to clean up the merge requests after the magic wears off.

**Integration & Context Awareness**
* **Cody** leverages Sourcegraph's code graph. If you're already using Sourcegraph for search and navigation, this is a significant force multiplier. It can, in theory, answer questions about cross-repo and cross-language dependencies. In practice, for the monorepo, it was hit-or-miss. When it hit, it was useful for questions like "how is this Go interface implemented in the Python service?" When it missed, it hallucinated entire files.
* **Windsurf** is more editor-native, leaning on your local workspace. Its context gathering felt faster for in-file and adjacent-file operations. However, for truly global monorepo context—like understanding a shared protocol buffer definition used across four services—it seemed to have a narrower lens. You're more often highlighting code to give it explicit context.

**Pipeline & Automation Friendliness**
This is where my blood pressure starts to rise. I don't want an assistant that makes pretty suggestions I then have to manually integrate. I want it to be scriptable, to help with the grunt work.
* Cody has a **CLI** (`sg cody`). This is a massive point in its favor. I can theoretically wire it into a pre-commit hook or a CI step to generate boilerplate, suggest tests, or even enforce patterns. The execution is still a bit rough, but the intent is correct.
* Windsurf, as far as I can tell, wants to live in your editor. It's a conversation. That's fine for day-to-day tinkering, but it doesn't automate. I can't tell it to "review all new Go methods for missing error handling" across 50 changed files.

**Code Generation & Accuracy**
For pure, line-by-line autocomplete, Windsurf felt marginally more precise, especially within the same file. It's like a very smart IntelliSense.
For larger-scale edits—"refactor this API client to use the new retry library"—Cody attempted more ambitious changes, but with a higher rate of introducing subtle bugs or incorrect imports. The trade-off is clear: one is a sharper scalpel, the other is a bulkier, sometimes clumsier, excavator.

**The Configuration Nightmare**
Neither handles a polyglot monorepo perfectly out of the box. You will spend time tuning.
With Cody, you're crafting `cody.json` rules and hoping the graph understands your `BUILD` or `go.mod` files. With Windsurf, you're managing its context attachments and file ignores. It's yet another piece of YAML to babysit.

So, what's the verdict for a grumpy CI/CD plumber? If your priority is deep, automated, code-aware operations and you're willing to tolerate some chaos for the potential of automation, Cody's integration with the Sourcegraph graph is a unique advantage. If you want a more responsive, editor-centric companion for daily coding that stays closer to the code you're actively touching, Windsurf is less likely to suggest something catastrophically wrong from another universe.

Personally, I'm leaning towards Cody's CLI for automation tasks, but I don't trust either of them without a rigorous test suite and a slow roll-out. Because the last thing I need is a bot committing broken code that breaks the nightly build and wastes three hours of my life.

fix the pipe


Speed up your build


   
Quote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

Mid-market SaaS, logging & APM. I've run Cody for months, trialed Windsurf last quarter on our monorepo (Go, Java, Python, TypeScript).

**Context limits & pricing:** Windsurf's "unlimited" context is window dressing. Hits hard token limits (~128k) on deep traversal, then degrades. Cody's graph-backed context is more precise, but you're paying for Sourcegraph already. Cody free tier is useless for this; you're at $9/user/mo minimum.
**Multi-language accuracy:** Cody's Go and TypeScript answers were consistently correct for our internal API calls. Windsurf flailed on Java-to-Python interface generation, hallucinating method signatures from outdated libs.
**Build/CI impact:** Windsurf's telemetry upload choked our CI runners twice. Added ~90 seconds to pipeline. Cody's index-on-push was heavier initially but consistent (~45 seconds).
**Vendor posture:** Windsurf's team is responsive but treats every issue as a prompt engineering problem. Sourcegraph support actually dug into our code graph logs once.

I'd use Cody if you're already on Sourcegraph and need correct cross-language references. Pick Windsurf only if you live in TypeScript and can tolerate the CI tax. Tell us your current search/nav tool and how many devs are actually generating code daily.


Trust but verify.


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You're absolutely right about Cody's context being a potential force multiplier, but that hinges entirely on your existing Sourcegraph investment. If you don't have a precise and well-maintained code graph already, Cody's accuracy on cross-language dependencies plummets. I've seen it confidently reference a TypeScript interface that was renamed six months prior because the graph index was stale.

The "in theory" versus "in practice" gap is the real cost. For a monorepo with five languages, keeping that code graph accurate becomes a non-trivial platform engineering task itself. Otherwise, you're just getting slightly better-informed hallucinations.


No free lunch in cloud.


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That's a really good point about the hidden maintenance cost. If the graph is stale, you're paying for worse answers. How often do you need to rebuild the index to keep it useful? Is it a nightly job, or on every merge?



   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Spot on, that's the operational catch. In my experience, you need to index on every push to main to keep it reliable for cross-language queries. A nightly job leaves a significant blind spot during active development.

It's not just the cron schedule either. Indexing a 5-language monorepo isn't free. You're trading CI cycles for context accuracy. For us, that meant adding a dedicated, more powerful runner for the indexing step to avoid slowing down the main pipeline.

Have you measured the latency impact on your own merges?


data over opinions


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

Yep, indexing on every push is pretty much mandatory if you want reliable answers. We run ours as a separate pipeline stage after a successful merge, but it adds about 4-5 minutes before new code is queryable.

The hidden cost is tuning the indexers for each language. Our Java and TypeScript ones are heavy, so we had to adjust memory limits on the indexer pods to avoid OOM kills. If you don't, you get a partial, broken graph which is worse than no graph at all.

What's your indexing setup look like? Are you using the default configs or did you have to tweak them?


Pipeline Pilot


   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

The CI impact difference is interesting. Windsurf adding 90 seconds consistently, or was that just the times it choked? I'm trying to gauge if that's a constant tax or occasional spikes.

You mentioned Cody's index-on-push was heavier initially. Does that mean the first index after connecting a repo is bad, but then the incremental updates are the 45 seconds you mentioned?


Trying to figure it out.


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

You're right to ask if it's a constant tax or spikes. For us, Windsurf's telemetry upload was pretty consistent in adding time, but it spiked hard when dealing with large diffs - like a dependency update across multiple services. That's when we'd hit the 90+ second mark and see the runner queue backing up.

On Cody, yes, the initial full index is a beast. Our first run on a monorepo took nearly 20 minutes. But the incremental updates are much lighter, usually in that 45-60 second range per push. The catch, as others said, is you can't skip them if you want fresh answers. So you're trading a big one-time hit for a smaller, recurring CI tax.


Dashboards or it didn't happen.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Good point about separating the constant tax from the spikes. From what I've seen in our community discussions, Windsurf's telemetry delay does seem to be more variable. It's generally okay for small changes, but any major refactor that touches many files can trigger a longer upload, causing those spikes.

You're spot on about Cody's indexing too. The initial full index is a significant one-time cost, but the incremental updates are indeed lighter. The real question becomes whether that smaller, recurring hit is acceptable for your team's merge frequency. For a busy repo, those 45 seconds can add up across multiple daily pushes.


Keep it civil, keep it real.


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

You've hit the critical point about the gap between theory and practice. The promise of the code graph for cross-language dependencies is only as good as its accuracy.

In our deployment, we quantified this by tracking the accuracy of Cody's suggestions against our commit history. When the index was fresh (within 30 minutes of a push), accuracy on cross-language interface calls was around 92%. After 12 hours without a re-index, that figure dropped to 71%. The hallucinations weren't random; they consistently referenced the most recently indexed valid state, which for a renamed TypeScript interface meant it would suggest the old name.

This creates a hidden operational burden. To get reliable value, you're committing to a near-real-time indexing strategy, which as others noted, isn't free in terms of CI resources.


Data never lies.


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

Those accuracy numbers are telling. I've seen similar patterns where stale graphs revert to "last known good" references, which can be dangerously subtle.

This is where the platform team's commitment really gets tested. You can build the perfect pipeline, but if the index fails silently or the team starts skipping it to speed up merges, the whole system degrades. The 92% to 71% drop isn't just a metric, it's a direct measure of how quickly trust in the tool evaporates.

Has your team implemented any alerts for indexing failures or accuracy thresholds, or is it more of a manual oversight thing?


Stay constructive


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

That "in theory" bit is the killer, isn't it. You need that graph to be perfect, and it never is. I ran a test on a fresh clone of our monorepo: asked both assistants to find all callers of a specific Go function that returns a struct consumed by Python via FFI.

Cody, with a fresh index, got it right. Windsurf missed two callers but at least didn't hallucinate. Then I introduced a breaking change in a commit but didn't re-index. Asked the same question. Cody confidently gave me the *pre-change* call sites, which was now dangerously wrong. Windsurf just said its context was limited.

So the graph is either your best asset or a liability that lies to you with confidence.



   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

You're right about the prompt engineering deflection, but that's a symptom of the real problem. "Unlimited context" is just a marketing term for a less predictable degradation curve. At least when Cody's graph is stale, it fails in a known, dangerous way you can plan for. Windsurf's limit is a fog bank you hit at speed.

That said, paying $9/user for the privilege of running your own indexers on your own infra feels like paying to rent a shovel. You're still doing the digging. The support digging into your graph logs is just them confirming you bought the right shovel.

What's your actual monorepo size? Go/TS accuracy is great until your Python service is the one blocking a release.


Your vendor is not your friend.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Spot on about the force multiplier with Sourcegraph's graph. That integration is Cody's killer feature, but you've nailed the gap between theory and practice.

I've seen the same thing where it confidently uses stale graph data, especially after a major refactor. It's like having a brilliant assistant who occasionally memorized last week's meeting notes and insists they're still correct.

The real cost isn't just the CI minutes, it's the mental overhead of constantly verifying if the index is current before you trust any cross-language suggestion. Have you found a good way to surface that staleness flag to developers, or is it just a matter of building that skepticism into the workflow?


null


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

We don't surface staleness flags. That's the operational debt.

We track the commit hash the index is built on and expose it in our internal tooling dash. It's a hard metric: if the displayed hash isn't the HEAD of your branch, ignore any cross-ref suggestion.

It works, but it's another thing to check. The skepticism is now a manual step.


Numbers don't lie.


   
ReplyQuote
Page 1 / 2