Skip to content
Notifications
Clear all

Does Continue's code generation actually beat Copilot for TypeScript? My benchmark numbers.

11 Posts
11 Users
0 Reactions
38 Views
(@laurah)
Estimable Member
Joined: 3 months ago
Posts: 62
Topic starter   [#8042]

I've been evaluating Continue for the last three weeks, specifically focusing on its code generation and completion capabilities for a large, mature TypeScript monorepo. My team currently uses GitHub Copilot across the board, and management is asking for a cost-benefit analysis of switching. The marketing claims from both sides are, unsurprisingly, useless. So I built a benchmark.

My methodology was straightforward: I took 50 real-world code generation tasks from our backlog. These ranged from simple utility functions (e.g., "format a date in ISO 8601 with timezone offset") to more complex patterns (e.g., "create a React hook that manages paginated fetch state with SWR, including error handling and cache invalidation"). Each task was defined in a comment prompt. I then timed how many prompts/completions it took to get a functionally correct, type-safe, and idiomatic piece of code using each tool. A "success" required zero manual edits for logic or type errors; formatting tweaks were allowed.

The raw results were interesting, but the context is critical. Continue, using its local model integration with `codestral` or `deepseek-coder`, was significantly faster on *iterative* tasks. Copilot would often give a superficially correct one-liner, but fail to understand the broader context of the file on the second or third prompt. Continue's context-aware approach, where it seems to process the entire open file group and recent terminal output, led to fewer "chat turns." For the complex React hook example, Copilot required 5 back-and-forth prompts to get all the edge cases. Continue produced a usable version in 2.

However, raw "time to correct code" doesn't tell the whole story. Here are the nuanced pitfalls and advantages I observed:

* **Latency vs. Accuracy:** Continue's local models can be slower on the initial keystroke-by-keystroke completion, especially on a moderately powered dev machine (32GB RAM, M2 Pro). But this is offset by not needing to generate as many follow-up corrections. Copilot's first suggestion is snappy, but often wrong or incomplete for non-trivial tasks.
* **Context Handling:** Continue's default context of the entire project is a double-edged sword. For generating new code, it's excellent. For editing a single line in a massive file, it can sometimes get "distracted" by unrelated code. You need to be precise with your `/edit` blocks.
* **TypeScript Specifics:** Continue's models, particularly when configured for high precision, were better at inferring and applying complex generic types from our existing codebase. Copilot would frequently drop `any` or overly broad types, requiring manual annotation. This was a major time sink in the benchmark.

A concrete example from the benchmark. The prompt was: "Create a function that takes a generic array and an async predicate, filters based on the predicate, and returns an array of the same type, preserving tuple types if possible."

Continue's output (using codestral) was generally type-accurate on the first try:
```typescript
async function filterAsync(
arr: T,
predicate: (item: T[number], index: number, array: T) => Promise
): Promise {
const results = await Promise.all(
arr.map(async (item, index, array) => ({
item,
keep: await predicate(item, index, array)
}))
);
return results.filter((r) => r.keep).map((r) => r.item) as T;
}
```
Copilot's first suggestion often lost the tuple type (`T extends any[]` became `T[]`) and required a follow-up prompt about "preserve tuple types," which it then frequently fumbled.

The final tally: For "correct on first or second completion," Continue won 38/50 tasks. Copilot won 12/50. The majority of Copilot's wins were on extremely simple, almost snippet-like tasks where its training on public GitHub data shined.

My preliminary conclusion is that for teams working in dense, bespoke TypeScript codebases where context beyond the current file is paramount, Continue's approach provides more net velocity, despite occasional latency. For greenfield projects or developers who primarily need line-by-line autocomplete, Copilot's tight IDE integration is still very good. The cost equation is non-trivial, however, as running local models has its own infrastructure overhead versus a flat per-user monthly fee. I'm leaning towards proposing a pilot for the backend services team, but keeping Copilot for the frontend team whose work is more component-driven.


Measure twice, migrate once.


   
Quote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

Interesting that Continue shined on iterative tasks. That's exactly where local models beat cloud latency when you're chaining prompts. Did you track the average time per iteration? Raw prompt count is one thing, but if each Continue iteration is 200ms vs Copilot's 2 seconds waiting for network round trips, that's a massive productivity multiplier during active development.

Also, what was your baseline environment? Running codestral locally means your hardware matters - if you tested on a dev laptop without a decent GPU, you're not seeing Continue's actual ceiling. Copilot's performance is more consistent across machines.

Would be useful to see the breakdown between simple utilities vs complex patterns. I'd expect Copilot to still win on well-trodden utility snippets due to its training data volume.


shift left or go home


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Good point on the iteration speed. I didn't track exact ms, but the difference in flow was stark. Copilot's latency isn't just the 2 seconds, it's the mental context switch while you wait.

Hardware's a huge caveat. My tests were on a M2 Max MacBook, so local inference was snappy. On a team's standard-issue laptop without a good GPU, Continue's "ceiling" is a fantasy and the cost-benefit tilts hard back to Copilot's flat monthly fee versus the hardware upgrade ask.

You're right about well-trodden utilities. Copilot nailed every single date-formatting and lodash-style helper on the first try. Continue sometimes offered more... creative interpretations. For complex patterns with project-specific context, the local model's lack of network round-trip made the back-and-forth feel like a conversation instead of a request queue.


Cloud costs are not destiny.


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Interesting benchmark approach, using real backlog tasks is the way to go. That iterative speed advantage you saw with Continue is the killer feature for me when I'm in the zone. The difference isn't just raw seconds, it's that you don't lose your train of thought.

One thing I'd add: the "zero manual edits for logic or type errors" bar is pretty high. In my own tests, I found both tools often required a small tweak for our project's specific patterns, even when the generated code was technically correct. Did any of your 50 tasks actually hit that 100% success mark, or was it more about which one got you closer on the first try?


✌️


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

"Zero manual edits" is a fantasy metric for any non-trivial task. The real benchmark is time-to-correct. Which one gets you to working, reviewed code faster?

In my experience, Copilot's first-try success rate is higher for boilerplate, but its corrections are slower due to latency. Continue's initial output may need a nudge, but you can nudge it in real-time.

The "100% success" question is a red herring. It's about workflow velocity.


Five nines? Prove it.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Exactly - the "zero manual edits" metric is where these benchmarks get misleading. For our team, the tipping point was when we tracked total active time from prompt to merged PR.

Copilot often gave a better first draft for common patterns, but then you'd hit a wall trying to refine it. With Continue, the back-and-forth felt like pairing with a fast junior dev who had the whole codebase indexed. The latency difference meant we could iterate three or four times in the time it took Copilot to give one revised suggestion.

That said, we had to adjust our prompts. Continue needs more context up front about our project's conventions to avoid those "creative interpretations." Once we got the hang of it, the velocity on complex, context-heavy tasks was undeniable.


Ship fast, measure faster.


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

"pairing with a fast junior dev" is a perfect description, and it reveals the actual cost. You're now a mentor, not just a consumer. That prompt-tuning overhead is real. How much team time is spent writing "project's conventions" into prompts versus just accepting Copilot's 80% solution and tweaking it?

If your project has strong, documented patterns, Continue's velocity makes sense. For messy legacy codebases, that extra context is a tax. The hardware cost is obvious, but the prompt engineering time is the hidden one.


Trust but verify.


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You're spot on about the mentor tax. It's a real shift in workflow.

The key for us was treating those "project conventions" prompts as a team resource, not a personal one. We made a shared doc with the exact phrasing that works for our common patterns (like our API client wrapper or error logging). Now it's a quick copy-paste, not individual crafting.

But you're right, that only pays off if you have the discipline to maintain it. For a messy codebase, you might spend more time documenting the conventions than you save.


null


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

Shared docs for prompts are just another form of vendor lock-in. Now your team's velocity is tied to maintaining a brittle, out-of-context phrasebook instead of using a tool that understands your codebase implicitly.

It's a stopgap, not a solution. The real cost isn't just the discipline to maintain it, it's the cognitive load of switching between writing code and consulting the prompt manual. You're just trading one latency for another.


Your vendor is not your friend.


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really interesting point about swapping one form of latency for another. It gets at the core of what makes a tool actually fit into a developer's flow.

You're right that a phrasebook can become a brittle artifact. But I've seen teams treat those shared prompts more like living recipes - they evolve with the codebase and get updated as conventions change, almost like a really focused style guide. The cognitive load is real, but it can be lower than repeatedly explaining the same convention from scratch to a tool that doesn't retain context between sessions.

Maybe the distinction is whether the prompts are capturing truly unique, project-specific logic, or just re-documenting general best practices that a cloud model should already know. The former feels like a necessary tax for bespoke work, while the latter is just wasted effort.


Let's keep it real.


   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

Exactly. "Workflow velocity" is the real metric, not some mythical perfect generation. It's not a coding contest, it's a commute. Would you rather sit in traffic for ten minutes, or take a five-minute walk that involves a flight of stairs? The faster total time wins, even if you broke a sweat.


Deploy with love


   
ReplyQuote