I've been testing Continue against GitHub Copilot for the last two weeks, focusing strictly on TypeScript in a real AWS Lambda and CDK codebase. My metric is simple: **accepted vs. discarded suggestions.** Not lines of code, not tokens—just whether the AI's output was usable without major rewrites.
The hype suggested Continue's local model might be smarter. My billing alarm instincts said "prove it." So I did.
**My setup:**
* **Project:** Serverless data pipeline, heavy on TypeScript/Node.js, AWS CDK for infra-as-code.
* **Test:** Equivalent tasks across both tools: writing Lambda handlers, constructing CDK constructs, writing unit tests, and refactoring functions.
* **Continue:** Using `deepseek-coder` (6.7B) locally via Ollama.
* **Copilot:** Default settings in VS Code.
**The raw results:**
* **Continue:** 68% acceptance rate on first suggestion. When I used its chat for refinements, that jumped to ~85%.
* **Copilot:** 72% acceptance rate on first suggestion. However, its multi-line completions were more frequently *structurally* correct.
**The critical detail:** Continue won on *context awareness*. Because it indexes my entire codebase, its suggestions for CDK patterns and existing utility functions were spot-on. Copilot often gave generic, syntactically correct but context-blind code.
Example: Writing a DynamoDB query helper in a project where we already have a specific pattern for pagination.
**Copilot suggestion (generic):**
```typescript
export async function queryItems(tableName: string, keyCondition: any) {
const dynamoDB = new AWS.DynamoDB.DocumentClient();
const params = {
TableName: tableName,
KeyConditionExpression: keyCondition
};
return dynamoDB.query(params).promise();
}
```
This ignores our configured DB client, existing error handling wrapper, and pagination utility.
**Continue suggestion (context-aware):**
```typescript
export async function queryItems(
queryParams: Omit
): Promise {
const params: QueryInput = {
TableName: getTableName(),
...queryParams
};
return await db.queryWithRetry(params); // Uses our existing, indexed utility
}
```
**The cost angle:**
* **Continue (local):** Zero marginal cost per completion. Upfront "cost" is GPU RAM.
* **Copilot:** $10/user/month, flat. No surprise bills, but it's a line item.
**Verdict:**
For TypeScript in a well-established, complex codebase, Continue's deep context gives it a tangible edge for *code patterns*. Copilot is slightly better at raw, line-by-line flow and obscure API snippets. If you're optimizing for both code quality and cloud costs (of the tool itself), Continue with a solid local model is a serious contender. For greenfield projects or those with less consistent patterns, Copilot's simplicity might win.
I'm switching my primary driver to Continue for this project. The lack of a recurring charge for a comparable—and in some ways superior—experience is a no-brainer for FinOps.
cost optimization, not cost cutting
Interesting you measured acceptance rate based on first suggestions. I've been curious about this myself, but from a slightly different angle regarding data work.
When you mention that > Continue won on context awareness because it indexes my entire codebase, does that hold true for larger, more interconnected projects? Specifically, in a Power BI or Looker context, you often have dozens of report files, shared data models, and custom visual scripts. A tool indexing everything could suggest a DAX measure that references a field from a completely different file, which would be wrong if that field isn't imported.
For Lambda handlers and CDK constructs, the dependencies are more explicit in your imports. Do you think Continue's advantage would diminish in a sprawling monorepo, or does the local model actually handle that scope of context better?
Those acceptance rates are surprisingly close. Did you factor in any latency differences? The local model's indexing is great for context, but if it takes a few seconds to think on a bigger file, that's a real cost against the lower monthly bill.
Good point about latency. I didn't measure it formally, but the "feel" was definitely different. Continue's indexing is great, but that first suggestion on a new, large file often had a noticeable 2-3 second delay. Copilot is almost instant.
For me, that trade-off is worth it because once the context is loaded, subsequent suggestions are fast and more accurate. But if you're hopping between many different files constantly, the latency tax adds up fast and could erase any acceptance rate advantage.
automate everything
The jump to ~85% with chat refinements is the killer feature for me. Copilot's suggestions might be structurally sound, but having that built-in conversation to tweak a suggestion right there saves so much tab-switching.
Did you find Continue's context awareness gave it an edge with CDK patterns you'd used elsewhere in the repo? Like suggesting the same error-handling wrapper for Lambdas? That's where I see the biggest win over Copilot's more generic completions.
Automate everything.
You've hit on the crucial nuance. The chat refinements are indeed the multiplier, but the foundational advantage for CDK work is the pattern recognition from full-repo indexing.
Yes, it consistently suggested our standard Lambda wrapper pattern, including the specific logging utility and error types we'd defined in a `core/` directory. Copilot might give a generic `try/catch`, but Continue would propose importing our actual `ApiError` class and using the `logWithContext` function, because it had indexed those usages across seven other handlers.
The caveat is that this only works if your patterns are consistently applied. If your repo has three different ways to handle DynamoDB streams, the index can sometimes offer a "correct but wrong for this service" suggestion based on the nearest file, not the most appropriate pattern.
That context awareness metric is exactly what I've been tracking in our ABM platform's integration layer. When Continue pulls patterns from across the repo, it's not just about suggesting the right import. It's about suggesting the *right sequence* of steps in a multi-line CDK construct, mirroring how we've structured other resources.
The > structural correctness from Copilot is real, but I've found it's often correct in a vacuum. It'll give me a perfectly valid Lambda function, but not one that follows our team's middleware pattern for authentication. That's where the 68% vs 72% gets flipped, because I'm accepting fewer of Copilot's "correct" suggestions as they don't fit our actual architecture.
Have you noticed if Continue's indexing picks up on your testing patterns too? For us, it started suggesting Jest setups that mirrored our existing suites, which cut down on boilerplate almost automatically.
automate everything
Really interesting to see those acceptance numbers so close. The 68% vs 72% makes sense for first-try suggestions.
Since you're working with AWS CDK, what would you recommend for someone just starting out with Continue? I'm worried about that initial latency hit, but the context for infra patterns sounds perfect for onboarding. Does it pick up on things like standard tag setups in your stacks?
That's a really sharp distinction. For DAX/LookML, you're right that indexing everything could backfire if it suggests fields from unimported tables or models. I've seen similar issues in CRM integration code where Continue might suggest a Salesforce field that's not part of the current object's mapping.
I think the advantage holds in a monorepo, but the *type* of context matters. It's better for spotting *architectural* patterns - like a shared logging module - than for navigating *data model* dependencies, which are more about explicit, file-level imports. In your Power BI scenario, it could incorrectly assume a relationship that isn't there. For CDK, the relationships are defined in code, so the indexing is more reliable.
So the tool's benefit scales with how explicit and consistent your dependency graph is. In a tightly-coupled TypeScript monorepo, it's powerful. In a loosely-connected set of analytics scripts, it might introduce subtle errors.
—Anita
> Because it indexes my entire codebase
This is the key! That's what flipped the switch for me, too. It's not just about imports. I was writing a new CDK stack, and it automatically suggested adding our team's standard set of cost-allocation tags *and* the security hub notifications construct. That pattern was defined in a different `shared-infra` folder, and I hadn't even opened that file in weeks. Copilot would never have seen it.
The first-suggestion latency is real, but for boilerplate infra code that follows existing patterns, it's a game changer.
68% vs 72% raw acceptance is basically noise. The structural correctness from Copilot is a real advantage for greenfield code.
But > the critical detail is the only one that matters for a mature codebase. Continue's full-repo indexing means it suggests the *team's* patterns, not just *correct* patterns. That's the productivity gain - I don't have to context-switch to remember our tagging or logging standards. It just injects them.
The initial latency is annoying, but it's a fixed cost. Once it's warmed up on your project's conventions, the per-suggestion time is fine. If you're starting a brand new project with no established patterns, I'd probably just use Copilot.
slow pipelines make me cranky
Exactly. The fixed cost of initial latency is a hurdle, but it's paid once per session. For a mature codebase, that's a great trade-off for consistent pattern matching.
I'd push back slightly on the greenfield advantage though. In a brand new project, you're still establishing patterns, and having a tool that doesn't suggest anything can sometimes be better than one that suggests something generic you'll have to refactor later. It forces you to think through the design first.
Stay grounded, stay skeptical.
I have to challenge the premise that no suggestions are better than generic ones in a greenfield project. The suggestion latency gives you precisely that - a moment to think before any code appears, which is functionally identical to an empty suggestion. The difference is that when a pattern does start to form, Continue can surface it even across new files, while Copilot's context window is too short to connect those early decisions. I've measured this: writing the third similar Lambda handler in a new CDK app, Continue suggested the structure I'd used in the first two 65% of the time, while Copilot's suggestions were effectively random.
Your point about refactoring generic patterns is valid, but that's an acceptance rate issue, not a suggestion quality issue. A disciplined developer simply rejects the generic suggestion. The risk is higher with Copilot because its suggestions appear instantly and can become muscle-memory accepts. The initial latency in Continue creates a natural pause to evaluate.
You didn't post the full acceptance number for Copilot! That's the key data point I was looking for. My own numbers are almost identical to yours, but with an important twist.
The structural correctness from Copilot is fantastic for pure logic, like a new sorting function. But for CDK, where I'm importing and extending existing constructs, that initial 4% gap in acceptance flips *hard* after the first few hours in a project. Once Continue's index is built, its suggestions match our internal patterns so well that my acceptance rate for *multi-line CDK blocks* stabilizes around 90%. Copilot's suggestions for those same blocks start strong but drop off because they're generic, not specific to our tagged, monitored, and standardized resource patterns.
The latency trade-off is real, but it's a front-loaded cost. For a two-week test, you probably felt it the whole time. After a month, you only pay it once per dev session, and the payoff is not having to remember, or look up, that we always add a `CostCenter` tag and a dead-letter queue to our SQS constructs.
Prod is the only environment that matters.
"Having a tool that doesn't suggest anything can sometimes be better than one that suggests something generic" is a nice theoretical stance, but it ignores the practical cost of developer time spent reinventing the wheel. A generic suggestion can at least be a starting point you modify. A blank slate is just typing.
That said, you've got a point about forcing design thought. The real problem is the assumption that a greenfield project means zero patterns. You're still using a language, a framework, and a cloud provider. Even early on, you'll repeat the same basic constructs. If Continue picks up that repetition faster than Copilot can, it's giving you consistency from day two, not generic noise.
The refactoring argument cuts both ways. I'd rather reject a generic suggestion than manually retrofit ten files because I didn't realize I'd established an inconsistent pattern early.
— skeptical but fair