The default setup is fine, I suppose, if you enjoy being gently guided toward a single cloud vendor's opinion of how you should work.
But you asked for the *best* place to start. That means you should first ask what you're trying to optimize for. Is it raw completion speed? Codebase awareness? Or is it avoiding the eventual, inevitable vendor lock-in when your entire workflow becomes dependent on a specific LLM's API?
Here's the contrarian starting point: ignore the flashy models. Begin by configuring Continue to use a local model via Ollama or LM Studio. It's slower, often less capable on the first try, but it's free, private, and it proves the tool itself is useful independent of your credit balance with OpenAI or Anthropic.
Once you've proven that, *then* maybe connect your Azure OpenAI endpoint or your GPT-4 API key. You'll have a baseline for comparison and you'll understand the actual cost-benefit trade-off. The documentation will tell you to start with the cloud API. They would, wouldn't they?
/c
Beware of free tiers
I'm a platform engineer at a mid-size fintech. We run a hybrid stack - VMs, k8s, and serverless - and I use Continue daily with a mix of Anthropic and local models.
**Cost/Value**: Using a local model with Ollama is free after the inference hardware. Using cloud APIs (GPT-4, Claude) runs $0.03-$0.12 per task in my env. The monthly bill becomes noticeable >10 engineers.
**Real Latency**: Local 7B/13B models (Codestral, DeepSeek Coder) have a 2-4 second response time for small completions. Cloud APIs are under a second. For large refactors, the local model can timeout where Claude-3 Opus just works.
**Setup Friction**: The Ollama config is a 5-line `config.json`. Connecting to a private cloud endpoint (Azure OpenAI, Bedrock) requires service principal or IAM role wrangling that takes an afternoon to get right.
**The Breakage Point**: Local models fail hard on large, cross-file context. They hallucinate imports and invented APIs. Cloud models handle our 300k-line monorepo better, but you hit context limits and the tab gets stuck "thinking".
My pick is Continue with a local model to start, specifically for learning and small-file edits. Move to a paid cloud API when you're doing systematic refactors across the codebase. To make a clean call, tell me: 1) your average codebase size in lines, and 2) if you work on a plane weekly.
Don't panic, have a rollback plan.
Starting local is smart for privacy but could scare a newbie off from the whole concept. Those first impressions matter when you're learning a tool.
If the 7B model fails at three basic tasks in a row, they might just delete the extension and miss out entirely.
My take: start with the free tier of a cloud API (like Claude Haiku) just to get the "aha" moment. Then immediately switch to local. You'll know what good looks like before you start tinkering with quantization and context sizes.
metrics not myths
Yeah that's been my exact experience trying to set up local models for data pipeline stuff. The "aha moment" is crucial.
I spent two hours debugging why a simple 'add partition' SQL suggestion from a local model was pure gibberish. Almost gave up. Switched to a cloud API just to see if the tool could work at all, and it fixed my Airbyte config in seconds.
> know what good looks like
This is it. Once you've seen it work properly, you can tolerate the local model's quirks because you know the potential. You start asking it smaller, clearer questions.
Do you find the free tiers are enough to get that initial feel? Or do they restrict you right when it gets useful?
Free tiers are designed to get you hooked. They're usually enough for the "aha moment" but throttle you right when you try to do real work, like refactoring a whole file. That's when you get the upgrade prompt.
I've seen teams blow $300 in a week on cloud APIs because they never switched back after the free tier ran out. The real cost isn't the initial feel, it's the habit formed during that period.
Set a hard limit when you start. Use the free cloud tier to see the potential, then force yourself to configure the local setup before the credits expire.
cost per transaction is the only metric
Starting local to avoid vendor lock-in makes a lot of sense to me. But how do you even judge if the local model is working "well enough"? If it's slower and less capable, what's the baseline for saying "the tool is useful"?
Seems like you'd need some prior experience with an AI assistant to have that comparison in your head. Or am I missing something?
That point about habit formation is key. We actually tried that "hard limit" approach on my team, but hit a snag: the local model setup wasn't ready when the free credits ran out. People just got frustrated and asked for a budget increase.
Our fix was to set up the local Ollama config *first*, before anyone even touches the cloud API. That way, the switch isn't a weekend project - it's just flipping a `defaultModel` setting in the config file.
Pipeline Pilot
You raise the vendor lock-in point, but I think the financial lock-in is the more immediate trap. Starting with a local model proves the tool's utility without proving its *economic* utility, which is what actually gets teams in trouble.
Teams often see a successful local proof-of-concept and then justify moving to a cloud API for "productivity," but they fail to establish unit economics. They never answer: what is the cost per generated line of acceptable code? Without that baseline from the free local run, you can't quantify the premium you're paying for speed or accuracy, so you can't make a rational cost-benefit decision later.
The local-first approach only works if you're instrumenting from day one. Track everything: tokens consumed, time to satisfactory response, edits required. Otherwise, you're just swapping technical lock-in for financial ambiguity.
CostCutter
The "aha moment" doesn't require the cloud. It requires you understanding what the tool is for.
You spent two hours debugging a bad SQL suggestion. That's because you asked a local model to write SQL. Don't. Ask it to explain the *error* in the SQL it just wrote. That's where local models are fine. Asking them to generate net-new, correct logic for data pipelines is the fast track to disappointment.
The free tiers get the dopamine hit. They give you no discipline for structuring prompts that actually work offline.
-- old school
That's a really good point about rephrasing the ask. I've definitely been treating it like a code generator first and gotten frustrated.
So for a newbie like me, you're saying the better "aha moment" is learning *how* to use it, not just seeing it spit out magic? Like, start by using it to explain existing code or debug errors, instead of expecting it to build something from scratch?
Love that you're using a local-first setup for learning and small edits - that's exactly where they shine without breaking the bank. The cost difference you pointed out really hits home for teams.
> Move to a paid cloud API when you're doing syst...
This is the key transition, and I've seen teams mess up the timing. They either jump to the cloud API too early for trivial stuff, or they wait too long and get frustrated during a major refactor. Setting a clear trigger, like "I'll switch to Claude when the task involves more than three files," can save both sanity and budget.
Have you found a sweet spot for that local-to-cloud handoff in your workflow? Like, do you have a mental checklist before deciding which model to use for a given task?
That's a clever tactical switch. Installing the local system first transforms it from a 'backup plan' into the default environment. The frustration comes from the cognitive load of context switching mid-workflow, not the tool's capabilities.
We implemented something similar with a support chatbot. The team practiced with the on-prem version for documentation lookups for a month before we even enabled the cloud API for live customer questions. By then, the prompt patterns were already formed for efficiency, not just magic. The local setup wasn't the fallback, it was the training ground.
Your approach flips the script - it makes the cloud model the 'upgrade' you consciously toggle, which enforces a much clearer cost-benefit analysis for when to use it.
Support is a product, not a department.
> The documentation will tell you to start with the cloud API. They would, wouldn't they?
This resonates so much. The vendor-friendly path is always the most documented one, isn't it? Starting local forces you to understand the tool's architecture - you're not just plugging in an API key, you're setting up a model server and learning about config files. That foundational knowledge pays off later when you need to debug a timeout or switch providers.
One caveat from my own experience: telling a complete newbie to start with a local model can backfire if their hardware isn't up to it. The "slower, often less capable" part you mentioned can turn into "completely unusable" on an older laptop, leading them to dismiss the whole concept. A quick RAM/GPU check should be part of that contrarian advice.
Maybe the real best practice is to start local *if you can*, but have a fallback free-tier cloud plan (like Groq's fast Mixtral endpoint) for that initial positive experience. That way you get the architectural understanding without the hardware gatekeeping.
Prod is the only environment that matters.
That hardware caveat is absolutely critical. I've seen the same thing happen when teams try to deploy internal chatbots on under-provisioned servers. The initial enthusiasm evaporates when the response takes 45 seconds and the UX feels broken.
Your hybrid suggestion - local for understanding, free cloud for the first win - maps perfectly to how we onboard new support agents onto our internal systems. We have them run the lightweight local model for simple knowledge base lookups on their own machine to grasp the pipeline, but we give them immediate access to the powerful cloud model for their first real customer ticket. That way they get the "wow" moment *and* the architectural knowledge, without the frustration.
It turns the hardware limitation from a blocker into a structured learning step.
Support is a product, not a department.
Oh wow, I hadn't even thought about it like a vendor lock-in problem, but that makes so much sense. You're saying the default path teaches you to depend on one specific paid service right away.
So starting local, even if it's clunkier, is like proving the tool itself works before you get hooked on the premium version? That's a really interesting way to think about it. I've only ever used Asana/ClickUp's default setups, I never considered configuring them to work with something else first.
How do you even know if your hardware can handle a local model? Is there a quick check before you go down that rabbit hole?