I've seen a lot of questions about getting Continue.dev to work with a self-managed LLM, specifically Ollama running on a separate machine. Most guides assume everything is on your local workstation. Running the model server remotely is a better use of hardware, but the setup isn't as obvious.
Here’s the straightforward process, focusing on the configuration that actually works. The key is in Continue's `config.json` file, located in `~/.continue/config.json` on macOS/Linux or `%USERPROFILE%.continueconfig.json` on Windows.
You need to add a model configuration for your remote Ollama instance. The critical part is the `apiBase` field. If your Ollama server is at `192.168.1.100` (just an example), your config should include a block like this:
```json
{
"models": [
{
"title": "My Remote Llama",
"provider": "ollama",
"model": "llama2",
"apiBase": "http://192.168.1.100:11434"
}
]
}
```
Important points:
* Use the machine's IP address or resolvable hostname. `localhost` or `127.0.0.1` will not work for a remote server.
* The default Ollama port is `11434`. Ensure this port is open on the remote server's firewall and accessible from your local machine.
* The `model` name must exactly match one you've pulled on the remote Ollama instance (e.g., `llama2`, `codellama`, `mistral`).
* Security note: This uses plain HTTP. Only do this on a trusted, private network. For wider reach, you'll need to set up a reverse proxy with HTTPS.
After saving the config, restart your IDE/editor with the Continue extension. The model should appear in the model dropdown. Test it. If it fails, the issue is almost always one of three things:
1. Network connectivity (firewall blocking port 11434).
2. Incorrect `apiBase` URL.
3. The specified model not existing on the remote Ollama server.
This setup decouples your development environment from your model hardware, which is the sane way to do it. Make sure your remote Ollama server has reliable uptime, or your coding flow will be interrupted.
—Chloe
SLA is not a suggestion.
This worked for me after I sorted out a firewall rule on the server. The part about the `apiBase` field being critical is spot on. I was stuck for a while because I initially tried using the server's local hostname, but my local machine couldn't resolve it. I ended up having to use the raw IP like you showed.
One thing I'd add: if you're on a Mac and using the Continue desktop app, you have to restart the entire app after editing the config.json file for the changes to take effect. Just closing and reopening the window wasn't enough for me.
This is really helpful. The apiBase field clarification is exactly what I was missing when I tried this last week. I was using the server's internal Docker IP, which obviously my laptop couldn't reach.
I'm curious about the model naming. If I have multiple remote instances, each with the same model name like "llama2," can I differentiate them in the Continue UI just by the title I give them in the config, or does it get confusing? Also, have you found any noticeable latency issues when the model server isn't on the same local network?
Glad the apiBase tip helped. The Docker IP trap is a classic one.
You can absolutely differentiate multiple remote instances with the same model name solely by the "title" field in your config. The Continue UI will show whatever title you give each entry, so you could have "llama2 - office server" and "llama2 - home lab" listed clearly.
On latency, it's definitely perceptible if you're working over a standard home internet link, mostly on the initial generation. The tokens can feel like they're dripping out. On a solid local network, it's usually fine. The bigger issue for me has been stability; if the remote server reboots or the VPN drops, Continue just hangs until it times out.
Trust the data, not the demo.
The stability point is key. I've found that timeout behavior is the main practical limitation for daily remote usage. Continue's default timeout seems quite long, so a dropped connection can freeze the interface for an extended period.
To manage multiple instances like you described, I maintain a primary and a fallback configuration. My config lists two entries for the same model type, one pointing to a remote server and one to a local Ollama instance as a backup. When the remote connection fails, I can switch models in the UI with a couple clicks, which is faster than editing the config. It's not seamless, but it reduces downtime.
Regarding latency over home internet, have you experimented with adjusting the 'context length' parameter in the model config for remote instances? Using a slightly lower value than the model's maximum can sometimes improve responsiveness, trading some capability for perceived speed.
Great walkthrough. The IP address tip is so important. I've had teammates spend an hour troubleshooting only to find they were using a localhost reference from a guide.
One extra caveat on port 11434: if you're setting this up on a cloud server, double-check your cloud provider's network security rules, not just the local firewall. I got bitten by that once - the instance firewall was open but the AWS security group wasn't.
Oh yeah, the full restart on Mac is a good tip! I thought just the window would do it too.
That hostname vs IP thing gets me all the time with Asana webhooks. My local machine never sees the server names. Did you have to add an entry to your hosts file, or was using the IP the only fix that worked for you?
Good to see the apiBase field getting the emphasis it needs. I've noticed people often miss the trailing slash on custom base URLs, which can cause silent failures. The Ollama default seems fine without it, but if you're routing through a reverse proxy with a path prefix, that's another thing to validate.
Your point on the firewall is the operational key. Beyond just opening port 11434, consider the inbound rule scope. If you're on an untrusted network, tunneling the connection over SSH is a more secure default than exposing the Ollama port directly.
Less spend, more headroom.
I hadn't considered the reverse proxy scenario. That's a good point. Does the `/api` path need to be appended to the apiBase when a proxy is involved, or does the proxy handle that mapping automatically?
The SSH tunnel suggestion is interesting for security, but wouldn't that require a persistent SSH connection to be maintained, separate from Continue's own process? That feels like another point of potential stability issues.
The double-layered firewall trap is a rite of passage in the cloud. Everyone gets burned by it once. It's the classic "my service is running, I can ping the port locally, why can't I connect?" scenario.
Beyond AWS security groups, the real fun begins with managed Kubernetes or container services. Their network policies and ingress controllers add a third or fourth layer to forget.