Continue
Continue is a VS Code extension that provides AI-powered chat, code editing, and autocomplete inside code-server. It connects to LiteLLM, so it uses the same remote Ollama model as the rest of the stack.
Features
- In-editor chat with context from open files and selections
- Inline code edit and apply: describe a change, preview it, apply in one keystroke
- Tab autocomplete powered by the same LiteLLM gateway as opencode
- Model switching without reconfiguring the extension (change the
alias in
config.yaml) - Persistent config in
.local/volumes/vscode-config/continue/ - Secrets decoupled from config via a local
.envfile
Functionalities
Configuration
The seeded config lives at:
.local/volumes/vscode-config/continue/config.yaml
The source template is at config/continue/config.yaml.
On first make init, the template is copied to the
volume.
models:
- name: RemoteLLM LiteLLM
provider: openai
model: ollama-default
apiBase: http://litellm:4000/v1
apiKey: ${{ secrets.OPENAI_API_KEY }}
roles: [chat, edit, apply]
defaultCompletionOptions:
temperature: 0.2
- name: RemoteLLM LiteLLM Autocomplete
provider: openai
model: ollama-default
apiBase: http://litellm:4000/v1
apiKey: ${{ secrets.OPENAI_API_KEY }}
roles: [autocomplete]
useLegacyCompletionsEndpoint: false
autocompleteOptions:
debounceDelay: 350
maxPromptTokens: 1024
onlyMyCode: true
API Key Secret
Continue resolves ${{ secrets.OPENAI_API_KEY }}
from:
.local/volumes/vscode-config/continue/.env
This file is created by make init with the value from
OPENAI_API_KEY (which maps to
AGENT_VIRTUAL_KEY or LITELLM_MASTER_KEY).
Update it after rotating the key:
# edit .local/volumes/vscode-config/continue/.env
OPENAI_API_KEY=sk-...
Switching Models
Change the model field in the Continue config to any
alias defined in config/litellm/config.yaml:
model: claude-default # use Anthropic via LiteLLM
model: bitnet-2b # use local BitNet (cpu-llm profile must be active)
No extension reinstall needed; the change takes effect on next Continue request.
Autocomplete Token Budget
The maxPromptTokens: 1024 default limits how much
context Continue sends for autocomplete. This keeps latency low with
large Ollama models over the VPN. Increase it for larger context but
expect slower suggestions:
maxPromptTokens: 2048
Installing the Extension
If the extension is not already installed in code-server, install it from the VS Code Open VSX marketplace within the browser IDE (search "Continue").
Limitations
- Requires code-server. Continue is a VS Code extension and only works inside the browser IDE.
- Autocomplete latency depends on model and VPN
throughput. Remote Ollama models typically take 1–5 seconds per
completion. The
debounceDelay: 350default reduces noise but cannot eliminate latency. - Single model for autocomplete. The seeded config
uses
ollama-defaultfor autocomplete. Switching to a purpose-built autocomplete model (e.g., a smaller fill-in-the-middle model) requires an additional entry in the LiteLLM config. maxPromptTokens: 1024is conservative. It may truncate context on large files. Increase only if the target model handles longer prompts well.- Config reset is manual. Customizations made in the
volume are not overwritten by
make init. If you want to reset to the seeded config, copy it manually:cp config/continue/config.yaml .local/volumes/vscode-config/continue/config.yaml.
Hardware Requirements
Continue itself has no independent hardware requirements — it runs as part of code-server. The autocomplete and chat workloads are sent to LiteLLM and then to the remote Ollama model. Hardware requirements are those of code-server and the upstream model server.