Architecture
RemoteLLM is a Docker Compose workspace for running coding agents behind a single LiteLLM gateway. Ollama is an optional model source: point it directly at a reachable Ollama instance, or route to one over your own WireGuard VPN. Neither is required — agents work through Anthropic (Claude) or the CPU-only BitNet profile without any Ollama or VPN setup.
Services
wireguard: optional VPN profile service that joins the VPN and owns the VPN network path.vpn-proxy: optional VPN profile service that shares the WireGuard network namespace and exposes a SOCKS5 proxy onwireguard:1080.agent: long-running coding-agent container with opencode, Codex home folders, Claude config folders, package safety tools, and project mounts.code-server: optional browser IDE at/code/.web-terminal: optional browser terminal at/terminal/.codex: optional one-shot Codex CLI profile using the same project mounts.claude: optional one-shot Claude Code profile using the same project mounts.hermes: optional one-shot Hermes Agent profile using the same project mounts.hermes-web: optional Hermes browser dashboard at portHERMES_WEB_PORT(default8645).litellm: OpenAI-compatible model gateway for routing calls to Ollama and Anthropic. Authenticated with virtual keys.litellm-postgres: PostgreSQL instance backing LiteLLM for virtual key persistence, usage logs, and team settings.langfuse-*: optional observability profile that stores LiteLLM traces, errors, and usage metrics.nginx: local reverse proxy bound to127.0.0.1:${NGINX_PORT:-8088}by default.
Network Flow
When the VPN profile is started manually, only LiteLLM uses the WireGuard proxy:
OLLAMA_PROXY=socks5h://wireguard:1080
NO_PROXY=localhost,127.0.0.1,agent,nginx,code-server,litellm,wireguard
The vpn-proxy service runs in the same network namespace
as wireguard, and both services stay off until
make up-vpn or the vpn Compose profile is
used. The agent and terminal do not set ALL_PROXY,
HTTP_PROXY, or HTTPS_PROXY; this keeps npm,
Bun, pip, Git, and opencode plugin installs off the VPN. Agents call
LiteLLM at http://litellm:4000/v1; LiteLLM can use the VPN
path when configured for wireguard:11434 or
OLLAMA_PROXY=socks5h://wireguard:1080. DNS resolution
happens through the SOCKS proxy when socks5h is used.
Volumes
Runtime state lives under .local/; it is machine-local,
generated by make init, and ignored by git.
.local/workspace/projects: editable project code..local/workspace/inputs: read-only input material..local/workspace/outputs: generated output and exports..local/volumes/agent-home: agent user home..local/volumes/opencode-home: opencode state..local/volumes/codex-home: Codex state and skills..local/volumes/claude-home: Claude state..local/volumes/hermes-home: Hermes state, skills, and dashboard database (kanban.db)..local/volumes/aide-memory: local agent memory..local/volumes/vscode-config/continue: Continue config and local secrets for code-server..local/volumes/wireguard: WireGuard runtime config and state..local/volumes/litellm: LiteLLM runtime state..local/volumes/litellm-postgres: LiteLLM PostgreSQL data (virtual keys, usage logs)..local/volumes/langfuse: optional Langfuse Postgres, ClickHouse, MinIO, and Redis state.
Reverse Proxy Paths
nginx exposes local browser entry points:
/code/->code-server:8080/terminal/->web-terminal:7681/api/litellm/->litellm:4000(prefix stripped before forwarding)/api/litellm/ui-> LiteLLM admin dashboard (login:LITELLM_UI_USERNAME/LITELLM_UI_PASSWORD)
The terminal and IDE both need WebSocket upgrade headers. nginx
preserves the full request URI so ttyd endpoints such as
/terminal/token and /terminal/ws reach ttyd
correctly.
Interactive shells source a managed RemoteLLM zsh profile from
/home/agent/.remotellm.zsh. The agent image includes Oh My
Zsh, a remotellm-agnoster theme based on
agnoster, and Powerline fonts so the browser terminal and
make shell use the same prompt and aliases. The custom
theme shows context, current directory, and git on the first line, uses
RemoteLLM color defaults, and places the > command
marker on a second line.
tmux is installed in the agent image and seeded to
/home/agent/.tmux.conf on first startup. The default tmux
configuration launches zsh -l, uses true-color terminal
settings, enables mouse support, and keeps pane/window creation in the
current directory.
Images
Custom Dockerfiles are multistage where useful to keep runtime images
smaller and avoid carrying build tooling into final stages. The agent
image pins key tool versions through build args in
.env.
The agent apt package list lives in
docker/agent/apt-packages.txt. Add common base applications
there, rebuild the image, and recreate the agent/terminal containers.
For one-off tools during development, use
sudo apt-get update and
sudo apt-get install ... inside the agent.
The agent and browser terminal allow passwordless sudo for
development installs. This intentionally removes
no-new-privileges from those two shell containers; other
services keep the stricter setting.
Memory And Skills
The entrypoint creates /workspace/.aide/memory,
/workspace/.aide/state, and
$HOME/.codex/skills on startup. Seed skills from the image
are copied only when missing, so edits in the mounted volume are
preserved.
Continue
The browser IDE mounts
.local/volumes/vscode-config/continue as
/home/coder/.continue. Host setup seeds
config.yaml from config/continue/config.yaml
and creates a local .env file with
OPENAI_API_KEY for Continue secret resolution. The seeded
config uses Continue's openai provider against LiteLLM at
http://litellm:4000/v1 and the ollama-default
model alias.
Optional Agents
The codex, claude, hermes, and
hermes-web Compose profiles are disabled by default. They
run wrapper commands from the agent image and share the normal
workspace, memory, and tool volumes. To keep builds reproducible, the
image only installs these CLIs when the corresponding variables are set
in .env.
Set INSTALL_CODEX_CLI=1 with
CODEX_CLI_PACKAGE=@openai/codex@<version> to include
Codex CLI. Set INSTALL_CLAUDE_CODE=1 with
CLAUDE_CODE_PACKAGE=@anthropic-ai/claude-code@<version>
to include Claude Code. Rebuild after changing either setting.
For Hermes, set INSTALL_HERMES_AGENT=1 and choose an
install method with HERMES_INSTALL_METHOD:
pypi— installs a pinned version from PyPI; requiresHERMES_AGENT_VERSION.official— runs the Nous Research install script at build time; installs the latest release;HERMES_AGENT_VERSIONis ignored.
The hermes profile runs the agent in one-shot mode
(run-hermes-agent). The hermes-web profile
starts the Hermes browser dashboard on port HERMES_WEB_PORT
(default 8645). The dashboard includes a Kanban board,
skills browser, and a live PTY terminal for the running agent.
Hermes state is persisted to .local/volumes/hermes-home
(mounted as /home/agent/.hermes inside the container).
Dependency Safety
The agent image includes:
safe-package-check: project-level package audit helper.safe-npm-install: npm install wrapper with preflight checks.safe-pip-install: pip install wrapper with preflight checks.pip-audit: Python dependency vulnerability scanning.osv-scanner: multi-ecosystem vulnerability scanning for npm, PyPI, Go, Rust, and more.scorecard: repository supply-chain posture checks.
Use these wrappers before adding dependencies from inside the agent container.
Resource Requirements
Figures below are approximate. "Idle" is the container running but not processing a request; "peak" is under active agent workload.
Per-service
| Service | Idle RAM | Peak RAM | CPU (idle) | CPU (peak) | Notes |
|---|---|---|---|---|---|
wireguard |
64 MB | 96 MB | <0.1 | 0.2 | Kernel module management only |
vpn-proxy |
16 MB | 32 MB | <0.1 | 0.1 | socat + microsocks |
agent |
256 MB | 4–8 GB | 0.1 | 4.0 | Peaks during build/test tasks |
nix-agent |
512 MB | 6–8 GB | 0.1 | 4.0 | +5–10 GB disk for /nix store on first run |
nginx |
16 MB | 64 MB | <0.1 | 0.2 | Negligible |
code-server |
512 MB | 2–4 GB | 0.5 | 4.0 | VS Code extensions add significant overhead |
web-terminal |
128 MB | 512 MB | <0.1 | 1.0 | ttyd + zsh |
litellm |
384 MB | 1 GB | 0.2 | 1.0 | Python proxy; includes Prisma ORM |
litellm-postgres |
128 MB | 384 MB | <0.1 | 0.3 | Small dataset — keys and usage logs only |
playwright-mcp |
512 MB | 2 GB | 0.5 | 2.0 | Chromium is memory-heavy |
langfuse-web |
512 MB | 1 GB | 0.3 | 1.0 | Next.js SSR |
langfuse-worker |
256 MB | 512 MB | 0.2 | 0.5 | Background trace processor |
langfuse-postgres |
256 MB | 1 GB | 0.2 | 1.0 | Grows with trace volume |
langfuse-clickhouse |
1 GB | 3 GB | 0.5 | 2.0 | Analytics engine; memory-hungry |
langfuse-minio |
128 MB | 256 MB | <0.1 | 0.3 | Object storage for media uploads |
langfuse-redis |
32 MB | 128 MB | <0.1 | 0.2 | Ephemeral cache |
By profile
| Profile | Services | Min RAM | Recommended RAM | Recommended CPU |
|---|---|---|---|---|
Core (make up) |
agent, litellm, litellm-postgres, nginx | 2 GB | 4 GB | 2 cores |
+ VPN (make up-vpn) |
+ wireguard, vpn-proxy | 2 GB | 4 GB | 2 cores |
+ IDE (make up-ide) |
+ code-server, web-terminal | 3 GB | 8 GB | 4 cores |
| + Observability | + all langfuse services | 6 GB | 16 GB | 6+ cores |
Full (make up-full) |
all of the above | 8 GB | 16 GB | 8+ cores |
ClickHouse (used by Langfuse) is the single largest consumer — 1 GB
minimum, often 2–3 GB under load. If memory is tight, skip the
observability profile and use make logs for basic
visibility instead.
Excluded Features
Docker Socket Access
The agent container does not mount the host Docker socket. No
confirmed use case requires DinD or host Docker access in this stack. If
a use case emerges (e.g., agent-driven container builds), it should be
added as an explicit docker-host profile with a separate
security review. Relevant tracking: remote-coding-agent-stack.md open
decisions.