BitNet CPU Inference
BitNet provides GPU-free local model inference using Microsoft's
BitNet.cpp engine. It runs under the cpu-llm Compose
profile and is proxied through LiteLLM as the bitnet-2b
model alias. It is designed as an offline fallback when the remote
Ollama GPU endpoint is unavailable.
Features
- CPU-only inference — no GPU required at all
- 1-bit (ternary) weight quantization: ~0.4 GB RAM for the 2B model (non-embedding)
- OpenAI-compatible API via
llama-server(BitNet.cpp's built-in server) - LiteLLM integration: agents use
bitnet-2blike any other model alias - Configurable thread count and context size via environment variables
- Health check on
http://localhost:8080/health - Resource limits: 4 CPU cores, 4 GB RAM (enforced in Compose)
- Works offline — no VPN or internet required once the model is downloaded
How It Works
BitNet.cpp uses 1-bit (ternary: -1, 0, +1) weights natively trained in this format. This is not post-hoc quantization of a float16 model; only models specifically trained in 1-bit format are compatible.
The bitnet service runs llama-server —
BitNet.cpp's built-in inference server that exposes an OpenAI-compatible
API at port 8080. LiteLLM proxies it via the openai
provider:
- model_name: bitnet-2b
litellm_params:
model: openai/bitnet-2b
api_base: http://bitnet:8080
api_key: "none"
Functionalities
Downloading a Model
Before starting the service, download a compatible GGUF model file. The recommended starting model:
# Install huggingface-cli if needed
pip install huggingface-hub
# Download the 2B model
huggingface-cli download microsoft/bitnet-b1.58-2B-4T-gguf \
--local-dir ./models/bitnet/
The model files land in ./models/bitnet/ and are mounted
read-only into the container.
Starting BitNet
docker compose --profile cpu-llm up -d bitnet
The make up-full target starts all profiles including
cpu-llm.
Configuration
BITNET_MODEL=/models/ggml-model-i2_s.gguf # path inside container
BITNET_CTX_SIZE=4096 # context window in tokens
BITNET_THREADS=4 # CPU threads for inference
Adjust BITNET_THREADS to match available CPU cores. More
threads → faster tokens/second, up to the core count.
Using BitNet from an Agent
From inside the agent container:
# Direct API test
curl -s http://bitnet:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"bitnet-2b","messages":[{"role":"user","content":"hello"}]}'
# Through LiteLLM
curl -s http://litellm:4000/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"bitnet-2b","messages":[{"role":"user","content":"hello"}]}'
To use BitNet in OpenCode, switch the model alias to
bitnet-2b. To use it in Continue, update the model field in
the Continue config.
Health Check
docker compose ps bitnet
curl -s http://127.0.0.1:4000/v1/models | jq '.data[].id'
The container takes up to 60 seconds on first load while
llama-server reads the model file.
Available Models
Only natively 1-bit-trained models are compatible with BitNet.cpp.
| Model | Params | CPU RAM | Notes |
|---|---|---|---|
microsoft/bitnet-b1.58-2B-4T-gguf |
2B | ~1–2 GB | Default; stable and fast |
HF1BitLLM/Llama3-8B-1.58-100B-tokens |
8B | ~4–5 GB | Best coding quality |
tiiuae/Falcon3-7B-Instruct-1.58bit |
7B | ~3–4 GB | Best reasoning option |
Start with the 2B model. Use 8B+ only if RAM allows.
Performance (approximate, modern CPU):
- 2B model: 5–20 tokens/second on 4–8 cores
- 8B model: 2–8 tokens/second on 4–8 cores
- AVX-512 (Intel) gives 2–6× speedup over AVX2
Limitations
- Only native 1-bit models. Arbitrary float16 models cannot be quantized into BitNet format. The model selection is small compared to Ollama.
- 4,096 token default context. The 2B model's context
is limited. Increase
BITNET_CTX_SIZEat the cost of more RAM (diminishing performance above 8k). - Single-user throughput. CPU inference is single-threaded per request. 5–20 t/s is fine for one agent; concurrent requests queue and slow down significantly.
- Build complexity. The BitNet Dockerfile uses Clang 18 (GCC is not supported). The build image is ~10 GB; the runtime image ~2 GB. First build takes 10–20 minutes.
- ARM64 known bug.
ggml-bitnet-mad.cppline 811 has a known ARM64 issue. The Dockerfile includes a workaround. Performance on ARM may be lower than the published benchmarks. - Limited non-English quality. Microsoft reports elevated error rates on non-English prompts, particularly for the 2B model.
- LiteLLM starts cleanly without BitNet. The
cpu-llmprofile is opt-in. Requests tobitnet-2bfail with a 502 when BitNet is not running; other model aliases are unaffected.
Hardware Requirements
| Requirement | Minimum | Recommended |
|---|---|---|
| CPU architecture | x86-64 with AVX2 | x86-64 with AVX-512 or ARM with NEON |
| CPU cores | 2 | 4–8 |
| System RAM (2B model) | 4 GB | 8 GB |
| System RAM (8B model) | 8 GB | 16 GB |
| Disk (model files) | 1 GB | 2 GB |
| GPU | Not needed | Not needed |
The Compose resource limits enforce cpus: "4" and
memory: 4g. Adjust in compose.yml for larger
models.