Why local deployment is the most reliable "free tokens"
Platform credits expire, rate limits change, and risk systems claw grants back. Open weights don't. As long as the license (Apache 2.0, MIT, or a permissive community license) allows self-hosting, once the weights are on your disk nobody meters your inference calls — you pay in electricity and GPU depreciation.
The 2026 reality: 7B–14B open models run at usable speed on consumer GPUs, and the two main paths split cleanly — Ollama for "install and go," vLLM for concurrency and long context.
Hardware thresholds: figure out what you can actually run
VRAM is the only hard constraint. Rough weight footprints at 4-bit quantization:
- 7B–8B: ~5–6 GB, fits an RTX 3060 12G / 4060 Ti 16G
- 14B: ~9–11 GB, fits 4060 Ti 16G or 4070 Ti Super 16G
- 32B: ~20–22 GB, needs a 4090 24G or dual cards
- 70B: 40 GB+, needs 2×24G or an A6000-class card
Remember KV cache on top of weights. Longer context and higher concurrency mean a bigger cache. A single request at 8K context typically adds 1–3 GB; budget multiples for 32K and beyond.
CPU-only works but is usually a few tokens per second on a 7B model — fine for offline batch jobs only. Apple Silicon unified memory is the exception: decent speed when memory is large enough, good for personal daily use.
Path 1: Ollama — first local inference in ten minutes
Steps
- Install from the official site (macOS/Windows) or the install script (Linux).
- Pull a model:
ollama pull qwen3:8b(orllama3.1:8b,gemma3:12b, depending on VRAM). - Chat directly:
ollama run qwen3:8b. - It already exposes an OpenAI-compatible API at
http://localhost:11434/v1/chat/completions. - Point your app's base_url to
http://localhost:11434/v1; any non-empty string works as the API key. - Verify with curl:
```bash
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3:8b","messages":[{"role":"user","content":"hi"}]}'
```
Key settings
OLLAMA_HOST=0.0.0.0:11434to serve other devices on your LAN.OLLAMA_KEEP_ALIVE=30mto keep the model resident and avoid cold starts.OLLAMA_NUM_PARALLELcontrols concurrent requests; set it too high and single-request speed suffers.- Context length is set with
PARAMETER num_ctxin a Modelfile. The default is often too small; raise it for long documents.
Path 2: vLLM — switch when you need throughput and long context
Ollama is fine for one person or a small team, but concurrency causes queueing. vLLM's PagedAttention and continuous batching multiply throughput and are the mainstream choice for self-hosted serving.
Steps
- Prefer Linux + NVIDIA GPU with matching CUDA drivers.
- Create a venv and
pip install vllm(follow the docs for the install matching your CUDA version). - Launch the OpenAI-compatible server:
```bash
vllm serve Qwen/Qwen3-8B \
--max-model-len 32768 \
--gpu-memory-utilization 0.90 \
--port 8000
```
- Point clients at
http://localhost:8000/v1. - For multi-GPU or tight VRAM, use
--tensor-parallel-size 2.
Key flags
--gpu-memory-utilization: defaults to 0.9; lower it on small cards or you'll OOM at startup.--max-model-len: directly drives KV cache size — set it explicitly for long context.--quantization: AWQ / GPTQ / FP8 supported; quantization cuts VRAM needs significantly.--max-num-seqs: concurrency cap, tune alongside VRAM.
Wiring local models into your workflow
- Code completion: VS Code extensions like Continue and Cline accept a custom OpenAI-compatible endpoint — just enter the local URL.
- CLI: any tool honoring
OPENAI_BASE_URLcan point at localhost. - Desktop clients: Cherry Studio, Chatbox and similar let you add a custom API base.
- Your own gateway: put a proxy like LiteLLM in front so local and cloud models share one route — sensitive tasks local, hard tasks remote.
Caveats
- Read the license. Llama models carry a community license; check terms before commercial use. Apache 2.0 / MIT models (Qwen, some Mistral releases) are more permissive.
- Quantization costs quality. 4-bit is acceptable for most tasks but degrades math and long-chain reasoning noticeably — use 8-bit or full precision for critical work.
- Security: local servers have no auth by default. Add authentication before exposing any port publicly.
- Power and thermals: long full-load runs on a consumer GPU mean real electricity and noise costs.
- Not everything belongs locally. For the hardest reasoning, an 8B local model still trails frontier cloud models — hybrid use is more realistic.
- Versions change. Model names and flags shift between releases; always check the official repo README.
When this fits
- High-frequency, short-request batch work: classification, extraction, translation, data cleaning — near-zero marginal cost locally.
- Data that cannot leave the network: contracts, medical records, internal code.
- Offline or poor-connectivity environments: field deployments, edge devices.
- Learning and debugging: understanding inference, quantization, and KV cache first-hand.
Bottom line
If your call volume is high or your data is sensitive, self-hosting open models is the most controllable "free tokens" route in 2026 — cost shifts from per-call billing to one-time hardware, and what's left is tuning.