Why local deployment is the most reliable "free tokens"

Platform credits expire, rate limits change, and risk systems claw grants back. Open weights don't. As long as the license (Apache 2.0, MIT, or a permissive community license) allows self-hosting, once the weights are on your disk nobody meters your inference calls — you pay in electricity and GPU depreciation.

The 2026 reality: 7B–14B open models run at usable speed on consumer GPUs, and the two main paths split cleanly — Ollama for "install and go," vLLM for concurrency and long context.

Hardware thresholds: figure out what you can actually run

VRAM is the only hard constraint. Rough weight footprints at 4-bit quantization:

  • 7B–8B: ~5–6 GB, fits an RTX 3060 12G / 4060 Ti 16G
  • 14B: ~9–11 GB, fits 4060 Ti 16G or 4070 Ti Super 16G
  • 32B: ~20–22 GB, needs a 4090 24G or dual cards
  • 70B: 40 GB+, needs 2×24G or an A6000-class card

Remember KV cache on top of weights. Longer context and higher concurrency mean a bigger cache. A single request at 8K context typically adds 1–3 GB; budget multiples for 32K and beyond.

CPU-only works but is usually a few tokens per second on a 7B model — fine for offline batch jobs only. Apple Silicon unified memory is the exception: decent speed when memory is large enough, good for personal daily use.

Path 1: Ollama — first local inference in ten minutes

Steps

  1. Install from the official site (macOS/Windows) or the install script (Linux).
  2. Pull a model: ollama pull qwen3:8b (or llama3.1:8b, gemma3:12b, depending on VRAM).
  3. Chat directly: ollama run qwen3:8b.
  4. It already exposes an OpenAI-compatible API at http://localhost:11434/v1/chat/completions.
  5. Point your app's base_url to http://localhost:11434/v1; any non-empty string works as the API key.
  6. Verify with curl:

```bash

curl http://localhost:11434/v1/chat/completions \

-H "Content-Type: application/json" \

-d '{"model":"qwen3:8b","messages":[{"role":"user","content":"hi"}]}'

```

Key settings

  • OLLAMA_HOST=0.0.0.0:11434 to serve other devices on your LAN.
  • OLLAMA_KEEP_ALIVE=30m to keep the model resident and avoid cold starts.
  • OLLAMA_NUM_PARALLEL controls concurrent requests; set it too high and single-request speed suffers.
  • Context length is set with PARAMETER num_ctx in a Modelfile. The default is often too small; raise it for long documents.

Path 2: vLLM — switch when you need throughput and long context

Ollama is fine for one person or a small team, but concurrency causes queueing. vLLM's PagedAttention and continuous batching multiply throughput and are the mainstream choice for self-hosted serving.

Steps

  1. Prefer Linux + NVIDIA GPU with matching CUDA drivers.
  2. Create a venv and pip install vllm (follow the docs for the install matching your CUDA version).
  3. Launch the OpenAI-compatible server:

```bash

vllm serve Qwen/Qwen3-8B \

--max-model-len 32768 \

--gpu-memory-utilization 0.90 \

--port 8000

```

  1. Point clients at http://localhost:8000/v1.
  2. For multi-GPU or tight VRAM, use --tensor-parallel-size 2.

Key flags

  • --gpu-memory-utilization: defaults to 0.9; lower it on small cards or you'll OOM at startup.
  • --max-model-len: directly drives KV cache size — set it explicitly for long context.
  • --quantization: AWQ / GPTQ / FP8 supported; quantization cuts VRAM needs significantly.
  • --max-num-seqs: concurrency cap, tune alongside VRAM.

Wiring local models into your workflow

  • Code completion: VS Code extensions like Continue and Cline accept a custom OpenAI-compatible endpoint — just enter the local URL.
  • CLI: any tool honoring OPENAI_BASE_URL can point at localhost.
  • Desktop clients: Cherry Studio, Chatbox and similar let you add a custom API base.
  • Your own gateway: put a proxy like LiteLLM in front so local and cloud models share one route — sensitive tasks local, hard tasks remote.

Caveats

  1. Read the license. Llama models carry a community license; check terms before commercial use. Apache 2.0 / MIT models (Qwen, some Mistral releases) are more permissive.
  2. Quantization costs quality. 4-bit is acceptable for most tasks but degrades math and long-chain reasoning noticeably — use 8-bit or full precision for critical work.
  3. Security: local servers have no auth by default. Add authentication before exposing any port publicly.
  4. Power and thermals: long full-load runs on a consumer GPU mean real electricity and noise costs.
  5. Not everything belongs locally. For the hardest reasoning, an 8B local model still trails frontier cloud models — hybrid use is more realistic.
  6. Versions change. Model names and flags shift between releases; always check the official repo README.

When this fits

  • High-frequency, short-request batch work: classification, extraction, translation, data cleaning — near-zero marginal cost locally.
  • Data that cannot leave the network: contracts, medical records, internal code.
  • Offline or poor-connectivity environments: field deployments, edge devices.
  • Learning and debugging: understanding inference, quantization, and KV cache first-hand.

Bottom line

If your call volume is high or your data is sensitive, self-hosting open models is the most controllable "free tokens" route in 2026 — cost shifts from per-call billing to one-time hardware, and what's left is tuning.