First, Be Honest About Where the “Free” Comes From

Running open models locally isn’t zero-cost — it swaps per-token billing for hardware and electricity. The genuinely free zone is: you already own a machine with a discrete GPU (or a spare inference card), and your call volume is high enough that the marginal cost per million tokens approaches your power bill.

So the goal here isn’t “how to install Ollama.” It’s how to wrap a locally served model into an OpenAI-compatible token supply layer, so existing code switches over without changes and the cloud only catches what local can’t handle.

Steps

Step 1: Pick model and quantization by VRAM

Rough inference VRAM formula:

```

VRAM ≈ params(B) × bytes-per-param + KV cache + framework overhead

```

Bytes per parameter by quantization:

| Quant | Bytes/param | 7B needs | 14B needs | 32B needs | 70B needs |

|---|---|---|---|---|---|

| FP16 | 2.0 | ~14 GB | ~28 GB | ~64 GB | ~140 GB |

| Q8 | 1.0 | ~7 GB | ~14 GB | ~32 GB | ~70 GB |

| Q4_K_M | ~0.55 | ~4 GB | ~8 GB | ~18 GB | ~40 GB |

Then leave headroom for KV cache: the longer the context and the higher the concurrency, the bigger it gets. A practical rule is 20%–40% of VRAM beyond the weights, more for long-context workloads.

Practical takeaways:

  • Single 8GB consumer card: 7B–8B at Q4, keep context under 8K, single stream is acceptable.
  • Single 16GB: 14B Q4, or 7B Q8 with light concurrency.
  • Single 24GB (3090/4090 class): 32B Q4 is the sweet spot, or 14B Q8 for higher concurrency.
  • Multi-GPU / 48GB+: 70B Q4 becomes realistic. Don’t force it otherwise.

Step 2: Ollama — fastest path to an OpenAI-compatible endpoint

Ollama exposes an OpenAI-compatible /v1 route on port 11434 by default. This is its most underrated feature.

```bash

# After install, pull models

ollama pull qwen3:8b

ollama pull llama3.1:8b

# Verify the OpenAI-compatible endpoint works

curl http://localhost:11434/v1/chat/completions \

-H "Content-Type: application/json" \

-d '{

"model": "qwen3:8b",

"messages": [{"role":"user","content":"Explain KV cache in one sentence"}]

}'

```

Note: Ollama’s /v1 endpoint does not validate API keys. Pass any non-empty string — many SDKs require the field to be present.

Bake system prompts and sampling params into a Modelfile so you don’t burn context re-sending a long system prompt every request:

```dockerfile

# Modelfile

FROM qwen3:8b

PARAMETER temperature 0.3

PARAMETER num_ctx 8192

PARAMETER num_predict 1024

SYSTEM """You are a rigorous technical assistant. Be concise. Say you don't know when unsure."""

```

```bash

ollama create my-assistant -f Modelfile

ollama run my-assistant

```

Make Ollama persistent and reachable from other hosts (it binds to localhost by default):

```bash

# On Linux with systemd

sudo systemctl edit ollama.service

# Add:

# [Service]

# Environment="OLLAMA_HOST=0.0.0.0:11434"

# Environment="OLLAMA_KEEP_ALIVE=-1" # keep model resident in VRAM

# Environment="OLLAMA_NUM_PARALLEL=4" # concurrent requests

```

OLLAMA_KEEP_ALIVE=-1 matters a lot: by default the model unloads after a few idle minutes, and the next request pays a reload that can push time-to-first-token into the tens of seconds.

Step 3: vLLM — switch when you need concurrency and throughput

Ollama suits single users and low concurrency. Once you’re serving a team or a service, vLLM’s PagedAttention and continuous batching typically deliver several times the throughput.

```bash

pip install vllm

# Start an OpenAI-compatible server

vllm serve Qwen/Qwen3-8B \

--served-model-name qwen3-8b \

--host 0.0.0.0 --port 8000 \

--max-model-len 8192 \

--gpu-memory-utilization 0.90 \

--max-num-seqs 32

```

Key flags, one by one:

  • --gpu-memory-utilization 0.90: lets vLLM claim 90% of VRAM for weights + KV cache pool. Too low wastes capacity; too high risks OOM.
  • --max-model-len: the most common footgun. If unset, vLLM reserves KV cache for the model’s nominal max context (possibly 32K/128K) and blows out VRAM. Set it to what you actually need.
  • --max-num-seqs: cap on concurrently processed sequences. Lower it if you’re VRAM-constrained but want high concurrency.
  • --tensor-parallel-size N: multi-GPU tensor parallelism, N = number of GPUs.
  • --quantization: specify when serving AWQ / GPTQ / FP8 quantized weights.

For quantized weights, grab an existing AWQ or GPTQ release on Hugging Face rather than quantizing locally.

Step 4: Local-first with cloud fallback via LiteLLM

This is the step that turns “unlimited local tokens” into an actually usable pipeline. LiteLLM Proxy puts local and cloud endpoints behind one OpenAI-compatible interface with ordered fallback.

```yaml

# litellm_config.yaml

model_list:

  • model_name: assistant

litellm_params:

model: openai/qwen3-8b

api_base: http://localhost:8000/v1

api_key: "not-needed"

  • model_name: assistant

litellm_params:

model: openai/gpt-4o-mini

api_key: os.environ/OPENAI_API_KEY

router_settings:

routing_strategy: simple-shuffle

num_retries: 2

fallbacks: [{"assistant": ["assistant"]}]

```

```bash

litellm --config litellm_config.yaml --port 4000

```

Callers only know http://localhost:4000 and model name assistant. If local is up, traffic goes local; if it dies or times out, it falls through to the cloud. Daily traffic hits local, and the cloud is triggered only on failure — keeping quota burn very low.

Step 5: Verify and load-test

Don’t guess whether it’s good enough. Run a concurrency probe:

```bash

for i in $(seq 1 10); do

curl -s http://localhost:8000/v1/chat/completions \

-H "Content-Type: application/json" \

-d '{"model":"qwen3-8b","messages":[{"role":"user","content":"Write a 100-word test paragraph"}],"max_tokens":128}' &

done

wait

```

Watch three numbers: time to first token (TTFT), output speed (tok/s), and the concurrency level at which requests start queuing. Those decide whether this is a production endpoint or a toy.

Caveats

  1. Quantization degrades quality, and not linearly. Q4_K_M is close to FP16 on many tasks, but long-chain reasoning, code generation and math suffer more. Don’t use the lowest tier for critical work.
  2. Context length is the VRAM killer. 32K context under high concurrency can make KV cache larger than the weights. Start with a small max-model-len, then scale up.
  3. Ollama’s concurrency is limited. Cranking OLLAMA_NUM_PARALLEL too high causes VRAM thrashing and lower throughput. Start at 2–4 per GPU.
  4. Never expose a bare endpoint publicly. Ollama’s /v1 doesn’t check keys, and vLLM has no auth by default. Put an authenticated gateway in front (LiteLLM virtual keys, Nginx + Basic Auth, or a private network).
  5. Read the model license. Open weights ≠ free commercial use. Llama, Qwen and Mistral all have different terms — check before shipping.
  6. Electricity and cooling are real costs. A 300W card running 24/7 is not cheap over a year. Do the math.
  7. Local ≠ frontier. Unless your task is narrow, don’t expect an 8B model to replace a frontier closed model. Its job is offloading high-volume, low-difficulty, tolerance-friendly traffic.

When This Fits

Good fit for a local supply layer:

  • High-volume, fixed-format tasks: classification, extraction, translation, summarization, log analysis.
  • Data that can’t leave the intranet: legal, medical, financial text.
  • Batch offline processing: cleaning or labeling tens of thousands of records is far cheaper locally than per-token.
  • Dev and test environments where you don’t want to burn quota debugging.

Poor fit:

  • Complex tasks needing frontier reasoning.
  • Very high concurrency without owned hardware (renting GPUs or buying API access wins).
  • Latency-critical real-time interaction — consumer-card TTFT usually loses to the cloud.

One-Line Summary

The value of local deployment isn’t “free” — it’s pulling predictable, high-volume traffic out of metered billing. Ollama to serve, vLLM to handle concurrency, LiteLLM to route and fall back: what you get is a token pipeline with marginal cost near your power bill, not a toy.