Where this approach fits
The paths covered earlier (official free tiers, gateway credits, student packs, cloud trial credits) are all "someone else meters your usage." Quantized local deployment is a different line: once the weights are on your disk, inference isn't billed per token — only electricity and hardware depreciation.
The preconditions are concrete:
- You have a discrete GPU with ≥ 8GB VRAM (NVIDIA is easiest; AMD via ROCm; Apple Silicon via Metal).
- Your tasks can run offline: code completion, summarization, batch translation, data cleaning, local RAG.
- You accept that the capability ceiling is set by your VRAM, not by a vendor's newest flagship.
If you need frontier-level reasoning, long context, or native multimodality, a local quantized model won't replace it — stay on the API route.
Estimate VRAM before you download
A rough but workable formula (4-bit, weights only):
```
weights (GB) ≈ params (B) × 0.6
```
So 7B ≈ 4.2GB, 14B ≈ 8.4GB, 32B ≈ 19GB. Then add:
- KV cache: scales with context length; typically 1–4GB at 8K, ballooning at long context.
- Runtime overhead: roughly 0.5–1.5GB.
Rules of thumb (4-bit, 8K context):
| VRAM | Comfortable range | Notes |
|---|---|---|
| 8GB | 7B–8B | Keep context short |
| 12GB | 8B–14B | Daily driver |
| 16GB | 14B | Balance of quality and speed |
| 24GB | 32B | Requires 4-bit and restrained context |
Note: for MoE architectures (large total params, small active params), VRAM usage follows total params while speed follows active params. Don't mix the two.
Steps
1. Pick a quantization format
Three practical tiers in 2026:
- GGUF (Q4_K_M / Q5_K_M): friendly to CPU+GPU hybrid inference, natively supported by llama.cpp and Ollama, broadest ecosystem. Quality loss at Q4_K_M is usually acceptable; Q5_K_M is safer but slower.
- AWQ / GPTQ (4-bit): GPU-oriented weight quantization, well supported by vLLM, noticeably higher throughput — good for batch offline jobs.
- FP8 / INT8: trade VRAM for quality. Comfortable for 14B on a 24GB card, but 32B generally won't fit.
Rule: GGUF for interactive single-machine use, AWQ + vLLM for batch throughput, higher bit-width whenever VRAM allows.
2. Ollama: the easiest start
```bash
# After installing, pull a quantized build directly
ollama pull qwen2.5:14b-instruct-q4_K_M
ollama run qwen2.5:14b-instruct-q4_K_M
```
Tune key parameters via Modelfile or environment variables:
```dockerfile
FROM qwen2.5:14b-instruct-q4_K_M
PARAMETER num_ctx 8192
PARAMETER num_gpu 99 # offload as many layers as possible
PARAMETER num_thread 8 # CPU threads
```
When VRAM is short, Ollama keeps some layers on CPU and speed drops to a few tokens/sec. Lower num_ctx or switch to a smaller model rather than pushing through.
3. llama.cpp: when you want fine control
```bash
# Build with CUDA enabled
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
# Start the server, exposing an OpenAI-compatible endpoint
./build/bin/llama-server \
-m ./models/qwen2.5-14b-instruct-q4_k_m.gguf \
-ngl 99 \
-c 8192 \
--host 127.0.0.1 --port 8080
```
-ngl (layers offloaded) is the core knob: walk it down from 99 until you stop hitting OOM, and keep the highest value the card tolerates.
4. vLLM: batch jobs and concurrency
```bash
vllm serve Qwen/Qwen2.5-14B-Instruct-AWQ \
--quantization awq \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--port 8000
```
--gpu-memory-utilization defaults to 0.9; drop it toward 0.8 when VRAM is tight. vLLM pre-allocates KV cache, so it looks "full" at startup — that's expected, not a leak.
5. Wire it into your toolchain
All three runtimes expose OpenAI-compatible endpoints (Ollama at http://localhost:11434/v1, llama.cpp at :8080/v1, vLLM at :8000/v1). Point your editor plugin or script's base_url at them; no call-site changes needed.
Caveats
- Quantization is lossy: Q4 degrades noticeably on math reasoning, long chains of logic, and code generation. Use Q5/Q6 or a smaller model if that's your workload.
- Context eats VRAM: going from 8K to 32K context can multiply KV cache several times over and cause OOM. Decide context first, then model size.
- Be realistic about speed: a 14B 4-bit model on a consumer card typically lands in the low tens of tokens/sec, depending on the card and context length. This is not an A100 experience.
- Power and thermals are real costs: sustained full-load inference draws meaningful wattage; laptops will throttle.
- Read the license: open weights do not automatically mean unrestricted commercial use; some models ship custom terms. Verify before commercial deployment.
- Trust the source: download weights from official orgs or reputable publishers; avoid unknown GGUF files.
Good fits
- Privacy-sensitive work: data never leaves the machine — medical, legal, internal codebases.
- High-frequency small calls: thousands of short requests per day cost less locally than per-token billing.
- Offline environments: inference on machines with no or restricted network.
- Batch processing: data cleaning, translation, labeling overnight.
Poor fits
- Frontier reasoning, very long context, or native multimodality.
- No discrete GPU — integrated graphics or CPU-only will run, but anything above 7B is effectively unusable.
- Production services that need elastic scaling; a single local box has no redundancy.