Where this approach fits

The paths covered earlier (official free tiers, gateway credits, student packs, cloud trial credits) are all "someone else meters your usage." Quantized local deployment is a different line: once the weights are on your disk, inference isn't billed per token — only electricity and hardware depreciation.

The preconditions are concrete:

  • You have a discrete GPU with ≥ 8GB VRAM (NVIDIA is easiest; AMD via ROCm; Apple Silicon via Metal).
  • Your tasks can run offline: code completion, summarization, batch translation, data cleaning, local RAG.
  • You accept that the capability ceiling is set by your VRAM, not by a vendor's newest flagship.

If you need frontier-level reasoning, long context, or native multimodality, a local quantized model won't replace it — stay on the API route.

Estimate VRAM before you download

A rough but workable formula (4-bit, weights only):

```

weights (GB) ≈ params (B) × 0.6

```

So 7B ≈ 4.2GB, 14B ≈ 8.4GB, 32B ≈ 19GB. Then add:

  • KV cache: scales with context length; typically 1–4GB at 8K, ballooning at long context.
  • Runtime overhead: roughly 0.5–1.5GB.

Rules of thumb (4-bit, 8K context):

| VRAM | Comfortable range | Notes |

|---|---|---|

| 8GB | 7B–8B | Keep context short |

| 12GB | 8B–14B | Daily driver |

| 16GB | 14B | Balance of quality and speed |

| 24GB | 32B | Requires 4-bit and restrained context |

Note: for MoE architectures (large total params, small active params), VRAM usage follows total params while speed follows active params. Don't mix the two.

Steps

1. Pick a quantization format

Three practical tiers in 2026:

  1. GGUF (Q4_K_M / Q5_K_M): friendly to CPU+GPU hybrid inference, natively supported by llama.cpp and Ollama, broadest ecosystem. Quality loss at Q4_K_M is usually acceptable; Q5_K_M is safer but slower.
  2. AWQ / GPTQ (4-bit): GPU-oriented weight quantization, well supported by vLLM, noticeably higher throughput — good for batch offline jobs.
  3. FP8 / INT8: trade VRAM for quality. Comfortable for 14B on a 24GB card, but 32B generally won't fit.

Rule: GGUF for interactive single-machine use, AWQ + vLLM for batch throughput, higher bit-width whenever VRAM allows.

2. Ollama: the easiest start

```bash

# After installing, pull a quantized build directly

ollama pull qwen2.5:14b-instruct-q4_K_M

ollama run qwen2.5:14b-instruct-q4_K_M

```

Tune key parameters via Modelfile or environment variables:

```dockerfile

FROM qwen2.5:14b-instruct-q4_K_M

PARAMETER num_ctx 8192

PARAMETER num_gpu 99 # offload as many layers as possible

PARAMETER num_thread 8 # CPU threads

```

When VRAM is short, Ollama keeps some layers on CPU and speed drops to a few tokens/sec. Lower num_ctx or switch to a smaller model rather than pushing through.

3. llama.cpp: when you want fine control

```bash

# Build with CUDA enabled

cmake -B build -DGGML_CUDA=ON

cmake --build build --config Release -j

# Start the server, exposing an OpenAI-compatible endpoint

./build/bin/llama-server \

-m ./models/qwen2.5-14b-instruct-q4_k_m.gguf \

-ngl 99 \

-c 8192 \

--host 127.0.0.1 --port 8080

```

-ngl (layers offloaded) is the core knob: walk it down from 99 until you stop hitting OOM, and keep the highest value the card tolerates.

4. vLLM: batch jobs and concurrency

```bash

vllm serve Qwen/Qwen2.5-14B-Instruct-AWQ \

--quantization awq \

--max-model-len 8192 \

--gpu-memory-utilization 0.90 \

--port 8000

```

--gpu-memory-utilization defaults to 0.9; drop it toward 0.8 when VRAM is tight. vLLM pre-allocates KV cache, so it looks "full" at startup — that's expected, not a leak.

5. Wire it into your toolchain

All three runtimes expose OpenAI-compatible endpoints (Ollama at http://localhost:11434/v1, llama.cpp at :8080/v1, vLLM at :8000/v1). Point your editor plugin or script's base_url at them; no call-site changes needed.

Caveats

  • Quantization is lossy: Q4 degrades noticeably on math reasoning, long chains of logic, and code generation. Use Q5/Q6 or a smaller model if that's your workload.
  • Context eats VRAM: going from 8K to 32K context can multiply KV cache several times over and cause OOM. Decide context first, then model size.
  • Be realistic about speed: a 14B 4-bit model on a consumer card typically lands in the low tens of tokens/sec, depending on the card and context length. This is not an A100 experience.
  • Power and thermals are real costs: sustained full-load inference draws meaningful wattage; laptops will throttle.
  • Read the license: open weights do not automatically mean unrestricted commercial use; some models ship custom terms. Verify before commercial deployment.
  • Trust the source: download weights from official orgs or reputable publishers; avoid unknown GGUF files.

Good fits

  • Privacy-sensitive work: data never leaves the machine — medical, legal, internal codebases.
  • High-frequency small calls: thousands of short requests per day cost less locally than per-token billing.
  • Offline environments: inference on machines with no or restricted network.
  • Batch processing: data cleaning, translation, labeling overnight.

Poor fits

  • Frontier reasoning, very long context, or native multimodality.
  • No discrete GPU — integrated graphics or CPU-only will run, but anything above 7B is effectively unusable.
  • Production services that need elastic scaling; a single local box has no redundancy.