Why go the self-hosted gateway route

The problem with public relay stations is that quota rules, model allowlists, and ban risk all sit with someone else. A self-hosted gateway puts control back in your hands: the upstreams are still official free tiers or trial credits, but you expose a single OpenAI-compatible endpoint and can swap upstreams, add keys, or change models at will.

By 2026 this pattern is mature: LiteLLM Proxy, One API / New API, and Portkey Gateway all support multiple upstreams, weighted routing, and retries.

Steps

1. Pick a gateway

  • LiteLLM Proxy: Python ecosystem, a config.yaml describes model_list, native support for 100+ upstreams (OpenAI / Anthropic / Gemini / Groq / Mistral), built-in /key/info usage tracking and virtual keys.
  • One API / New API: single Go binary, one-command Docker start, friendly Chinese UI, good for handing sub-keys to teammates with per-key budgets.
  • Portkey Gateway: leans toward observability, useful once you have traffic and want latency/cost breakdowns.

For individuals and small teams, start with LiteLLM or New API — the former is more config-driven, the latter has a more visual dashboard.

2. Collect upstream keys first, then write config

The order matters: gather every quota you can legitimately get, then wire them together. Typical sources:

  • Vendor official free tiers (usually with RPM/TPM caps, refreshed daily or monthly);
  • Model services opened with cloud trial credits (e.g. Bedrock, Vertex AI side);
  • Free inference quota on open-model hosting platforms (e.g. Groq, Cerebras — usually rate-limited);
  • Your own local Ollama / vLLM as a backstop.

Record each key as one row: upstream, model, rate limit, quota period, and whether production use is permitted.

3. Configure multiple upstreams and failover

In LiteLLM, define several deployments under the same logical model name, describing each one's capacity with rpm, tpm, and weight; the gateway handles balancing and retries:

```yaml

model_list:

  • model_name: gpt-4o-mini

litellm_params:

model: openai/gpt-4o-mini

api_key: os.environ/OPENAI_KEY_1

rpm: 3

  • model_name: gpt-4o-mini

litellm_params:

model: groq/llama-3.3-70b

api_key: os.environ/GROQ_KEY_1

rpm: 30

router_settings:

routing_strategy: usage-based-routing-v2

num_retries: 2

fallbacks: [{"gpt-4o-mini": ["local-llama"]}]

```

Setting rpm to real values matters: too high and you will keep hitting upstream 429s; slightly conservative values give more stable overall throughput.

4. Isolate callers with virtual keys

Issue a virtual key per app or teammate with max_budget and rpm_limit. If one key is abused, only its own budget is affected, not the whole free pool.

5. Add caching and degradation

  • Use LiteLLM's Redis cache or your own semantic cache for repeated prompts — it saves a lot of quota;
  • Configure local vLLM as a fallback so you still get answers (lower quality) when every upstream returns 429;
  • Enable /metrics and watch per-upstream failure rates in Prometheus; drop unstable keys promptly.

Caveats

  • Compliance first: free tiers typically state "development and testing only". Wiring free quota into a paid product may violate terms; use paid keys or self-hosted models for production traffic.
  • Do not farm quotas with multiple accounts: scripted signups on the same platform are commonly treated as abuse and can get your IP or payment method banned.
  • Rate limits are hard constraints: free-tier RPM/TPM is low; a gateway smooths traffic, it does not create quota.
  • Key security: keep all keys in environment variables or a secret manager — never commit them into config.yaml.
  • Quotas change: model allowlists and rate limits shift often; re-verify upstream availability monthly.
  • Regional limits: some upstreams are unavailable in certain regions; confirm before deploying.

When this fits

  • Solo developers doing prototypes, evals, or small tools — low volume but wanting multi-model comparison;
  • Small-team internal tools (support drafts, doc summarization, code completion) that need scattered quotas managed centrally;
  • A/B testing across models without writing an adapter per upstream;
  • You already own local GPUs and want self-hosted models mixed into the same endpoint as a backstop.

A minimal validation flow

  1. pip install 'litellm[proxy]' and write the config.yaml above;
  2. litellm --config config.yaml --port 4000;
  3. Hit /v1/chat/completions with curl and confirm normal responses;
  4. Deliberately break one upstream key and verify fallback kicks in;
  5. Open /ui and check that usage is tracked per virtual key.

Once these five steps pass, you have a token pipeline that depends on no single relay station.