docs(guides): reorganize documentation into dedicated usage manuals

Expand the documentation set with standalone guides for free-tier
providers, proxy configuration, and PWA installation while refreshing
the README badge layout and navigation structure.

Remove the docs ignore allowlist so the full documentation tree can be
tracked directly, and drop the outdated context-relay feature page.
This commit is contained in:
Antigravity Assistant
2026-05-01 13:18:27 -03:00
parent fa950a8a49
commit abe8cae083
6 changed files with 1277 additions and 187 deletions

40
.gitignore vendored
View File

@@ -68,46 +68,6 @@ antigravity-manager-analysis/
.sisyphus/
.plans/
# docs (allow specific tracked files)
docs/*
!docs/ARCHITECTURE.md
!docs/CODEBASE_DOCUMENTATION.md
!docs/CONTRIBUTING.md
!docs/USER_GUIDE.md
!docs/API_REFERENCE.md
!docs/TERMUX_GUIDE.md
!docs/TROUBLESHOOTING.md
!docs/EXECUTION_CONTEXT_PROVIDER_SYNC.md
!docs/TASK_NEBIUS_BACKEND_ENABLEMENT.md
!docs/frontend-backend-provider-gap-report.md
!docs/openapi.yaml
!docs/RELEASE_CHECKLIST.md
!docs/PLANO-IMPLANTACAO.md
!docs/TASKS.md
!docs/FASE-*.md
!docs/adr/
!docs/cli-tools/
!docs/planning/
!docs/improvement-plans/
!docs/api/
!docs/VM_DEPLOYMENT_GUIDE.md
!docs/FEATURES.md
!docs/screenshots/
!docs/i18n/
!docs/i18n/**
!docs/features/
!docs/features/**
!docs/A2A-SERVER.md
!docs/AUTO-COMBO.md
!docs/MCP-SERVER.md
!docs/CLI-TOOLS.md
!docs/COVERAGE_PLAN.md
!docs/ENVIRONMENT.md
!docs/UNINSTALL.md
!docs/I18N.md
!docs/FLY_IO_DEPLOYMENT_GUIDE.md
# open-sse tests
open-sse/test/*

105
README.md
View File

@@ -1,5 +1,3 @@
[![MseeP.ai Security Assessment Badge](https://mseep.net/pr/diegosouzapw-omniroute-badge.png)](https://mseep.ai/app/diegosouzapw-omniroute)
# 🚀 OmniRoute — The Free AI Gateway
### Never stop coding. Smart routing to **FREE & low-cost AI models** with automatic fallback.
@@ -12,30 +10,39 @@ _Your universal API proxy — one endpoint, 160+ providers, zero downtime. Now w
<div align="center">
<!-- Package & Distribution -->
[![npm version](https://img.shields.io/npm/v/omniroute?color=cb3837&logo=npm)](https://www.npmjs.com/package/omniroute)
[![Docker Hub](https://img.shields.io/docker/v/diegosouzapw/omniroute?label=Docker%20Hub&logo=docker&color=2496ED)](https://hub.docker.com/r/diegosouzapw/omniroute)
[![tag](https://custom-icon-badges.demolab.com/github/v/tag/diegosouzapw/OmniRoute?logo=tag&logoColor=white)](https://github.com/diegosouzapw/OmniRoute/tags)
[![license](https://custom-icon-badges.demolab.com/github/license/diegosouzapw/OmniRoute?logo=law)](https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE)
![NPM Downloads](https://img.shields.io/npm/dw/omniroute?label=npm%20down%20week&color=red)
![NPM Downloads](https://img.shields.io/npm/dm/omniroute?label=npm%20down%20month&color=red)
<!-- Downloads -->
![NPM Downloads](https://img.shields.io/npm/d18m/omniroute?label=npm%20down%20year&color=red)
![Docker Pulls](https://img.shields.io/docker/pulls/diegosouzapw/omniroute)
![GitHub Downloads (all assets, all releases)](https://img.shields.io/github/downloads/diegosouzapw/omniroute/total?style=flat&label=eletron%20donwloads&color=blue)
![NPM Weekly](https://img.shields.io/npm/dw/omniroute?label=npm/week&color=cb3837&logo=npm)
![NPM Monthly](https://img.shields.io/npm/dm/omniroute?label=npm/month&color=cb3837&logo=npm)
![NPM Yearly](https://img.shields.io/npm/d18m/omniroute?label=npm/year&color=cb3837&logo=npm)
![Docker Pulls](https://img.shields.io/docker/pulls/diegosouzapw/omniroute?label=docker%20pulls&logo=docker&color=2496ED)
![Electron Downloads](https://img.shields.io/github/downloads/diegosouzapw/omniroute/total?style=flat&label=electron%20downloads&logo=electron&color=47848F)
<!-- Repository Health -->
[![stars](https://custom-icon-badges.demolab.com/github/stars/diegosouzapw/OmniRoute?logo=star&style=flat)](https://github.com/diegosouzapw/OmniRoute/stargazers)
[![open issues](https://custom-icon-badges.demolab.com/github/issues-raw/diegosouzapw/OmniRoute?logo=issue)](https://github.com/diegosouzapw/OmniRoute/issues)
[![license](https://custom-icon-badges.demolab.com/github/license/diegosouzapw/OmniRoute?logo=law)](https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE)
[![last commit](https://custom-icon-badges.demolab.com/github/last-commit/diegosouzapw/OmniRoute?logo=history&logoColor=white)](https://github.com/diegosouzapw/OmniRoute/commits/main)
[![total contributions](https://custom-icon-badges.demolab.com/badge/dynamic/json?logo=graph&logoColor=fff&color=blue&label=total%20contributions&query=%24.totalContributions&url=https%3A%2F%2Fstreak-stats.demolab.com%2F%3Fuser%3Ddiegosouzapw%26type%3Djson)](https://github.com/diegosouzapw)
[![code size](https://custom-icon-badges.demolab.com/github/languages/code-size/diegosouzapw/OmniRoute?logo=file-code&logoColor=white)](https://github.com/diegosouzapw/OmniRoute)
[![pr closed](https://custom-icon-badges.demolab.com/github/issues-pr-closed/diegosouzapw/OmniRoute?color=purple&logo=git-pull-request&logoColor=white)](https://github.com/diegosouzapw/OmniRoute/pulls?q=is%3Apr+is%3Aclosed)
[![tag](https://custom-icon-badges.demolab.com/github/v/tag/diegosouzapw/OmniRoute?logo=tag&logoColor=white)](https://github.com/diegosouzapw/OmniRoute/tags)
[![github streak](https://custom-icon-badges.demolab.com/badge/dynamic/json?logo=fire&logoColor=fff&color=orange&label=github%20streak&query=%24.currentStreak.length&suffix=%20days&url=https%3A%2F%2Fstreak-stats.demolab.com%2F%3Fuser%3Ddiegosouzapw%26type%3Djson)](https://github.com/diegosouzapw)
[![followers](https://custom-icon-badges.demolab.com/github/followers/diegosouzapw?logo=person-add)](https://github.com/diegosouzapw?tab=followers)
[![fork](https://custom-icon-badges.demolab.com/github/forks/diegosouzapw/OmniRoute?logo=fork)](https://github.com/diegosouzapw/OmniRoute/network/members)
[![watch](https://custom-icon-badges.demolab.com/github/watchers/diegosouzapw/OmniRoute?logo=eye)](https://github.com/diegosouzapw/OmniRoute/watchers)
[![open issues](https://custom-icon-badges.demolab.com/github/issues-raw/diegosouzapw/OmniRoute?logo=issue)](https://github.com/diegosouzapw/OmniRoute/issues)
[![pr closed](https://custom-icon-badges.demolab.com/github/issues-pr-closed/diegosouzapw/OmniRoute?color=purple&logo=git-pull-request&logoColor=white)](https://github.com/diegosouzapw/OmniRoute/pulls?q=is%3Apr+is%3Aclosed)
[![last commit](https://custom-icon-badges.demolab.com/github/last-commit/diegosouzapw/OmniRoute?logo=history&logoColor=white)](https://github.com/diegosouzapw/OmniRoute/commits/main)
[![code size](https://custom-icon-badges.demolab.com/github/languages/code-size/diegosouzapw/OmniRoute?logo=file-code&logoColor=white)](https://github.com/diegosouzapw/OmniRoute)
<!-- Community & Social -->
[![total contributions](https://custom-icon-badges.demolab.com/badge/dynamic/json?logo=graph&logoColor=fff&color=blue&label=total%20contributions&query=%24.totalContributions&url=https%3A%2F%2Fstreak-stats.demolab.com%2F%3Fuser%3Ddiegosouzapw%26type%3Djson)](https://github.com/diegosouzapw)
[![github streak](https://custom-icon-badges.demolab.com/badge/dynamic/json?logo=fire&logoColor=fff&color=orange&label=github%20streak&query=%24.currentStreak.length&suffix=%20days&url=https%3A%2F%2Fstreak-stats.demolab.com%2F%3Fuser%3Ddiegosouzapw%26type%3Djson)](https://github.com/diegosouzapw)
[![followers](https://custom-icon-badges.demolab.com/github/followers/diegosouzapw?logo=person-add)](https://github.com/diegosouzapw?tab=followers)
<!-- Links -->
[![License](https://img.shields.io/github/license/diegosouzapw/OmniRoute)](https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE)
[![Website](https://img.shields.io/badge/Website-omniroute.online-blue?logo=google-chrome&logoColor=white)](https://omniroute.online)
[![WhatsApp](https://img.shields.io/badge/WhatsApp-Community-25D366?logo=whatsapp&logoColor=white)](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
@@ -161,6 +168,42 @@ _Connect any AI-powered IDE or CLI tool through OmniRoute — free API gateway f
---
## 📺 OmniRoute in Action — Video Guides
<div align="center">
<table>
<tr>
<td align="center" width="320">
<a href="https://www.youtube.com/watch?v=Rxdc36yUyOQ">
<img src="https://img.youtube.com/vi/Rxdc36yUyOQ/maxresdefault.jpg" alt="OmniRoute — Guia em Português" width="300"/>
</a><br/>
<b>🇧🇷 Português</b><br/>
<sub>Guia completo do OmniRoute</sub>
</td>
<td align="center" width="320">
<a href="https://www.youtube.com/watch?v=CMzyOiUyEVc">
<img src="https://img.youtube.com/vi/CMzyOiUyEVc/maxresdefault.jpg" alt="OmniRoute — English Guide" width="300"/>
</a><br/>
<b>🇺🇸 English</b><br/>
<sub>Complete OmniRoute walkthrough</sub>
</td>
<td align="center" width="320">
<a href="https://www.youtube.com/watch?v=il_5Ii6v4-Y">
<img src="https://img.youtube.com/vi/il_5Ii6v4-Y/maxresdefault.jpg" alt="OmniRoute — Руководство на русском" width="300"/>
</a><br/>
<b>🇷🇺 Русский</b><br/>
<sub>Полное руководство по OmniRoute</sub>
</td>
</tr>
</table>
</div>
> 🎬 **Made a video about OmniRoute?** We'd love to feature it here! Open an [issue](https://github.com/diegosouzapw/OmniRoute/issues/new) or [discussion](https://github.com/diegosouzapw/OmniRoute/discussions) with the link and we'll add it to this showcase.
---
## 🤔 Why OmniRoute?
**Stop wasting money and hitting limits:**
@@ -1617,6 +1660,33 @@ Available free: `qwen3-235b-a22b-instruct-2507` (Qwen3 235B!), `llama-3.1-70b-in
> Cerebras (cerebras/) → Llama/Qwen world-fastest — 1M tok/day
> ```
---
## 🌐 Free API Provider Directory — 25+ Providers, 500+ Models, $0
> **We analyzed 6 community repositories** aggregating free LLM API providers and consolidated everything into one definitive reference. This is the most comprehensive free-tier directory available.
| Provider | Best Free Model | RPM | RPD | Tokens | Speed |
| ----------------- | ---------------- | ----- | ------------ | ------------- | --------- |
| **Groq** | Llama 3.3 70B | 30 | 14,400 | 6K TPM | 🟢 Fast |
| **Cerebras** | Qwen3 235B | 30 | 14,400 | 1M TPD | 🟢 Fast |
| **Mistral AI** | Mistral Large 3 | 60 | Unlimited | 1B/month | 🟡 Medium |
| **Google Gemini** | Gemini 2.5 Flash | 515 | 201,500 | 250K TPM | 🟢 Fast |
| **NVIDIA NIM** | 129 models | 40 | — | — | 🟡 Medium |
| **OpenRouter** | 35+ :free models | 20 | 501,000 | — | 🟡 Medium |
| **GitHub Models** | GPT-4.1, GPT-5 | 1015 | 50150 | 8K/4K per req | 🟡 Medium |
| **Cloudflare AI** | 50+ models | — | 10K neurons | — | 🟡 Medium |
| **Pollinations** | Text+Image+Video | — | Hourly reset | — | 🟡 Medium |
| **SiliconFlow** | Qwen3-8B | 1,000 | — | 50K TPM | 🟡 Medium |
**Combined free capacity across all providers: ~31,000+ RPD · ~32B+ tokens/month · 500+ models · $0 forever.**
The full directory includes 25+ providers with detailed rate limits, base URLs, model tables, trial credit providers (Baseten $30, AI21 $10, SambaNova $5, etc.), China-specific platforms (ModelScope, Volcengine, Tencent Hunyuan), and step-by-step OmniRoute combo configuration.
📖 **Complete free provider directory with all models, quotas, and integration guide:** [`docs/FREE_TIERS.md`](docs/FREE_TIERS.md)
---
## 🎙️ Free Transcription Combo
> Transcribe any audio/video for **$0** — Deepgram leads with $200 free, AssemblyAI $50 fallback, Groq Whisper as unlimited emergency backup.
@@ -2602,6 +2672,7 @@ Se não quiser criar credenciais próprias agora, ainda é possível usar o flux
| [Release Checklist](docs/RELEASE_CHECKLIST.md) | Pre-release validation steps |
| [PWA Guide](docs/PWA_GUIDE.md) | Progressive Web App install, caching, architecture |
| [Proxy Guide](docs/PROXY_GUIDE.md) | Proxy system, 1proxy marketplace, registry CRUD |
| [Free Tiers](docs/FREE_TIERS.md) | 25+ free API providers consolidated directory |
| [Termux Guide](docs/TERMUX_GUIDE.md) | Run OmniRoute on Android via Termux |
---

407
docs/FREE_TIERS.md Normal file
View File

@@ -0,0 +1,407 @@
# 🆓 Free LLM API Providers — Consolidated Directory
> **The ultimate aggregated reference for all permanently free LLM API providers.**
> Consolidated from 6 community repositories. Use with OmniRoute to route through 25+ free providers simultaneously.
_Last consolidated: May 2026 · Sources: awesome-free-llm-apis, awesome-free-llm-apis2, free-llm-api-resources, Free-LLM-Collection, FREE-LLM-API-Provider, gpt4free_
---
## Table of Contents
- [Quick Comparison](#quick-comparison)
- [Provider APIs (First-Party)](#provider-apis-first-party)
- [Inference Providers (Third-Party)](#inference-providers-third-party)
- [China-Based Providers](#china-based-providers)
- [Trial Credit Providers](#trial-credit-providers)
- [Using with OmniRoute](#using-with-omniroute)
- [Glossary](#glossary)
---
## Quick Comparison
All free providers at a glance, sorted by generosity of free tier:
| Provider | Type | Best Free Model | RPM | RPD | Tokens | OpenAI Compat | Speed |
| ----------------- | --------- | ---------------------- | ------- | ------------ | ---------------- | --------------- | --------- |
| **Groq** | Inference | Llama 3.3 70B | 30 | 14,400 | 6K TPM | ✅ | 🟢 Fast |
| **Cerebras** | Inference | Qwen3 235B | 30 | 14,400 | 1M TPD | ✅ | 🟢 Fast |
| **Mistral AI** | Provider | Mistral Large 3 | 60 | Unlimited | 1B/month | ✅ | 🟡 Medium |
| **Google Gemini** | Provider | Gemini 2.5 Flash | 515 | 201,500 | 250K TPM | ✅ | 🟢 Fast |
| **NVIDIA NIM** | Inference | 129 models | 40 | — | — | ✅ | 🟡 Medium |
| **Ollama Cloud** | Inference | 400+ models | — | — | Session limits | ❌ (Ollama API) | 🟡 Medium |
| **OpenRouter** | Inference | 35+ free models | 20 | 501,000 | — | ✅ | 🟡 Medium |
| **GitHub Models** | Inference | GPT-4.1, GPT-5 | 1015 | 50150 | 8K in/4K out | ✅ | 🟡 Medium |
| **Cloudflare AI** | Inference | 50+ models | — | 10K neurons | — | ⚠️ Partial | 🟡 Medium |
| **Hugging Face** | Inference | Thousands | — | — | $0.10/mo credits | ✅ | 🔴 Slow |
| **Cohere** | Provider | Command A (111B) | 20 | — | 1K calls/month | ⚠️ Partial | 🟡 Medium |
| **Pollinations** | Inference | Text+Image+Video+Audio | — | Hourly reset | — | ✅ | 🟡 Medium |
| **Z.AI (Zhipu)** | Provider | GLM-4.7-Flash | — | — | Undocumented | ✅ | 🟡 Medium |
| **SiliconFlow** | Inference | Qwen3-8B | 1,000 | — | 50K TPM | ✅ | 🟡 Medium |
| **Kilo Code** | Inference | Free auto-router | ~200/hr | — | — | ✅ | 🟡 Medium |
| **LLM7.io** | Inference | 30+ models | 1530 | — | — | ✅ | 🟡 Medium |
| **Kluster AI** | Inference | DeepSeek-R1 | — | — | Undocumented | ✅ | 🟡 Medium |
| **ModelScope** | Inference | Qwen, DeepSeek | — | 2,000 | ≤500/model/day | ✅ | 🟡 Medium |
| **IBM watsonx** | Provider | Granite models | 2/sec | — | 300K/month | ❌ | 🟡 Medium |
---
## Provider APIs (First-Party)
APIs from the companies that train or fine-tune the models.
### Google Gemini 🇺🇸
🔗 [Get API Key](https://aistudio.google.com/app/apikey) · Base URL: `https://generativelanguage.googleapis.com/v1beta`
> ⚠️ Free tier NOT available in EU/UK/Switzerland. Prompts may be used by Google to improve products.
| Model | Context | Max Output | Modality | RPM | RPD |
| --------------------------------- | ------- | ---------- | ---------------------- | --- | ------ |
| Gemini 2.5 Flash / Gemini 3 Flash | 1M | 65K | Text+Image+Audio+Video | 5 | 20 |
| Gemini 2.5 Flash-Lite | 1M | 65K | Text+Image+Audio+Video | 10 | 20 |
| Gemini 3.1 Flash-Lite | 1M | 65K | Text+Image+Audio+Video | 15 | 1,500 |
| Gemma 4 26B/31B | — | — | Text | 15 | 1,500 |
| Gemma 3 (1B/4B/12B/27B) | — | — | Text | 30 | 14,400 |
### Mistral AI 🇫🇷
🔗 [Get API Key](https://console.mistral.ai/api-keys) · Base URL: `https://api.mistral.ai/v1`
Free "Experiment" plan, no credit card. ~1B tokens/month. Requires phone verification.
| Model | Context | Max Output | Modality | Rate Limit |
| ------------------ | ------- | ---------- | --------------- | --------------- |
| Mistral Small 4 | 256K | 256K | Text+Image+Code | 1 RPS, 500K TPM |
| Mistral Medium 3 | 128K | 128K | Text | 1 RPS, 500K TPM |
| Mistral Large 3 | 256K | 256K | Text | 1 RPS, 500K TPM |
| Mistral Nemo (12B) | 128K | 128K | Text | 1 RPS, 500K TPM |
| Codestral | 256K | 256K | Code | 30 RPM, 2K RPD |
| Pixtral Large | 128K | 128K | Text+Image | 1 RPS, 500K TPM |
### Cohere 🇨🇦
🔗 [Get API Key](https://dashboard.cohere.com/api-keys) · Base URL: `https://api.cohere.com/v2`
Free "Trial" key. 1,000 API calls/month. Non-commercial use only. 20 RPM.
| Model | Context | Max Output | Modality |
| ------------------- | ------- | ---------- | ----------------------- |
| Command A (111B) | 256K | 4K | Text |
| Command A Reasoning | 256K | 4K | Text (reasoning) |
| Command A Vision | 256K | 4K | Text+Image |
| Command A Translate | 256K | 4K | Translation |
| Command R+ | 128K | 4K | Text |
| Command R | 128K | 4K | Text |
| Command R7B | 128K | 4K | Text |
| Embed 4 | — | — | Embeddings (Text+Image) |
| Rerank 3.5 | — | — | Reranking |
### Z.AI (Zhipu AI) 🇨🇳
🔗 [Get API Key](https://open.bigmodel.cn/usercenter/apikeys) · Base URL: `https://open.bigmodel.cn/api/paas/v4`
Permanent free models, no credit card. No published rate limits.
| Model | Context | Max Output | Modality |
| --------------- | ------- | ---------- | ---------- |
| GLM-4.7-Flash | 200K | 128K | Text |
| GLM-4.5-Flash | 128K | ~8K | Text |
| GLM-4.6V-Flash | 128K | ~4K | Text+Image |
| GLM-5 / GLM-5.1 | — | — | Text |
### IBM watsonx 🇺🇸
🔗 [Pricing](https://www.ibm.com/products/watsonx-ai/pricing)
Free tier: 2 RPS, 300K tokens/month. Granite foundation models.
---
## Inference Providers (Third-Party)
### Groq 🇺🇸
🔗 [Get API Key](https://console.groq.com/keys) · Base URL: `https://api.groq.com/openai/v1`
Ultra-fast LPU inference (~300500 tok/s). No credit card required.
| Model | RPM | RPD | TPM | Modality |
| ---------------------------------- | --- | ------ | --- | ---------------- |
| llama-3.3-70b-versatile | 30 | 1,000 | 12K | Text |
| llama-3.1-8b-instant | 30 | 14,400 | 6K | Text |
| llama-4-scout-17b-16e-instruct | 30 | 1,000 | 30K | Text+Vision |
| llama-4-maverick-17b-128e-instruct | 30 | 1,000 | 6K | Text+Vision |
| qwen3-32b | 60 | 1,000 | 6K | Text |
| kimi-k2-instruct | 60 | 1,000 | 10K | Text |
| gpt-oss-120b / gpt-oss-20b | 30 | 1,000 | 8K | Text |
| deepseek-r1-distill-70b | 30 | 14,400 | — | Text (reasoning) |
| whisper-large-v3 / v3-turbo | 20 | 2,000 | — | Audio→Text |
### Cerebras 🇺🇸
🔗 [Get API Key](https://cloud.cerebras.ai/) · Base URL: `https://api.cerebras.ai/v1`
Wafer-scale chip inference (~2,600 tok/s). 1M tokens/day cap.
| Model | RPM | RPH | RPD | TPM | TPD |
| ------------------------------ | --- | --- | ------ | --- | --- |
| gpt-oss-120b | 30 | 900 | 14,400 | 64K | 1M |
| llama3.1-8b | 30 | 900 | 14,400 | 60K | 1M |
| qwen-3-235b-a22b-instruct-2507 | 30 | 900 | 14,400 | 60K | 1M |
| zai-glm-4.7 | 10 | 100 | 100 | 60K | 1M |
### NVIDIA NIM 🇺🇸
🔗 [Explore Models](https://build.nvidia.com/explore/discover) · Base URL: `https://integrate.api.nvidia.com/v1`
Free with NVIDIA Developer Program. **129 models**, 40 RPM. Phone verification required.
**Notable models:** DeepSeek-R1, DeepSeek-V3.2, Nemotron Ultra 253B, Llama 3.1 405B, Qwen3 Coder 480B, Mistral Large 3, Kimi K2, GLM-5.1, MiniMax M2.7, Gemma 4 31B, + 100 more.
### OpenRouter 🇺🇸
🔗 [Get API Key](https://openrouter.ai/keys) · Base URL: `https://openrouter.ai/api/v1`
35+ free models (suffix `:free`). 20 RPM.
| Credits Purchased | RPD |
| ----------------- | ----- |
| < $10 | 50 |
| ≥ $10 (one-time) | 1,000 |
**Notable free models:** DeepSeek R1, DeepSeek V3, Qwen3 Coder 480B, Llama 4 Scout/Maverick, GPT-OSS 120B, Nemotron 3 Super 120B, MiniMax M2.5, Gemma 4 31B, Devstral, + 23 more.
### GitHub Models 🇺🇸
🔗 [Marketplace](https://github.com/marketplace/models) · Base URL: `https://models.inference.ai.azure.com`
Free for all GitHub users. 45+ models including frontier models.
| Tier | RPM | RPD | Tokens/Request |
| ----------------------- | --- | --- | -------------- |
| Low tier models | 15 | 150 | 8K in / 4K out |
| High tier models | 10 | 50 | 8K in / 4K out |
| DeepSeek-R1 / MAI-DS-R1 | 1 | 8 | 4K in / 4K out |
| Grok-3 | 1 | 15 | 4K in / 4K out |
**Notable models:** GPT-4.1, GPT-4o, GPT-5, GPT-5-mini, o3-mini, o4-mini, DeepSeek-R1, Llama 4 Scout/Maverick, Codestral, Mistral Medium 3, Phi-4, Grok-3.
### Cloudflare Workers AI 🇺🇸
🔗 [Get Token](https://dash.cloudflare.com/profile/api-tokens) · 10,000 Neurons/day free. 50+ models.
**Notable models:** Llama 3.3 70B, Llama 4 Scout, Qwen3 30B-A3B, QwQ 32B, DeepSeek R1 Distill, Gemma 4 26B, GLM 4.7 Flash, Nemotron 3 120B, Kimi K2.5/K2.6, Mistral Small 3.1, GPT-OSS 120B/20B, + 40 more.
### Hugging Face 🇺🇸
🔗 [Get Token](https://huggingface.co/settings/tokens) · Base URL: `https://api-inference.huggingface.co/v1`
$0.10/month free credits (auto-replenished). Thousands of models. Serverless limited to <10GB models.
### Ollama Cloud 🇺🇸
🔗 [Get Key](https://ollama.com/settings/keys) · Base URL: `https://api.ollama.com`
400+ models. Session/weekly limits (unpublished). NOT OpenAI SDK-compatible.
**Notable models:** GPT-OSS 120B, DeepSeek V3.2/V4, Kimi K2/K2.5/K2.6, GLM-5/5.1, Qwen3 Coder 480B, Gemini 3 Flash, MiniMax M2.7, Cogito 2.1 671B, Nemotron 3 Super 120B.
### Pollinations AI 🇩🇪
🔗 [Get Key](https://enter.pollinations.ai) · Base URL: `https://gen.pollinations.ai/v1`
No sign-up required for basic use. Unique: **text + image + video + audio** all free.
**Text models:** openai, openai-large, openai-reasoning, gemini, mistral, llama.
**Image models:** flux, gpt-image, seedream, kontext.
**Video:** wan-fast. **Audio:** tts-1, 30+ ElevenLabs voices.
### SiliconFlow 🇨🇳
🔗 [Get Key](https://cloud.siliconflow.cn/account/ak) · Base URL: `https://api.siliconflow.cn/v1`
14 CNY signup credits. Permanently free models: 1,000 RPM, 50K TPM.
| Model | Context | Modality |
| --------------------------- | ------- | ---------------- |
| Qwen/Qwen3-8B | 131K | Text |
| DeepSeek-R1-0528-Qwen3-8B | ~33K | Text (reasoning) |
| DeepSeek-R1-Distill-Qwen-7B | 131K | Text (reasoning) |
| THUDM/glm-4-9b-chat | 32K | Text |
| THUDM/GLM-4.1V-9B-Thinking | 66K | Vision+Text |
| DeepSeek-OCR | — | Vision (OCR) |
### Kilo Code 🇺🇸
🔗 [Get Key](https://kilo.ai) · Base URL: `https://api.kilo.ai/api/gateway`
Free models, no credit card. ~200 req/hr. Auto-router `kilo-auto/free`.
### LLM7.io 🇬🇧
🔗 [Get Token](https://token.llm7.io) · Base URL: `https://api.llm7.io/v1`
30+ models. 15 RPM (30 RPM with free token). No registration for basic access.
### Kluster AI 🇺🇸
🔗 [Get Key](https://platform.kluster.ai/apikeys) · DeepSeek-R1, Llama 4 Maverick, Qwen3-235B + more.
### OpenCode Zen
🔗 [Docs](https://opencode.ai/docs/zen/) · Free models (Big Pickle Stealth, MiniMax M2.5 Free, Arcee Large).
### Vercel AI Gateway
🔗 [Docs](https://vercel.com/docs/ai-gateway) · $5/month free credits. Routes to various providers.
---
## China-Based Providers
### ModelScope (魔搭社区) 🇨🇳
🔗 [Get Token](https://modelscope.cn/my/myaccesstoken) · Base URL: `https://api-inference.modelscope.cn/v1`
2,000 req/day total, ≤500/model/day. Requires Alibaba Cloud account + real-name verification.
**Models:** DeepSeek V4 Pro/Flash, DeepSeek V3.2, GLM-5/5.1, MiniMax M2.5, Qwen3-235B, Qwen3 Coder 480B, Ling-2.6-1T.
### Tencent Hunyuan (腾讯混元)
Hunyuan-Lite: free. Other models: 100M tokens free (1-year expiry).
### Volcengine (火山引擎)
500 resource points/day. Tongyi Qwen free (100 calls/day). Doubao models with tiered pricing.
### ChatAnywhere
🔗 Base URL: `https://api.chatanywhere.tech` · GPT-5.4-mini, DeepSeek-V4, and more.
### InternAI (书生)
🔗 Base URL: `https://chat.intern-ai.org.cn/api/v1` · 10 RPM. Keys valid 6 months.
**Models:** intern-latest, intern-s1-pro, internvl3.5-241b-a28b.
### Bigmodel (智谱)
🔗 Base URL: `https://open.bigmodel.cn/api/paas/v4/` · 30 concurrent requests.
**Models:** GLM-4-Flash, GLM-4V-Flash, GLM-4.1V-Thinking-Flash, GLM-4.6V-Flash, GLM-4.7-Flash.
---
## Trial Credit Providers
These offer one-time or time-limited credits (not permanent free tiers):
| Provider | Credits | Expiry | Notable Models |
| ---------------------------------------------------------- | ---------------- | -------- | ----------------------------- |
| [Baseten](https://app.baseten.co/) | $30 | — | Any model (pay by compute) |
| [NLP Cloud](https://nlpcloud.com) | $15 | — | Various open models |
| [AI21](https://studio.ai21.com/) | $10 | 3 months | Jamba family |
| [Upstage](https://console.upstage.ai/) | $10 | 3 months | Solar Pro/Mini |
| [Modal](https://modal.com) | $5/mo | Monthly | Any model (compute time) |
| [SambaNova](https://cloud.sambanova.ai/) | $5 | 3 months | Llama 3.3, Qwen3, DeepSeek R1 |
| [Scaleway](https://console.scaleway.com/generative-api) | 1M tokens | One-time | Llama 3.3, Gemma 3, GPT-OSS |
| [Alibaba Cloud](https://bailian.console.alibabacloud.com/) | 1M tokens/model | — | Qwen family |
| [Fireworks](https://fireworks.ai/) | $1 | — | Various open models |
| [Nebius](https://tokenfactory.nebius.com/) | $1 | — | Various open models |
| [Inference.net](https://inference.net) | $1 (+$25 survey) | — | Various open models |
| [Hyperbolic](https://app.hyperbolic.ai/) | $1 | — | DeepSeek V3, Llama 3.3 |
| [Novita](https://novita.ai/) | $0.50 | 1 year | Various open models |
---
## Using with OmniRoute
OmniRoute supports **all providers listed above** as connections. Here's how to maximize free usage:
### 1. Add Multiple Free Providers
```
Dashboard → Providers → Add Connection
```
Add API keys for Groq, Cerebras, Mistral, Google Gemini, OpenRouter, GitHub Models, etc.
### 2. Create a Free-Tier Combo
```
Dashboard → Combos → Create Combo → Add all free providers as targets
```
Use the **"priority"** or **"round-robin"** strategy to distribute load across free tiers.
### 3. Recommended Free Combo Strategy
| Priority | Provider | Why |
| -------- | ----------------- | --------------------------------------------- |
| 1 | **Groq** | Fastest inference, 14,400 RPD on small models |
| 2 | **Cerebras** | 1M TPD, fast wafer-scale chips |
| 3 | **Mistral** | 1B tokens/month, large model selection |
| 4 | **Google Gemini** | 1M context, multimodal |
| 5 | **NVIDIA NIM** | 129 models, 40 RPM |
| 6 | **OpenRouter** | 35+ free models as final fallback |
### 4. Environment Variables
```bash
# These providers work out of the box with OmniRoute:
GROQ_API_KEY=your-key
CEREBRAS_API_KEY=your-key
MISTRAL_API_KEY=your-key
GOOGLE_AI_API_KEY=your-key
NVIDIA_API_KEY=your-key
OPENROUTER_API_KEY=your-key
GITHUB_TOKEN=your-token
CLOUDFLARE_API_TOKEN=your-token
COHERE_API_KEY=your-key
SILICONFLOW_API_KEY=your-key
```
### 5. Estimated Free Capacity
With all top-6 providers combined in a combo:
| Metric | Combined Free Capacity |
| -------------------- | ---------------------- |
| **Requests/Day** | ~31,000+ RPD |
| **Tokens/Month** | ~32B+ tokens |
| **Models Available** | 200+ unique models |
| **Cost** | $0.00 |
---
## Glossary
| Term | Meaning |
| ----------- | ------------------------------------------- |
| **RPM** | Requests per minute |
| **RPD** | Requests per day |
| **RPH** | Requests per hour |
| **RPS** | Requests per second |
| **TPM** | Tokens per minute |
| **TPD** | Tokens per day |
| **Neurons** | Cloudflare's compute unit (~1 output token) |
---
## Sources
This document consolidates data from 6 community repositories:
| Repository | Focus |
| -------------------------------------------------------------------------- | ------------------------------------------------ |
| [awesome-free-llm-apis](https://github.com/mnfst/awesome-free-llm-apis) | Curated list with detailed model tables |
| [awesome-free-llm-apis2](https://github.com/) | Extended list with speed tiers and code snippets |
| [free-llm-api-resources](https://github.com/) | Auto-generated model lists with trial credits |
| [Free-LLM-Collection](https://github.com/for-the-zero/Free-LLM-Collection) | Chinese + global providers with rate limits |
| [FREE-LLM-API-Provider](https://github.com/CYBIRD-D/FREE-LLM-API-Provider) | Deep provider analysis with CN platforms |
| [gpt4free](https://github.com/xtekky/gpt4free) | Config-based routing with quota awareness |
> ⚠️ **Disclaimer:** Rate limits change frequently. Always verify with the provider's official documentation before relying on specific limits. Trial credits and time-limited promotions are separated from permanent free tiers.

596
docs/PROXY_GUIDE.md Normal file
View File

@@ -0,0 +1,596 @@
# 🌐 OmniRoute Proxy Guide
> **Bypass geographic blocks, protect your identity, and route AI traffic through any proxy — with zero configuration complexity.**
OmniRoute includes a full-featured proxy management system that lets you route upstream AI provider traffic through HTTP, HTTPS, or SOCKS5 proxies. Whether you're in a blocked region, need IP rotation, or want stealth fingerprinting — this guide covers everything.
---
## Table of Contents
- [Why Use Proxies?](#why-use-proxies)
- [Architecture Overview](#architecture-overview)
- [3-Level Proxy System](#3-level-proxy-system)
- [Proxy Registry (CRUD)](#proxy-registry-crud)
- [1proxy Free Marketplace](#1proxy-free-proxy-marketplace)
- [Proxy Rotation](#proxy-rotation)
- [Anti-Detection & Stealth](#anti-detection--stealth)
- [Upstream Proxy Modes](#upstream-proxy-modes)
- [Dashboard UI](#dashboard-ui)
- [API Reference](#api-reference)
- [Environment Variables](#environment-variables)
- [Troubleshooting](#troubleshooting)
---
## Why Use Proxies?
Many AI providers restrict access by geographic region. Developers in **Russia, China, Iran, Cuba, Turkey**, and other countries encounter errors like:
```
unsupported_country_region_territory
```
Even outside blocked regions, proxies are useful for:
| Use Case | Description |
| --------------------- | --------------------------------------------------------------- |
| **Geographic bypass** | Access OpenAI, Anthropic, Codex, Copilot from blocked countries |
| **IP rotation** | Distribute requests across multiple IPs to avoid rate limiting |
| **Privacy** | Hide your real IP from upstream providers |
| **Compliance** | Route traffic through specific jurisdictions |
| **Testing** | Simulate requests from different regions |
---
## Architecture Overview
```
┌───────────────────────────────────────────────────────────────┐
│ OmniRoute Server │
│ │
│ ┌─────────────┐ ┌──────────────┐ ┌──────────────────┐ │
│ │ Proxy │ │ Proxy │ │ Proxy │ │
│ │ Registry │───▶│ Dispatcher │───▶│ Fetch (undici) │ │
│ │ (SQLite) │ │ (cached) │ │ │ │
│ └─────────────┘ └──────────────┘ └────────┬─────────┘ │
│ ▲ │ │
│ │ ▼ │
│ ┌──────┴──────┐ ┌──────────────────┐ │
│ │ 1proxy Sync │ │ Upstream │ │
│ │ (free pool) │ │ Provider API │ │
│ └─────────────┘ └──────────────────┘ │
└───────────────────────────────────────────────────────────────┘
```
### Key Components
| Component | File | Role |
| -------------------- | -------------------------------------------- | ---------------------------------------------------------- |
| **Proxy Registry** | `src/lib/db/proxies.ts` | CRUD for proxy entries + scope assignments |
| **Proxy Dispatcher** | `open-sse/utils/proxyDispatcher.ts` | Creates `undici` ProxyAgent/SOCKS dispatchers with caching |
| **Proxy Fetch** | `open-sse/utils/proxyFetch.ts` | Wraps `fetch()` with proxy dispatcher injection |
| **Settings Route** | `src/app/api/settings/proxy/route.ts` | Legacy proxy config API (GET/PUT/DELETE) |
| **Management Route** | `src/app/api/v1/management/proxies/route.ts` | Registry CRUD API (GET/POST/PATCH/DELETE) |
| **1proxy DB** | `src/lib/db/oneproxy.ts` | Free proxy marketplace persistence |
| **1proxy Sync** | `src/lib/oneproxySync.ts` | Fetches proxies from 1proxy API |
| **1proxy Rotator** | `src/lib/oneproxyRotator.ts` | Rotation strategies (quality/random/sequential) |
---
## 3-Level Proxy System
OmniRoute supports proxy configuration at **four independent scopes**, resolved in priority order:
```
Priority Resolution Order (highest → lowest):
1. 🔵 Account/Connection Proxy → per API key / OAuth connection
2. 🟡 Provider Proxy → per provider (e.g., all OpenAI traffic)
3. 🟠 Combo Proxy → per combo/routing configuration
4. 🟢 Global Proxy → all traffic, all providers
```
### How Resolution Works
When OmniRoute sends a request to an upstream provider, it calls `resolveProxyForConnectionFromRegistry()` which checks each level in order:
1. **Account-level** — Is there a proxy assigned to this specific connection ID?
2. **Provider-level** — Is there a proxy assigned to this provider (e.g., `openai`)?
3. **Global-level** — Is there a global proxy configured?
4. **No proxy** — Direct connection to the provider.
The first match wins. This means you can set a global proxy as a fallback but override it for specific providers or connections.
### What Gets Proxied
| Traffic Type | Proxied? | Notes |
| -------------------- | -------- | --------------------------------------------- |
| Chat completions | ✅ | All `/v1/chat/completions` requests |
| Embeddings | ✅ | `/v1/embeddings` |
| Image generation | ✅ | `/v1/images/generations` |
| Audio (TTS/STT) | ✅ | `/v1/audio/*` |
| OAuth token exchange | ✅ | Solves `unsupported_country_region_territory` |
| Connection tests | ✅ | "Test Connection" button uses proxy |
| Token refresh | ✅ | Background OAuth renewal |
| Model sync | ✅ | Model listing and discovery |
---
## Proxy Registry (CRUD)
The proxy registry is a SQLite table (`proxy_registry`) that stores all your proxies. Each proxy has:
| Field | Type | Description |
| ---------- | ------- | ----------------------------------- |
| `id` | UUID | Unique identifier |
| `name` | String | Human-readable label |
| `type` | String | Protocol: `http`, `https`, `socks5` |
| `host` | String | Proxy hostname or IP |
| `port` | Integer | Port number |
| `username` | String | Auth username (encrypted at rest) |
| `password` | String | Auth password (encrypted at rest) |
| `region` | String | Geographic region label |
| `notes` | String | Free-text notes |
| `status` | String | `active` or `inactive` |
| `source` | String | `manual` or `oneproxy` |
### Creating a Proxy
**Via Dashboard:**
1. Go to **Settings → Proxy**
2. Click **Add Proxy**
3. Fill in the type, host, port, and optional auth credentials
4. Save
**Via API:**
```bash
curl -X POST http://localhost:20128/api/v1/management/proxies \
-H "Content-Type: application/json" \
-d '{
"name": "US Proxy",
"type": "http",
"host": "proxy.example.com",
"port": 8080,
"username": "user",
"password": "pass",
"region": "US"
}'
```
### Updating a Proxy
```bash
curl -X PATCH http://localhost:20128/api/v1/management/proxies \
-H "Content-Type: application/json" \
-d '{
"id": "proxy-uuid-here",
"host": "new-proxy.example.com",
"port": 9090
}'
```
> **Note:** Credentials are preserved unless you explicitly send non-empty replacements. Sending empty strings for `username`/`password` will keep the stored values.
### Deleting a Proxy
```bash
# Fails if proxy is assigned to any scope
curl -X DELETE "http://localhost:20128/api/v1/management/proxies?id=proxy-uuid"
# Force delete (removes assignments too)
curl -X DELETE "http://localhost:20128/api/v1/management/proxies?id=proxy-uuid&force=1"
```
### Listing Proxies
```bash
curl "http://localhost:20128/api/v1/management/proxies?limit=50&offset=0"
```
### Assigning Proxies to Scopes
```bash
# Assign to global scope
curl -X PUT http://localhost:20128/api/settings/proxy \
-H "Content-Type: application/json" \
-d '{"level": "global", "proxy": {"type":"http","host":"proxy.example.com","port":8080}}'
# Assign to a specific provider
curl -X PUT http://localhost:20128/api/settings/proxy \
-H "Content-Type: application/json" \
-d '{"level": "provider", "id": "openai", "proxy": {"type":"socks5","host":"socks.example.com","port":1080}}'
# Assign to a specific connection/key
curl -X PUT http://localhost:20128/api/settings/proxy \
-H "Content-Type: application/json" \
-d '{"level": "key", "id": "connection-uuid", "proxy": {"type":"http","host":"key-proxy.com","port":3128}}'
```
### Resolving Effective Proxy
Check which proxy would be used for a given connection:
```bash
curl "http://localhost:20128/api/settings/proxy?resolve=connection-uuid"
```
Returns the resolved proxy with its level (`account`, `provider`, or `global`) and source.
### Bulk Assignment
Assign one proxy to multiple providers or connections at once:
```bash
curl -X POST http://localhost:20128/api/v1/management/proxies/bulk-assign \
-H "Content-Type: application/json" \
-d '{
"scope": "provider",
"scopeIds": ["openai", "anthropic", "codex"],
"proxyId": "proxy-uuid"
}'
```
### Import/Export
Proxies are included in the **Backup/Restore** system. When you export your OmniRoute configuration:
1. Go to **Dashboard → Settings → Backup**
2. Click **Export** — proxy registry and assignments are included
3. To restore, click **Import** and upload the backup file
The proxy registry also supports **upsert by host+port** — if you import a proxy that already exists (same host and port), it updates instead of creating a duplicate.
### Legacy Migration
If you configured proxies in an older version (pre-registry), OmniRoute automatically migrates them:
```
Legacy key_value store → proxy_registry + proxy_assignments
```
This happens once on first startup after upgrade. Use `migrateLegacyProxyConfigToRegistry({ force: true })` to re-run.
---
## 1proxy Free Proxy Marketplace
> 🆕 **Contributed by [@oyi77](https://github.com/oyi77)** — PR [#1847](https://github.com/diegosouzapw/OmniRoute/pull/1847) (Issue [#1788](https://github.com/diegosouzapw/OmniRoute/issues/1788))
OmniRoute integrates with the **[1proxy](https://1proxy-api.aitradepulse.com)** community platform to provide access to **hundreds of free, validated proxies** from around the world. This is perfect for users who don't have their own proxy infrastructure.
### How It Works
```
┌─────────────┐ Sync ┌─────────────────┐ Rotate ┌──────────┐
│ 1proxy API │ ────────────▶ │ proxy_registry │ ────────────▶ │ Provider │
│ (external) │ up to 500 │ source=oneproxy │ by quality │ API │
└─────────────┘ proxies └─────────────────┘ └──────────┘
```
1. **Sync** — OmniRoute fetches validated proxies from the 1proxy API
2. **Store** — Proxies are saved in the same `proxy_registry` table with `source = 'oneproxy'`
3. **Filter** — Filter by protocol, country, quality score
4. **Rotate** — Pick the best proxy using quality, random, or sequential strategies
5. **Auto-degrade** — Failed proxies get their quality score reduced; below threshold → marked inactive
### Syncing Proxies
**Via Dashboard:**
1. Go to **Settings → 1proxy** tab
2. Click **"Sync Now"**
3. View stats: total proxies, active count, average quality, by-country breakdown
**Via API:**
```bash
# Trigger sync
curl -X POST http://localhost:20128/api/settings/oneproxy \
-H "Content-Type: application/json" \
-d '{}'
# Response:
# { "success": true, "added": 127, "updated": 45, "failed": 2, "total": 172 }
```
### Filtering Proxies
```bash
# Filter by protocol
curl "http://localhost:20128/api/settings/oneproxy?protocol=socks5"
# Filter by country
curl "http://localhost:20128/api/settings/oneproxy?countryCode=US"
# Filter by minimum quality score
curl "http://localhost:20128/api/settings/oneproxy?minQuality=80"
# Combine filters
curl "http://localhost:20128/api/settings/oneproxy?protocol=http&countryCode=DE&minQuality=70"
```
### Proxy Quality Scores
Each 1proxy proxy comes with metadata:
| Field | Description |
| --------------- | -------------------------------------------- |
| `qualityScore` | 0-100 rating from 1proxy validation |
| `latencyMs` | Measured network latency |
| `anonymity` | `transparent`, `anonymous`, or `elite` |
| `googleAccess` | Whether the proxy can access Google services |
| `countryCode` | Two-letter ISO country code |
| `lastValidated` | Timestamp of last validation |
Quality scores are dynamically adjusted:
- **Failed requests** reduce the score by 10 points
- **Score drops to ≤10** → proxy is marked `inactive`
- Inactive proxies are excluded from rotation
### Rotation Strategies
```bash
# Rotate by quality (best proxy first) — default
curl -X POST http://localhost:20128/api/settings/oneproxy/rotate \
-H "Content-Type: application/json" \
-d '{"strategy": "quality"}'
# Random rotation
curl -X POST http://localhost:20128/api/settings/oneproxy/rotate \
-d '{"strategy": "random"}'
# Sequential (least recently validated first)
curl -X POST http://localhost:20128/api/settings/oneproxy/rotate \
-d '{"strategy": "sequential"}'
```
### Circuit Breaker
The 1proxy sync has a built-in circuit breaker:
- After **5 consecutive sync failures**, further sync attempts are blocked
- Reset with: `resetOneproxyCircuitBreaker()` or restart the server
- Sync status is available at `GET /api/settings/oneproxy?action=status`
### Clearing 1proxy Proxies
```bash
# Delete a single 1proxy proxy
curl -X DELETE "http://localhost:20128/api/settings/oneproxy?id=proxy-uuid"
# Clear ALL 1proxy proxies (manual proxies are untouched)
curl -X DELETE "http://localhost:20128/api/settings/oneproxy?clearAll=1"
```
---
## Anti-Detection & Stealth
OmniRoute doesn't just route traffic through a proxy — it makes the traffic look legitimate:
### TLS Fingerprint Spoofing
Uses `wreq-js` to generate browser-like TLS fingerprints, bypassing bot detection systems that flag non-browser TLS handshakes.
### CLI Fingerprint Matching
The **CLI Fingerprint Toggle** (`Settings → Security`) reorders HTTP headers and JSON body fields to match the exact signature of native CLI binaries (Claude Code, Codex, etc.). This works **on top of** the proxy:
```
Your IP (blocked) → Proxy IP (US) → Provider API
+ TLS spoof
+ CLI fingerprint
```
You get both **IP masking** and **request authenticity** simultaneously.
### Proxy IP Preservation
Color-coded badges in the dashboard show which proxy level is active:
| Badge | Level | Meaning |
| ----- | ---------- | ----------------------------------------- |
| 🟢 | Global | All traffic goes through this proxy |
| 🟡 | Provider | Only this provider's traffic is proxied |
| 🔵 | Connection | This specific key/account uses this proxy |
The badge also shows the resolved proxy IP for verification.
---
## Upstream Proxy Modes
For providers that use the CLIProxyAPI pattern, OmniRoute supports three upstream proxy modes:
| Mode | Description |
| ------------- | -------------------------------------------------- |
| `native` | OmniRoute handles proxy routing directly (default) |
| `cliproxyapi` | Delegates to an external CLIProxyAPI instance |
| `fallback` | Tries native first, falls back to CLIProxyAPI |
Configure per-provider:
```bash
curl -X PUT "http://localhost:20128/api/upstream-proxy/openai" \
-H "Content-Type: application/json" \
-d '{"mode": "native", "enabled": true}'
```
---
## Dashboard UI
### Settings → Proxy Tab
- **Global proxy** configuration (set once for all traffic)
- **Per-provider proxy** overrides
- **Per-connection proxy** assignments
- **Connection test** through configured proxy
- **Color-coded badges** showing active proxy level
### Settings → 1proxy Tab
- **Sync Now** button to fetch free proxies
- **Stats cards**: Total, Active, Avg Quality, Last Sync
- **Filters**: Protocol, Country Code, Min Quality
- **Proxy table** with host, protocol, country, quality score, latency, anonymity, Google access
- **Sync status** panel with success/failure tracking and consecutive failure count
- **Clear All** to remove all 1proxy entries
---
## API Reference
### Proxy Settings API
| Method | Endpoint | Description |
| -------- | ---------------------------------------------- | ----------------------- |
| `GET` | `/api/settings/proxy` | Get full proxy config |
| `GET` | `/api/settings/proxy?level=global` | Get global proxy |
| `GET` | `/api/settings/proxy?level=provider&id=openai` | Get provider proxy |
| `GET` | `/api/settings/proxy?resolve=connectionId` | Resolve effective proxy |
| `PUT` | `/api/settings/proxy` | Update proxy config |
| `DELETE` | `/api/settings/proxy?level=provider&id=openai` | Remove proxy at level |
### Proxy Registry API
| Method | Endpoint | Description |
| -------- | ------------------------------------------------- | --------------------- |
| `GET` | `/api/v1/management/proxies` | List all proxies |
| `GET` | `/api/v1/management/proxies?id=uuid` | Get proxy by ID |
| `GET` | `/api/v1/management/proxies?id=uuid&where_used=1` | Get proxy assignments |
| `POST` | `/api/v1/management/proxies` | Create proxy |
| `PATCH` | `/api/v1/management/proxies` | Update proxy |
| `DELETE` | `/api/v1/management/proxies?id=uuid` | Delete proxy |
| `DELETE` | `/api/v1/management/proxies?id=uuid&force=1` | Force delete |
| `POST` | `/api/v1/management/proxies/bulk-assign` | Bulk assign |
| `GET` | `/api/v1/management/proxies/assignments` | List assignments |
| `GET` | `/api/v1/management/proxies/health` | Proxy health stats |
### 1proxy API
| Method | Endpoint | Description |
| -------- | -------------------------------------- | ----------------------- |
| `GET` | `/api/settings/oneproxy` | List 1proxy proxies |
| `GET` | `/api/settings/oneproxy?action=stats` | Get stats + sync status |
| `GET` | `/api/settings/oneproxy?action=status` | Get sync status only |
| `POST` | `/api/settings/oneproxy` | Trigger sync |
| `POST` | `/api/settings/oneproxy/rotate` | Rotate to next proxy |
| `DELETE` | `/api/settings/oneproxy?id=uuid` | Delete one |
| `DELETE` | `/api/settings/oneproxy?clearAll=1` | Clear all |
### Upstream Proxy API
| Method | Endpoint | Description |
| -------- | --------------------------------- | ---------------------------- |
| `GET` | `/api/upstream-proxy/:providerId` | Get upstream proxy config |
| `PUT` | `/api/upstream-proxy/:providerId` | Set upstream proxy mode |
| `DELETE` | `/api/upstream-proxy/:providerId` | Remove upstream proxy config |
---
## Environment Variables
| Variable | Default | Description |
| -------------------------------- | ------------------------------------- | ------------------------------- |
| `ENABLE_SOCKS5_PROXY` | `false` | Enable SOCKS5 proxy support |
| `ONEPROXY_ENABLED` | `true` | Enable 1proxy integration |
| `ONEPROXY_API_URL` | `https://1proxy-api.aitradepulse.com` | 1proxy API endpoint |
| `ONEPROXY_MAX_PROXIES` | `500` | Maximum proxies to sync |
| `ONEPROXY_MIN_QUALITY_THRESHOLD` | `50` | Minimum quality score to import |
---
## Troubleshooting
### "SOCKS5 proxy is disabled"
Set `ENABLE_SOCKS5_PROXY=true` in your `.env` file and restart.
### "socket hang up" errors through proxy
This is normal with cheap proxies that drop idle connections. OmniRoute already handles this by:
- Disabling keep-alive on proxy connections (`keepAliveTimeout: 1`)
- Disabling pipelining (`pipelining: 0`)
- Caching dispatchers to avoid repeated handshakes
If it persists, try a different proxy or use the 1proxy rotation feature.
### "unsupported_country_region_territory" during OAuth
Make sure the proxy is configured **before** starting the OAuth flow. OmniRoute routes OAuth token exchange through the configured proxy. Set a global or provider-level proxy first, then connect.
### Proxy not being used
Check the resolution order:
1. Verify with `GET /api/settings/proxy?resolve=your-connection-id`
2. Check if the proxy `status` is `active` (not `inactive`)
3. Ensure the proxy assignment scope matches your connection
### 1proxy sync failing
Check the sync status:
```bash
curl "http://localhost:20128/api/settings/oneproxy?action=status"
```
If `consecutiveFailures >= 5`, the circuit breaker has tripped. Restart the server to reset, or wait for manual reset.
---
## Database Schema
### `proxy_registry` Table
```sql
CREATE TABLE proxy_registry (
id TEXT PRIMARY KEY,
name TEXT NOT NULL,
type TEXT NOT NULL DEFAULT 'http',
host TEXT NOT NULL,
port INTEGER NOT NULL,
username TEXT DEFAULT '',
password TEXT DEFAULT '',
region TEXT,
notes TEXT,
status TEXT DEFAULT 'active',
source TEXT NOT NULL DEFAULT 'manual', -- 'manual' or 'oneproxy'
quality_score INTEGER, -- 0-100 (1proxy only)
latency_ms INTEGER, -- milliseconds (1proxy only)
anonymity TEXT, -- transparent/anonymous/elite
google_access INTEGER DEFAULT 0, -- can access Google? (1proxy)
last_validated TEXT, -- ISO timestamp (1proxy)
country_code TEXT, -- ISO 2-letter code (1proxy)
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL
);
```
### `proxy_assignments` Table
```sql
CREATE TABLE proxy_assignments (
id INTEGER PRIMARY KEY AUTOINCREMENT,
proxy_id TEXT NOT NULL REFERENCES proxy_registry(id),
scope TEXT NOT NULL, -- 'global', 'provider', 'account', 'combo'
scope_id TEXT, -- provider ID, connection ID, or combo ID
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL,
UNIQUE(scope, scope_id)
);
```
---
> 📖 **Related documentation:**
>
> - [User Guide](USER_GUIDE.md) — General setup and configuration
> - [API Reference](API_REFERENCE.md) — Full API documentation
> - [Environment Config](ENVIRONMENT.md) — All environment variables

186
docs/PWA_GUIDE.md Normal file
View File

@@ -0,0 +1,186 @@
# Progressive Web App (PWA) Guide
OmniRoute ships as a fully installable Progressive Web App. When you access the dashboard from any mobile browser — Android (Chrome) or iOS (Safari) — you can "Add to Home Screen" and get a native app-like experience with no app store required.
## What Is a PWA?
A Progressive Web App turns the OmniRoute web dashboard into something that looks and feels like a native mobile app. Once installed, it:
- Launches from your home screen with its own icon
- Opens fullscreen — no browser address bar or tab UI
- Works offline with a dedicated connectivity page
- Caches static assets for faster loading
- Supports both portrait and landscape orientations
## Installation
### Android (Chrome)
1. Open the OmniRoute dashboard in Chrome: `http://YOUR_IP:20128`
2. Chrome will show an **"Add OmniRoute to Home screen"** banner automatically, or:
- Tap the **⋮** menu (three dots) → **"Add to Home screen"** or **"Install app"**
3. Confirm the prompt
4. OmniRoute appears on your home screen as a standalone app
### iOS (Safari)
1. Open the OmniRoute dashboard in Safari: `http://YOUR_IP:20128`
2. Tap the **Share** button (box with arrow)
3. Scroll down and tap **"Add to Home Screen"**
4. Name it (defaults to "OmniRoute") and tap **Add**
5. OmniRoute appears on your home screen with the app icon
### Desktop (Chrome / Edge)
1. Open the OmniRoute dashboard
2. Click the **install icon** in the address bar (or ⋮ → "Install OmniRoute...")
3. Confirm the prompt
4. OmniRoute opens as a standalone window — no tabs, no address bar
## Features
### Fullscreen Experience
The manifest is configured with `display: "fullscreen"`, which means the installed app uses the entire screen — no browser chrome, no status bar overlap. This makes the dashboard feel truly native.
### Offline Support
OmniRoute includes a service worker (`sw.js`) that provides intelligent caching:
| Asset Type | Strategy | Behavior |
| ------------------------------------------------------- | ---------------------------------- | ---------------------------------------------------------------------------- |
| **App Shell** | Cache-first | `/`, `/offline`, manifest, and icons are pre-cached on install |
| **Static assets** (CSS, JS, images, fonts) | Network-first with cache fallback | Fetches fresh from the network; falls back to cache if offline |
| **Next.js bundles** (`/_next/`) | Network-first with cache update | Fetches from network and updates cache; serves cached version if offline |
| **Navigation requests** | Network-only with offline fallback | Always fetches from network; shows `/offline` page if network is unavailable |
| **API routes** (`/api/`, `/a2a`, `/dashboard/endpoint`) | Bypass (never cached) | Always goes directly to the server — never intercepted by the service worker |
### Offline Page
When the network is unavailable and a user navigates to a new page, the service worker serves a dedicated `/offline` page that:
- Displays a clear **"Connectivity Issue"** message
- Shows a live **online/offline status indicator** that updates in real time
- Provides a **"Retry Connection"** button to reload when connectivity returns
- Links to the **Status Page** for diagnostics
### App Icons
OmniRoute provides icons optimized for each platform:
| File | Size | Used By |
| ---------------------- | ---------------- | ------------------------------------- |
| `icon-512.png` | 512×512 | Android install prompt, splash screen |
| `apple-touch-icon.png` | 180×180 | iOS home screen icon |
| `icon-192.svg` | 192×192 (vector) | Android adaptive icon |
| `apple-touch-icon.svg` | 180×180 (vector) | Apple fallback |
| `favicon.svg` | Vector | Browser tabs |
| `favicon.ico` | Multi-size | Legacy browsers |
### Automatic Registration
The service worker is registered automatically via the `<PwaRegister />` component in the root layout. No user action is needed — the app becomes installable as soon as the browser detects the valid manifest and service worker.
## Technical Architecture
### Web App Manifest (`manifest.webmanifest`)
Generated by Next.js via `src/app/manifest.ts`:
```json
{
"name": "OmniRoute",
"short_name": "OmniRoute",
"description": "OmniRoute is an AI gateway for multi-provider LLMs. One endpoint for all your AI providers.",
"start_url": "/",
"scope": "/",
"display": "fullscreen",
"orientation": "any",
"background_color": "#0b0f1a",
"theme_color": "#0b0f1a",
"icons": [
{ "src": "/icon-512.png", "sizes": "512x512", "type": "image/png", "purpose": "any maskable" },
{ "src": "/apple-touch-icon.png", "sizes": "180x180", "type": "image/png" }
]
}
```
### Service Worker (`public/sw.js`)
A vanilla service worker (no framework dependencies) with:
- **Install phase**: Pre-caches the app shell (root, offline page, manifest, icons)
- **Activate phase**: Cleans up old cache versions and claims all clients
- **Fetch phase**: Intelligent routing based on request type (navigation, static asset, API)
- **Cache versioning**: `omniroute-pwa-v2` — bump this to force a fresh cache on update
### Layout Metadata (`src/app/layout.tsx`)
The root layout provides all the meta tags required for PWA compliance:
- `manifest` link to `/manifest.webmanifest`
- `apple-web-app-capable: true` for iOS standalone mode
- `apple-web-app-status-bar-style: black-translucent`
- `mobile-web-app-capable: yes` for Android Chrome
- `theme-color: #0b0f1a`
- `viewport-fit: cover` for edge-to-edge rendering
### Component: `PwaRegister`
Located at `src/shared/components/PwaRegister.tsx`, this client component:
1. Runs on mount (client-side only)
2. Checks for `serviceWorker` support in the browser
3. Registers `/sw.js` silently (errors are swallowed to avoid blocking the app)
4. Renders nothing (`return null`) — it's a side-effect-only component
## Use With Termux (Android)
When running OmniRoute on Android via Termux, the PWA works seamlessly:
1. Start OmniRoute in Termux: `npx omniroute`
2. Open Chrome on the same phone: `http://localhost:20128`
3. Install the PWA via "Add to Home Screen"
4. The PWA connects to the local Termux server — everything runs on-device
This combination means your Android phone is both the **server** (Termux) and the **client** (PWA) — a complete self-contained AI gateway.
## Use From Other Devices
Install the PWA on any device that has browser access to your OmniRoute server:
- **Another phone/tablet**: Navigate to `http://PHONE_IP:20128` and install the PWA
- **Laptop**: Open Chrome/Edge and install it as a desktop PWA
- **Smart TV with browser**: Access the dashboard fullscreen
## Customization
### Instance Name
The PWA title respects the **Instance Name** setting from `Dashboard → Settings`. If you rename your instance to "My AI Gateway", the installed PWA will show that name.
### Custom Favicon
If you upload a custom favicon via `Dashboard → Settings`, the PWA icon on desktop will reflect the custom icon. Mobile home screen icons use the pre-built `icon-512.png` and `apple-touch-icon.png` files.
## Limitations
- **No push notifications** — The service worker does not implement the Push API. Notifications are handled by the Electron app instead.
- **No background sync** — Offline actions are not queued for replay. The PWA is primarily a dashboard viewer.
- **iOS restrictions** — Safari on iOS does not support all PWA features (e.g., install prompts are manual, and background service workers are limited).
- **Cache size** — The service worker caches static assets only. Large response payloads from `/api/` routes are never cached.
- **Custom icons on mobile** — Changing the favicon in settings does not update the home screen icon on mobile (this requires regenerating the PWA icons).
## Files Reference
| File | Purpose |
| --------------------------------------- | ---------------------------------------------------------------- |
| `src/app/manifest.ts` | Next.js manifest route (generates `manifest.webmanifest`) |
| `public/sw.js` | Service worker with caching logic |
| `src/shared/components/PwaRegister.tsx` | Client component that registers the service worker |
| `src/app/offline/page.tsx` | Offline fallback page with live status indicator |
| `src/app/layout.tsx` | Root layout with PWA metadata (apple-web-app, theme-color, etc.) |
| `public/icon-512.png` | 512×512 PNG icon (Android, splash screen) |
| `public/apple-touch-icon.png` | 180×180 PNG icon (iOS home screen) |
| `public/icon-192.svg` | 192×192 SVG icon (Android adaptive) |
| `public/apple-touch-icon.svg` | 180×180 SVG icon (Apple fallback) |

View File

@@ -1,130 +0,0 @@
# Context Relay
🌐 **Languages:** 🇺🇸 [English](context-relay.md) · 🇪🇸 [es](../i18n/es/docs/features/context-relay.md) · 🇫🇷 [fr](../i18n/fr/docs/features/context-relay.md) · 🇩🇪 [de](../i18n/de/docs/features/context-relay.md) · 🇮🇹 [it](../i18n/it/docs/features/context-relay.md) · 🇷🇺 [ru](../i18n/ru/docs/features/context-relay.md) · 🇨🇳 [zh-CN](../i18n/zh-CN/docs/features/context-relay.md) · 🇯🇵 [ja](../i18n/ja/docs/features/context-relay.md) · 🇰🇷 [ko](../i18n/ko/docs/features/context-relay.md) · 🇸🇦 [ar](../i18n/ar/docs/features/context-relay.md) · 🇮🇳 [hi](../i18n/hi/docs/features/context-relay.md) · 🇮🇳 [in](../i18n/in/docs/features/context-relay.md) · 🇹🇭 [th](../i18n/th/docs/features/context-relay.md) · 🇻🇳 [vi](../i18n/vi/docs/features/context-relay.md) · 🇮🇩 [id](../i18n/id/docs/features/context-relay.md) · 🇲🇾 [ms](../i18n/ms/docs/features/context-relay.md) · 🇳🇱 [nl](../i18n/nl/docs/features/context-relay.md) · 🇵🇱 [pl](../i18n/pl/docs/features/context-relay.md) · 🇸🇪 [sv](../i18n/sv/docs/features/context-relay.md) · 🇳🇴 [no](../i18n/no/docs/features/context-relay.md) · 🇩🇰 [da](../i18n/da/docs/features/context-relay.md) · 🇫🇮 [fi](../i18n/fi/docs/features/context-relay.md) · 🇵🇹 [pt](../i18n/pt/docs/features/context-relay.md) · 🇷🇴 [ro](../i18n/ro/docs/features/context-relay.md) · 🇭🇺 [hu](../i18n/hu/docs/features/context-relay.md) · 🇧🇬 [bg](../i18n/bg/docs/features/context-relay.md) · 🇸🇰 [sk](../i18n/sk/docs/features/context-relay.md) · 🇺🇦 [uk-UA](../i18n/uk-UA/docs/features/context-relay.md) · 🇮🇱 [he](../i18n/he/docs/features/context-relay.md) · 🇵🇭 [phi](../i18n/phi/docs/features/context-relay.md) · 🇧🇷 [pt-BR](../i18n/pt-BR/docs/features/context-relay.md) · 🇨🇿 [cs](../i18n/cs/docs/features/context-relay.md) · 🇹🇷 [tr](../i18n/tr/docs/features/context-relay.md)
---
`context-relay` is a combo strategy that keeps session continuity when the active account
rotates before the conversation is finished.
The current runtime behaves like priority routing for model selection, then adds a
handoff layer on top:
- before the active account is exhausted, OmniRoute generates a compact structured summary
- after authentication selects a different account for the same session, OmniRoute injects
that summary as a system message into the next request
- once the handoff is consumed successfully, it is removed from storage
## When To Use It
Use `context-relay` when all of the following are true:
- the combo is expected to rotate between multiple accounts of the same provider
- losing short-term conversational continuity would hurt task quality
- the provider exposes enough quota information to predict an approaching account limit
This is most useful for long-running coding or research sessions that may outlive a single
account window.
## Runtime Flow
The current behavior is intentionally split across two runtime layers.
### 0% to 84% quota used
No handoff is generated. Requests behave like normal priority routing.
### 85% to 94% quota used
If the active provider is enabled in `handoffProviders`, OmniRoute generates a structured
handoff summary in the background before the account is fully exhausted.
Important details:
- the default warning threshold is `0.85`
- the hard stop for generation is `0.95`
- only one in-flight handoff generation is allowed per `sessionId + comboName`
- if an active handoff already exists for that session/combo, no duplicate summary is generated
### 95% or more quota used
No new handoff is generated. At this point the system is already in or near exhaustion and
the runtime avoids scheduling another summary request.
### After account rotation
When the next request for the same session resolves to a different authenticated account,
OmniRoute prepends the stored handoff as a system message. Injection happens only after the
real account switch is known.
## Handoff Payload
The persisted handoff payload is stored in `context_handoffs` and includes:
- `sessionId`
- `comboName`
- `fromAccount`
- `summary`
- `keyDecisions`
- `taskProgress`
- `activeEntities`
- `messageCount`
- `model`
- `warningThresholdPct`
- `generatedAt`
- `expiresAt`
The summary model is instructed to return a JSON object with this structure:
```json
{
"summary": "Dense summary of what matters for continuity",
"keyDecisions": ["Decision 1", "Decision 2"],
"taskProgress": "What is done, what is pending, and the next step",
"activeEntities": ["fileA.ts", "feature X", "provider Y"]
}
```
At injection time, OmniRoute converts that payload into a `<context_handoff>` system
message so the next account can continue with the correct local context.
## Configuration
`context-relay` supports these config fields:
- `handoffThreshold`: warning threshold for summary generation, default `0.85`
- `handoffModel`: optional model override used only for summary generation
- `handoffProviders`: allowlist of providers allowed to trigger handoff generation
Global defaults can be configured in Settings, and combo-specific values can override them
in the Combos page.
## Architectural Note
The current implementation does not use a standalone `handleContextRelayCombo` handler.
Instead:
- `open-sse/services/combo.ts` decides whether a successful turn should generate a handoff
- `src/sse/handlers/chat.ts` injects the handoff only after authentication resolves the
actual account used for the request
This split is intentional in the current codebase because the combo loop alone does not know
whether the request stayed on the same account or actually switched accounts.
## Limitations
- Effective runtime support is currently centered on `codex` quota rotation.
- `handoffProviders` is already modeled as a config surface, but real handoff generation
still depends on provider-specific quota plumbing.
- The summary is intentionally compact and recent-history based; it is not a full transcript
replay mechanism.
- Handoffs are scoped by `sessionId + comboName` and expire automatically.
- If the session does not switch accounts, the stored handoff is not injected.
## Recommended Usage Pattern
- use multiple accounts from the same provider
- keep stable `sessionId` values across the session
- set `handoffThreshold` early enough to leave room for the background summary request
- treat the feature as continuity assistance, not as a replacement for persistent memory