Rebased onto the current release/v3.8.51 tip as part of a combined provider-retirement/provenance merge batch (Designer Web, Felo Web, Runtime, GPL-derived removal, Qwen Web already landed). Large conflict set (this is the biggest PR in the batch — the common ChatGPT Web provider touches chat, images, count-tokens, session leases, and combos). Conflicts resolved: - `open-sse/config/providers/registry/chatgpt-web/*`, `open-sse/executors/chatgpt-web*`, `open-sse/handlers/imageGeneration/providers/chatgptWeb.ts`, and their tests: kept deleted, matching the PR's stated scope. - `open-sse/config/providers/registry/minimax/web/index.ts`, `open-sse/handlers/imageGeneration/providers/geminiWeb.ts`, `open-sse/executors/gemini-web.ts`'s stale image-mode branch: base-drift collisions against already-merged sibling retirements (#11691, #11708) — kept deleted / dropped the dead code, since this PR's own branch forked before those merged. - `src/shared/constants/reservedProviderPrefixes.ts`, `open-sse/executors/index.ts`, `executorProxy.ts`, `virtualFactory.ts`, `autoStrategy.ts`, `src/lib/db/providers.ts`, `src/sse/handlers/chat.ts`: combined the Designer + Runtime (Felo/Qwen) + common-ChatGPT-Web retirement guard calls at each shared chokepoint — compute-once-then-OR pattern, consistent with prior combinations in this batch. - `src/sse/services/model.ts` / `src/sse/handlers/chatHelpers.ts`: adopted this PR's new `getModelInfoOrRetirementResponse()` central wrapper (a real improvement over ad-hoc try/catch), and extended it to also catch the Designer + Runtime retirement errors it didn't originally cover, so the consolidation doesn't regress the other two mechanisms. - `src/app/api/v1/images/edits/route.ts`: this PR moved the retirement check earlier (before `enforceApiKeyPolicy`) but left the old later call+catch block in place from base drift — removed the now-redundant duplicate `resolveImageRouteModel()` call and merged the Designer catch into the earlier one. - `open-sse/config/imageRegistry.ts`, `tests/snapshots/executors/executor-map.json` (`keyCount` recomputed to 133), `tests/snapshots/provider/translate-path.json`: same "both sides inserted a different retired provider at the same slot" pattern — resolved by dropping both. - `tests/unit/chatcore-executor-proxy.test.ts`, `provider-node-reserved-prefix.test.ts`, `combo-auto-candidate-expansion.test.ts`, `messages-count-tokens-route.test.ts`, `virtual-auto-combo.test.ts`: split into independent per-mechanism test blocks (established pattern); `virtual-auto-combo.test.ts`'s old "includes cookie web-session providers" positive-inclusion test (which used chatgpt-web as its example) was retired along with the provider and replaced by this PR's negative-exclusion test for the same slot. - `docs/architecture/ARCHITECTURE.md`, `CODEBASE_DOCUMENTATION.md` (+ 4 i18n mirrors), `README.md`, `FREE-TIERS-GUIDE.md`, `docs/diagrams/free-tier-budget.svg`, `docs/screenshots/free-tier-budget-card.svg`, `docs/reference/PROVIDER_REFERENCE.md`: recomputed every stale count from the real merged state — 104 executors (`countFiles` gate logic), 351 providers (regenerated via `gen:provider-reference`), 152/351 `hasFree` entries, 445/438/7 free-tier catalog rows, 13 ToS-avoid providers, budget-card regenerated via its real generator script. One doc conflict (`oauth/` module list) needed picking HEAD's side specifically — theirs still listed the already-removed `raycast` module instead of the real `openference`. - `config/quality/test-masking-allowlist.json`: additive merge of the PR's 17 `_deletedWithReplacement` entries alongside the batch's existing ones (one real duplicate-key mistake in my first pass, caught and fixed via a `object_pairs_hook` duplicate-key check before finalizing). Also fixed two real, unrelated-to-my-merge issues surfaced by the focused suite: - `tests/unit/resolve-web-provider-host.test.ts`: the PR's own test had a typo — it asserted `perplexity-web`'s resolved host as `"perplexity.ai"`, but the provider's registered `website` is `"https://www.perplexity.ai"` and the resolver returns the URL's `host` verbatim (no www-stripping), so the correct value is `"www.perplexity.ai"` (consistent with the same test's own `url` assertion). - `tests/unit/hard-session-lease-bypass-inventory.test.ts`: this golden call-site inventory was already stale on the pristine post-#11713 tip (confirmed via a throwaway probe worktree) — `src/lib/db/providers.ts`'s 3 connection-fallback sites and a third `src/app/api/providers/route.ts` site were never added to the golden list by the earlier-merged #11698/#11720 PRs. Updated it to the real current inventory (dated inline comments explain each delta and which PR introduced it), plus this PR's own legitimate deltas (image-edits duplicate-call removal, `ChatGptWebExecutor.execute()` site removed). Focused suite green (433/433 across executor-proxy, reserved-prefix, hard-session-lease-bypass-inventory, resolve-web-provider-host, retirement/runtime-block/source-retirement/management-retirement/image-handler-retirement, migration-168, combo-auto-candidate-expansion, virtual-auto-combo, executor-map-golden and siblings), plus `typecheck:core`, `check-file-size`, and `check-changelog-integrity` clean. Thanks for the thorough provenance-hold retirement work — appreciated.
13 KiB
Free Tiers Guide: Understand and Combine Free AI Access
TL;DR: OmniRoute registers 351 provider IDs, with 152 provider-catalog entries marked
hasFree. The stricter audited free-model catalog covers 39 recurring pool keys / 445 entries (438 active + 7 discontinued). Connect several suitable providers for broader fallback capacity; every quota, approval rule, privacy policy, and paid-overage condition still applies.
What Are Free Tiers?
Many AI providers offer some form of free access. Depending on the provider, that may mean a no-auth endpoint, recurring quota, rate-limited uncapped access, a signup grant, manual approval, or a temporary promotion. Some options require an account, API key, credit card, KYC, or acceptance of provider-specific terms.
OmniRoute aggregates these free tiers into one endpoint. Instead of signing up for 10 different services, you connect them all to OmniRoute and use model: "auto" to automatically pick the best free option for each request.
Representative Free-Access Providers
Recurring, Keyless, or Uncapped Access
These providers have a recurring, keyless, or uncapped free-access path in the audited catalog. “Uncapped” means no published token cap; rate, concurrency, account, regional, and policy limits can still apply:
| Provider | Models | Quota | How to Connect |
|---|---|---|---|
| Kiro AI | Claude Sonnet 4.5, Haiku 4.5, DeepSeek V3.2, and others | Audited catalog estimates a 25K-token shared monthly pool | OAuth/account flow; ToS flagged avoid in the catalog |
| OpenCode Free | Current *-free model set in the provider registry |
Keyless; no published token cap | No provider credential; ToS flagged avoid |
| Pollinations | Current keyless model set; some former models are discontinued or key-required | Keyless; no published token cap | No provider credential for the keyless models |
| Logfare | kimi-k3, deepseek-v4-pro, glm-5.2, gpt-5.6-luna, minimax-m3, and more | Free API key (no rate limits, no card); every request is logged for research (opt out at logfare.ai/consent) | Instant key at logfare.ai/register; ToS/privacy at logfare.ai/tos and logfare.ai/privacy |
| Cloudflare AI | Workers AI catalog | Audited pool estimates ~30M tokens/month from published usage units | Cloudflare account and API credentials |
| Gemini | Gemini Flash family | Audited pool estimates ~60M tokens/month | Google AI Studio API key; rate limits apply |
| Groq | Llama, GPT-OSS, and Qwen models | Audited pool estimates ~15M tokens/month | Groq API key; rate limits apply |
| Cerebras | GLM 4.7 and GPT-OSS 120B | Audited pool estimates ~30M tokens/month | Cerebras API key; rate limits apply |
Signup Grants and Provider-Specific Credits
These providers give you free credits when you sign up:
| Provider | Free Credits | Models | How to Get |
|---|---|---|---|
| DeepSeek | 5M free tokens | DeepSeek V4 | Sign up at platform.deepseek.com |
| LongCat | 10M-token one-time grant | LongCat 2.0 | API key + KYC; pay-as-you-go after the grant |
| Together | $25 signup credit represented as ~25M tokens in the budget model | Provider catalog | Sign up and verify current terms |
| Vertex AI | $300 signup credit represented as ~300M tokens in the budget model | Gemini and partner models | Google Cloud account; billing and eligibility rules apply |
Other Limited Access
These providers have free tiers with specific limits:
| Provider | Free Limit | Models | Best For |
|---|---|---|---|
| GitHub Models | Audited shared pool estimates ~18M tokens/month | Broad model evaluation | |
| Hugging Face | Small recurring monthly pool | Experiments and model variety | |
| OpenRouter free models | Shared request-limited pool; optional one-time top-up increases the recurring allowance | Broad fallback catalog | |
| AI Horde | Keyless community capacity; availability varies | Opportunistic distributed inference |
How to Stack Free Tiers
The magic of OmniRoute is stacking free tiers. Instead of relying on one provider, you connect multiple free providers and let OmniRoute automatically pick the best one for each request.
Example: Broader Free-Tier Coverage
Connect several providers to reduce dependence on any single quota:
- Gemini — recurring API-key quota
- Groq — recurring API-key quota
- Pollinations — keyless, rate-limited access
- LongCat — one-time signup grant (requires KYC)
Then use model: "auto" and OmniRoute will:
- Try the highest-ranked eligible connection first
- If its quota or health check fails → try the next configured provider
- If the keyless provider is unavailable → continue through the remaining targets
- If all fail → use LongCat as backup
Result: broader free-tier coverage with automatic fallback — not a guarantee of unlimited capacity.
How to Connect Free Providers
Step 1: Open the Dashboard
Go to http://localhost:20128 in your browser.
Step 2: Go to Providers
Click Providers in the sidebar.
Step 3: Click Add Provider
Click the + Add Provider button.
Step 4: Select a Free Provider
Browse the catalog and inspect each provider's current hasFree, auth, quota, privacy,
and ToS metadata. The provider card and the
Free Tiers Reference distinguish recurring pools,
uncapped/keyless access, signup credits, discontinued entries, and higher-risk sources.
Step 5: Click Connect
For a NOAUTH provider, no credential is required. OAuth and API-key providers must be
connected through their documented account flow.
Step 6: Repeat
Connect several providers whose terms and privacy model fit your use case.
Reading the Catalog Correctly
NOAUTHmeans OmniRoute does not ask you for a provider credential; it does not guarantee uptime, privacy, or unlimited capacity.hasFreeis discovery metadata. It can represent a recurring quota, keyless access, signup credit, approval program, or promotion.recurring-uncappedmeans no published token ceiling was available; rate and concurrency limits still apply.one-time-initialdoes not recur after the signup grant is consumed.tos: avoidis a warning to review provider terms and account risk before use.- Entries marked
discontinuedremain historical evidence and must not be presented as currently free.
How OmniRoute Makes Free Tiers Better
1. Automatic Fallback
If one free provider is busy or down, OmniRoute automatically tries the next one. You don't need to do anything.
2. Smart Routing
OmniRoute picks the best free provider for each request based on:
- Speed — Which provider is fastest right now?
- Quality — Which provider is best for this task?
- Capacity — Which provider has quota remaining?
3. Token Savings
OmniRoute's compression pipeline can reduce eligible prompt and tool-output tokens. The actual savings depend on content, selected engines, provider accounting, and fidelity settings; compression does not multiply every provider quota by a fixed amount.
4. Multi-Account Support
If provider terms permit multiple accounts or credentials, OmniRoute can treat each connection as a separate routing candidate. Do not create extra accounts to evade a provider's quota or access policy.
Free Tier Math
The live, pool-deduplicated catalog currently reports:
| Metric | Current audited value | Interpretation |
|---|---|---|
| Recurring quantified grant | ~1.51B tokens/month | Shared pools counted once; excludes uncapped providers from the sum |
| First month with signup grants | ~2.13B tokens | Recurring total plus one-time and recurring credits |
| Audited free-model inventory | 39 recurring pool keys / 445 catalog entries | 438 active + 7 discontinued; distinct from the 351-provider catalog |
| Recurring/keyless free-forever providers represented | 55 | Unique providers across recurring daily/monthly/credit/uncapped and keyless catalog types |
Provider catalog entries marked hasFree |
152 / 351 | Broader provider metadata; not all have a quantifiable recurring quota |
These values are computed from open-sse/config/freeModelCatalog.ts; see the
Free Tiers Reference for pool deduplication, ToS flags,
discontinued entries, and signup-credit methodology.
Common Questions
"Is this really free?"
The catalog records provider-published terms and project research, but offers can change. Verify the provider's current pricing, quota, privacy policy, and eligibility before use.
"Will the free tier run out?"
Every provider can rate-limit, change models, suspend access, or go offline. Multiple connections improve fallback coverage but do not guarantee an available free route.
"Can I use free providers for production?"
Only if the provider's SLA, data handling, limits, and terms meet your production requirements. Critical workloads should have monitored, contractually suitable fallback.
"What's the catch?"
Tradeoffs may include strict limits, waitlists, KYC, credit-card verification, training on prompts, weaker privacy, no SLA, model churn, geographic restrictions, paid overage, or account-policy risk. OmniRoute surfaces the available metadata; you choose what to enable.
"How do I get more free quota?"
- Connect more free providers
- Enable appropriate compression engines and measure savings for your workload
- Use
auto/cheapto prioritize free/cheap providers - Add additional permitted providers or credentials without violating provider terms
"Do free providers have worse quality?"
Not necessarily. Some providers expose the same model families available through paid routes, but limits, latency, privacy, reliability, and model versions can differ. Use the Free Provider Rankings page as a quality signal and verify the actual model served.
What's Next?
- Auto-Combo Guide — Let OmniRoute pick the best AI for you
- Providers Guide — Connect more providers
- Troubleshooting — Fix common issues
- Free Tiers Reference — Full list of free tiers