refactor(providers): remove retired GitHub Models (#9023)

This commit is contained in:
backryun
2026-08-11 19:19:32 +09:00
committed by GitHub
parent ed2f3cd067
commit 668ed4813f
23 changed files with 24 additions and 581 deletions

View File

@@ -0,0 +1 @@
- **refactor(providers):** removed the retired GitHub Models provider and its catalog, discovery, embedding, free-tier, UI, and documentation surfaces; upgrades now run a durable, idempotent purge of its stored credentials, usage state, structured configuration, and call-log artifacts while preserving GitHub Copilot and live members of mixed configurations ([#9023](https://github.com/diegosouzapw/OmniRoute/pull/9023))

View File

@@ -128,7 +128,7 @@ Content-Type: application/json
}
```
Available providers: Nebius, OpenAI, Mistral, Together AI, Fireworks, NVIDIA, **OpenRouter**, **GitHub Models**.
Available providers: Nebius, OpenAI, Mistral, Together AI, Fireworks, NVIDIA, **OpenRouter**.
Registry models that advertise multimodal support also accept up to 32 provider-neutral structured
items. Media item types are `text`, `image`, `audio`, `video`, and `document`. Their media `source`

View File

@@ -1,7 +1,7 @@
---
title: "Free Tiers & Free-Token Budget"
version: 3.8.40
lastUpdated: 2026-06-28
lastUpdated: 2026-07-31
---
# Free Tiers & Free-Token Budget
@@ -15,19 +15,19 @@ lastUpdated: 2026-06-28
| Metric | Tokens / month | Meaning |
| ------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Documented recurring grant (steady)** | **~1.53B** | Free-tier **pools** (per-model catalog), each shared pool counted **once**. The live source behind `/api/free-tier/summary` and the dashboard's Free-Tier Budget page. **Use this number.** |
| **+ first month with signup credits** | **~2.15B** | Steady + one-time signup credits (Together $25, Z.AI 20M, DeepSeek 5M, …), deduped per account. **First month only** — does not recur. |
| **Documented recurring grant (steady)** | **~1.51B** | Free-tier **pools** (per-model catalog), each shared pool counted **once**. The live source behind `/api/free-tier/summary` and the dashboard's Free-Tier Budget page. **Use this number.** |
| **+ first month with signup credits** | **~2.13B** | Steady + one-time signup credits (Together $25, Z.AI 20M, DeepSeek 5M, …), deduped per account. **First month only** — does not recur. |
| **+ permanently free, no published cap** | _un-quantifiable_ | `siliconflow`, `glm-cn` (GLM-4-Flash), `tencent`, `baidu`, `kilo-gateway`, `opencode-zen` — real recurring access, rate/concurrency-limited, **no token cap to count**. Listed, never summed (counting them at `RPM×24/7` is the inflation we reject). |
| **+ deposit-unlock boost** | **+~24M** | A one-time **$10** OpenRouter top-up raises its free pool from 50 → 1000 req/day. Reported separately so it never inflates the steady number. |
| Theoretical ceiling (all rate limits, 24/7) | ~10B | Sum of every provider rate limit extrapolated to non-stop use. **Not a guarantee** — do not headline this. |
**Honest headline:** _OmniRoute aggregates **~1.53B documented free tokens per month** (up to ~2.15B in your first month with signup credits) across 43 free-tier pools — plus a long tail of permanently-free, no-cap providers — and RTK + Caveman compression (1595% token savings) stretches that further._
**Honest headline:** _OmniRoute aggregates **~1.51B documented free tokens per month** (up to ~2.13B in your first month with signup credits) across 42 free-tier pools — plus a long tail of permanently-free, no-cap providers — and RTK + Caveman compression (1595% token savings) stretches that further._
> **Why this dropped from the previous ~1.94B.** The 2026-06-17 refresh is an honesty correction, not a loss: `gemini` is now pool-deduped (was inflated by counting each Flash variant separately, 462M → 60M), `cloudflare-ai` corrected to its real 10k-Neurons/day (122M → 30M), `doubao` reclassified as a one-time signup credit (not recurring), and shut-down tiers removed (`github-models` closed to new signups, `chutes`/`phind`/`kluster` discontinued). Partly offset by `llm7` (correct 5M/day → 150M) and new free providers (Kilo, OpenCode Zen, Z.AI GLM-Flash).
> **Why this dropped from the previous ~1.94B.** The 2026-06-17 refresh is an honesty correction, not a loss: `gemini` is now pool-deduped (was inflated by counting each Flash variant separately, 462M → 60M), `cloudflare-ai` corrected to its real 10k-Neurons/day (122M → 30M), `doubao` reclassified as a one-time signup credit (not recurring), and shut-down tiers removed (`chutes`/`phind`/`kluster` discontinued). Partly offset by `llm7` (correct 5M/day → 150M) and new free providers (Kilo, OpenCode Zen, Z.AI GLM-Flash).
>
> **Further corrected to ~1.37B in v3.8.42:** `longcat` was reclassified from a 150M/mo recurring grant to a one-time 10M signup credit after its free preview ended. Same honesty rule — no provider was dropped by mistake.
>
> **Updated to ~1.53B in v3.8.49:** the pool count grew from 39 to 43 after mapping free tiers that were documented upstream but missing from the catalog (`requesty`, `ovhcloud`, `agnes`, `glm`) plus new providers `navy` and `aihorde` (#7840). This is the live, CI-gated number (`check:docs-counts` fails the build if this drifts from `computeFreeModelTotals()`).
> **Updated to ~1.51B after removing a retired provider:** the pool count is now 42 after mapping free tiers that were documented upstream but missing from the catalog (`requesty`, `ovhcloud`, `agnes`, `glm`) plus new providers `navy` and `aihorde` (#7840). This is the live, CI-gated number (`check:docs-counts` fails the build if this drifts from `computeFreeModelTotals()`).
Biggest **documented** contributors: `mistral` 1.00B, `llm7` 150M, `groq` 117M, `gemini` 60M, `cerebras` 30M, `cloudflare-ai` 30M, `sambanova` 30M. (`longcat` is excluded — its 10M LongCat-2.0 grant is a one-time, KYC-gated signup credit, not a recurring monthly budget.)
@@ -40,7 +40,6 @@ Biggest **documented** contributors: `mistral` 1.00B, `llm7` 150M, `groq` 117M,
A 50-agent web-research pass (official docs + last-7-days news, adversarially verified) refreshed the whole catalog. Highlights:
- **Removed / no free tier (2026):** `chutes` (free tier ended 2026-03), `phind` (company shut down 2026-01), `kluster` (sunset 2026-06-09 → MITO), `gitlawb` + `gitlawb-gmi` (MiMo free revoked 2026-05-24, Nemotron promo ended 2026-06 — re-verified 2026-06-18), `aimlapi` (free tier paused — re-verified 2026-06-18), `yi` (Yi-Light retired, pay-as-you-go — re-verified 2026-06-18), `theoldllm` / `featherless-ai` (no current free tier). `iflytek` / `sparkdesk` stay listed but carry a ToS-caution note (Spark Lite is free; the ToS restricts proxy/relay use).
- **GitHub Models** — closed to **new** customers on 2026-06-16; existing accounts keep API/playground access, so it stays in the catalog with a note (not removed).
- **Gemini** — `2.0 Flash` / `2.0 Flash-Lite` shut down 2026-06-01 and `2.5 Pro` left the free tier (2026-04); free tier is now **Flash-family only** (2.5/3/3.1/3.5 Flash + Gemma). The catalog now **pools** the Flash family (was inflated by counting each variant separately: 462M → 60M).
- **Corrected numbers:** `cloudflare-ai` 122M → **30M** (real 10k-Neurons/day), `doubao` reclassified as a one-time signup credit (not recurring), `llm7` 4M → **150M** (documented 5M tokens/day), `together` "-Free" endpoints discontinued → only the **$25** signup credit remains, `longcat` Preview ended + Flash models retired → **LongCat-2.0** only, reclassified as a one-time **10M**-token signup credit (KYC-gated, not recurring).
- **New free providers discovered:** ⭐ **Kilo Code** (`kilo-gateway` — rotating "Auto Free" set: NVIDIA Nemotron 3 family, StepFun, Poolside, Nex-N2-Pro), ⭐ **OpenCode Zen** (`opencode-zen` — 6 rotating free coding models), ⭐ **Z.AI / Zhipu** (`glm-cn` — GLM-4-Flash / 4.5-Flash / 4.7-Flash permanently free + 20M signup bonus), and `arcee-ai` Trinity Large Preview.
@@ -118,7 +117,6 @@ A 50-agent web-research pass (official docs + last-7-days news, adversarially ve
| `exa-search` | caution | No explicit "no proxy" or "evaluation only" clauses found; Exa actively offers a reseller partner program allowing API … |
| `firecrawl` | caution | Cloud API ToS has no explicit personal-proxy prohibition found, but the open-source self-hosted version is AGPL-3.0 (re… |
| `gemini` | caution | ToS explicitly states the free tier is for "developers building with Google AI models for professional or business purp… |
| `github-models` | caution | GitHub's Acceptable Use Policy prohibits reselling/proxying the service; GitHub Models ToS delegates to each model's ho… |
| `groq` | caution | Services Agreement §6.3 prohibits reselling, sublicensing, or distributing API access; §3.2 bars reselling/leasing acco… |
| `hackclub` | caution | Service is explicitly scoped to Hack Club teen members building projects/learning; no public ToS found explicitly permi… |
| `huggingchat` | caution | Hugging Face ToS does not explicitly ban personal self-hosted proxies, but supplemental terms (referenced but not fully… |
@@ -183,7 +181,6 @@ A 50-agent web-research pass (official docs + last-7-days news, adversarially ve
| `cloudflare-ai` | recurring | ~30M | — | caution | 6 |
| `api-airforce` | recurring | ~24M | — | caution | 7 |
| `ollama-cloud` | recurring | ~20M | — | ambiguous | 8 |
| `github-models` | recurring | ~18M | — | caution | 14 |
| `groq` | recurring | ~15M | — | caution | 5 |
| `bluesminds` | recurring | ~7M | — | ambiguous | 22 |
| `sambanova` | recurring | ~6M | — | caution | 5 |
@@ -280,7 +277,6 @@ A 50-agent web-research pass (official docs + last-7-days news, adversarially ve
- **`freemodel-dev`** — Our shipped freeNote is "(none)" — this was likely a placeholder meaning the provider was not yet cataloged. In reality the provider does have a $300 one-time trial credit offer. However, this is a o…
- **`friendliai`** — The shipped freeNote ("Free tier for serverless inference") is partially accurate but misleading. There is free access via Tier 0 and free-designated models, but the rate limits are undefined and ada…
- **`gemini`** — The shipped freeNote says "1,500 req/day for Gemini 2.5 Flash" — this was accurate before December 2025. Google cut free-tier limits by 50-80% in December 2025, reducing Gemini 2.5 Flash from 1,500 R…
- **`github-models`** — Catalog note "Free GPT-5, o-series, DeepSeek-R1, Llama 4, Grok 3" is directionally correct about model availability but omits the daily rate limits (50 RPD for high-tier models, 150 RPD for low-tier)…
- **`gitlawb`** — The shipped freeNote "Free tier available" is effectively stale. The original free MiMo access was removed in May 2026; the only remaining "free" option is a temporary promotional model (Nemotron 3 U…
- **`gitlawb-gmi`** — Partially still accurate — free tier exists but is now narrowed to a single model (Nemotron 3 Ultra) after MiMo free access was revoked in late May 2026. The shipped note "Free tier available" unders…
- **`groq`** — The shipped freeNote "30 RPM / 14.4K RPD" is accurate only for llama-3.1-8b-instant. Most other models (including llama-3.3-70b-versatile) have a much lower 1K RPD cap. The note omits model-specific …

View File

@@ -287,36 +287,6 @@ export const EMBEDDING_PROVIDERS: Record<string, EmbeddingProvider> = {
],
},
"github-models": {
id: "github-models",
baseUrl: "https://models.github.ai/inference/embeddings",
authType: "apikey",
authHeader: "bearer",
models: [
{
id: "openai/text-embedding-3-large",
name: "OpenAI Text Embedding 3 (large)",
dimensions: 3_072,
},
{
id: "openai/text-embedding-3-small",
name: "OpenAI Text Embedding 3 (small)",
dimensions: 1_536,
},
],
},
github: {
id: "github",
baseUrl: "https://models.inference.ai.azure.com/embeddings",
authType: "apikey",
authHeader: "bearer",
models: [
{ id: "text-embedding-3-small", name: "Text Embedding 3 Small (GitHub)", dimensions: 1536 },
{ id: "text-embedding-3-large", name: "Text Embedding 3 Large (GitHub)", dimensions: 3072 },
],
},
"jina-ai": {
id: "jina-ai",
structuredInputProtocol: "jina-v1",

View File

@@ -188,27 +188,6 @@ export const FREE_MODEL_BUDGETS: FreeModelBudget[] = [
{ provider: "gemini", modelId: "gemini-3-flash-preview", displayName: "Gemini 3 Flash Preview", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-daily", poolKey: "gemini-free", tos: "caution" },
{ provider: "gemini", modelId: "gemini-3.1-flash-lite", displayName: "Gemini 3.1 Flash-Lite", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-daily", poolKey: "gemini-free", tos: "caution" },
{ provider: "gemini", modelId: "gemini-3.5-flash", displayName: "Gemini 3.5 Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-daily", poolKey: "gemini-free", tos: "caution" },
{ provider: "github-models", modelId: "cohere/cohere-command-a", displayName: "Cohere Command A (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "deepseek/deepseek-r1-0528", displayName: "DeepSeek-R1-0528 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "deepseek/deepseek-v3-0324", displayName: "DeepSeek-V3-0324 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "meta/llama-4-maverick-17b-128e-instruct-fp8", displayName: "Llama 4 Maverick 17B 128E Instruct FP8 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "meta/llama-3.3-70b-instruct", displayName: "Llama-3.3-70B-Instruct (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "meta/llama-4-scout-17b-16e-instruct", displayName: "Llama 4 Scout 17B 16E Instruct (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "microsoft/phi-4-multimodal-instruct", displayName: "Phi-4-multimodal-instruct (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "microsoft/phi-4-reasoning", displayName: "Phi-4-reasoning (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "mistral-ai/codestral-2501", displayName: "Codestral 25.01 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "mistral-ai/mistral-medium-2505", displayName: "Mistral Medium 3 (25.05) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-4.1", displayName: "OpenAI GPT-4.1 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-4.1-mini", displayName: "OpenAI GPT-4.1-mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-4o", displayName: "OpenAI GPT-4o (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-4o-mini", displayName: "OpenAI GPT-4o mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-5", displayName: "OpenAI gpt-5 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-5-chat", displayName: "OpenAI gpt-5-chat (preview) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/gpt-5-mini", displayName: "OpenAI gpt-5-mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/o3", displayName: "OpenAI o3 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/o4-mini", displayName: "OpenAI o4-mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/text-embedding-3-large", displayName: "OpenAI Text Embedding 3 (large) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "github-models", modelId: "openai/text-embedding-3-small", displayName: "OpenAI Text Embedding 3 (small) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
{ provider: "glm-cn", modelId: "glm-4-flash", displayName: "GLM-4-Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-uncapped", poolKey: "zhipu-flash-free", tos: "ok" },
{ provider: "glm-cn", modelId: "glm-4.5-flash", displayName: "GLM-4.5-Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-uncapped", poolKey: "zhipu-flash-free", tos: "ok" },
{ provider: "glm-cn", modelId: "glm-4.7-flash", displayName: "GLM-4.7-Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-uncapped", poolKey: "zhipu-flash-free", tos: "ok" },

View File

@@ -19,7 +19,6 @@ export const FREE_TIER_BUDGETS: Record<string, number> = {
cerebras: 30_000_000,
"api-airforce": 24_000_000,
"ollama-cloud": 20_000_000,
"github-models": 18_000_000,
groq: 15_000_000,
bluesminds: 7_200_000,
sambanova: 6_000_000,

View File

@@ -27,7 +27,6 @@ import { raycastProvider } from "./registry/raycast/index.ts";
import { muse_spark_webProvider } from "./registry/muse-spark-web/index.ts";
import { lmarenaProvider } from "./registry/lmarena/index.ts";
import { kilocodeProvider } from "./registry/kilocode/index.ts";
import { github_modelsProvider } from "./registry/github/models/index.ts";
import { githubProvider } from "./registry/github/index.ts";
import { gheCopilotProvider } from "./registry/ghe-copilot/index.ts";
import { difyProvider } from "./registry/dify/index.ts";
@@ -256,7 +255,6 @@ export const REGISTRY: Record<string, RegistryEntry> = {
"muse-spark-web": muse_spark_webProvider,
lmarena: lmarenaProvider,
kilocode: kilocodeProvider,
"github-models": github_modelsProvider,
github: githubProvider,
"ghe-copilot": gheCopilotProvider,
dify: difyProvider,

View File

@@ -1,187 +0,0 @@
import type { RegistryEntry } from "../../../shared.ts";
export const github_modelsProvider: RegistryEntry = {
id: "github-models",
alias: "ghm",
format: "openai",
executor: "default",
baseUrl: "https://models.github.ai/inference/chat/completions",
modelsUrl: "https://models.github.ai/catalog/models",
authType: "apikey",
authHeader: "Authorization",
authPrefix: "Bearer",
headers: {
"X-GitHub-Api-Version": "2026-03-10",
Accept: "application/vnd.github+json",
},
defaultContextLength: 128000,
models: [
{
id: "cohere/cohere-command-a",
name: "Cohere Command A",
contextLength: 131_072,
maxInputTokens: 131_072,
maxOutputTokens: 4_096,
},
{
id: "deepseek/deepseek-r1-0528",
name: "DeepSeek-R1-0528",
contextLength: 128_000,
maxInputTokens: 128_000,
maxOutputTokens: 4_096,
supportsReasoning: true,
toolCalling: true,
},
{
id: "deepseek/deepseek-v3-0324",
name: "DeepSeek-V3-0324",
contextLength: 128_000,
maxInputTokens: 128_000,
maxOutputTokens: 4_096,
toolCalling: true,
},
{
id: "meta/llama-4-maverick-17b-128e-instruct-fp8",
name: "Llama 4 Maverick 17B 128E Instruct FP8",
contextLength: 1_000_000,
maxInputTokens: 1_000_000,
maxOutputTokens: 4_096,
supportsVision: true,
toolCalling: true,
},
{
id: "meta/llama-3.3-70b-instruct",
name: "Llama-3.3-70B-Instruct",
contextLength: 128_000,
maxInputTokens: 128_000,
maxOutputTokens: 4_096,
},
{
id: "meta/llama-4-scout-17b-16e-instruct",
name: "Llama 4 Scout 17B 16E Instruct",
contextLength: 10_000_000,
maxInputTokens: 10_000_000,
maxOutputTokens: 4_096,
supportsVision: true,
toolCalling: true,
},
{
id: "microsoft/phi-4-multimodal-instruct",
name: "Phi-4-multimodal-instruct",
contextLength: 128_000,
maxInputTokens: 128_000,
maxOutputTokens: 4_096,
supportsVision: true,
},
{
id: "microsoft/phi-4-reasoning",
name: "Phi-4-reasoning",
contextLength: 32_768,
maxInputTokens: 32_768,
maxOutputTokens: 4_096,
supportsReasoning: true,
},
{
id: "mistral-ai/codestral-2501",
name: "Codestral 25.01",
contextLength: 256_000,
maxInputTokens: 256_000,
maxOutputTokens: 4_096,
},
{
id: "mistral-ai/mistral-medium-2505",
name: "Mistral Medium 3 (25.05)",
contextLength: 128_000,
maxInputTokens: 128_000,
maxOutputTokens: 4_096,
supportsVision: true,
toolCalling: true,
},
{
id: "openai/gpt-4.1",
name: "OpenAI GPT-4.1",
contextLength: 1_048_576,
maxInputTokens: 1_048_576,
maxOutputTokens: 32_768,
supportsVision: true,
toolCalling: true,
},
{
id: "openai/gpt-4.1-mini",
name: "OpenAI GPT-4.1-mini",
contextLength: 1_048_576,
maxInputTokens: 1_048_576,
maxOutputTokens: 32_768,
supportsVision: true,
toolCalling: true,
},
{
id: "openai/gpt-4o",
name: "OpenAI GPT-4o",
contextLength: 131_072,
maxInputTokens: 131_072,
maxOutputTokens: 16_384,
supportsVision: true,
toolCalling: true,
},
{
id: "openai/gpt-4o-mini",
name: "OpenAI GPT-4o mini",
contextLength: 131_072,
maxInputTokens: 131_072,
maxOutputTokens: 4_096,
supportsVision: true,
toolCalling: true,
},
{
id: "openai/gpt-5",
name: "OpenAI gpt-5",
contextLength: 200_000,
maxInputTokens: 200_000,
maxOutputTokens: 100_000,
supportsVision: true,
supportsReasoning: true,
toolCalling: true,
},
{
id: "openai/gpt-5-chat",
name: "OpenAI gpt-5-chat (preview)",
contextLength: 200_000,
maxInputTokens: 200_000,
maxOutputTokens: 100_000,
supportsVision: true,
supportsReasoning: true,
toolCalling: true,
},
{
id: "openai/gpt-5-mini",
name: "OpenAI gpt-5-mini",
contextLength: 200_000,
maxInputTokens: 200_000,
maxOutputTokens: 100_000,
supportsVision: true,
supportsReasoning: true,
toolCalling: true,
},
{
id: "openai/o3",
name: "OpenAI o3",
contextLength: 200_000,
maxInputTokens: 200_000,
maxOutputTokens: 100_000,
supportsVision: true,
supportsReasoning: true,
toolCalling: true,
},
{
id: "openai/o4-mini",
name: "OpenAI o4-mini",
contextLength: 200_000,
maxInputTokens: 200_000,
maxOutputTokens: 100_000,
supportsVision: true,
supportsReasoning: true,
toolCalling: true,
},
],
};

View File

@@ -1,14 +1,7 @@
import { getRegistryEntry } from "@omniroute/open-sse/config/providerRegistry.ts";
import type { ProviderModelsConfigEntry } from "./discovery/providerModelsConfig";
const DISCOVERY_EXCLUDED_MODEL_IDS: Readonly<Record<string, ReadonlySet<string>>> = {
"github-models": new Set([
"meta/llama-3.2-11b-vision-instruct",
"meta/llama-3.2-90b-vision-instruct",
]),
};
function parseRegistryModelsResponse(provider: string, data: unknown): unknown[] {
function parseRegistryModelsResponse(data: unknown): unknown[] {
const response = data as { data?: unknown; models?: unknown } | null;
const models = Array.isArray(data)
? data
@@ -17,15 +10,7 @@ function parseRegistryModelsResponse(provider: string, data: unknown): unknown[]
: Array.isArray(response?.models)
? response.models
: [];
const excludedIds = DISCOVERY_EXCLUDED_MODEL_IDS[provider];
if (!excludedIds) return models;
return models.filter((model) => {
if (!model || typeof model !== "object" || !("id" in model)) return true;
const id = (model as { id?: unknown }).id;
return typeof id !== "string" || !excludedIds.has(id.toLowerCase());
});
return models;
}
/**
@@ -47,7 +32,7 @@ export function deriveConfigFromRegistryModelsUrl(
authHeader: "Authorization",
authPrefix: "Bearer ",
headers: { "Content-Type": "application/json" },
parseResponse: (data) => parseRegistryModelsResponse(provider, data),
parseResponse: parseRegistryModelsResponse,
};
}
return undefined;

View File

@@ -5897,7 +5897,6 @@
"friendliai": "Free tier for serverless inference — no credit card required",
"gemini": "Free forever: 1,500 req/day for Gemini 2.5 Flash — no credit card, get key at aistudio.google.com",
"gigachat": "Connect GigaChat (Sber) with an API key.",
"github-models": "Create a GitHub PAT with 'models: read' scope at github.com/settings/tokens",
"gitlab": "GitLab personal access token for the public Code Suggestions API. Configure a self-hosted base URL when not using gitlab.com.",
"gitlawb-gmi": "Get your API key from Gitlawb Opengateway dashboard.",
"gitlawb": "Get your API key from Gitlawb Opengateway dashboard.",

View File

@@ -22,9 +22,7 @@ import { getTrustedLocalRateLimitError } from "@omniroute/open-sse/services/rate
const INTERNAL_ORIGIN = "http://omniroute.internal";
export const DEFAULT_MODEL_TEST_TIMEOUT_MS = 30_000;
const DOLA_PRO_TEST_TIMEOUT_MS = 90_000;
const GITHUB_PHI_REASONING_TEST_TIMEOUT_MS = 60_000;
const DOUBAO_WEB_PROVIDER_ID = "doubao-web";
const GITHUB_MODELS_PROVIDER_ID = "github-models";
const SLOW_WEB_TEST_MODELS = new Set(["dola-pro"]);
const STREAMING_CHAT_TEST_MAX_TOKENS = 64;
@@ -105,10 +103,6 @@ export function resolveModelTestTimeoutMs(
return Math.max(requestedTimeoutMs, DOLA_PRO_TEST_TIMEOUT_MS);
}
if (normalizedProviderId === GITHUB_MODELS_PROVIDER_ID && modelLeafId === "phi-4-reasoning") {
return Math.max(requestedTimeoutMs, GITHUB_PHI_REASONING_TEST_TIMEOUT_MS);
}
return requestedTimeoutMs;
}

View File

@@ -137,7 +137,6 @@ export const MODELS_DEV_PROVIDER_MAP: Record<string, string[]> = {
// OAuth / special providers
bedrock: ["kiro", "kr"], // kr = Kiro (AWS Bedrock)
"github-copilot": ["github", "gh"],
"github-models": ["github", "gh"],
kilo: ["kilocode", "kc", "kilo-gateway"],
kilocode: ["kilocode", "kc", "kilo-gateway"],
"kimi-for-coding": ["kimi-coding", "kmc", "kimi-coding-apikey", "kmca"],

View File

@@ -369,7 +369,6 @@ const LOBE_PROVIDER_ALIASES = {
"gemini-web": "Gemini",
"gemini-business": "Gemini",
github: "GithubCopilot",
"github-models": "Github",
"github-copilot": "GithubCopilot",
"ghe-copilot": "GithubCopilot",
glm: "Zhipu",

View File

@@ -146,18 +146,6 @@ export const APIKEY_PROVIDERS_INFERENCE = {
hasFree: true,
freeNote: "Free Inference API for thousands of models (Whisper, VITS, SDXL…)",
},
"github-models": {
id: "github-models",
alias: "ghm",
name: "GitHub Models",
icon: "code",
color: "#238636",
textIcon: "GH",
website: "https://github.com/marketplace/models",
hasFree: true,
freeNote: "Free GPT-5, o-series, DeepSeek-R1, Llama 4, Grok 3 — GitHub account only.",
authHint: "Create a GitHub PAT with 'models: read' scope at github.com/settings/tokens",
},
deepinfra: {
id: "deepinfra",
alias: "deepinfra",

View File

@@ -2162,33 +2162,6 @@
"stream": "https://api.githubcopilot.com/chat/completions"
}
},
"github-models": {
"format": "openai",
"headers": {
"apiKey": {
"Accept": "text/event-stream",
"Authorization": "Bearer <TOK>",
"Content-Type": "application/json",
"X-GitHub-Api-Version": "2026-03-10"
},
"nonStream": {
"Accept": "application/vnd.github+json",
"Authorization": "Bearer <TOK>",
"Content-Type": "application/json",
"X-GitHub-Api-Version": "2026-03-10"
},
"oauth": {
"Accept": "text/event-stream",
"Authorization": "Bearer <TOK>",
"Content-Type": "application/json",
"X-GitHub-Api-Version": "2026-03-10"
}
},
"url": {
"nonStream": "https://models.github.ai/inference/chat/completions",
"stream": "https://models.github.ai/inference/chat/completions"
}
},
"gitlab-duo": {
"format": "openai",
"headers": {

View File

@@ -7,7 +7,7 @@ import {
} from "../../open-sse/config/freeTierCatalog.ts";
test("FREE_TIER_BUDGETS holds positive integer monthly-token budgets", () => {
assert.ok(Object.keys(FREE_TIER_BUDGETS).length >= 20);
assert.ok(Object.keys(FREE_TIER_BUDGETS).length >= 19);
for (const [id, tokens] of Object.entries(FREE_TIER_BUDGETS)) {
assert.ok(Number.isInteger(tokens) && tokens > 0, `${id} must be a positive integer`);
}
@@ -27,7 +27,7 @@ test("FREE_TIER_TOS marks proxy-prohibited providers as avoid", () => {
test("computeFreeTierTotals sums the documented budgets", () => {
const t = computeFreeTierTotals();
assert.equal(t.providerCount, 20);
assert.equal(t.providerCount, 19);
assert.ok(t.documentedMonthlyTokens >= 1_350_000_000);
assert.ok(t.documentedMonthlyTokens <= 1_450_000_000);
assert.equal(typeof t.headline, "string");
@@ -38,5 +38,5 @@ test("computeFreeTierTotals can exclude ToS-avoid providers", () => {
const all = computeFreeTierTotals();
const clean = computeFreeTierTotals({ excludeTosAvoid: true });
assert.equal(all.documentedMonthlyTokens - clean.documentedMonthlyTokens, 25_000);
assert.equal(clean.providerCount, 19);
assert.equal(clean.providerCount, 18);
});

View File

@@ -5,8 +5,7 @@ import { getRegistryEntry } from "../../open-sse/config/providerRegistry.ts";
// Regression guard: the GitHub Copilot (`gh`) provider must expose `gpt-4o-mini`
// alongside the curated date-pinned GPT-4o model. Copilot serves the cheaper mini variant via
// chat/completions, so apps that hard-code `gpt-4o-mini` should resolve to the
// Copilot provider — not only to the separate github-models (`ghm`) marketplace
// entry, which lists it under the `openai/` prefix (`openai/gpt-4o-mini`).
// Copilot provider.
//
// Ported from upstream decolua/9router (add GPT-4o mini to GitHub Copilot).
@@ -27,20 +26,3 @@ test("github (Copilot) provider exposes gpt-4o-mini next to curated GPT-4o", ()
assert.equal(mini?.contextLength, 128000);
assert.equal(ids.includes("gpt-4o"), false, "bare gpt-4o is not in the curated list");
});
test("Copilot gpt-4o-mini is distinct from the github-models openai/gpt-4o-mini", () => {
const copilot = getRegistryEntry("github");
const marketplace = getRegistryEntry("github-models");
const copilotIds = (copilot?.models ?? []).map((m) => m.id);
const marketplaceIds = (marketplace?.models ?? []).map((m) => m.id);
// The two providers reference the same upstream model under different ids:
// Copilot uses the bare `gpt-4o-mini`; the marketplace uses `openai/gpt-4o-mini`.
assert.ok(copilotIds.includes("gpt-4o-mini"));
assert.ok(marketplaceIds.includes("openai/gpt-4o-mini"));
assert.ok(
!copilotIds.includes("openai/gpt-4o-mini"),
"Copilot should not carry the marketplace-prefixed id"
);
});

View File

@@ -1,116 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import { FREE_MODEL_BUDGETS } from "../../open-sse/config/freeModelCatalog.data.ts";
import { getEmbeddingProvider } from "../../open-sse/config/embeddingRegistry.ts";
import { REGISTRY } from "../../open-sse/config/providerRegistry.ts";
import { deriveConfigFromRegistryModelsUrl } from "../../src/app/api/providers/[id]/models/discoveryConfig.ts";
import { getStaticModelsForProvider } from "../../src/lib/providers/staticModels.ts";
const EXPECTED_CHAT_IDS = [
"cohere/cohere-command-a",
"deepseek/deepseek-r1-0528",
"deepseek/deepseek-v3-0324",
"meta/llama-4-maverick-17b-128e-instruct-fp8",
"meta/llama-3.3-70b-instruct",
"meta/llama-4-scout-17b-16e-instruct",
"microsoft/phi-4-multimodal-instruct",
"microsoft/phi-4-reasoning",
"mistral-ai/codestral-2501",
"mistral-ai/mistral-medium-2505",
"openai/gpt-4.1",
"openai/gpt-4.1-mini",
"openai/gpt-4o",
"openai/gpt-4o-mini",
"openai/gpt-5",
"openai/gpt-5-chat",
"openai/gpt-5-mini",
"openai/o3",
"openai/o4-mini",
] as const;
const EXPECTED_EMBEDDING_IDS = [
"openai/text-embedding-3-large",
"openai/text-embedding-3-small",
] as const;
test("github-models curates the exact chat roster and preserves live discovery", () => {
const provider = REGISTRY["github-models"];
const ids = provider.models.map((model) => model.id);
assert.deepEqual(ids, EXPECTED_CHAT_IDS);
assert.equal(provider.modelsUrl, "https://models.github.ai/catalog/models");
assert.ok(!ids.includes("xai/grok-3"));
assert.ok(!ids.includes("openai/o1"));
assert.ok(!ids.includes("meta/meta-llama-3.1-405b-instruct"));
assert.ok(!ids.includes("meta/llama-3.2-11b-vision-instruct"));
assert.ok(!ids.includes("meta/llama-3.2-90b-vision-instruct"));
const scout = provider.models.find((model) => model.id === "meta/llama-4-scout-17b-16e-instruct");
assert.deepEqual(scout, {
id: "meta/llama-4-scout-17b-16e-instruct",
name: "Llama 4 Scout 17B 16E Instruct",
contextLength: 10_000_000,
maxInputTokens: 10_000_000,
maxOutputTokens: 4_096,
supportsVision: true,
toolCalling: true,
});
const gpt5 = provider.models.find((model) => model.id === "openai/gpt-5");
assert.equal(gpt5?.contextLength, 200_000);
assert.equal(gpt5?.maxInputTokens, 200_000);
assert.equal(gpt5?.maxOutputTokens, 100_000);
assert.equal(gpt5?.supportsVision, true);
assert.equal(gpt5?.supportsReasoning, true);
assert.equal(gpt5?.toolCalling, true);
});
test("github-models registers embedding-only models outside the chat roster", () => {
const chatIds = REGISTRY["github-models"].models.map((model) => model.id);
const provider = getEmbeddingProvider("github-models");
assert.ok(provider);
assert.equal(provider.baseUrl, "https://models.github.ai/inference/embeddings");
assert.equal(provider.authType, "apikey");
assert.equal(provider.authHeader, "bearer");
assert.deepEqual(
provider.models.map((model) => model.id),
EXPECTED_EMBEDDING_IDS
);
for (const id of EXPECTED_EMBEDDING_IDS) assert.ok(!chatIds.includes(id));
const specialty = getStaticModelsForProvider("github-models") || [];
assert.deepEqual(
specialty.map((model) => ({
id: model.id,
apiFormat: model.apiFormat,
supportedEndpoints: model.supportedEndpoints,
})),
EXPECTED_EMBEDDING_IDS.map((id) => ({
id,
apiFormat: "embeddings",
supportedEndpoints: ["embeddings"],
}))
);
});
test("github-models free metadata matches the combined curated roster", () => {
const freeIds = FREE_MODEL_BUDGETS.filter((entry) => entry.provider === "github-models").map(
(entry) => entry.modelId
);
assert.deepEqual(freeIds, [...EXPECTED_CHAT_IDS, ...EXPECTED_EMBEDDING_IDS]);
assert.equal(new Set(freeIds).size, 21);
});
test("github-models live discovery filters confirmed dead Llama 3.2 vision models", () => {
const config = deriveConfigFromRegistryModelsUrl("github-models");
assert.ok(config);
assert.equal(config.url, "https://models.github.ai/catalog/models");
assert.deepEqual(
config.parseResponse([
{ id: "meta/llama-3.2-11b-vision-instruct", name: "dead-11b" },
{ id: "meta/llama-3.2-90b-vision-instruct", name: "dead-90b" },
{ id: "meta/llama-3.3-70b-instruct", name: "live" },
]),
[{ id: "meta/llama-3.3-70b-instruct", name: "live" }]
);
});

View File

@@ -1,52 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import { DefaultExecutor } from "../../open-sse/executors/default.ts";
test("github-models uses max_completion_tokens for namespaced recent OpenAI models", () => {
const executor = new DefaultExecutor("github-models");
for (const model of [
"openai/gpt-5",
"openai/gpt-5-chat",
"openai/gpt-5-mini",
"openai/o4-mini",
]) {
const transformed = executor.transformRequest(
model,
{
model,
messages: [{ role: "user", content: "hi" }],
max_tokens: 2_048,
stream: false,
},
false,
{ providerSpecificData: {} }
) as Record<string, unknown>;
assert.equal(transformed.max_tokens, undefined, `${model} must not send max_tokens`);
assert.equal(
transformed.max_completion_tokens,
2_048,
`${model} must send max_completion_tokens`
);
}
});
test("github-models leaves legacy namespaced OpenAI models on max_tokens", () => {
const executor = new DefaultExecutor("github-models");
const transformed = executor.transformRequest(
"openai/gpt-4.1",
{
model: "openai/gpt-4.1",
messages: [{ role: "user", content: "hi" }],
max_tokens: 2_048,
stream: false,
},
false,
{ providerSpecificData: {} }
) as Record<string, unknown>;
assert.equal(transformed.max_tokens, 2_048);
assert.equal(transformed.max_completion_tokens, undefined);
});

View File

@@ -293,22 +293,7 @@ test("runSingleModelTest preserves slow timeout after chatCore converts AbortErr
});
test("resolveModelTestTimeoutMs defaults ordinary model checks to 30 seconds", () => {
assert.equal(resolveModelTestTimeoutMs("github-models", "openai/gpt-4.1"), 30_000);
});
test("resolveModelTestTimeoutMs gives GitHub Phi-4 Reasoning up to 60 seconds", () => {
assert.equal(
resolveModelTestTimeoutMs("github-models", "microsoft/phi-4-reasoning", 30_000),
60_000
);
assert.equal(
resolveModelTestTimeoutMs("github-models", "github-models/microsoft/phi-4-reasoning", 30_000),
60_000
);
assert.equal(
resolveModelTestTimeoutMs("github-models", "microsoft/phi-4-reasoning", 90_000),
90_000
);
assert.equal(resolveModelTestTimeoutMs("openai", "gpt-4.1"), 30_000);
});
// ---------------------------------------------------------------------------

View File

@@ -7,7 +7,7 @@ import {
} from "../../open-sse/config/embeddingRegistry.ts";
import { getImageProvider, parseImageModel } from "../../open-sse/config/imageRegistry.ts";
describe("OpenRouter & GitHub registry entries (#960)", () => {
describe("OpenRouter registry entries (#960)", () => {
// ── Embedding Registry ──────────────────────────────────────────────────
describe("embeddingRegistry — openrouter", () => {
@@ -36,26 +36,6 @@ describe("OpenRouter & GitHub registry entries (#960)", () => {
});
});
describe("embeddingRegistry — github", () => {
it("resolves github provider config", () => {
const p = getEmbeddingProvider("github");
assert.ok(p, "github should be in EMBEDDING_PROVIDERS");
assert.equal(p.baseUrl, "https://models.inference.ai.azure.com/embeddings");
assert.equal(p.authHeader, "bearer");
});
it("github has at least 2 models", () => {
const p = getEmbeddingProvider("github");
assert.ok(p.models.length >= 2, `Expected ≥2 models, got ${p.models.length}`);
});
it("parses github/text-embedding-3-small correctly", () => {
const result = parseEmbeddingModel("github/text-embedding-3-small");
assert.equal(result.provider, "github");
assert.equal(result.model, "text-embedding-3-small");
});
});
// ── Image Registry ───────────────────────────────────────────────────────
describe("imageRegistry — openrouter", () => {

View File

@@ -17,7 +17,8 @@
// sarvam+plamo in regional) to 193, then #8170 (inception/typhoon — inception in frontier-labs,
// typhoon in regional) to 195, then Firecrawl dual search+fetch under SEARCH_PROVIDERS.firecrawl
// (removed specialty-media duplicate) to 194, #8861 (Xiaomi MiMo Token Plan, regional) to 195, and
// the Cheaper Inference gateway (OSS-sponsor reseller, gateways family) to 198 (UnoRouter #9009, Raycast Pro #8895).
// the Cheaper Inference gateway (OSS-sponsor reseller, gateways family) to 198 (UnoRouter #9009,
// Raycast Pro #8895), then later additions to 199; retiring GitHub Models brings it to 198.
import { test } from "node:test";
import assert from "node:assert/strict";
@@ -46,12 +47,12 @@ test("barrel still exports every catalog + key helpers", () => {
}
});
test("APIKEY_PROVIDERS merges the 6 family files into 199 entries (no loss / no dup)", async () => {
test("APIKEY_PROVIDERS merges the 6 family files into 198 entries (no loss / no dup)", async () => {
const keys = Object.keys((P as Record<string, object>).APIKEY_PROVIDERS);
assert.equal(keys.length, 199);
assert.equal(new Set(keys).size, 199, "duplicate keys after spread-merge");
assert.equal(keys.length, 198);
assert.equal(new Set(keys).size, 198, "duplicate keys after spread-merge");
// the merged object's entry-count equals the sum of the 6 semantic family files; families are a
// strict partition (every provider in exactly one), so the sum must be exactly 199.
// strict partition (every provider in exactly one), so the sum must be exactly 198.
const families: [string, string][] = [
["gateways", "APIKEY_PROVIDERS_GATEWAYS"],
["frontier-labs", "APIKEY_PROVIDERS_FRONTIER"],
@@ -71,7 +72,7 @@ test("APIKEY_PROVIDERS merges the 6 family files into 199 entries (no loss / no
seen.add(k);
}
}
assert.equal(famTotal, 199, "families must partition all 199 providers");
assert.equal(famTotal, 198, "families must partition all 198 providers");
});
test("AI_PROVIDERS Proxy aggregates all sections; lookups resolve", () => {

View File

@@ -15,8 +15,6 @@ const embeddingsRoute = await import("../../src/app/api/v1/embeddings/route.ts")
const videoRoute = await import("../../src/app/api/v1/videos/generations/route.ts");
const musicRoute = await import("../../src/app/api/v1/music/generations/route.ts");
const v1ModelsCatalog = await import("../../src/app/api/v1/models/catalog.ts");
const providerModelsRoute =
await import("../../src/app/api/v1/providers/[provider]/models/route.ts");
async function resetStorage() {
core.resetDbInstance();
@@ -98,34 +96,6 @@ test("embedding catalog GET hides providers without active credentials", async (
assert.ok(!ids.includes("openai/text-embedding-3-small"));
});
test("provider-scoped GitHub Models catalog preserves nested embedding IDs", async () => {
await seedConnection("github-models");
const response = await providerModelsRoute.GET(
new Request("http://localhost/v1/providers/github-models/models"),
{ params: Promise.resolve({ provider: "github-models" }) }
);
const body = (await response.json()) as {
data?: Array<{ id: string; root?: string; type?: string }>;
};
assert.equal(response.status, 200);
const embeddings = (body.data || []).filter((model) => model.type === "embedding");
assert.deepEqual(
embeddings.map((model) => ({ id: model.id, root: model.root })),
[
{
id: "openai/text-embedding-3-large",
root: "openai/text-embedding-3-large",
},
{
id: "openai/text-embedding-3-small",
root: "openai/text-embedding-3-small",
},
]
);
});
test("video catalog GET hides credential-backed providers without credentials", async () => {
await seedConnection("kie");