mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-17 20:52:15 +03:00
refactor(providers): remove retired GitHub Models (#9023)
This commit is contained in:
@@ -0,0 +1 @@
|
||||
- **refactor(providers):** removed the retired GitHub Models provider and its catalog, discovery, embedding, free-tier, UI, and documentation surfaces; upgrades now run a durable, idempotent purge of its stored credentials, usage state, structured configuration, and call-log artifacts while preserving GitHub Copilot and live members of mixed configurations ([#9023](https://github.com/diegosouzapw/OmniRoute/pull/9023))
|
||||
@@ -128,7 +128,7 @@ Content-Type: application/json
|
||||
}
|
||||
```
|
||||
|
||||
Available providers: Nebius, OpenAI, Mistral, Together AI, Fireworks, NVIDIA, **OpenRouter**, **GitHub Models**.
|
||||
Available providers: Nebius, OpenAI, Mistral, Together AI, Fireworks, NVIDIA, **OpenRouter**.
|
||||
|
||||
Registry models that advertise multimodal support also accept up to 32 provider-neutral structured
|
||||
items. Media item types are `text`, `image`, `audio`, `video`, and `document`. Their media `source`
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
title: "Free Tiers & Free-Token Budget"
|
||||
version: 3.8.40
|
||||
lastUpdated: 2026-06-28
|
||||
lastUpdated: 2026-07-31
|
||||
---
|
||||
|
||||
# Free Tiers & Free-Token Budget
|
||||
@@ -15,19 +15,19 @@ lastUpdated: 2026-06-28
|
||||
|
||||
| Metric | Tokens / month | Meaning |
|
||||
| ------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| **Documented recurring grant (steady)** | **~1.53B** | Free-tier **pools** (per-model catalog), each shared pool counted **once**. The live source behind `/api/free-tier/summary` and the dashboard's Free-Tier Budget page. **Use this number.** |
|
||||
| **+ first month with signup credits** | **~2.15B** | Steady + one-time signup credits (Together $25, Z.AI 20M, DeepSeek 5M, …), deduped per account. **First month only** — does not recur. |
|
||||
| **Documented recurring grant (steady)** | **~1.51B** | Free-tier **pools** (per-model catalog), each shared pool counted **once**. The live source behind `/api/free-tier/summary` and the dashboard's Free-Tier Budget page. **Use this number.** |
|
||||
| **+ first month with signup credits** | **~2.13B** | Steady + one-time signup credits (Together $25, Z.AI 20M, DeepSeek 5M, …), deduped per account. **First month only** — does not recur. |
|
||||
| **+ permanently free, no published cap** | _un-quantifiable_ | `siliconflow`, `glm-cn` (GLM-4-Flash), `tencent`, `baidu`, `kilo-gateway`, `opencode-zen` — real recurring access, rate/concurrency-limited, **no token cap to count**. Listed, never summed (counting them at `RPM×24/7` is the inflation we reject). |
|
||||
| **+ deposit-unlock boost** | **+~24M** | A one-time **$10** OpenRouter top-up raises its free pool from 50 → 1000 req/day. Reported separately so it never inflates the steady number. |
|
||||
| Theoretical ceiling (all rate limits, 24/7) | ~10B | Sum of every provider rate limit extrapolated to non-stop use. **Not a guarantee** — do not headline this. |
|
||||
|
||||
**Honest headline:** _OmniRoute aggregates **~1.53B documented free tokens per month** (up to ~2.15B in your first month with signup credits) across 43 free-tier pools — plus a long tail of permanently-free, no-cap providers — and RTK + Caveman compression (15–95% token savings) stretches that further._
|
||||
**Honest headline:** _OmniRoute aggregates **~1.51B documented free tokens per month** (up to ~2.13B in your first month with signup credits) across 42 free-tier pools — plus a long tail of permanently-free, no-cap providers — and RTK + Caveman compression (15–95% token savings) stretches that further._
|
||||
|
||||
> **Why this dropped from the previous ~1.94B.** The 2026-06-17 refresh is an honesty correction, not a loss: `gemini` is now pool-deduped (was inflated by counting each Flash variant separately, 462M → 60M), `cloudflare-ai` corrected to its real 10k-Neurons/day (122M → 30M), `doubao` reclassified as a one-time signup credit (not recurring), and shut-down tiers removed (`github-models` closed to new signups, `chutes`/`phind`/`kluster` discontinued). Partly offset by `llm7` (correct 5M/day → 150M) and new free providers (Kilo, OpenCode Zen, Z.AI GLM-Flash).
|
||||
> **Why this dropped from the previous ~1.94B.** The 2026-06-17 refresh is an honesty correction, not a loss: `gemini` is now pool-deduped (was inflated by counting each Flash variant separately, 462M → 60M), `cloudflare-ai` corrected to its real 10k-Neurons/day (122M → 30M), `doubao` reclassified as a one-time signup credit (not recurring), and shut-down tiers removed (`chutes`/`phind`/`kluster` discontinued). Partly offset by `llm7` (correct 5M/day → 150M) and new free providers (Kilo, OpenCode Zen, Z.AI GLM-Flash).
|
||||
>
|
||||
> **Further corrected to ~1.37B in v3.8.42:** `longcat` was reclassified from a 150M/mo recurring grant to a one-time 10M signup credit after its free preview ended. Same honesty rule — no provider was dropped by mistake.
|
||||
>
|
||||
> **Updated to ~1.53B in v3.8.49:** the pool count grew from 39 to 43 after mapping free tiers that were documented upstream but missing from the catalog (`requesty`, `ovhcloud`, `agnes`, `glm`) plus new providers `navy` and `aihorde` (#7840). This is the live, CI-gated number (`check:docs-counts` fails the build if this drifts from `computeFreeModelTotals()`).
|
||||
> **Updated to ~1.51B after removing a retired provider:** the pool count is now 42 after mapping free tiers that were documented upstream but missing from the catalog (`requesty`, `ovhcloud`, `agnes`, `glm`) plus new providers `navy` and `aihorde` (#7840). This is the live, CI-gated number (`check:docs-counts` fails the build if this drifts from `computeFreeModelTotals()`).
|
||||
|
||||
Biggest **documented** contributors: `mistral` 1.00B, `llm7` 150M, `groq` 117M, `gemini` 60M, `cerebras` 30M, `cloudflare-ai` 30M, `sambanova` 30M. (`longcat` is excluded — its 10M LongCat-2.0 grant is a one-time, KYC-gated signup credit, not a recurring monthly budget.)
|
||||
|
||||
@@ -40,7 +40,6 @@ Biggest **documented** contributors: `mistral` 1.00B, `llm7` 150M, `groq` 117M,
|
||||
A 50-agent web-research pass (official docs + last-7-days news, adversarially verified) refreshed the whole catalog. Highlights:
|
||||
|
||||
- **Removed / no free tier (2026):** `chutes` (free tier ended 2026-03), `phind` (company shut down 2026-01), `kluster` (sunset 2026-06-09 → MITO), `gitlawb` + `gitlawb-gmi` (MiMo free revoked 2026-05-24, Nemotron promo ended 2026-06 — re-verified 2026-06-18), `aimlapi` (free tier paused — re-verified 2026-06-18), `yi` (Yi-Light retired, pay-as-you-go — re-verified 2026-06-18), `theoldllm` / `featherless-ai` (no current free tier). `iflytek` / `sparkdesk` stay listed but carry a ToS-caution note (Spark Lite is free; the ToS restricts proxy/relay use).
|
||||
- **GitHub Models** — closed to **new** customers on 2026-06-16; existing accounts keep API/playground access, so it stays in the catalog with a note (not removed).
|
||||
- **Gemini** — `2.0 Flash` / `2.0 Flash-Lite` shut down 2026-06-01 and `2.5 Pro` left the free tier (2026-04); free tier is now **Flash-family only** (2.5/3/3.1/3.5 Flash + Gemma). The catalog now **pools** the Flash family (was inflated by counting each variant separately: 462M → 60M).
|
||||
- **Corrected numbers:** `cloudflare-ai` 122M → **30M** (real 10k-Neurons/day), `doubao` reclassified as a one-time signup credit (not recurring), `llm7` 4M → **150M** (documented 5M tokens/day), `together` "-Free" endpoints discontinued → only the **$25** signup credit remains, `longcat` Preview ended + Flash models retired → **LongCat-2.0** only, reclassified as a one-time **10M**-token signup credit (KYC-gated, not recurring).
|
||||
- **New free providers discovered:** ⭐ **Kilo Code** (`kilo-gateway` — rotating "Auto Free" set: NVIDIA Nemotron 3 family, StepFun, Poolside, Nex-N2-Pro), ⭐ **OpenCode Zen** (`opencode-zen` — 6 rotating free coding models), ⭐ **Z.AI / Zhipu** (`glm-cn` — GLM-4-Flash / 4.5-Flash / 4.7-Flash permanently free + 20M signup bonus), and `arcee-ai` Trinity Large Preview.
|
||||
@@ -118,7 +117,6 @@ A 50-agent web-research pass (official docs + last-7-days news, adversarially ve
|
||||
| `exa-search` | caution | No explicit "no proxy" or "evaluation only" clauses found; Exa actively offers a reseller partner program allowing API … |
|
||||
| `firecrawl` | caution | Cloud API ToS has no explicit personal-proxy prohibition found, but the open-source self-hosted version is AGPL-3.0 (re… |
|
||||
| `gemini` | caution | ToS explicitly states the free tier is for "developers building with Google AI models for professional or business purp… |
|
||||
| `github-models` | caution | GitHub's Acceptable Use Policy prohibits reselling/proxying the service; GitHub Models ToS delegates to each model's ho… |
|
||||
| `groq` | caution | Services Agreement §6.3 prohibits reselling, sublicensing, or distributing API access; §3.2 bars reselling/leasing acco… |
|
||||
| `hackclub` | caution | Service is explicitly scoped to Hack Club teen members building projects/learning; no public ToS found explicitly permi… |
|
||||
| `huggingchat` | caution | Hugging Face ToS does not explicitly ban personal self-hosted proxies, but supplemental terms (referenced but not fully… |
|
||||
@@ -183,7 +181,6 @@ A 50-agent web-research pass (official docs + last-7-days news, adversarially ve
|
||||
| `cloudflare-ai` | recurring | ~30M | — | caution | 6 |
|
||||
| `api-airforce` | recurring | ~24M | — | caution | 7 |
|
||||
| `ollama-cloud` | recurring | ~20M | — | ambiguous | 8 |
|
||||
| `github-models` | recurring | ~18M | — | caution | 14 |
|
||||
| `groq` | recurring | ~15M | — | caution | 5 |
|
||||
| `bluesminds` | recurring | ~7M | — | ambiguous | 22 |
|
||||
| `sambanova` | recurring | ~6M | — | caution | 5 |
|
||||
@@ -280,7 +277,6 @@ A 50-agent web-research pass (official docs + last-7-days news, adversarially ve
|
||||
- **`freemodel-dev`** — Our shipped freeNote is "(none)" — this was likely a placeholder meaning the provider was not yet cataloged. In reality the provider does have a $300 one-time trial credit offer. However, this is a o…
|
||||
- **`friendliai`** — The shipped freeNote ("Free tier for serverless inference") is partially accurate but misleading. There is free access via Tier 0 and free-designated models, but the rate limits are undefined and ada…
|
||||
- **`gemini`** — The shipped freeNote says "1,500 req/day for Gemini 2.5 Flash" — this was accurate before December 2025. Google cut free-tier limits by 50-80% in December 2025, reducing Gemini 2.5 Flash from 1,500 R…
|
||||
- **`github-models`** — Catalog note "Free GPT-5, o-series, DeepSeek-R1, Llama 4, Grok 3" is directionally correct about model availability but omits the daily rate limits (50 RPD for high-tier models, 150 RPD for low-tier)…
|
||||
- **`gitlawb`** — The shipped freeNote "Free tier available" is effectively stale. The original free MiMo access was removed in May 2026; the only remaining "free" option is a temporary promotional model (Nemotron 3 U…
|
||||
- **`gitlawb-gmi`** — Partially still accurate — free tier exists but is now narrowed to a single model (Nemotron 3 Ultra) after MiMo free access was revoked in late May 2026. The shipped note "Free tier available" unders…
|
||||
- **`groq`** — The shipped freeNote "30 RPM / 14.4K RPD" is accurate only for llama-3.1-8b-instant. Most other models (including llama-3.3-70b-versatile) have a much lower 1K RPD cap. The note omits model-specific …
|
||||
|
||||
@@ -287,36 +287,6 @@ export const EMBEDDING_PROVIDERS: Record<string, EmbeddingProvider> = {
|
||||
],
|
||||
},
|
||||
|
||||
"github-models": {
|
||||
id: "github-models",
|
||||
baseUrl: "https://models.github.ai/inference/embeddings",
|
||||
authType: "apikey",
|
||||
authHeader: "bearer",
|
||||
models: [
|
||||
{
|
||||
id: "openai/text-embedding-3-large",
|
||||
name: "OpenAI Text Embedding 3 (large)",
|
||||
dimensions: 3_072,
|
||||
},
|
||||
{
|
||||
id: "openai/text-embedding-3-small",
|
||||
name: "OpenAI Text Embedding 3 (small)",
|
||||
dimensions: 1_536,
|
||||
},
|
||||
],
|
||||
},
|
||||
|
||||
github: {
|
||||
id: "github",
|
||||
baseUrl: "https://models.inference.ai.azure.com/embeddings",
|
||||
authType: "apikey",
|
||||
authHeader: "bearer",
|
||||
models: [
|
||||
{ id: "text-embedding-3-small", name: "Text Embedding 3 Small (GitHub)", dimensions: 1536 },
|
||||
{ id: "text-embedding-3-large", name: "Text Embedding 3 Large (GitHub)", dimensions: 3072 },
|
||||
],
|
||||
},
|
||||
|
||||
"jina-ai": {
|
||||
id: "jina-ai",
|
||||
structuredInputProtocol: "jina-v1",
|
||||
|
||||
@@ -188,27 +188,6 @@ export const FREE_MODEL_BUDGETS: FreeModelBudget[] = [
|
||||
{ provider: "gemini", modelId: "gemini-3-flash-preview", displayName: "Gemini 3 Flash Preview", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-daily", poolKey: "gemini-free", tos: "caution" },
|
||||
{ provider: "gemini", modelId: "gemini-3.1-flash-lite", displayName: "Gemini 3.1 Flash-Lite", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-daily", poolKey: "gemini-free", tos: "caution" },
|
||||
{ provider: "gemini", modelId: "gemini-3.5-flash", displayName: "Gemini 3.5 Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-daily", poolKey: "gemini-free", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "cohere/cohere-command-a", displayName: "Cohere Command A (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "deepseek/deepseek-r1-0528", displayName: "DeepSeek-R1-0528 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "deepseek/deepseek-v3-0324", displayName: "DeepSeek-V3-0324 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "meta/llama-4-maverick-17b-128e-instruct-fp8", displayName: "Llama 4 Maverick 17B 128E Instruct FP8 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "meta/llama-3.3-70b-instruct", displayName: "Llama-3.3-70B-Instruct (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "meta/llama-4-scout-17b-16e-instruct", displayName: "Llama 4 Scout 17B 16E Instruct (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "microsoft/phi-4-multimodal-instruct", displayName: "Phi-4-multimodal-instruct (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "microsoft/phi-4-reasoning", displayName: "Phi-4-reasoning (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "mistral-ai/codestral-2501", displayName: "Codestral 25.01 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "mistral-ai/mistral-medium-2505", displayName: "Mistral Medium 3 (25.05) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-4.1", displayName: "OpenAI GPT-4.1 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-4.1-mini", displayName: "OpenAI GPT-4.1-mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-4o", displayName: "OpenAI GPT-4o (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-4o-mini", displayName: "OpenAI GPT-4o mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-5", displayName: "OpenAI gpt-5 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-5-chat", displayName: "OpenAI gpt-5-chat (preview) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/gpt-5-mini", displayName: "OpenAI gpt-5-mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/o3", displayName: "OpenAI o3 (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/o4-mini", displayName: "OpenAI o4-mini (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/text-embedding-3-large", displayName: "OpenAI Text Embedding 3 (large) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "github-models", modelId: "openai/text-embedding-3-small", displayName: "OpenAI Text Embedding 3 (small) (Free)", monthlyTokens: 18000000, creditTokens: 0, freeType: "recurring-daily", poolKey: "github-models", tos: "caution" },
|
||||
{ provider: "glm-cn", modelId: "glm-4-flash", displayName: "GLM-4-Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-uncapped", poolKey: "zhipu-flash-free", tos: "ok" },
|
||||
{ provider: "glm-cn", modelId: "glm-4.5-flash", displayName: "GLM-4.5-Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-uncapped", poolKey: "zhipu-flash-free", tos: "ok" },
|
||||
{ provider: "glm-cn", modelId: "glm-4.7-flash", displayName: "GLM-4.7-Flash", monthlyTokens: 0, creditTokens: 0, freeType: "recurring-uncapped", poolKey: "zhipu-flash-free", tos: "ok" },
|
||||
|
||||
@@ -19,7 +19,6 @@ export const FREE_TIER_BUDGETS: Record<string, number> = {
|
||||
cerebras: 30_000_000,
|
||||
"api-airforce": 24_000_000,
|
||||
"ollama-cloud": 20_000_000,
|
||||
"github-models": 18_000_000,
|
||||
groq: 15_000_000,
|
||||
bluesminds: 7_200_000,
|
||||
sambanova: 6_000_000,
|
||||
|
||||
@@ -27,7 +27,6 @@ import { raycastProvider } from "./registry/raycast/index.ts";
|
||||
import { muse_spark_webProvider } from "./registry/muse-spark-web/index.ts";
|
||||
import { lmarenaProvider } from "./registry/lmarena/index.ts";
|
||||
import { kilocodeProvider } from "./registry/kilocode/index.ts";
|
||||
import { github_modelsProvider } from "./registry/github/models/index.ts";
|
||||
import { githubProvider } from "./registry/github/index.ts";
|
||||
import { gheCopilotProvider } from "./registry/ghe-copilot/index.ts";
|
||||
import { difyProvider } from "./registry/dify/index.ts";
|
||||
@@ -256,7 +255,6 @@ export const REGISTRY: Record<string, RegistryEntry> = {
|
||||
"muse-spark-web": muse_spark_webProvider,
|
||||
lmarena: lmarenaProvider,
|
||||
kilocode: kilocodeProvider,
|
||||
"github-models": github_modelsProvider,
|
||||
github: githubProvider,
|
||||
"ghe-copilot": gheCopilotProvider,
|
||||
dify: difyProvider,
|
||||
|
||||
@@ -1,187 +0,0 @@
|
||||
import type { RegistryEntry } from "../../../shared.ts";
|
||||
|
||||
export const github_modelsProvider: RegistryEntry = {
|
||||
id: "github-models",
|
||||
alias: "ghm",
|
||||
format: "openai",
|
||||
executor: "default",
|
||||
baseUrl: "https://models.github.ai/inference/chat/completions",
|
||||
modelsUrl: "https://models.github.ai/catalog/models",
|
||||
authType: "apikey",
|
||||
authHeader: "Authorization",
|
||||
authPrefix: "Bearer",
|
||||
headers: {
|
||||
"X-GitHub-Api-Version": "2026-03-10",
|
||||
Accept: "application/vnd.github+json",
|
||||
},
|
||||
defaultContextLength: 128000,
|
||||
models: [
|
||||
{
|
||||
id: "cohere/cohere-command-a",
|
||||
name: "Cohere Command A",
|
||||
contextLength: 131_072,
|
||||
maxInputTokens: 131_072,
|
||||
maxOutputTokens: 4_096,
|
||||
},
|
||||
{
|
||||
id: "deepseek/deepseek-r1-0528",
|
||||
name: "DeepSeek-R1-0528",
|
||||
contextLength: 128_000,
|
||||
maxInputTokens: 128_000,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsReasoning: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "deepseek/deepseek-v3-0324",
|
||||
name: "DeepSeek-V3-0324",
|
||||
contextLength: 128_000,
|
||||
maxInputTokens: 128_000,
|
||||
maxOutputTokens: 4_096,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "meta/llama-4-maverick-17b-128e-instruct-fp8",
|
||||
name: "Llama 4 Maverick 17B 128E Instruct FP8",
|
||||
contextLength: 1_000_000,
|
||||
maxInputTokens: 1_000_000,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "meta/llama-3.3-70b-instruct",
|
||||
name: "Llama-3.3-70B-Instruct",
|
||||
contextLength: 128_000,
|
||||
maxInputTokens: 128_000,
|
||||
maxOutputTokens: 4_096,
|
||||
},
|
||||
{
|
||||
id: "meta/llama-4-scout-17b-16e-instruct",
|
||||
name: "Llama 4 Scout 17B 16E Instruct",
|
||||
contextLength: 10_000_000,
|
||||
maxInputTokens: 10_000_000,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "microsoft/phi-4-multimodal-instruct",
|
||||
name: "Phi-4-multimodal-instruct",
|
||||
contextLength: 128_000,
|
||||
maxInputTokens: 128_000,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsVision: true,
|
||||
},
|
||||
{
|
||||
id: "microsoft/phi-4-reasoning",
|
||||
name: "Phi-4-reasoning",
|
||||
contextLength: 32_768,
|
||||
maxInputTokens: 32_768,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsReasoning: true,
|
||||
},
|
||||
{
|
||||
id: "mistral-ai/codestral-2501",
|
||||
name: "Codestral 25.01",
|
||||
contextLength: 256_000,
|
||||
maxInputTokens: 256_000,
|
||||
maxOutputTokens: 4_096,
|
||||
},
|
||||
{
|
||||
id: "mistral-ai/mistral-medium-2505",
|
||||
name: "Mistral Medium 3 (25.05)",
|
||||
contextLength: 128_000,
|
||||
maxInputTokens: 128_000,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-4.1",
|
||||
name: "OpenAI GPT-4.1",
|
||||
contextLength: 1_048_576,
|
||||
maxInputTokens: 1_048_576,
|
||||
maxOutputTokens: 32_768,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-4.1-mini",
|
||||
name: "OpenAI GPT-4.1-mini",
|
||||
contextLength: 1_048_576,
|
||||
maxInputTokens: 1_048_576,
|
||||
maxOutputTokens: 32_768,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-4o",
|
||||
name: "OpenAI GPT-4o",
|
||||
contextLength: 131_072,
|
||||
maxInputTokens: 131_072,
|
||||
maxOutputTokens: 16_384,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-4o-mini",
|
||||
name: "OpenAI GPT-4o mini",
|
||||
contextLength: 131_072,
|
||||
maxInputTokens: 131_072,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-5",
|
||||
name: "OpenAI gpt-5",
|
||||
contextLength: 200_000,
|
||||
maxInputTokens: 200_000,
|
||||
maxOutputTokens: 100_000,
|
||||
supportsVision: true,
|
||||
supportsReasoning: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-5-chat",
|
||||
name: "OpenAI gpt-5-chat (preview)",
|
||||
contextLength: 200_000,
|
||||
maxInputTokens: 200_000,
|
||||
maxOutputTokens: 100_000,
|
||||
supportsVision: true,
|
||||
supportsReasoning: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/gpt-5-mini",
|
||||
name: "OpenAI gpt-5-mini",
|
||||
contextLength: 200_000,
|
||||
maxInputTokens: 200_000,
|
||||
maxOutputTokens: 100_000,
|
||||
supportsVision: true,
|
||||
supportsReasoning: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/o3",
|
||||
name: "OpenAI o3",
|
||||
contextLength: 200_000,
|
||||
maxInputTokens: 200_000,
|
||||
maxOutputTokens: 100_000,
|
||||
supportsVision: true,
|
||||
supportsReasoning: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
{
|
||||
id: "openai/o4-mini",
|
||||
name: "OpenAI o4-mini",
|
||||
contextLength: 200_000,
|
||||
maxInputTokens: 200_000,
|
||||
maxOutputTokens: 100_000,
|
||||
supportsVision: true,
|
||||
supportsReasoning: true,
|
||||
toolCalling: true,
|
||||
},
|
||||
],
|
||||
};
|
||||
@@ -1,14 +1,7 @@
|
||||
import { getRegistryEntry } from "@omniroute/open-sse/config/providerRegistry.ts";
|
||||
import type { ProviderModelsConfigEntry } from "./discovery/providerModelsConfig";
|
||||
|
||||
const DISCOVERY_EXCLUDED_MODEL_IDS: Readonly<Record<string, ReadonlySet<string>>> = {
|
||||
"github-models": new Set([
|
||||
"meta/llama-3.2-11b-vision-instruct",
|
||||
"meta/llama-3.2-90b-vision-instruct",
|
||||
]),
|
||||
};
|
||||
|
||||
function parseRegistryModelsResponse(provider: string, data: unknown): unknown[] {
|
||||
function parseRegistryModelsResponse(data: unknown): unknown[] {
|
||||
const response = data as { data?: unknown; models?: unknown } | null;
|
||||
const models = Array.isArray(data)
|
||||
? data
|
||||
@@ -17,15 +10,7 @@ function parseRegistryModelsResponse(provider: string, data: unknown): unknown[]
|
||||
: Array.isArray(response?.models)
|
||||
? response.models
|
||||
: [];
|
||||
const excludedIds = DISCOVERY_EXCLUDED_MODEL_IDS[provider];
|
||||
|
||||
if (!excludedIds) return models;
|
||||
|
||||
return models.filter((model) => {
|
||||
if (!model || typeof model !== "object" || !("id" in model)) return true;
|
||||
const id = (model as { id?: unknown }).id;
|
||||
return typeof id !== "string" || !excludedIds.has(id.toLowerCase());
|
||||
});
|
||||
return models;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -47,7 +32,7 @@ export function deriveConfigFromRegistryModelsUrl(
|
||||
authHeader: "Authorization",
|
||||
authPrefix: "Bearer ",
|
||||
headers: { "Content-Type": "application/json" },
|
||||
parseResponse: (data) => parseRegistryModelsResponse(provider, data),
|
||||
parseResponse: parseRegistryModelsResponse,
|
||||
};
|
||||
}
|
||||
return undefined;
|
||||
|
||||
@@ -5897,7 +5897,6 @@
|
||||
"friendliai": "Free tier for serverless inference — no credit card required",
|
||||
"gemini": "Free forever: 1,500 req/day for Gemini 2.5 Flash — no credit card, get key at aistudio.google.com",
|
||||
"gigachat": "Connect GigaChat (Sber) with an API key.",
|
||||
"github-models": "Create a GitHub PAT with 'models: read' scope at github.com/settings/tokens",
|
||||
"gitlab": "GitLab personal access token for the public Code Suggestions API. Configure a self-hosted base URL when not using gitlab.com.",
|
||||
"gitlawb-gmi": "Get your API key from Gitlawb Opengateway dashboard.",
|
||||
"gitlawb": "Get your API key from Gitlawb Opengateway dashboard.",
|
||||
|
||||
@@ -22,9 +22,7 @@ import { getTrustedLocalRateLimitError } from "@omniroute/open-sse/services/rate
|
||||
const INTERNAL_ORIGIN = "http://omniroute.internal";
|
||||
export const DEFAULT_MODEL_TEST_TIMEOUT_MS = 30_000;
|
||||
const DOLA_PRO_TEST_TIMEOUT_MS = 90_000;
|
||||
const GITHUB_PHI_REASONING_TEST_TIMEOUT_MS = 60_000;
|
||||
const DOUBAO_WEB_PROVIDER_ID = "doubao-web";
|
||||
const GITHUB_MODELS_PROVIDER_ID = "github-models";
|
||||
const SLOW_WEB_TEST_MODELS = new Set(["dola-pro"]);
|
||||
const STREAMING_CHAT_TEST_MAX_TOKENS = 64;
|
||||
|
||||
@@ -105,10 +103,6 @@ export function resolveModelTestTimeoutMs(
|
||||
return Math.max(requestedTimeoutMs, DOLA_PRO_TEST_TIMEOUT_MS);
|
||||
}
|
||||
|
||||
if (normalizedProviderId === GITHUB_MODELS_PROVIDER_ID && modelLeafId === "phi-4-reasoning") {
|
||||
return Math.max(requestedTimeoutMs, GITHUB_PHI_REASONING_TEST_TIMEOUT_MS);
|
||||
}
|
||||
|
||||
return requestedTimeoutMs;
|
||||
}
|
||||
|
||||
|
||||
@@ -137,7 +137,6 @@ export const MODELS_DEV_PROVIDER_MAP: Record<string, string[]> = {
|
||||
// OAuth / special providers
|
||||
bedrock: ["kiro", "kr"], // kr = Kiro (AWS Bedrock)
|
||||
"github-copilot": ["github", "gh"],
|
||||
"github-models": ["github", "gh"],
|
||||
kilo: ["kilocode", "kc", "kilo-gateway"],
|
||||
kilocode: ["kilocode", "kc", "kilo-gateway"],
|
||||
"kimi-for-coding": ["kimi-coding", "kmc", "kimi-coding-apikey", "kmca"],
|
||||
|
||||
@@ -369,7 +369,6 @@ const LOBE_PROVIDER_ALIASES = {
|
||||
"gemini-web": "Gemini",
|
||||
"gemini-business": "Gemini",
|
||||
github: "GithubCopilot",
|
||||
"github-models": "Github",
|
||||
"github-copilot": "GithubCopilot",
|
||||
"ghe-copilot": "GithubCopilot",
|
||||
glm: "Zhipu",
|
||||
|
||||
@@ -146,18 +146,6 @@ export const APIKEY_PROVIDERS_INFERENCE = {
|
||||
hasFree: true,
|
||||
freeNote: "Free Inference API for thousands of models (Whisper, VITS, SDXL…)",
|
||||
},
|
||||
"github-models": {
|
||||
id: "github-models",
|
||||
alias: "ghm",
|
||||
name: "GitHub Models",
|
||||
icon: "code",
|
||||
color: "#238636",
|
||||
textIcon: "GH",
|
||||
website: "https://github.com/marketplace/models",
|
||||
hasFree: true,
|
||||
freeNote: "Free GPT-5, o-series, DeepSeek-R1, Llama 4, Grok 3 — GitHub account only.",
|
||||
authHint: "Create a GitHub PAT with 'models: read' scope at github.com/settings/tokens",
|
||||
},
|
||||
deepinfra: {
|
||||
id: "deepinfra",
|
||||
alias: "deepinfra",
|
||||
|
||||
@@ -2162,33 +2162,6 @@
|
||||
"stream": "https://api.githubcopilot.com/chat/completions"
|
||||
}
|
||||
},
|
||||
"github-models": {
|
||||
"format": "openai",
|
||||
"headers": {
|
||||
"apiKey": {
|
||||
"Accept": "text/event-stream",
|
||||
"Authorization": "Bearer <TOK>",
|
||||
"Content-Type": "application/json",
|
||||
"X-GitHub-Api-Version": "2026-03-10"
|
||||
},
|
||||
"nonStream": {
|
||||
"Accept": "application/vnd.github+json",
|
||||
"Authorization": "Bearer <TOK>",
|
||||
"Content-Type": "application/json",
|
||||
"X-GitHub-Api-Version": "2026-03-10"
|
||||
},
|
||||
"oauth": {
|
||||
"Accept": "text/event-stream",
|
||||
"Authorization": "Bearer <TOK>",
|
||||
"Content-Type": "application/json",
|
||||
"X-GitHub-Api-Version": "2026-03-10"
|
||||
}
|
||||
},
|
||||
"url": {
|
||||
"nonStream": "https://models.github.ai/inference/chat/completions",
|
||||
"stream": "https://models.github.ai/inference/chat/completions"
|
||||
}
|
||||
},
|
||||
"gitlab-duo": {
|
||||
"format": "openai",
|
||||
"headers": {
|
||||
|
||||
@@ -7,7 +7,7 @@ import {
|
||||
} from "../../open-sse/config/freeTierCatalog.ts";
|
||||
|
||||
test("FREE_TIER_BUDGETS holds positive integer monthly-token budgets", () => {
|
||||
assert.ok(Object.keys(FREE_TIER_BUDGETS).length >= 20);
|
||||
assert.ok(Object.keys(FREE_TIER_BUDGETS).length >= 19);
|
||||
for (const [id, tokens] of Object.entries(FREE_TIER_BUDGETS)) {
|
||||
assert.ok(Number.isInteger(tokens) && tokens > 0, `${id} must be a positive integer`);
|
||||
}
|
||||
@@ -27,7 +27,7 @@ test("FREE_TIER_TOS marks proxy-prohibited providers as avoid", () => {
|
||||
|
||||
test("computeFreeTierTotals sums the documented budgets", () => {
|
||||
const t = computeFreeTierTotals();
|
||||
assert.equal(t.providerCount, 20);
|
||||
assert.equal(t.providerCount, 19);
|
||||
assert.ok(t.documentedMonthlyTokens >= 1_350_000_000);
|
||||
assert.ok(t.documentedMonthlyTokens <= 1_450_000_000);
|
||||
assert.equal(typeof t.headline, "string");
|
||||
@@ -38,5 +38,5 @@ test("computeFreeTierTotals can exclude ToS-avoid providers", () => {
|
||||
const all = computeFreeTierTotals();
|
||||
const clean = computeFreeTierTotals({ excludeTosAvoid: true });
|
||||
assert.equal(all.documentedMonthlyTokens - clean.documentedMonthlyTokens, 25_000);
|
||||
assert.equal(clean.providerCount, 19);
|
||||
assert.equal(clean.providerCount, 18);
|
||||
});
|
||||
|
||||
@@ -5,8 +5,7 @@ import { getRegistryEntry } from "../../open-sse/config/providerRegistry.ts";
|
||||
// Regression guard: the GitHub Copilot (`gh`) provider must expose `gpt-4o-mini`
|
||||
// alongside the curated date-pinned GPT-4o model. Copilot serves the cheaper mini variant via
|
||||
// chat/completions, so apps that hard-code `gpt-4o-mini` should resolve to the
|
||||
// Copilot provider — not only to the separate github-models (`ghm`) marketplace
|
||||
// entry, which lists it under the `openai/` prefix (`openai/gpt-4o-mini`).
|
||||
// Copilot provider.
|
||||
//
|
||||
// Ported from upstream decolua/9router (add GPT-4o mini to GitHub Copilot).
|
||||
|
||||
@@ -27,20 +26,3 @@ test("github (Copilot) provider exposes gpt-4o-mini next to curated GPT-4o", ()
|
||||
assert.equal(mini?.contextLength, 128000);
|
||||
assert.equal(ids.includes("gpt-4o"), false, "bare gpt-4o is not in the curated list");
|
||||
});
|
||||
|
||||
test("Copilot gpt-4o-mini is distinct from the github-models openai/gpt-4o-mini", () => {
|
||||
const copilot = getRegistryEntry("github");
|
||||
const marketplace = getRegistryEntry("github-models");
|
||||
|
||||
const copilotIds = (copilot?.models ?? []).map((m) => m.id);
|
||||
const marketplaceIds = (marketplace?.models ?? []).map((m) => m.id);
|
||||
|
||||
// The two providers reference the same upstream model under different ids:
|
||||
// Copilot uses the bare `gpt-4o-mini`; the marketplace uses `openai/gpt-4o-mini`.
|
||||
assert.ok(copilotIds.includes("gpt-4o-mini"));
|
||||
assert.ok(marketplaceIds.includes("openai/gpt-4o-mini"));
|
||||
assert.ok(
|
||||
!copilotIds.includes("openai/gpt-4o-mini"),
|
||||
"Copilot should not carry the marketplace-prefixed id"
|
||||
);
|
||||
});
|
||||
|
||||
@@ -1,116 +0,0 @@
|
||||
import assert from "node:assert/strict";
|
||||
import test from "node:test";
|
||||
|
||||
import { FREE_MODEL_BUDGETS } from "../../open-sse/config/freeModelCatalog.data.ts";
|
||||
import { getEmbeddingProvider } from "../../open-sse/config/embeddingRegistry.ts";
|
||||
import { REGISTRY } from "../../open-sse/config/providerRegistry.ts";
|
||||
import { deriveConfigFromRegistryModelsUrl } from "../../src/app/api/providers/[id]/models/discoveryConfig.ts";
|
||||
import { getStaticModelsForProvider } from "../../src/lib/providers/staticModels.ts";
|
||||
|
||||
const EXPECTED_CHAT_IDS = [
|
||||
"cohere/cohere-command-a",
|
||||
"deepseek/deepseek-r1-0528",
|
||||
"deepseek/deepseek-v3-0324",
|
||||
"meta/llama-4-maverick-17b-128e-instruct-fp8",
|
||||
"meta/llama-3.3-70b-instruct",
|
||||
"meta/llama-4-scout-17b-16e-instruct",
|
||||
"microsoft/phi-4-multimodal-instruct",
|
||||
"microsoft/phi-4-reasoning",
|
||||
"mistral-ai/codestral-2501",
|
||||
"mistral-ai/mistral-medium-2505",
|
||||
"openai/gpt-4.1",
|
||||
"openai/gpt-4.1-mini",
|
||||
"openai/gpt-4o",
|
||||
"openai/gpt-4o-mini",
|
||||
"openai/gpt-5",
|
||||
"openai/gpt-5-chat",
|
||||
"openai/gpt-5-mini",
|
||||
"openai/o3",
|
||||
"openai/o4-mini",
|
||||
] as const;
|
||||
|
||||
const EXPECTED_EMBEDDING_IDS = [
|
||||
"openai/text-embedding-3-large",
|
||||
"openai/text-embedding-3-small",
|
||||
] as const;
|
||||
|
||||
test("github-models curates the exact chat roster and preserves live discovery", () => {
|
||||
const provider = REGISTRY["github-models"];
|
||||
const ids = provider.models.map((model) => model.id);
|
||||
assert.deepEqual(ids, EXPECTED_CHAT_IDS);
|
||||
assert.equal(provider.modelsUrl, "https://models.github.ai/catalog/models");
|
||||
assert.ok(!ids.includes("xai/grok-3"));
|
||||
assert.ok(!ids.includes("openai/o1"));
|
||||
assert.ok(!ids.includes("meta/meta-llama-3.1-405b-instruct"));
|
||||
assert.ok(!ids.includes("meta/llama-3.2-11b-vision-instruct"));
|
||||
assert.ok(!ids.includes("meta/llama-3.2-90b-vision-instruct"));
|
||||
|
||||
const scout = provider.models.find((model) => model.id === "meta/llama-4-scout-17b-16e-instruct");
|
||||
assert.deepEqual(scout, {
|
||||
id: "meta/llama-4-scout-17b-16e-instruct",
|
||||
name: "Llama 4 Scout 17B 16E Instruct",
|
||||
contextLength: 10_000_000,
|
||||
maxInputTokens: 10_000_000,
|
||||
maxOutputTokens: 4_096,
|
||||
supportsVision: true,
|
||||
toolCalling: true,
|
||||
});
|
||||
|
||||
const gpt5 = provider.models.find((model) => model.id === "openai/gpt-5");
|
||||
assert.equal(gpt5?.contextLength, 200_000);
|
||||
assert.equal(gpt5?.maxInputTokens, 200_000);
|
||||
assert.equal(gpt5?.maxOutputTokens, 100_000);
|
||||
assert.equal(gpt5?.supportsVision, true);
|
||||
assert.equal(gpt5?.supportsReasoning, true);
|
||||
assert.equal(gpt5?.toolCalling, true);
|
||||
});
|
||||
|
||||
test("github-models registers embedding-only models outside the chat roster", () => {
|
||||
const chatIds = REGISTRY["github-models"].models.map((model) => model.id);
|
||||
const provider = getEmbeddingProvider("github-models");
|
||||
assert.ok(provider);
|
||||
assert.equal(provider.baseUrl, "https://models.github.ai/inference/embeddings");
|
||||
assert.equal(provider.authType, "apikey");
|
||||
assert.equal(provider.authHeader, "bearer");
|
||||
assert.deepEqual(
|
||||
provider.models.map((model) => model.id),
|
||||
EXPECTED_EMBEDDING_IDS
|
||||
);
|
||||
for (const id of EXPECTED_EMBEDDING_IDS) assert.ok(!chatIds.includes(id));
|
||||
|
||||
const specialty = getStaticModelsForProvider("github-models") || [];
|
||||
assert.deepEqual(
|
||||
specialty.map((model) => ({
|
||||
id: model.id,
|
||||
apiFormat: model.apiFormat,
|
||||
supportedEndpoints: model.supportedEndpoints,
|
||||
})),
|
||||
EXPECTED_EMBEDDING_IDS.map((id) => ({
|
||||
id,
|
||||
apiFormat: "embeddings",
|
||||
supportedEndpoints: ["embeddings"],
|
||||
}))
|
||||
);
|
||||
});
|
||||
|
||||
test("github-models free metadata matches the combined curated roster", () => {
|
||||
const freeIds = FREE_MODEL_BUDGETS.filter((entry) => entry.provider === "github-models").map(
|
||||
(entry) => entry.modelId
|
||||
);
|
||||
assert.deepEqual(freeIds, [...EXPECTED_CHAT_IDS, ...EXPECTED_EMBEDDING_IDS]);
|
||||
assert.equal(new Set(freeIds).size, 21);
|
||||
});
|
||||
|
||||
test("github-models live discovery filters confirmed dead Llama 3.2 vision models", () => {
|
||||
const config = deriveConfigFromRegistryModelsUrl("github-models");
|
||||
assert.ok(config);
|
||||
assert.equal(config.url, "https://models.github.ai/catalog/models");
|
||||
assert.deepEqual(
|
||||
config.parseResponse([
|
||||
{ id: "meta/llama-3.2-11b-vision-instruct", name: "dead-11b" },
|
||||
{ id: "meta/llama-3.2-90b-vision-instruct", name: "dead-90b" },
|
||||
{ id: "meta/llama-3.3-70b-instruct", name: "live" },
|
||||
]),
|
||||
[{ id: "meta/llama-3.3-70b-instruct", name: "live" }]
|
||||
);
|
||||
});
|
||||
@@ -1,52 +0,0 @@
|
||||
import assert from "node:assert/strict";
|
||||
import test from "node:test";
|
||||
|
||||
import { DefaultExecutor } from "../../open-sse/executors/default.ts";
|
||||
|
||||
test("github-models uses max_completion_tokens for namespaced recent OpenAI models", () => {
|
||||
const executor = new DefaultExecutor("github-models");
|
||||
|
||||
for (const model of [
|
||||
"openai/gpt-5",
|
||||
"openai/gpt-5-chat",
|
||||
"openai/gpt-5-mini",
|
||||
"openai/o4-mini",
|
||||
]) {
|
||||
const transformed = executor.transformRequest(
|
||||
model,
|
||||
{
|
||||
model,
|
||||
messages: [{ role: "user", content: "hi" }],
|
||||
max_tokens: 2_048,
|
||||
stream: false,
|
||||
},
|
||||
false,
|
||||
{ providerSpecificData: {} }
|
||||
) as Record<string, unknown>;
|
||||
|
||||
assert.equal(transformed.max_tokens, undefined, `${model} must not send max_tokens`);
|
||||
assert.equal(
|
||||
transformed.max_completion_tokens,
|
||||
2_048,
|
||||
`${model} must send max_completion_tokens`
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test("github-models leaves legacy namespaced OpenAI models on max_tokens", () => {
|
||||
const executor = new DefaultExecutor("github-models");
|
||||
const transformed = executor.transformRequest(
|
||||
"openai/gpt-4.1",
|
||||
{
|
||||
model: "openai/gpt-4.1",
|
||||
messages: [{ role: "user", content: "hi" }],
|
||||
max_tokens: 2_048,
|
||||
stream: false,
|
||||
},
|
||||
false,
|
||||
{ providerSpecificData: {} }
|
||||
) as Record<string, unknown>;
|
||||
|
||||
assert.equal(transformed.max_tokens, 2_048);
|
||||
assert.equal(transformed.max_completion_tokens, undefined);
|
||||
});
|
||||
@@ -293,22 +293,7 @@ test("runSingleModelTest preserves slow timeout after chatCore converts AbortErr
|
||||
});
|
||||
|
||||
test("resolveModelTestTimeoutMs defaults ordinary model checks to 30 seconds", () => {
|
||||
assert.equal(resolveModelTestTimeoutMs("github-models", "openai/gpt-4.1"), 30_000);
|
||||
});
|
||||
|
||||
test("resolveModelTestTimeoutMs gives GitHub Phi-4 Reasoning up to 60 seconds", () => {
|
||||
assert.equal(
|
||||
resolveModelTestTimeoutMs("github-models", "microsoft/phi-4-reasoning", 30_000),
|
||||
60_000
|
||||
);
|
||||
assert.equal(
|
||||
resolveModelTestTimeoutMs("github-models", "github-models/microsoft/phi-4-reasoning", 30_000),
|
||||
60_000
|
||||
);
|
||||
assert.equal(
|
||||
resolveModelTestTimeoutMs("github-models", "microsoft/phi-4-reasoning", 90_000),
|
||||
90_000
|
||||
);
|
||||
assert.equal(resolveModelTestTimeoutMs("openai", "gpt-4.1"), 30_000);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
@@ -7,7 +7,7 @@ import {
|
||||
} from "../../open-sse/config/embeddingRegistry.ts";
|
||||
import { getImageProvider, parseImageModel } from "../../open-sse/config/imageRegistry.ts";
|
||||
|
||||
describe("OpenRouter & GitHub registry entries (#960)", () => {
|
||||
describe("OpenRouter registry entries (#960)", () => {
|
||||
// ── Embedding Registry ──────────────────────────────────────────────────
|
||||
|
||||
describe("embeddingRegistry — openrouter", () => {
|
||||
@@ -36,26 +36,6 @@ describe("OpenRouter & GitHub registry entries (#960)", () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe("embeddingRegistry — github", () => {
|
||||
it("resolves github provider config", () => {
|
||||
const p = getEmbeddingProvider("github");
|
||||
assert.ok(p, "github should be in EMBEDDING_PROVIDERS");
|
||||
assert.equal(p.baseUrl, "https://models.inference.ai.azure.com/embeddings");
|
||||
assert.equal(p.authHeader, "bearer");
|
||||
});
|
||||
|
||||
it("github has at least 2 models", () => {
|
||||
const p = getEmbeddingProvider("github");
|
||||
assert.ok(p.models.length >= 2, `Expected ≥2 models, got ${p.models.length}`);
|
||||
});
|
||||
|
||||
it("parses github/text-embedding-3-small correctly", () => {
|
||||
const result = parseEmbeddingModel("github/text-embedding-3-small");
|
||||
assert.equal(result.provider, "github");
|
||||
assert.equal(result.model, "text-embedding-3-small");
|
||||
});
|
||||
});
|
||||
|
||||
// ── Image Registry ───────────────────────────────────────────────────────
|
||||
|
||||
describe("imageRegistry — openrouter", () => {
|
||||
|
||||
@@ -17,7 +17,8 @@
|
||||
// sarvam+plamo in regional) to 193, then #8170 (inception/typhoon — inception in frontier-labs,
|
||||
// typhoon in regional) to 195, then Firecrawl dual search+fetch under SEARCH_PROVIDERS.firecrawl
|
||||
// (removed specialty-media duplicate) to 194, #8861 (Xiaomi MiMo Token Plan, regional) to 195, and
|
||||
// the Cheaper Inference gateway (OSS-sponsor reseller, gateways family) to 198 (UnoRouter #9009, Raycast Pro #8895).
|
||||
// the Cheaper Inference gateway (OSS-sponsor reseller, gateways family) to 198 (UnoRouter #9009,
|
||||
// Raycast Pro #8895), then later additions to 199; retiring GitHub Models brings it to 198.
|
||||
import { test } from "node:test";
|
||||
import assert from "node:assert/strict";
|
||||
|
||||
@@ -46,12 +47,12 @@ test("barrel still exports every catalog + key helpers", () => {
|
||||
}
|
||||
});
|
||||
|
||||
test("APIKEY_PROVIDERS merges the 6 family files into 199 entries (no loss / no dup)", async () => {
|
||||
test("APIKEY_PROVIDERS merges the 6 family files into 198 entries (no loss / no dup)", async () => {
|
||||
const keys = Object.keys((P as Record<string, object>).APIKEY_PROVIDERS);
|
||||
assert.equal(keys.length, 199);
|
||||
assert.equal(new Set(keys).size, 199, "duplicate keys after spread-merge");
|
||||
assert.equal(keys.length, 198);
|
||||
assert.equal(new Set(keys).size, 198, "duplicate keys after spread-merge");
|
||||
// the merged object's entry-count equals the sum of the 6 semantic family files; families are a
|
||||
// strict partition (every provider in exactly one), so the sum must be exactly 199.
|
||||
// strict partition (every provider in exactly one), so the sum must be exactly 198.
|
||||
const families: [string, string][] = [
|
||||
["gateways", "APIKEY_PROVIDERS_GATEWAYS"],
|
||||
["frontier-labs", "APIKEY_PROVIDERS_FRONTIER"],
|
||||
@@ -71,7 +72,7 @@ test("APIKEY_PROVIDERS merges the 6 family files into 199 entries (no loss / no
|
||||
seen.add(k);
|
||||
}
|
||||
}
|
||||
assert.equal(famTotal, 199, "families must partition all 199 providers");
|
||||
assert.equal(famTotal, 198, "families must partition all 198 providers");
|
||||
});
|
||||
|
||||
test("AI_PROVIDERS Proxy aggregates all sections; lookups resolve", () => {
|
||||
|
||||
@@ -15,8 +15,6 @@ const embeddingsRoute = await import("../../src/app/api/v1/embeddings/route.ts")
|
||||
const videoRoute = await import("../../src/app/api/v1/videos/generations/route.ts");
|
||||
const musicRoute = await import("../../src/app/api/v1/music/generations/route.ts");
|
||||
const v1ModelsCatalog = await import("../../src/app/api/v1/models/catalog.ts");
|
||||
const providerModelsRoute =
|
||||
await import("../../src/app/api/v1/providers/[provider]/models/route.ts");
|
||||
|
||||
async function resetStorage() {
|
||||
core.resetDbInstance();
|
||||
@@ -98,34 +96,6 @@ test("embedding catalog GET hides providers without active credentials", async (
|
||||
assert.ok(!ids.includes("openai/text-embedding-3-small"));
|
||||
});
|
||||
|
||||
test("provider-scoped GitHub Models catalog preserves nested embedding IDs", async () => {
|
||||
await seedConnection("github-models");
|
||||
|
||||
const response = await providerModelsRoute.GET(
|
||||
new Request("http://localhost/v1/providers/github-models/models"),
|
||||
{ params: Promise.resolve({ provider: "github-models" }) }
|
||||
);
|
||||
const body = (await response.json()) as {
|
||||
data?: Array<{ id: string; root?: string; type?: string }>;
|
||||
};
|
||||
assert.equal(response.status, 200);
|
||||
|
||||
const embeddings = (body.data || []).filter((model) => model.type === "embedding");
|
||||
assert.deepEqual(
|
||||
embeddings.map((model) => ({ id: model.id, root: model.root })),
|
||||
[
|
||||
{
|
||||
id: "openai/text-embedding-3-large",
|
||||
root: "openai/text-embedding-3-large",
|
||||
},
|
||||
{
|
||||
id: "openai/text-embedding-3-small",
|
||||
root: "openai/text-embedding-3-small",
|
||||
},
|
||||
]
|
||||
);
|
||||
});
|
||||
|
||||
test("video catalog GET hides credential-backed providers without credentials", async () => {
|
||||
await seedConnection("kie");
|
||||
|
||||
|
||||
Reference in New Issue
Block a user