From b0f5f92f1a19a9cb1f939cf06eb69db9d96e59bc Mon Sep 17 00:00:00 2001 From: diegosouzapw Date: Sat, 14 Mar 2026 11:04:09 -0300 Subject: [PATCH] =?UTF-8?q?feat(release):=20v2.4.2=20=E2=80=94=20task-awar?= =?UTF-8?q?e=20routing,=20HuggingFace/Vertex=20providers,=20streaming=20fi?= =?UTF-8?q?xes,=20token=20tracking,=20playground=20uploads?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - feat: Task-Aware Smart Routing (T05) — auto-select model by task type - feat: HuggingFace and Vertex AI provider support - feat: Playground audio/image file uploads for transcription and vision - feat: ModelSelectModal shows ✓ for already-added models (#180) - fix: Claude Haiku routed to OpenAI without provider prefix (#73) - fix: Token counts always 0 for Antigravity/Claude streaming (#74) - fix: OpenAI SDK stream=False drops tool_calls (#302) - fix: Media page generation errors — inline rendering for images/transcription - fix: Round-robin state management for excluded accounts (#349) - fix: Qwen user agent and CLI fingerprint compatibility (#352) - deps: undici→7.24.2, dompurify→3.3.3, docker actions v4 - docs: CHANGELOG 2.4.2 with full feature/fix list - docs: README with Task-Aware Routing table entry --- CHANGELOG.md | 48 ++++++++++++++- README.md | 152 +++++++++++++++++++++++----------------------- docs/openapi.yaml | 2 +- package.json | 2 +- 4 files changed, 125 insertions(+), 79 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 118b76cb2d..20fd4c490a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,53 @@ ## [Unreleased] +## [2.4.2] - 2026-03-14 + +> Multiple improvements from community issue analysis, new provider support, bug fixes for token tracking, model routing, and streaming reliability. + +### ✨ New Features + +- **Task-Aware Smart Routing (T05)**: Automatic model selection based on request content type — coding → deepseek-chat, analysis → gemini-2.5-pro, vision → gpt-4o, summarization → gemini-2.5-flash. Configurable via Settings. New `GET/PUT/POST /api/settings/task-routing` API. +- **HuggingFace Provider**: Added HuggingFace Router as an OpenAI-compatible provider with Llama 3.1 70B/8B, Qwen 2.5 72B, Mistral 7B, Phi-3.5 Mini. +- **Vertex AI Provider**: Added Vertex AI (Google Cloud) provider with Gemini 2.5 Pro/Flash, Gemma 2 27B, Claude via Vertex. +- **Playground File Uploads**: Audio upload for transcription, image upload for vision models (auto-detect by model name), inline image rendering for image generation results. +- **Model Select Visual Feedback**: Already-added models in combo picker now show ✓ green badge — prevents duplicate confusion. +- **Qwen Compatibility (PR #352)**: Updated User-Agent and CLI fingerprint settings for Qwen provider compatibility. +- **Round-Robin State Management (PR #349)**: Enhanced round-robin logic to handle excluded accounts and maintain rotation state correctly. +- **Clipboard UX (PR #360)**: Hardened clipboard operations with fallback for non-secure contexts; Claude tool normalization improvements. + +### 🐛 Bug Fixes + +- **Fix #302 — OpenAI SDK stream=False drops tool_calls**: T01 Accept header negotiation no longer forces streaming when `body.stream` is explicitly `false`. Was causing tool_calls to be silently dropped when using the OpenAI Python SDK in non-streaming mode. +- **Fix #73 — Claude Haiku routed to OpenAI without provider prefix**: `claude-*` models sent without a provider prefix now correctly route to the `antigravity` (Anthropic) provider. Added `gemini-*`/`gemma-*` → `gemini` heuristic as well. +- **Fix #74 — Token counts always 0 for Antigravity/Claude streaming**: The `message_start` SSE event which carries `input_tokens` was not being parsed by `extractUsage()`, causing all input token counts to drop. Input/output token tracking now works correctly for streaming responses. +- **Fix #180 — Model import duplicates with no feedback**: `ModelSelectModal` now shows ✓ green highlight for models already in the combo, making it obvious they're already added. +- **Media page generation errors**: Image results now render as `` tags instead of raw JSON. Transcription results shown as readable text. Credential errors show an amber banner instead of silent failure. +- **Token refresh button on provider page**: Manual token refresh UI added for OAuth providers. + +### 🔧 Improvements + +- **Provider Registry**: HuggingFace and Vertex AI added to `providerRegistry.ts` and `providers.ts` (frontend). +- **Read Cache**: New `src/lib/db/readCache.ts` for efficient DB read caching. +- **Quota Cache**: Improved quota cache with TTL-based eviction. + +### 📦 Dependencies + +- `dompurify` → 3.3.3 (PR #347) +- `undici` → 7.24.2 (PR #348, #361) +- `docker/setup-qemu-action` → v4 (PR #342) +- `docker/setup-buildx-action` → v4 (PR #343) + +### 📁 New Files + +| File | Purpose | +| --------------------------------------------- | --------------------------------------- | +| `open-sse/services/taskAwareRouter.ts` | Task-aware routing logic (7 task types) | +| `src/app/api/settings/task-routing/route.ts` | Task routing config API | +| `src/app/api/providers/[id]/refresh/route.ts` | Manual OAuth token refresh | +| `src/lib/db/readCache.ts` | Efficient DB read cache | +| `src/shared/utils/clipboard.ts` | Hardened clipboard with fallback | + ## [2.4.1] - 2026-03-13 ### 🐛 Fix @@ -40,7 +87,6 @@ ## [2.3.14] - 2026-03-13 - ### 🐛 Bug Fixes - **iFlow OAuth (#339)**: Restored the valid default `clientSecret` — was previously an empty string, causing "Bad client credentials" on every connect attempt. The public credential is now the default fallback (overridable via `IFLOW_OAUTH_CLIENT_SECRET` env var). diff --git a/README.md b/README.md index 6ae1fa0c9a..b48c28b98f 100644 --- a/README.md +++ b/README.md @@ -706,19 +706,18 @@ Outcome: deep fallback depth for deadline-critical workloads > Setup AI coding in minutes at **$0/month**. Connect these free accounts and use the built-in **Free Stack** combo. -| Step | Action | Providers Unlocked | -|---|---|---| -| 1 | Connect **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 — **unlimited** | -| 2 | Connect **iFlow** (Google OAuth) | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1... — **unlimited** | -| 3 | Connect **Qwen** (Device Code) | qwen3-coder-plus, qwen3-coder-flash... — **unlimited** | -| 4 | Connect **Gemini CLI** (Google OAuth) | gemini-3-flash, gemini-2.5-pro — **180K/mo free** | -| 5 | `/dashboard/combos` → **Free Stack ($0)** template | Round-robin all free providers automatically | +| Step | Action | Providers Unlocked | +| ---- | -------------------------------------------------- | ------------------------------------------------------------------ | +| 1 | Connect **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 — **unlimited** | +| 2 | Connect **iFlow** (Google OAuth) | kimi-k2-thinking, qwen3-coder-plus, deepseek-r1... — **unlimited** | +| 3 | Connect **Qwen** (Device Code) | qwen3-coder-plus, qwen3-coder-flash... — **unlimited** | +| 4 | Connect **Gemini CLI** (Google OAuth) | gemini-3-flash, gemini-2.5-pro — **180K/mo free** | +| 5 | `/dashboard/combos` → **Free Stack ($0)** template | Round-robin all free providers automatically | **Point any IDE/CLI to:** `http://localhost:20128/v1` · API Key: `any-string` · Done. > **Optional extra coverage (also free):** Groq API key (30 RPM free), NVIDIA NIM (40 RPM free, 70+ models), Cerebras (1M tok/day). - ## ⚡ Quick Start ### 1) Install and run @@ -899,25 +898,25 @@ When minimized, OmniRoute lives in your system tray with quick actions: ## 💰 Pricing at a Glance -| Tier | Provider | Cost | Quota Reset | Best For | -| ------------------- | ----------------- | ----------------------- | ---------------- | -------------------- | -| **💳 SUBSCRIPTION** | Claude Code (Pro) | $20/mo | 5h + weekly | Already subscribed | -| | Codex (Plus/Pro) | $20-200/mo | 5h + weekly | OpenAI users | -| | Gemini CLI | **FREE** | 180K/mo + 1K/day | Everyone! | -| | GitHub Copilot | $10-19/mo | Monthly | GitHub users | -| **🔑 API KEY** | NVIDIA NIM | **FREE** (dev forever) | ~40 RPM | 70+ open models | -| | Cerebras | **FREE** (1M tok/day) | 60K TPM / 30 RPM | World's fastest | -| | Groq | **FREE** (30 RPM) | 14.4K RPD | Ultra-fast Llama/Gemma | -| | DeepSeek | Pay-per-use | None | Best price/quality | -| | xAI (Grok) | Pay-per-use | None | Grok models | -| | Mistral | Free trial + paid | Rate limited | European AI | -| | OpenRouter | Pay-per-use | None | 100+ models aggr. | -| **💰 CHEAP** | GLM-4.7 | $0.6/1M | Daily 10AM | Budget backup | -| | MiniMax M2.1 | $0.2/1M | 5-hour rolling | Cheapest option | -| | Kimi K2 | $9/mo flat | 10M tokens/mo | Predictable cost | -| **🆓 FREE** | iFlow | **$0** | Unlimited | 5 models unlimited | -| | Qwen | **$0** | Unlimited | 4 models unlimited | -| | Kiro | **$0** | Unlimited | Claude (AWS Builder ID) | +| Tier | Provider | Cost | Quota Reset | Best For | +| ------------------- | ----------------- | ---------------------- | ---------------- | ----------------------- | +| **💳 SUBSCRIPTION** | Claude Code (Pro) | $20/mo | 5h + weekly | Already subscribed | +| | Codex (Plus/Pro) | $20-200/mo | 5h + weekly | OpenAI users | +| | Gemini CLI | **FREE** | 180K/mo + 1K/day | Everyone! | +| | GitHub Copilot | $10-19/mo | Monthly | GitHub users | +| **🔑 API KEY** | NVIDIA NIM | **FREE** (dev forever) | ~40 RPM | 70+ open models | +| | Cerebras | **FREE** (1M tok/day) | 60K TPM / 30 RPM | World's fastest | +| | Groq | **FREE** (30 RPM) | 14.4K RPD | Ultra-fast Llama/Gemma | +| | DeepSeek | Pay-per-use | None | Best price/quality | +| | xAI (Grok) | Pay-per-use | None | Grok models | +| | Mistral | Free trial + paid | Rate limited | European AI | +| | OpenRouter | Pay-per-use | None | 100+ models aggr. | +| **💰 CHEAP** | GLM-4.7 | $0.6/1M | Daily 10AM | Budget backup | +| | MiniMax M2.1 | $0.2/1M | 5-hour rolling | Cheapest option | +| | Kimi K2 | $9/mo flat | 10M tokens/mo | Predictable cost | +| **🆓 FREE** | iFlow | **$0** | Unlimited | 5 models unlimited | +| | Qwen | **$0** | Unlimited | 4 models unlimited | +| | Kiro | **$0** | Unlimited | Claude (AWS Builder ID) | **💡 $0 Combo Stack:** Gemini CLI (180K/mo) → iFlow (unlimited: kimi-k2-thinking, qwen3-coder-plus, deepseek-r1) → Kiro (Claude for free) → Qwen (4 models, unlimited) — **Zero cost, never stops coding.** When Gemini quota runs out, OmniRoute auto-falls back to iFlow or Kiro with zero config. @@ -931,63 +930,64 @@ When minimized, OmniRoute lives in your system tray with quick actions: ### 🔵 CLAUDE MODELS (via Kiro — AWS Builder ID) -| Model | Prefix | Limit | Rate Limit | -|---|---|---|---| -| `claude-sonnet-4.5` | `kr/` | **Unlimited** | No reported daily cap | -| `claude-haiku-4.5` | `kr/` | **Unlimited** | No reported daily cap | -| `claude-opus-4.6` | `kr/` | **Unlimited** | Latest Opus via Kiro | +| Model | Prefix | Limit | Rate Limit | +| ------------------- | ------ | ------------- | --------------------- | +| `claude-sonnet-4.5` | `kr/` | **Unlimited** | No reported daily cap | +| `claude-haiku-4.5` | `kr/` | **Unlimited** | No reported daily cap | +| `claude-opus-4.6` | `kr/` | **Unlimited** | Latest Opus via Kiro | ### 🟢 IFLOW MODELS (Free OAuth — No Credit Card) -| Model | Prefix | Limit | Rate Limit | -|---|---|---|---| -| `kimi-k2-thinking` | `if/` | **Unlimited** | No reported cap | -| `qwen3-coder-plus` | `if/` | **Unlimited** | No reported cap | -| `deepseek-r1` | `if/` | **Unlimited** | No reported cap | -| `minimax-m2.1` | `if/` | **Unlimited** | No reported cap | -| `kimi-k2` | `if/` | **Unlimited** | No reported cap | +| Model | Prefix | Limit | Rate Limit | +| ------------------ | ------ | ------------- | --------------- | +| `kimi-k2-thinking` | `if/` | **Unlimited** | No reported cap | +| `qwen3-coder-plus` | `if/` | **Unlimited** | No reported cap | +| `deepseek-r1` | `if/` | **Unlimited** | No reported cap | +| `minimax-m2.1` | `if/` | **Unlimited** | No reported cap | +| `kimi-k2` | `if/` | **Unlimited** | No reported cap | ### 🟡 QWEN MODELS (Device Code Auth) -| Model | Prefix | Limit | Rate Limit | -|---|---|---|---| -| `qwen3-coder-plus` | `qw/` | **Unlimited** | No reported cap | -| `qwen3-coder-flash` | `qw/` | **Unlimited** | No reported cap | -| `qwen3-coder-next` | `qw/` | **Unlimited** | No reported cap | -| `vision-model` | `qw/` | **Unlimited** | Multimodal (images) | +| Model | Prefix | Limit | Rate Limit | +| ------------------- | ------ | ------------- | ------------------- | +| `qwen3-coder-plus` | `qw/` | **Unlimited** | No reported cap | +| `qwen3-coder-flash` | `qw/` | **Unlimited** | No reported cap | +| `qwen3-coder-next` | `qw/` | **Unlimited** | No reported cap | +| `vision-model` | `qw/` | **Unlimited** | Multimodal (images) | ### 🟣 GEMINI CLI (Google OAuth) -| Model | Prefix | Limit | Rate Limit | -|---|---|---|---| -| `gemini-3-flash-preview` | `gc/` | **180K tok/month** + 1K/day | Monthly reset | -| `gemini-2.5-pro` | `gc/` | 180K/month (shared pool) | High quality | +| Model | Prefix | Limit | Rate Limit | +| ------------------------ | ------ | --------------------------- | ------------- | +| `gemini-3-flash-preview` | `gc/` | **180K tok/month** + 1K/day | Monthly reset | +| `gemini-2.5-pro` | `gc/` | 180K/month (shared pool) | High quality | ### ⚫ NVIDIA NIM (Free API Key — build.nvidia.com) -| Tier | Daily Limit | Rate Limit | Notes | -|---|---|---|---| +| Tier | Daily Limit | Rate Limit | Notes | +| ---------- | ------------ | ----------- | ------------------------------------------------------ | | Free (Dev) | No token cap | **~40 RPM** | 70+ models; transitioning to pure rate limits mid-2025 | Popular free models: `moonshotai/kimi-k2.5` (Kimi K2.5), `z-ai/glm4.7` (GLM 4.7), `deepseek-ai/deepseek-v3.2` (DeepSeek V3.2), `nvidia/llama-3.3-70b-instruct`, `deepseek/deepseek-r1` ### ⚪ CEREBRAS (Free API Key — inference.cerebras.ai) -| Tier | Daily Limit | Rate Limit | Notes | -|---|---|---|---| +| Tier | Daily Limit | Rate Limit | Notes | +| ---- | ----------------- | ---------------- | ------------------------------------------- | | Free | **1M tokens/day** | 60K TPM / 30 RPM | World's fastest LLM inference; resets daily | Available free: `llama-3.3-70b`, `llama-3.1-8b`, `deepseek-r1-distill-llama-70b` ### 🔴 GROQ (Free API Key — console.groq.com) -| Tier | Daily Limit | Rate Limit | Notes | -|---|---|---|---| +| Tier | Daily Limit | Rate Limit | Notes | +| ---- | ------------- | ---------------- | ----------------------------------------- | | Free | **14.4K RPD** | 30 RPM per model | No credit card; 429 on limit, not charged | Available free: `llama-3.3-70b-versatile`, `gemma2-9b-it`, `mixtral-8x7b`, `whisper-large-v3` > **💡 The Ultimate Free Stack:** +> > ``` > Kiro (Claude, unlimited) > → iFlow (5 models, unlimited) @@ -997,18 +997,18 @@ Available free: `llama-3.3-70b-versatile`, `gemma2-9b-it`, `mixtral-8x7b`, `whis > → Groq (14.4K req/day) > → NVIDIA NIM (40 RPM, 70+ models) > ``` +> > Configure this as an OmniRoute combo and you'll never pay for AI again. - ## 🎙️ Free Transcription Combo > Transcribe any audio/video for **$0** — Deepgram leads with $200 free, AssemblyAI $50 fallback, Groq Whisper as unlimited emergency backup. -| Provider | Free Credits | Best Model | Rate Limit | -|---|---|---|---| -| 🟢 **Deepgram** | **$200 free** (signup) | `nova-3` — best accuracy, 30+ languages | No RPM limit on free credits | -| 🔵 **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` — chapters, sentiment, PII | No RPM limit on free credits | -| 🔴 **Groq** | **Free forever** | `whisper-large-v3` — OpenAI Whisper | 30 RPM (rate limited) | +| Provider | Free Credits | Best Model | Rate Limit | +| ----------------- | ---------------------- | -------------------------------------------- | ---------------------------- | +| 🟢 **Deepgram** | **$200 free** (signup) | `nova-3` — best accuracy, 30+ languages | No RPM limit on free credits | +| 🔵 **AssemblyAI** | **$50 free** (signup) | `universal-3-pro` — chapters, sentiment, PII | No RPM limit on free credits | +| 🔴 **Groq** | **Free forever** | `whisper-large-v3` — OpenAI Whisper | 30 RPM (rate limited) | **Suggested combo in `/dashboard/combos`:** @@ -1023,7 +1023,6 @@ Nodes: Then in `/dashboard/media` → **Transcription** tab: upload any audio or video file → select your combo endpoint → get transcription in supported formats. - ## 💡 Key Features OmniRoute v2.0 is built as an operational platform, not just a relay proxy. @@ -1058,20 +1057,21 @@ OmniRoute v2.0 is built as an operational platform, not just a relay proxy. ### 🧠 Routing & Intelligence -| Feature | What It Does | -| ---------------------------------- | --------------------------------------------------------------------- | -| 🎯 **Smart 4-Tier Fallback** | Auto-route: Subscription → API Key → Cheap → Free | -| 📊 **Real-Time Quota Tracking** | Live token count + reset countdown per provider | -| 🔄 **Format Translation** | OpenAI ↔ Claude ↔ Gemini ↔ Responses with schema-safe conversions | -| 👥 **Multi-Account Support** | Multiple accounts per provider with intelligent selection | -| 🔄 **Auto Token Refresh** | OAuth tokens refresh automatically with retry | -| 🎨 **Custom Combos** | 6 balancing strategies + fallback chain control | -| 🌐 **Wildcard Router** | `provider/*` dynamic routing | -| 🧠 **Thinking Budget Controls** | Passthrough, auto, custom, and adaptive reasoning limits | -| 🔀 **Model Aliases** | Built-in + custom model aliasing and migration safety | -| ⚡ **Background Degradation** | Route low-priority background tasks to cheaper models | -| 💬 **System Prompt Injection** | Global behavior controls applied consistently | -| 📄 **Responses API Compatibility** | Full `/v1/responses` support for Codex and advanced agentic workflows | +| Feature | What It Does | +| ---------------------------------- | ------------------------------------------------------------------------ | +| 🎯 **Smart 4-Tier Fallback** | Auto-route: Subscription → API Key → Cheap → Free | +| 📊 **Real-Time Quota Tracking** | Live token count + reset countdown per provider | +| 🔄 **Format Translation** | OpenAI ↔ Claude ↔ Gemini ↔ Responses with schema-safe conversions | +| 👥 **Multi-Account Support** | Multiple accounts per provider with intelligent selection | +| 🔄 **Auto Token Refresh** | OAuth tokens refresh automatically with retry | +| 🎨 **Custom Combos** | 6 balancing strategies + fallback chain control | +| 🌐 **Wildcard Router** | `provider/*` dynamic routing | +| 🧠 **Thinking Budget Controls** | Passthrough, auto, custom, and adaptive reasoning limits | +| 🔀 **Model Aliases** | Built-in + custom model aliasing and migration safety | +| ⚡ **Background Degradation** | Route low-priority background tasks to cheaper models | +| 🧪 **Task-Aware Smart Routing** | Auto-select model by content type (coding/vision/analysis/summarization) | +| 💬 **System Prompt Injection** | Global behavior controls applied consistently | +| 📄 **Responses API Compatibility** | Full `/v1/responses` support for Codex and advanced agentic workflows | ### 🎵 Multi-Modal APIs diff --git a/docs/openapi.yaml b/docs/openapi.yaml index 378acc556a..0bd39ab028 100644 --- a/docs/openapi.yaml +++ b/docs/openapi.yaml @@ -1,7 +1,7 @@ openapi: 3.1.0 info: title: OmniRoute API - version: 2.4.1 + version: 2.4.2 description: | OmniRoute is a local-first AI API proxy router. It provides an OpenAI-compatible endpoint that routes requests to multiple AI providers with load balancing, diff --git a/package.json b/package.json index f87a24956c..12355f4e0b 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "omniroute", - "version": "2.4.1", + "version": "2.4.2", "description": "Smart AI Router with auto fallback — route to FREE & cheap models, zero downtime. Works with Cursor, Cline, Claude Desktop, Codex, and any OpenAI-compatible tool.", "type": "module", "bin": {