mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-07-26 09:52:11 +03:00
* fix routed target request parameters * chore: rerun CI * test(chatcore): align PR #7323 Codex-routing test with #7533 verbosity gating The "Codex Responses routing keeps reasoning effort while dropping GPT-only verbosity" test translated a Responses-shape request with credentials=null and only the positional `provider` arg set to "opencode-go". #7533's verbosity carry-over (Responses `text.verbosity` -> Chat `verbosity`) reads the destination from `credentials.provider`, which the source->openai translation step never threads from the positional provider arg — so `translated.verbosity` came back undefined instead of "low", failing before prepareUpstreamBody's sanitizer was even reached. The test's intent (per its own name/comment) is a combo/fallback reroute: translate while still addressed at Codex (an #7533-allowlisted OpenAI-param destination, so verbosity legitimately survives translateRequest), then resolve the final upstream target to opencode-go/GLM so prepareUpstreamBody's sanitizeRequestForResolvedTarget (#7050/#7533) strips the GPT-only verbosity for that concrete target while preserving reasoning_effort. Fixed by passing `credentials: { provider: "codex" }` to the first translateRequest call to match how production actually carries the destination provider, instead of relying on the provider positional argument. No production code changed — #7050 and #7533's sanitization are intentional and protected; only the test's setup was misaligned with them. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
4.3 KiB
4.3 KiB
open-sse/services/ — Routing Engine & Cross-Cutting Services
Purpose: 134 service modules (top-level) powering request routing, rate limiting, quota management, token refresh, fallback strategies, and runtime state. The combo routing engine (combo.ts) is the core; supporting services handle resilience, accounting, and decision-making.
Live count: ls open-sse/services/*.ts | wc -l (currently 134). More including sub-dirs like autoCombo/ and compression/.
Combo Routing Engine
combo.ts— Entry point for multi-model routing.handleComboChat()iterates through targets in order until success or all fail.resolveComboTargets()expands combo config into orderedResolvedComboTarget[](provider + model + account + credentials).- Strategies (17):
priority,weighted,fill-first,round-robin,P2C,random,least-used,reset-aware,reset-window,cost-optimized,strict-random,auto,lkgp,context-optimized,context-relay,headroom,fusion. Source:ROUTING_STRATEGY_VALUESinsrc/shared/constants/routingStrategies.ts. - Each target calls
handleSingleModel()which wrapshandleChatCore()with per-target error handling and circuit breaker checks.
Key Services
Quota & Rate Limiting
rateLimitManager.ts— Token bucket per API key + provider combo. Rejects before dispatch.usage.ts— Per-request token/cost consumption tracking.quotaCache.ts— In-memory quota snapshots, pre-loaded at startup.
Account & Token Management
tokenRefresh.ts— OAuth token expiration detection and refresh.accountFallback.ts— Account switching on quota/rate-limit. Also houses model lockout.sessionManager.ts— Request session state across retries.
Request Routing & Intelligence
wildcardRouter.ts— Wildcard route matching in combo configs.intentClassifier.ts— Request intent classification for intelligent routing.taskAwareRouter.ts— Task-type-based routing (reasoning → o1, code-gen → Cursor).targetRequestSanitizer.ts— Final provider/model-aware parameter sanitation after routing resolution and before executor dispatch.thinkingBudget.ts— Thinking token allocation for o1/o3 models.contextManager.ts— Routing context injection (system prompts, memory).
Model Lifecycle & Fallback
modelDeprecation.ts— Deprecated model detection and successor routing.modelFamilyFallback.ts— T5 intra-family fallback chains.emergencyFallback.ts— Last-resort fallback to stable free providers.
State & Detection
workflowFSM.ts— Multi-turn workflow state machine.backgroundTaskDetector.ts— Long-running task detection for batch routing.ipFilter.ts— IP-based routing rules.signatureCache.ts— Request signature caching for deduplication.volumeDetector.ts— Volume spike detection for rate-limit escalation.contextHandoff.ts— Session context serialization for A2A handoff.
Prompt Compression Pipeline (compression/)
strategySelector.ts— Compression mode selection (off/lite/standard/aggressive/ultra/rtk/stacked).lite.ts— 5 lite techniques (whitespace, dedup, tool results, redundant removal, image URLs).caveman.ts/cavemanRules.ts— Caveman-style semantic condensation with rule packs.engines/rtk/— RTK tool-output compression (command detection, JSON filters, dedup, truncation).engines/registry.ts— Engine registry for standalone and stacked pipelines.stats.ts— Per-request compression stats.types.ts— Shared types (CompressionMode,CompressionConfig,CompressionStats).
Adding a New Service
- Create
open-sse/services/[serviceName].ts - Export main handler function
- Add unit tests in
tests/unit/services/ - Integrate into
handlers/chatCore.ts(if routing-related) orcombo.ts - Document in this file
Anti-Patterns
- Synchronous DB calls in
combo.tshot path — pre-compute and cache - Retry logic in handlers — use
retry()from resilience service - Direct provider config access — use
providerRegistrygetter functions - Hardcoded fallback chains — define in
modelFamilyFallback.ts - State mutations across concurrent requests — use request-scoped context only