mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-08 00:02:20 +03:00
* chore(release): open v3.8.35 development cycle
* fix db vacuum scheduler settings (#4726)
Scheduled VACUUM now follows Storage page settings (scheduledVacuum/vacuumHour) as single source of truth; env-flag control path removed. 11/11 vacuum-scheduler tests pass against release/v3.8.35 tip; no orphaned env refs. Integrated into release/v3.8.35.
* fix(tier): noAuth providers count as free; free filter returns empty … (#4753)
noAuth providers now classified free (union of legacy list + NOAUTH_PROVIDERS chat-tier derivation), -free arena_elo alias, and auto/<cat>:free returns an empty pool when no free candidate matches (opt-in legacy fallback via OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL). New env var documented in .env.example + ENVIRONMENT.md; CHANGELOG bullet added (maintainer co-author). 46/46 node + 56/56 vitest tests pass on release tip; env-doc-sync, docs-sync, typecheck:core, lint, file-size all green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai 11 helpers de nível superior para 6 leaves puros (#3501) (#4571)
chatCore god-file decomposition (#3501): extract 6 pure leaves (cacheUsageMeta, executorClientHeaders, nonStreamingResponseBody, skillsFormat, streamErrorResult, streamFinalize) from chatCore.ts. Rebased onto release/v3.8.35 tip (resolved single chatCore.ts conflict — removed now-extracted inline buildExecutorClientHeaders). 265/265 chatcore tests, 26/26 new leaf tests, typecheck:core, cycles, file-size all green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai resolveExecutorWithProxy + getExecutionCredentials para leaves (#3501) (#4646)
chatCore #3501: extract resolveExecutorWithProxy + getExecutionCredentials to leaves (executorProxy.ts, executionCredentials.ts). Clean cherry-pick onto release tip post-#4571. 12/12 new leaf tests, typecheck:core, cycles, file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai transforms de mensagens Claude p/ leaf (#3501) (#4708)
chatCore #3501: extract Claude upstream-message transforms to leaf (claudeUpstreamMessages.ts + claudeMessageTypes.ts). Clean cherry-pick post-#4646. 8/8 new leaf tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai persistAttemptLogs para leaf (#3501) (#4717)
chatCore #3501: extract persistAttemptLogs to leaf (attemptLogging.ts). Rebased onto release tip post-#4708 (resolved imports conflict: kept tip's resolveCompressionHeader from compression Phase 3, dropped now-unused logTruncation import moved into the leaf). 288/288 chatcore tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai stageTrace + compressionUsageReceipt para leaves (#3501) (#4721)
chatCore #3501: extract stageTrace + compressionUsageReceipt to leaves. Clean cherry-pick post-#4717. 6/6 new leaf tests, typecheck/cycles/file-size green. Integrated into release/v3.8.35.
* refactor(chatCore): extrai prepareUpstreamBody (1ª sub-fatia do executeProviderRequest, #3501) (#4730)
chatCore #3501: extract prepareUpstreamBody (first sub-slice of executeProviderRequest) to leaf (upstreamBody.ts). Clean cherry-pick post-#4721. 7/7 new leaf tests, full 301/301 chatcore suite, typecheck/cycles/file-size green. Completes the 6-PR chatCore decomposition stack into release/v3.8.35.
* fix(db): make db-backup import size cap configurable (#4719) (#4757)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* chore(quality): expand check:release-green to the FULL release-PR gate set (#4758)
The release-green pre-flight (Solution C) previously covered only a subset of the
gates that run exclusively on the release PR (PR→main), so reds still accrued
silently on release/** and surfaced in ~40-min layers at release time (v3.8.34:
3 CI rounds — CodeQL sanitization, then the fail-fast Quality Ratchet revealing
openapi then cyclomatic-complexity one push at a time, plus zizmor/integration).
Now check:release-green reproduces the COMPLETE release-PR gate set and reports
EVERY red in one pass (collected, not fail-fast):
- New DRIFT ratchets (report-only, rebaselined at release, never block):
cyclomatic complexity, dead-code, type-coverage, compression-budget,
openapi-coverage, workflow-lint (zizmor), codeql-ratchet.
- New HARD gates (real defects): docs-all (fabricated-docs strict + i18n mirror
sync) and the integration test suite (gated behind !--quick).
The only release-PR gates it still cannot reproduce locally are GitHub-side CodeQL
semantic analysis and SonarQube/SonarCloud (external services).
The nightly-release-green workflow and /green-prs inherit the expanded coverage
automatically (they invoke this script), so cycle drift is now surfaced
continuously and the release PR is green on its first CI run.
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* fix(dashboard): add missing onboarding.tiers step title (#4698) (#4755)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* feat(compression): Output Styles registry + D0 telemetry (Phase 4A) (#4694)
Phase 4A: Output Styles registry + D0 telemetry. Integrated into release/v3.8.35.
* feat(compression): SLM tier for ultra (Phase 4B) [stacked on #4694] (#4707)
Phase 4B: SLM tier for ultra. Integrated into release/v3.8.35.
* feat(compression): context-budget adaptive compression (Phase 4C) [stacked on #4707] (#4716)
Phase 4C: adaptive context-budget compression. Integrated into release/v3.8.35.
* feat(compression): offline evaluation harness (Phase 4 D1) [stacked on #4716] (#4720)
Phase 4 D1: offline evaluation harness. Integrated into release/v3.8.35.
* fix(sse): deepseek-web folds role:tool results into prompt transcript (#4712) (#4756)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* fix(dashboard): remove dead unconditional useLiveRequests call in HomePageClient (#4759, #4745, #4596) (#4761)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* fix(dashboard): dedupe provider nodes by id on compatible-provider add (#4746) (#4768)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* chore(db): re-export compressionRunTelemetry from localDb to satisfy db-rules (#4775)
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
* docs(security): add canonical STRIDE-based threat model (#4783)
Canonical STRIDE threat model. Integrated into release/v3.8.35.
* test(dashboard): add smoke test for home client dashboard (#4793)
Smoke test guarding the dashboard home client render (regression #4745/#4759). Code fix already landed via #4761; this PR's jsdom smoke test is the net-new regression guard. Integrated into release/v3.8.35.
* fix(combos): auto-promote zeroLatencyOptimizationsEnabled so legacy configs (pre-3.8.33 fallbackCompressionMode="lite") round-trip on the first GUI edit (#4774)
Auto-promote zeroLatencyOptimizationsEnabled + strip v3.8.31-era removed keys so legacy combo configs round-trip through PUT /api/combos/{id} on first GUI edit (closes #4382 followup). Pre-merge: rewrote the now-stale reject test to assert auto-promotion + added passthrough/round-trip regression guards; reconciled combos/page.tsx file-size baseline. Integrated into release/v3.8.35.
* refactor(chatCore): extrai parse + usage-stats não-streaming do executeProviderRequest (#3501) (#4762)
chatCore #3501: extract parseNonStreamingResponseBody + recordNonStreamingUsageStats. Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordContextEditingTelemetryHook (#3501) (#4779)
chatCore #3501: extract recordContextEditingTelemetryHook. Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordCompressionCacheStats (#3501) (#4792)
chatCore #3501: extract recordCompressionCacheStats. Integrated into release/v3.8.35.
* refactor(chatCore): extrai writeCavemanOutputAnalytics (#3501) (#4794)
chatCore #3501: extract writeCavemanOutputAnalytics. Integrated into release/v3.8.35.
* refactor(chatCore): extrai scheduleQuotaShareConsumption (POST-hook não-streaming, #3501) (#4780)
chatCore #3501: extract scheduleQuotaShareConsumption (non-streaming POST-hook). Integrated into release/v3.8.35.
* refactor(chatCore): extrai emitRequestGamificationEvent (helper compartilhado DRY, #3501) (#4776)
chatCore #3501: extract emitRequestGamificationEvent (DRY streaming/non-streaming). Integrated into release/v3.8.35.
* refactor(chatCore): extrai runPluginOnResponseHook (#3501) (#4782)
chatCore #3501: extract runPluginOnResponseHook. Integrated into release/v3.8.35.
* refactor(chatCore): extrai scheduleStreamingQuotaShareConsumption (POST-hook streaming, #3501) (#4784)
chatCore #3501: extract scheduleStreamingQuotaShareConsumption (streaming POST-hook). Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordStreamingUsageStats (analytics de usage streaming, #3501) (#4791)
chatCore #3501: extract recordStreamingUsageStats. Integrated into release/v3.8.35.
* refactor(chatCore): extrai recordStreamingCost (custo por-request streaming, #3501) (#4790)
chatCore #3501: extract recordStreamingCost (per-request streaming cost). Integrated into release/v3.8.35.
* docs(readme): credit ponytail + OmniCompress; restore env-doc-sync release-green (#4799)
README compression credits (ponytail/OmniCompress) + env-doc-sync ignore for eval-only OMNIROUTE_EVAL_CREDENTIALS (restores release-green after #4720). Integrated into release/v3.8.35.
* chore(quality): trim combo-config.test.ts comments under file-size cap (#4774 follow-up) (#4800)
Restore file-size release-green. Integrated into release/v3.8.35.
* feat(api-docs): Redoc-rendered /api/docs + consolidate OpenAPI spec to docs/openapi.yaml (#4781)
Redoc /api/docs + OpenAPI spec consolidated to docs/openapi.yaml (canonical 201-path complete spec; old path → legacy fallback). All refs/gates/tests/CI updated. Integrated into release/v3.8.35.
* docs(compression): declare Phase 4 layers — Output Styles, adaptive dial, per-request control (#4801)
The README compression section listed the 9 input engines but not the Phase 4
layers now in production:
- Output Styles (output-axis steering: terse-prose / less-code / terse-cjk, lite/full/ultra)
- adaptive context-budget dial (reserve-output|percentage|absolute · floor|replace-autotrigger|off)
- per-request x-omniroute-compression precedence + the offline eval harness
Also bumped the highlights range to v3.8.35, expanded the compression feature bullet,
and marked the GUIDE's Phase 4 row Shipped (was 'Planned' — it's merged on v3.8.35).
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(release): finalize v3.8.35 CHANGELOG + docs reconciliation
- CHANGELOG: complete 3.8.35 section (all 35 commits since v3.8.34,
contributor attribution: @rdself @megamen32 @KooshaPari @JxnLexn)
- docs(security): align THREAT_MODEL.md refs with real code
(routeGuard.ts, tokenLimits.ts, /api/monitoring/health) — fabricated-docs gate
- check:fabricated-docs: skip docs/superpowers/specs (dated research reports)
- i18n: sync 3.8.35 section into 41 CHANGELOG mirrors (docs-sync size gate)
- ratchet rebaseline: cyclomatic 1916->1920, eslintWarnings 3907->3912
(inherited cycle drift; release-finalize diff is docs-only)
* fix(release): resolve inherited base-reds surfaced by v3.8.35 release CI
Cycle base-reds that only run on PR→main (not the PR→release fast-path):
- test(autoCombo): suffixComposition-4517 used node:test in a vitest-only dir
(#4753) → vitest found no suite. Switch to the vitest API. (Vitest job)
- test(agentSkills): openapiParser fixture wrote docs/reference/openapi.yaml;
parser reads docs/openapi.yaml since #4781 → point fixture at the new path.
(Unit/Coverage/Node24/Node26 shard 4)
- test(integration): proxy-pipeline source-scan expected inline streaming-cost
code that #4790/#3501 extracted to the recordStreamingCost leaf → assert the
delegation instead. (Integration 1/2)
- fix(chatCore): derive the log trace id from crypto, not Math.random
(CodeQL js/insecure-randomness — log-correlation id, not a secret).
- test(resilience): circuit-breaker invalid-cooldown fallback asserted t>29000,
flaking on slow CI where ~1.6s elapsed gave t=28401 → tolerate wall-clock
drift (t>25000). (Unit 6/8)
* fix(usage): derive pending-request id from crypto, not Math.random
CodeQL js/insecure-randomness (#669): the pending-request id generated in
trackPendingRequest (usageHistory.ts) flows into attempt logging and was flagged
as insecure randomness in a security context. It's a log-correlation id, not a
secret — switch to crypto RNG to clear the alert. Pairs with the chatCore traceId
fix in 37c49781a (same sink).
---------
Co-authored-by: Diego Rodrigues de Sa e Souza <souzamiriamrodrigues790@gmail.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: Demiurge The Single <megamen932@gmail.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: Jan Leon <Jan.gaschler@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
414 lines
14 KiB
TypeScript
414 lines
14 KiB
TypeScript
import { AutoComboConfig } from "./engine";
|
|
import { MODE_PACKS } from "./modePacks";
|
|
import { DEFAULT_WEIGHTS, ScoringWeights } from "./scoring";
|
|
import { AutoVariant } from "./autoPrefix";
|
|
import { getProviderConnections } from "@/lib/db/providers";
|
|
import { getSettings } from "@/lib/db/settings";
|
|
import { getProviderRegistry } from "./providerRegistryAccessor";
|
|
import type { ConnectionFields } from "@/lib/db/encryption";
|
|
import { NOAUTH_PROVIDERS } from "@/shared/constants/providers";
|
|
import { hasUsableWebSessionCredential } from "@/shared/providers/webSessionCredentials";
|
|
import { defaultLogger as log } from "@omniroute/open-sse/utils/logger";
|
|
import { getTokenLimit } from "../contextManager";
|
|
import { getResolvedModelCapabilities } from "@/lib/modelCapabilities";
|
|
import {
|
|
buildAutoCandidateFilter,
|
|
tierToWeightVariant,
|
|
type AutoCategory,
|
|
type AutoTier,
|
|
} from "./suffixComposition";
|
|
import { getHiddenModelsByProvider } from "@/models";
|
|
|
|
/** #4235 Phase B: optional category/tier overlay for `auto/<category>:<tier>` combos. */
|
|
export interface AutoComboSpec {
|
|
category?: AutoCategory;
|
|
tier?: AutoTier;
|
|
}
|
|
|
|
/** Minimal connection shape needed for virtual auto-combo factory */
|
|
interface VirtualFactoryConn extends ConnectionFields {
|
|
id: string;
|
|
provider: string;
|
|
defaultModel?: string;
|
|
expiresAt?: number | string | null;
|
|
tokenExpiresAt?: number | string | null;
|
|
providerSpecificData?: Record<string, unknown> | null;
|
|
}
|
|
|
|
type NoAuthProviderDefinition = {
|
|
id?: string;
|
|
alias?: string;
|
|
noAuth?: boolean;
|
|
serviceKinds?: string[];
|
|
};
|
|
|
|
export interface VirtualAutoComboCandidate {
|
|
provider: string;
|
|
connectionId: string;
|
|
model: string;
|
|
modelStr: string; // e.g., 'openai/gpt-4o'
|
|
costPer1MTokens: number; // from providerRegistry
|
|
}
|
|
|
|
type VirtualAutoCombo = AutoComboConfig & {
|
|
strategy: "auto";
|
|
models: Array<{
|
|
id: string;
|
|
kind: "model";
|
|
model: string;
|
|
providerId: string;
|
|
connectionId: string;
|
|
weight: number;
|
|
label: string;
|
|
}>;
|
|
/** MAX of candidates' context windows — safe to advertise because the
|
|
* auto-combo context pre-filter routes oversized requests to large-window
|
|
* candidates. null when the pool is empty. */
|
|
advertisedContextLength: number | null;
|
|
advertisedMaxOutputTokens: number | null;
|
|
autoConfig: {
|
|
candidatePool: string[];
|
|
weights: ScoringWeights;
|
|
explorationRate: number;
|
|
routerStrategy: string;
|
|
};
|
|
config: {
|
|
auto: {
|
|
candidatePool: string[];
|
|
weights: ScoringWeights;
|
|
explorationRate: number;
|
|
routerStrategy: string;
|
|
};
|
|
};
|
|
};
|
|
|
|
function toExpiryMs(value: unknown): number | null {
|
|
if (value === null || value === undefined || value === "") return null;
|
|
|
|
const parsed =
|
|
typeof value === "number"
|
|
? value
|
|
: typeof value === "string" && value.trim().length > 0
|
|
? Number(value)
|
|
: Number.NaN;
|
|
|
|
if (Number.isFinite(parsed) && parsed > 0) {
|
|
return parsed < 10_000_000_000 ? parsed * 1000 : parsed;
|
|
}
|
|
|
|
if (typeof value === "string") {
|
|
const timestamp = new Date(value).getTime();
|
|
return Number.isFinite(timestamp) ? timestamp : null;
|
|
}
|
|
|
|
return null;
|
|
}
|
|
|
|
function hasUsableOAuthToken(conn: VirtualFactoryConn): boolean {
|
|
if (typeof conn.accessToken !== "string" || conn.accessToken.trim().length === 0) return false;
|
|
|
|
const expiryMs = toExpiryMs(conn.tokenExpiresAt) ?? toExpiryMs(conn.expiresAt);
|
|
|
|
return expiryMs === null || expiryMs > Date.now();
|
|
}
|
|
|
|
function hasProviderSpecificSessionData(conn: VirtualFactoryConn): boolean {
|
|
return hasUsableWebSessionCredential(conn.provider, conn.providerSpecificData);
|
|
}
|
|
|
|
function hasUsableConnectionCredential(conn: VirtualFactoryConn): boolean {
|
|
const hasApiKey = typeof conn.apiKey === "string" && conn.apiKey.trim().length > 0;
|
|
return hasApiKey || hasUsableOAuthToken(conn) || hasProviderSpecificSessionData(conn);
|
|
}
|
|
|
|
const SYNTHETIC_NOAUTH_CONNECTION_ID = "noauth";
|
|
|
|
function isChatAutoComboNoAuthProvider(providerDef: NoAuthProviderDefinition): boolean {
|
|
if (providerDef.noAuth !== true) return false;
|
|
if (!Array.isArray(providerDef.serviceKinds) || providerDef.serviceKinds.length === 0)
|
|
return true;
|
|
return providerDef.serviceKinds.includes("llm");
|
|
}
|
|
|
|
function getNoAuthCandidates(
|
|
excludedProviders: Set<string>,
|
|
blockedProviders: Set<string>
|
|
): VirtualAutoComboCandidate[] {
|
|
const registry = getProviderRegistry();
|
|
const candidates: VirtualAutoComboCandidate[] = [];
|
|
|
|
for (const providerDef of Object.values(NOAUTH_PROVIDERS) as NoAuthProviderDefinition[]) {
|
|
if (!isChatAutoComboNoAuthProvider(providerDef)) continue;
|
|
|
|
const providerId = providerDef.id;
|
|
if (!providerId || excludedProviders.has(providerId)) continue;
|
|
if (
|
|
blockedProviders.has(providerId) ||
|
|
(typeof providerDef.alias === "string" && blockedProviders.has(providerDef.alias))
|
|
)
|
|
continue;
|
|
|
|
const providerInfo = registry[providerId];
|
|
const registryModels = Array.isArray(providerInfo?.models) ? providerInfo.models : [];
|
|
if (registryModels.length === 0) continue;
|
|
|
|
// No-auth providers do not have provider_connections rows. Use the same
|
|
// synthetic connection id returned by getProviderCredentials() so the
|
|
// downstream combo path can still carry a stable target/account identity.
|
|
// Prefer provider aliases because some canonical provider IDs are reserved
|
|
// for credentialed tiers with different routing semantics.
|
|
const registryAlias =
|
|
typeof providerInfo?.alias === "string" && providerInfo.alias.trim().length > 0
|
|
? providerInfo.alias
|
|
: null;
|
|
const routingPrefix = providerDef.alias || registryAlias || providerId;
|
|
|
|
for (const model of registryModels) {
|
|
const modelId = typeof model?.id === "string" && model.id.trim().length > 0 ? model.id : null;
|
|
if (!modelId) continue;
|
|
candidates.push({
|
|
provider: providerId,
|
|
connectionId: SYNTHETIC_NOAUTH_CONNECTION_ID,
|
|
model: modelId,
|
|
modelStr: `${routingPrefix}/${modelId}`,
|
|
costPer1MTokens: 0,
|
|
});
|
|
}
|
|
}
|
|
|
|
return candidates;
|
|
}
|
|
|
|
/**
|
|
* Creates a virtual AutoCombo configuration dynamically based on connected providers and a specified variant.
|
|
* This combo is not persisted in the DB.
|
|
*/
|
|
/**
|
|
* Aggregate the context window / max output to ADVERTISE for an auto combo.
|
|
*
|
|
* MAX across candidates (not min): the auto-combo context pre-filter
|
|
* (combo.ts::filterTargetsByRequestCompatibility + the estimated-tokens
|
|
* pre-filter) already routes oversized requests away from small-window
|
|
* candidates, so advertising the largest window lets clients (e.g. opencode)
|
|
* keep their smart auto-compaction calibrated to the best candidate instead
|
|
* of compacting prematurely — or, worse, receiving 0 and disabling
|
|
* compaction entirely (the "agent keeps forgetting things" bug).
|
|
*
|
|
* Unknown candidates resolve through getTokenLimit()'s fallback chain, so a
|
|
* non-empty pool always yields a positive contextLength.
|
|
*/
|
|
export function computeAdvertisedLimits(candidates: Array<{ provider: string; model: string }>): {
|
|
contextLength: number | null;
|
|
maxOutputTokens: number | null;
|
|
} {
|
|
if (!Array.isArray(candidates) || candidates.length === 0) {
|
|
return { contextLength: null, maxOutputTokens: null };
|
|
}
|
|
|
|
let contextLength: number | null = null;
|
|
let maxOutputTokens: number | null = null;
|
|
for (const candidate of candidates) {
|
|
const limit = getTokenLimit(candidate.provider, candidate.model);
|
|
if (Number.isFinite(limit) && limit > 0) {
|
|
contextLength = contextLength === null ? limit : Math.max(contextLength, limit);
|
|
}
|
|
const output = getResolvedModelCapabilities({
|
|
provider: candidate.provider,
|
|
model: candidate.model,
|
|
}).maxOutputTokens;
|
|
if (typeof output === "number" && Number.isFinite(output) && output > 0) {
|
|
maxOutputTokens = maxOutputTokens === null ? output : Math.max(maxOutputTokens, output);
|
|
}
|
|
}
|
|
return { contextLength, maxOutputTokens };
|
|
}
|
|
|
|
export async function createVirtualAutoCombo(
|
|
variant: AutoVariant | undefined,
|
|
spec?: AutoComboSpec
|
|
): Promise<VirtualAutoCombo> {
|
|
const [connections, settings] = await Promise.all([
|
|
getProviderConnections({ isActive: true }) as Promise<VirtualFactoryConn[]>,
|
|
getSettings().catch(() => ({}) as Record<string, unknown>),
|
|
]);
|
|
const blockedProviders = new Set(
|
|
Array.isArray(settings.blockedProviders) ? (settings.blockedProviders as string[]) : []
|
|
);
|
|
const hiddenModelsMap = getHiddenModelsByProvider();
|
|
|
|
const validConnections = connections.filter(hasUsableConnectionCredential);
|
|
|
|
const candidatePool: VirtualAutoComboCandidate[] = [];
|
|
for (const conn of validConnections) {
|
|
const providerInfo = getProviderRegistry()[conn.provider];
|
|
if (!providerInfo) continue; // Skip unknown providers
|
|
|
|
let modelId: string | undefined = conn.defaultModel;
|
|
if (!modelId) {
|
|
const firstModel = providerInfo.models[0];
|
|
modelId = firstModel?.id;
|
|
}
|
|
if (!modelId) continue; // Skip providers without a model
|
|
|
|
// Skip models that the user has hidden in the dashboard
|
|
const hiddenModels = hiddenModelsMap.get(conn.provider);
|
|
if (hiddenModels?.has(modelId)) continue;
|
|
|
|
candidatePool.push({
|
|
provider: conn.provider,
|
|
connectionId: conn.id,
|
|
model: modelId,
|
|
modelStr: `${conn.provider}/${modelId}`,
|
|
costPer1MTokens: 0, // Not used in virtual auto-combo (LKGP uses session stickiness)
|
|
});
|
|
}
|
|
|
|
candidatePool.push(
|
|
...getNoAuthCandidates(new Set(validConnections.map((conn) => conn.provider)), blockedProviders)
|
|
);
|
|
|
|
if (candidatePool.length === 0) {
|
|
log.warn("AUTO", "No connected providers with valid credentials for virtual auto-combo");
|
|
const emptyPool: string[] = [];
|
|
const autoConfig = {
|
|
candidatePool: emptyPool,
|
|
weights: { ...DEFAULT_WEIGHTS },
|
|
explorationRate: 0.05,
|
|
routerStrategy: "lkgp",
|
|
};
|
|
return {
|
|
id: `virtual-auto-${variant || "default"}`,
|
|
name: `Auto ${variant || "Default"}`,
|
|
type: "auto" as const,
|
|
strategy: "auto",
|
|
models: [],
|
|
candidatePool: emptyPool,
|
|
weights: autoConfig.weights,
|
|
explorationRate: autoConfig.explorationRate,
|
|
routerStrategy: autoConfig.routerStrategy,
|
|
autoConfig,
|
|
config: { auto: autoConfig },
|
|
advertisedContextLength: null,
|
|
advertisedMaxOutputTokens: null,
|
|
};
|
|
}
|
|
|
|
// #4235 Phase B: narrow the pool by the `auto/<category>:<tier>` overlay
|
|
// (vision/reasoning capability, free/premium model tier).
|
|
//
|
|
// Default behavior: when the filter yields zero candidates, return an EMPTY
|
|
// pool — never silently fall back to the full pool. This makes
|
|
// `auto/coding:free` actually mean "free tier only" and prevents a paid
|
|
// expensive model from being picked just because no free provider is
|
|
// connected. Operators who want the old "never break routing, lose the bias"
|
|
// behavior can opt back in via the env var below.
|
|
let effectivePool = candidatePool;
|
|
const candidateFilter = spec ? buildAutoCandidateFilter(spec.category, spec.tier) : null;
|
|
if (candidateFilter) {
|
|
const narrowed = candidatePool.filter((c) =>
|
|
candidateFilter({ provider: c.provider, model: c.model })
|
|
);
|
|
if (narrowed.length > 0) {
|
|
effectivePool = narrowed;
|
|
} else if (
|
|
process.env.OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL === "true" ||
|
|
process.env.OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL === "1"
|
|
) {
|
|
// Opt-in legacy behavior: warn loudly, then keep the full pool.
|
|
log.warn(
|
|
"AUTO",
|
|
`auto/${spec?.category ?? ""}${spec?.tier ? `:${spec.tier}` : ""} matched no connected models; falling back to the full pool (OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL=true)`
|
|
);
|
|
} else {
|
|
log.warn(
|
|
"AUTO",
|
|
`auto/${spec?.category ?? ""}${spec?.tier ? `:${spec.tier}` : ""} matched no connected models; returning an empty pool. Set OMNIROUTE_AUTO_FREE_FALLBACK_TO_FULL_POOL=true to restore the legacy "use full pool" behavior.`
|
|
);
|
|
effectivePool = [];
|
|
}
|
|
}
|
|
|
|
let weights: ScoringWeights = { ...DEFAULT_WEIGHTS };
|
|
let explorationRate = 0.05; // Default exploration rate
|
|
let routerStrategy = "lkgp"; // All auto variants use LKGP
|
|
|
|
switch (variant) {
|
|
case "coding":
|
|
weights = { ...MODE_PACKS["quality-first"] };
|
|
break;
|
|
case "fast":
|
|
weights = { ...MODE_PACKS["ship-fast"] };
|
|
break;
|
|
case "cheap":
|
|
weights = { ...MODE_PACKS["cost-saver"] };
|
|
break;
|
|
case "offline":
|
|
weights = { ...MODE_PACKS["offline-friendly"] };
|
|
break;
|
|
case "smart":
|
|
weights = { ...MODE_PACKS["quality-first"] };
|
|
explorationRate = 0.1; // Override default exploration rate
|
|
break;
|
|
case "lkgp":
|
|
// LKGP is default for all auto variants, this variant just explicitly names it.
|
|
// Use default weights.
|
|
break;
|
|
case undefined: // Default auto
|
|
// Use default weights
|
|
break;
|
|
}
|
|
|
|
// #4235 Phase B: category/tier weight overlay. A non-chat category leans
|
|
// quality-first; the tier then refines toward latency (fast), cost (cheap/floor)
|
|
// or availability (reliable). free/pro keep the base weights — their bias is the
|
|
// candidate filter above (free → free-tier models, pro → premium models).
|
|
if (spec) {
|
|
if (spec.category && spec.category !== "chat") {
|
|
weights = { ...MODE_PACKS["quality-first"] };
|
|
}
|
|
const weightVariant = tierToWeightVariant(spec.tier);
|
|
if (weightVariant === "fast") {
|
|
weights = { ...MODE_PACKS["ship-fast"] };
|
|
} else if (weightVariant === "cheap") {
|
|
weights = { ...MODE_PACKS["cost-saver"] };
|
|
} else if (weightVariant === "reliability") {
|
|
weights = { ...MODE_PACKS["reliability-first"] };
|
|
}
|
|
}
|
|
|
|
const providerPool = [...new Set(effectivePool.map((c) => c.provider))];
|
|
const models = effectivePool.map((candidate, index) => ({
|
|
id: `virtual-auto-${variant || "default"}-${index + 1}-${candidate.provider}`,
|
|
kind: "model" as const,
|
|
model: candidate.modelStr,
|
|
providerId: candidate.provider,
|
|
connectionId: candidate.connectionId,
|
|
weight: 1,
|
|
label: candidate.provider,
|
|
}));
|
|
const autoConfig = {
|
|
candidatePool: providerPool,
|
|
weights,
|
|
explorationRate,
|
|
routerStrategy,
|
|
};
|
|
|
|
const advertisedLimits = computeAdvertisedLimits(effectivePool);
|
|
|
|
return {
|
|
id: `virtual-auto-${variant || "default"}`,
|
|
name: `Auto ${variant || "Default"}`,
|
|
type: "auto",
|
|
strategy: "auto",
|
|
models,
|
|
candidatePool: providerPool,
|
|
weights,
|
|
explorationRate,
|
|
routerStrategy,
|
|
autoConfig,
|
|
config: { auto: autoConfig },
|
|
advertisedContextLength: advertisedLimits.contextLength,
|
|
advertisedMaxOutputTokens: advertisedLimits.maxOutputTokens,
|
|
};
|
|
}
|