The weight table said stability accounts for "low latency stdDev / error rate". Grep errorRate in scoring.ts and you find it declared on ProviderCandidate and read nowhere — while combo.ts pulls 24 hours of usage history behind a ten-sample floor, falls back to real-time metrics, and hands every candidate an errorRate the scorer ignores. Two candidates, one failing 1% of calls and one failing 99%, scored identically at 0.459486. This declares reliability as a sixteenth factor: 1 - failureRate, using the same formula, field precedence and rate-bounding speedRanking.ts already applies, so a corrupt reading means "nothing observed" rather than "fails every call". It ships at weight 0, leaving the ranking unchanged to the digit — the honest default, since which weight this deserves is a product call backed by traffic the author does not have. Two declared-but-silent factors already ship (cacheAffinity, resetWindowAffinity), so the pattern is not new. The stability row now describes what that factor actually computes: latency variance. The rest is the mechanical 15 → 16 across nineteen documents and the forty-two llm.txt mirrors — sourced from check:docs-counts rather than a grep, the first real use of the gate #12316 extended. Protected-surface note: this PR touches AGENTS.md, llm.txt and its 42 mirrors, and skills/omni-combos-routing/SKILL.md. Every changed line in those 45 files is a digit substitution and nothing else — masking all digits makes the removed and added lines identical, with no sentence added, removed or reworded. Reviewed and approved on that basis before merging. Verified on the author's rebased head: check:docs-counts green (the gate that now enforces the count this PR moves), typecheck:core clean, and 71/71 focused tests across scoring-reliability-factor, combo-scoring-weights-schema-coverage, check-docs-counts-sync, lkgp-enabled-context, intelligent-routing-options and the combo-matrix auto integration suite. Thanks @maxmad64bis — shipping the factor at weight 0 and saying plainly that the weight is someone else's call is the right way to land this.
4.3 KiB
title, version, lastUpdated
| title | version | lastUpdated |
|---|---|---|
| OmniRoute Tiers — User Guide | 3.8.40 | 2026-06-28 |
OmniRoute Tiers — User Guide
OmniRoute organizes the 352 supported providers into 3 economic tiers. Each request travels through them in order until one returns successfully — you get the cheapest viable response without ever writing fallback code.
Tier 1 — Subscription
Providers you already pay for. OmniRoute uses every drop of quota before it expires.
| Provider | Why Tier 1 |
|---|---|
| Claude Code OAuth | Anthropic Pro/Team — flat-rate, often unused |
| OpenAI Codex (ChatGPT subscription) | Plus/Team includes Codex quota |
| GitHub Copilot | Per-seat — quota resets monthly |
| Cursor IDE | Pro plan quota |
| Antigravity / Devin Desktop | Built-in quotas |
Strategy: route here first for every request that fits the model's
strengths. The quota tracker monitors approaching resets, and the reset-aware
combo strategy prioritizes accordingly. To route Tier 1 first and only step out
to paid tiers as quota runs out, use the auto/thrifty id — or auto/subscription
to stay on plan-included capacity and fail closed instead. See
Subscription-first routing.
Tier 2 — Cheap
Pay-per-token providers under $1/1M tokens. Reserved for high-volume work or after Tier 1 quotas hit limits.
| Provider | Price (input/output) | Strengths |
|---|---|---|
| DeepSeek V4 Pro | $0.27 / $1.10 per 1M | Code, reasoning |
| GLM-4.5 | $0.60 / $2.20 per 1M | Long context |
| MiniMax M1 | $0.20 / $1.10 per 1M | Speed |
| Qwen Coder | $0.30 / $1.20 per 1M | Code |
| OpenRouter (price-optimized) | varies | 100+ models, dynamic |
Strategy: combo cost-optimized picks lowest $/token model that meets
the task's capability filter (vision, JSON mode, tools, max-context).
Tier 3 — Free
Zero-cost providers — free tiers, credit programs, OAuth daily quotas.
| Provider | Free quota / credits |
|---|---|
| Kiro AI | Free Claude tier (generous fair-use) |
| OpenCode Free | No auth, generous rate limits |
| Qoder | Free OAuth |
| Google Vertex AI | $300 new-account credits |
| Amazon Q | Free tier for AWS users |
| Pollinations | Open public API |
| Cloudflare AI | Workers AI free tier |
Strategy: combo auto with budget cap routes here when Tier 1+2 fail
or when useFreeOnly=true is set. Free providers often have weaker
rate limits — circuit breaker recovers them on backoff.
Configuring tiers
Dashboard → Tiers → assign your providers. Defaults (from tierDefaults.json) are
sensible; edit when you have specific subscriptions to prioritize or providers to exclude.
Auto-Combo's 16-factor scoring also considers tier. See
docs/routing/AUTO-COMBO.md.
Telemetry
Dashboard → Usage shows tokens spent per tier per day. Use this to:
- Confirm Tier 1 is utilized fully (otherwise you're wasting subscription value)
- Identify which Tier 2 models are picked most (consolidate to 1-2)
- Verify Tier 3 saves money on test/exploration workloads
Common patterns
Pure-free workload
{
"strategy": "auto",
"config": { "auto": { "weights": { "costInv": 0.5, "tierPriority": 0.3 } } }
}
Forces strongly towards Tier 3; only uses Tier 2 if Tier 3 is unavailable.
Subscription-first with cheap fallback
{
"strategy": "priority",
"targets": [
{ "provider": "claude-code-oauth", "weight": 1 },
{ "provider": "deepseek", "weight": 1 },
{ "provider": "kiro", "weight": 1 }
]
}
Explicit ordered list matching Tier 1 → Tier 2 → Tier 3.