* feat(routing): self-hosted unified OpenAI-compatible entry (RIC-738) Divert /v1/chat/completions through the self-hosted provider adapters when OMNIROUTE_SELF_HOSTED_PROVIDERS / OMNIROUTE_SELF_HOSTED_PROVIDERS_FILE is set: one OpenAI-compatible contract in, auto-route to the selected provider (x-omniroute-provider header, provider/model prefix, or first provider), standard OpenAI error shape out. Optional OMNIROUTE_SELF_HOSTED_API_KEY guards the entry (D5 reserved); unset = open loopback route. Upstream credentials stay runtime-only and are stripped from echoed responses. Brings in the provider-adapters baseline from sibling branch (RIC-737) that this entry depends on. Includes 21 passing unit tests (provider selection, model-prefix forwarding, header hygiene, auth, error normalization, SSE passthrough, fall-through/misconfig), docs, env example, changelog fragment. * feat(routing): deterministic routing strategies for self-hosted entry (RIC-740) Add the M2 deterministic routing strategy engine (D3 可审计路由) to the self-hosted unified entry: a declarative `strategy:` block expressing five explainable, non-predictive policies — blacklist/whitelist hard filters, cooldown circuit breaker, cost-priority, latency-aware ordering, and an explicit fallback chain. The ordered candidate list is the fallback chain: a failed primary (network or non-2xx) falls through to the next candidate and each failure feeds the breaker. Every response carries an x-omniroute-route-decision header answering "why this model / why not that one". A pinned provider rejected by a hard filter returns 400 (never a silent re-route); no eligible providers returns 503 with the full explainable decision. No ML/predict dependency. Covers the RIC-740 acceptance: 5 strategy types with unit tests + HTTP fault-injection tests (primary down -> fallback works), config matching docs, and no predict/ML deps. Adds docs, .env.example entries, and a changelog fragment. * refactor(routing): reduce complexity-ratchet violations in new self-hosted routing files Extract cost/id validation, pin-blocked resolution, ordering, and env/file source resolution into small helpers so routingStrategies.ts and selfHostedEntry.ts stay under the complexity-ratchets cap. No behavior change — the same 51 routing-strategies/self-hosted-entry/provider-adapters tests pass unmodified. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * docs(routing): document the 5 self-hosted env vars in ENVIRONMENT.md check:env-doc-sync failed because OMNIROUTE_SELF_HOSTED_PROVIDERS(_FILE), OMNIROUTE_SELF_HOSTED_API_KEY and OMNIROUTE_SELF_HOSTED_STRATEGY(_FILE) were present in .env.example but missing from docs/reference/ENVIRONMENT.md. Add them under "6. Tool & Routing Policies". Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Ant Rich <ant@richants.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: luyuehm <luyuehm@users.noreply.github.com>
3.5 KiB
title
| title |
|---|
| Self-Hosted OpenAI-Compatible Entry |
Self-hosted unified OpenAI-compatible entry
When enabled, OmniRoute's existing /v1/chat/completions (and the OpenAI-compatible
contract it serves) becomes a self-hosted gateway: one OpenAI-compatible request
in, auto-routed to the provider of your choice through the provider adapters, with a
standard OpenAI error shape out. Client code does not change.
This is the D4 (接入即用) differentiator from the RIC-697 design: the same /v1 path
the OpenAI SDK already targets, backed by your own providers instead of a single
vendored catalog.
Enable
Set either env var (see .env.example for both):
OMNIROUTE_SELF_HOSTED_PROVIDERS— inline YAML document (runtime-only, not logged).OMNIROUTE_SELF_HOSTED_PROVIDERS_FILE— path to a YAML file.
# providers.yaml
providers:
- id: local
kind: openai # openai | anthropic | local
baseUrl: http://127.0.0.1:11434/v1
model: llama3
# apiKey: sk-... # optional, runtime-only
- id: claude
kind: anthropic
baseUrl: http://127.0.0.1:8080
model: claude-sonnet
When either var is set, the unified entry is active for every request to
/v1/chat/completions. Config present but unparseable returns 500 (it never
silently falls through to cloud routing). When neither is set, the route behaves
exactly as before.
Provider auto-route
Precedence (deterministic, no predictive model):
x-omniroute-provider: <id>header — exact provider id.modelprefix:provider/model(slash) orprovider::model(double colon).- First configured provider.
The routing prefix is stripped before forwarding — upstream receives the bare model
(claude-sonnet, not claude/claude-sonnet).
# via header
curl http://localhost:20128/v1/chat/completions \
-H "x-omniroute-provider: claude" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet","messages":[{"role":"user","content":"hi"}]}'
# or via model prefix
curl ... -d '{"model":"claude/claude-sonnet","messages":[...]}'
Auth scaffold (D5 reserved)
Optional OMNIROUTE_SELF_HOSTED_API_KEY. When set, requests must carry
Authorization: Bearer <key>. Unset = open route (loopback / trusted network),
mirroring how OmniRoute's existing local providers work. The D5 quota/quota-自治
key system is expected to take over this header.
Failure contract
- Upstream non-2xx body is normalized through
parseUpstreamError+buildErrorBodyinto the standard OpenAIerrorshape (Hard Rule #12 — never raw upstream text). - Network-level failures (connection refused / DNS / TLS) return a normalized
502witherror.message: "Upstream provider unreachable: …". - Credential / session headers echoed upstream are stripped from every response.
Tests
tests/unit/self-hosted-entry.test.ts covers provider selection, model-prefix
forwarding, header hygiene, auth, upstream-error normalization, SSE streaming
passthrough, and the fall-through / misconfig paths — all over real HTTP against a
local upstream.
Notes
- Self-hosted models bypass the cloud-only retirement / alias machinery by design:
the divert happens before those cloud checks, so ids like
local/llama3never trip cloud-peer 410s or alias rewrites. - Deterministic routing strategies (fallback / cooldown / cost / latency / blacklist)
are M2's scope (
RIC-740) and inject into the gateway layer here. Configure them with astrategy:block — see DETERMINISTIC_ROUTING.md.