From 99279b037a883a4e6252e11f97c3c4e4fce1a196 Mon Sep 17 00:00:00 2001
From: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Date: Tue, 28 Jul 2026 17:16:31 -0300
Subject: [PATCH] docs(release): v3.8.49 feature-documentation sync
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Phase 1 step 6b. Swept the cycle's 284 New Features bullets against the existing
docs before writing anything: nearly every large theme (Kimi, xAI OAuth, session
affinity, bun:sqlite, Firecrawl, Opus 5, omniglyph, GCF v3.2, homologation suite)
was already covered. Six real gaps were left undocumented by the PRs that shipped
them, each verified in source before being written up:
- CredentialMaskerGuardrail (#7683) is registered in guardrails/registry.ts but the
GUARDRAILS table listed only 3 of the 4 guardrails
- the cacheAffinity scoring factor and the cache-optimized combo strategy (#8008):
the docs still said 12 factors / 18 strategies, the code has 13 / 19
- the optional dashboard OIDC login gate (#6973) — /api/auth/oidc/{login,callback}
had no mention in AUTHZ_GUIDE
- GET /api/usage/cache-health (#8827) and GET /api/usage/model-latency-stats (#6873)
were missing from the API reference
README "What's New" gains one bullet (routing transparency) and merges two others
rather than growing a second changelog. PROVIDER_REFERENCE regenerated with the
generator (Firecrawl reclassified to Search, Xiaomi MiMo added by #8861).
check:docs-all green: 134 docs, 813 internal links, no fabricated API/env/CLI
references. Known pre-existing drift left alone and reported: stale nominal counts
in ARCHITECTURE/CODEBASE_DOCUMENTATION (soft), the 9-factor mentions scattered in
AUTO-COMBO, and the auto-combo diagram SVG (the renderer needs a browser this
environment does not have — the .mmd source is updated and the .md says so).
---
README.md | 5 ++--
docs/architecture/AUTHZ_GUIDE.md | 26 ++++++++++++++++
docs/architecture/RESILIENCE_GUIDE.md | 2 +-
docs/diagrams/auto-combo-12factor.mmd | 9 +++---
docs/reference/API_REFERENCE.md | 2 ++
docs/reference/PROVIDER_REFERENCE.md | 11 +++----
docs/routing/AUTO-COMBO.md | 12 ++++----
docs/security/GUARDRAILS.md | 43 +++++++++++++++++++++++----
8 files changed, 87 insertions(+), 23 deletions(-)
diff --git a/README.md b/README.md
index 00394b5fc0..b4f5d9a292 100644
--- a/README.md
+++ b/README.md
@@ -418,11 +418,12 @@ OmniRoute is free and open source, built and maintained in the open. If it saves
- **🕵️ Transparent MITM decrypt (TPROXY)** — capture CLIs that ignore proxy env vars, with a per-SNI CA + trust-store installer. → [MITM/TPROXY](docs/security/MITM-TPROXY-DECRYPT.md)
- **💸 Cost telemetry everywhere** — `X-OmniRoute-*` cost/usage headers on every endpoint, cache-HIT savings header, per-key USD spend quotas. → [API Reference](docs/reference/API_REFERENCE.md)
- **🧠 Memory you control** — off by default, opt-in int8 vector quantization + typed decay, per-request `x-omniroute-no-memory`. → [Memory](docs/frameworks/MEMORY.md)
-- **🛡️ Security** — prompt-injection guard on every LLM route (red-team suite) + free DuckDuckGo last-resort web search. → [Guardrails](docs/security/GUARDRAILS.md)
+- **🛡️ Security** — prompt-injection guard on every LLM route (red-team suite), opt-in credential-masking guardrail (redacts leaked API keys/secrets in both directions), free DuckDuckGo last-resort web search, and an optional OIDC login gate for the dashboard (password login always stays available). → [Guardrails](docs/security/GUARDRAILS.md)
- **🖼️ New endpoints** — `/v1/ocr` (Mistral OCR) and `/v1/audio/translations` (Whisper-style) round out the media surface. → [API Reference](docs/reference/API_REFERENCE.md)
- **🎨 Image / video / audio generation** — one API for media: xAI Grok Imagine & Novita AI video, ComfyUI, Freepik, Adobe Firefly, Microsoft Designer, Google Imagen, Segmind, EdgeTTS. → [API Reference](docs/reference/API_REFERENCE.md)
- **🌍 Deployment & ops** — reverse-proxy `basePath`, browser-language auto-detect, per-key device tracking, root-less MITM trust, zh-TW localization. → [Environment](docs/reference/ENVIRONMENT.md)
-- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI), Ollama first-class card, Claude Sonnet 5, Zed, Requesty, SenseNova, Yuanbao, Agnes AI… and a refreshed **290-provider catalog**. → [Providers](docs/reference/PROVIDER_REFERENCE.md)
+- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI) with browser + OAuth login, Ollama first-class card, Claude Opus 5 & Sonnet 5, Kimi official partnership (Code/Web/Moonshot), Zed, Requesty, SenseNova, Yuanbao, Agnes AI… and a refreshed **290-provider catalog**. → [Providers](docs/reference/PROVIDER_REFERENCE.md)
+- **📡 Routing transparency** — every response carries an `X-OmniRoute-Decision` header naming the strategy/provider/latency that served it, a new `cache-optimized` combo strategy + Auto-Combo `cacheAffinity` factor route repeat requests back to the connection holding the cached prefix, and a read-only `/v1/auto-combo/{channel}/candidates` endpoint exposes an `auto/*` channel's live candidate pool. → [Auto-Combo](docs/routing/AUTO-COMBO.md)
- **⚡ Local performance & infra** — one-click local Redis, Cloudflare Workers / Deno Deploy relay deployers, Bifrost & Mux as supervised embedded services. → [Embedded Services](docs/frameworks/EMBEDDED-SERVICES.md)
diff --git a/docs/architecture/AUTHZ_GUIDE.md b/docs/architecture/AUTHZ_GUIDE.md
index 4f10555901..c12efe8c4c 100644
--- a/docs/architecture/AUTHZ_GUIDE.md
+++ b/docs/architecture/AUTHZ_GUIDE.md
@@ -39,6 +39,32 @@ Verified by `isDashboardSessionAuthenticated()` in `src/shared/utils/apiAuth.ts`
Some management routes accept **either** mode: cookie OR `Bearer ` when the API key has the `manage` (or `admin`) scope. This is what enables the "configurable via API calls" workflow added in v3.8.
+#### Optional OIDC login gate (#6973)
+
+The dashboard admin login also supports an **opt-in** OIDC (OpenID Connect) flow
+alongside the default password login — password login is never removed, only
+supplemented:
+
+- Disabled unless `settings.oidcEnabled === true` **and** `oidcIssuer` /
+ `oidcClientId` / `oidcClientSecret` are all configured (Settings → Auth).
+ `GET /api/auth/oidc/login` returns `400` otherwise.
+- `GET /api/auth/oidc/login` discovers the `authorization_endpoint` from the
+ issuer's `/.well-known/openid-configuration` (falls back to
+ `/authorize`), builds the redirect URI from the incoming request
+ (`x-forwarded-proto`-aware), and redirects to the IdP with a random `state`
+ stored in an `httpOnly` `oidc_state` cookie.
+- `GET /api/auth/oidc/callback` validates `state`, exchanges the authorization
+ code, and verifies the ID token's signature via the issuer's JWKS
+ (`jose`'s `createRemoteJWKSet`, cached per JWKS URI) with `issuer`/`audience`
+ checks. An optional `oidcAllowedSubjects` allowlist matches the token's
+ `sub` claim or its `email` claim — the email claim is only honored when
+ `email_verified === true`, so an unverified email at the IdP can never pass
+ the gate.
+- On success it mints the **exact same** 30-day `auth_token` JWT the password
+ login issues (`src/app/api/auth/login/route.ts`), so the rest of the
+ dashboard session pipeline (auto-refresh, cookie flags) is unchanged —
+ OIDC only replaces how the cookie gets minted, not what it grants.
+
## Route Classes
`src/server/authz/types.ts` defines three classes; any route that cannot be classified deterministically falls back to `MANAGEMENT`.
diff --git a/docs/architecture/RESILIENCE_GUIDE.md b/docs/architecture/RESILIENCE_GUIDE.md
index 5c04f29723..060e93c247 100644
--- a/docs/architecture/RESILIENCE_GUIDE.md
+++ b/docs/architecture/RESILIENCE_GUIDE.md
@@ -262,7 +262,7 @@ it is unit-testable without a real Bottleneck limiter.
## Other Resilience Features
-- **18 routing strategies** (priority, weighted, round-robin, context-relay, fill-first, p2c, random, least-used, cost-optimized, reset-aware, reset-window, headroom, strict-random, auto, lkgp, context-optimized, fusion, pipeline) — see [AUTO-COMBO.md](../routing/AUTO-COMBO.md).
+- **19 routing strategies** (priority, weighted, round-robin, context-relay, fill-first, p2c, random, least-used, cost-optimized, reset-aware, reset-window, headroom, strict-random, auto, lkgp, context-optimized, cache-optimized, fusion, pipeline) — see [AUTO-COMBO.md](../routing/AUTO-COMBO.md).
- **Reset-aware routing** (v3.8.0) — prioritizes connections by quota reset time.
- **Background mode degradation** — Responses API `background: true` degraded to sync with warning.
- **Dynamic tool limit detection** — backs off providers when tool count limits hit.
diff --git a/docs/diagrams/auto-combo-12factor.mmd b/docs/diagrams/auto-combo-12factor.mmd
index 0cff9e8d28..3c7f967534 100644
--- a/docs/diagrams/auto-combo-12factor.mmd
+++ b/docs/diagrams/auto-combo-12factor.mmd
@@ -1,11 +1,11 @@
-%% Auto-Combo 12-factor scoring
+%% Auto-Combo 13-factor scoring
%% Reflects: open-sse/services/autoCombo/scoring.ts (DEFAULT_WEIGHTS, sum = 1.0)
-%% v3.8.40
+%% v3.8.49
flowchart TB
Request["Incoming request"] --> Candidates["Eligible candidates
(provider × model × account)"]
Candidates --> Score["Compute composite score
per candidate"]
- subgraph Factors["12-factor scoring weights (sum = 1.0)"]
+ subgraph Factors["13-factor scoring weights (sum = 1.0)"]
f1["health (0.20)"]
f2["quota (0.15)"]
f3["costInv (0.15)"]
@@ -17,7 +17,8 @@ flowchart TB
f9["specificityMatch (0.05)"]
f10["contextAffinity (0.05)"]
f11["connectionDensity (0.05)"]
- f12["resetWindowAffinity (0.00)"]
+ f12["cacheAffinity (0.00)"]
+ f13["resetWindowAffinity (0.00)"]
end
Score --> Factors
diff --git a/docs/reference/API_REFERENCE.md b/docs/reference/API_REFERENCE.md
index fbf41ea286..3f3e6467a0 100644
--- a/docs/reference/API_REFERENCE.md
+++ b/docs/reference/API_REFERENCE.md
@@ -532,6 +532,8 @@ Response example:
| `/api/usage/request-logs` | GET | Request-level logs |
| `/api/usage/[connectionId]` | GET | Per-connection usage |
| `/api/usage/token-limits` | GET/POST/DELETE | Per-API-key token-limit budgets |
+| `/api/usage/model-latency-stats` | GET | Rolling per-provider/model latency aggregate (avg/p50/p95/p99, success rate); filters: `windowHours`/`minSamples`/`maxRows`/`provider`/`model` (#6873) |
+| `/api/usage/cache-health` | GET | Prompt-cache health summary over `call_logs` — write/read ratio, p50/p90/p99 write-size distribution, heavy-write concentration, per-model split, and a `healthy`/`degraded`/`thrash`/`no-data` verdict; query params `range` (`1h`\|`24h`\|`7d`\|`30d`, default `24h`) and optional `model` (#8827) |
### Settings
diff --git a/docs/reference/PROVIDER_REFERENCE.md b/docs/reference/PROVIDER_REFERENCE.md
index 6e25e6b8e0..fa597a551d 100644
--- a/docs/reference/PROVIDER_REFERENCE.md
+++ b/docs/reference/PROVIDER_REFERENCE.md
@@ -1,14 +1,14 @@
---
title: "Provider Reference"
version: 3.8.49
-lastUpdated: 2026-07-23
+lastUpdated: 2026-07-28
---
# Provider Reference
> **Auto-generated** from `src/shared/constants/providers.ts` — do not edit by hand.
> Regenerate with: `npm run gen:provider-reference`
-> **Last generated:** 2026-07-23
+> **Last generated:** 2026-07-28
Total providers: **290**. See category breakdown below.
@@ -49,7 +49,7 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each
| `devin-cli` | `dv` | Devin CLI (Official) | OAuth | [link](https://cli.devin.ai) | Requires the Devin CLI binary. Run `devin auth login` to authenticate, or provide your WINDSURF_API_KEY. Install: https://cli.devin.ai |
| `ghe-copilot` | `ghe-copilot` | GitHub Enterprise Copilot | OAuth | — | Enter your GHE instance URL (e.g., https://ghe.company.com) in provider settings, then authenticate via device flow. |
| `github` | `gh` | GitHub Copilot | OAuth | — | — |
-| `gitlab-duo` | `gitlab-duo` | GitLab Duo | OAuth | [link](https://docs.gitlab.com/user/duo_agent_platform/code_suggestions/) | OAuth application with ai_features + read_user scopes. Configure GITLAB_DUO_OAUTH_CLIENT_ID and optionally GITLAB_DUO_OAUTH_CLIENT_SECRET on this OmniRoute instance. |
+| `gitlab-duo` | `gitlab-duo` | GitLab Duo | OAuth | [link](https://docs.gitlab.com/user/duo_agent_platform/code_suggestions/) | GitLab Duo OAuth is not configured. Register an OAuth application at https://gitlab.com/-/profile/applications with redirect URI http://localhost:20128/callback and scopes "ai_features read_user", then set GITLAB_DUO_OAUTH_CLIENT_ID (and optionally GITLAB_DUO_OAUTH_CLIENT_SECRET) and restart. |
| `grok-cli` | `gc` | Grok Build | OAuth | — | Sign in with your browser, or paste your ~/.grok/auth.json (or the JWT access token) from the Grok Build CLI; refresh_token is rotated automatically either way. |
| `kilocode` | `kc` | Kilo Code | OAuth | — | — |
| `kimi-coding` | `kmc` | Kimi Code CLI | OAuth | [link](https://www.kimi.com/code?aff=omniroute) | Sign in with the same Kimi account used by Kimi Code CLI. OmniRoute uses the CLI OAuth flow and Kimi Coding Plan endpoints. |
@@ -155,7 +155,6 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each
| `fal-ai` | `fal` | Fal.ai | API key, image | [link](https://fal.ai) | — |
| `featherless-ai` | `featherless` | Featherless AI | API key | [link](https://featherless.ai) | Free tier available — no credit card required |
| `fenayai` | `fenayai` | FenayAI | API key, aggregator | [link](https://fenayai.com) | Bearer API key for the FenayAI OpenAI-compatible gateway. |
-| `firecrawl` | `fc` | Firecrawl | API key | [link](https://firecrawl.dev) | — |
| `fireworks` | `fireworks` | Fireworks AI | API key | [link](https://fireworks.ai) | $1 free starter credits on signup for API testing |
| `freeaiapikey` | `faik` | FreeAIAPIKey | API key | [link](https://freeaiapikey.com) | — |
| `freemodel-dev` | `fmd` | FreeModel.dev | API key | [link](https://freemodel.dev) | $300 free credits on signup — no credit card required. Access GPT-5.4 and GPT-5.5 (OpenAI's latest flagship models) through an OpenAI-compatible API. |
@@ -293,6 +292,7 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each
| `x5lab` | `x5lab` | X5Lab | API key | [link](https://x5lab.dev) | Use your X5Lab API key (x5-...) in Authorization: Bearer . Fully OpenAI-compatible. API base URL: https://api.x5lab.dev/v1. |
| `xai` | `xai` | xAI (Grok) | API key | [link](https://x.ai) | — |
| `xiaomi-mimo` | `mimo` | Xiaomi MiMo | API key | [link](https://mimo.mi.com) | — |
+| `xiaomi-mimo-token-plan` | `mimotp` | Xiaomi MiMo Token Plan | API key | [link](https://mimo.mi.com) | — |
| `yi` | `yi` | Yi (01.AI) | API key | [link](https://01.ai) | Get API key at platform.lingyiwanwu.com |
| `zai` | `zai` | Z.AI | API key | [link](https://open.bigmodel.cn) | — |
| `zenmux` | `zm` | ZenMux | API key | [link](https://zenmux.ai) | Use your ZenMux API key in Authorization: Bearer . ZenMux is fully OpenAI-compatible. Base URL: https://zenmux.ai/api/v1. |
@@ -314,12 +314,13 @@ Use the dashboard at `/dashboard/providers` to enable, configure, and test each
| `vllm` | `vllm` | vLLM | Local, self-hosted | [link](https://github.com/vllm-project/vllm) | API key optional. Configure the local vLLM OpenAI-compatible base URL (default: http://localhost:8000/v1). |
| `xinference` | `xinference` | XInference | Local, self-hosted | [link](https://inference.readthedocs.io) | API key optional. Configure the local XInference OpenAI-compatible base URL (default: http://localhost:9997/v1). |
-## Search Providers (11)
+## Search Providers (12)
| ID | Alias | Name | Tags | Website | Notes |
|----|-------|------|------|---------|-------|
| `brave-search` | `brave-search` | Brave Search | Search | [link](https://brave.com/search/api) | Subscription token from Brave Search API dashboard |
| `exa-search` | `exa-search` | Exa Search | Search | [link](https://exa.ai) | API key from dashboard.exa.ai |
+| `firecrawl` | `fc` | Firecrawl | Search | [link](https://firecrawl.dev) | — |
| `google-pse-search` | `google-pse` | Google Programmable Search | Search | [link](https://developers.google.com/custom-search/v1/overview) | Requires a Google API key and your Programmable Search Engine ID (cx) |
| `linkup-search` | `linkup` | Linkup Search | Search | [link](https://docs.linkup.so) | Bearer API key from the Linkup dashboard |
| `ollama-search` | `ollama-search` | Ollama Search | Search | [link](https://ollama.com/settings/keys) | Same API key as Ollama Cloud (from ollama.com/settings/keys) |
diff --git a/docs/routing/AUTO-COMBO.md b/docs/routing/AUTO-COMBO.md
index 424926507a..948cd255fa 100644
--- a/docs/routing/AUTO-COMBO.md
+++ b/docs/routing/AUTO-COMBO.md
@@ -161,11 +161,11 @@ The detection helper lives in `src/lib/combos/modelNameCollision.ts`.
## How It Works (Persisted Auto-Combos)
-The Auto-Combo Engine dynamically selects the best provider/model for each request using a **12-factor scoring function** (defined in `open-sse/services/autoCombo/scoring.ts` → `DEFAULT_WEIGHTS`). All weights sum to **1.0**.
+The Auto-Combo Engine dynamically selects the best provider/model for each request using a **13-factor scoring function** (defined in `open-sse/services/autoCombo/scoring.ts` → `DEFAULT_WEIGHTS`). All weights sum to **1.0**.

-> Source: [diagrams/auto-combo-12factor.mmd](../diagrams/auto-combo-12factor.mmd) (regenerate via `npm run docs:render-diagrams`).
+> Source: [diagrams/auto-combo-12factor.mmd](../diagrams/auto-combo-12factor.mmd) (regenerate via `npm run docs:render-diagrams`). Diagram/filename predate the `cacheAffinity` factor added by #8008 and still show 12 factors.
| Factor | Default Weight | Description |
| :-------------------- | :------------- | :------------------------------------------------------------------------------------------------- |
@@ -180,9 +180,10 @@ The Auto-Combo Engine dynamically selects the best provider/model for each reque
| `specificityMatch` | 0.05 | Match between request specificity (manifest hint) and model tier |
| `contextAffinity` | 0.05 | Affinity between the request's context-window need and the model's context window |
| `connectionDensity` | 0.05 | Spreads load across connections of the same provider (anti-concentration) |
+| `cacheAffinity` | 0.00 | Rendezvous-hash affinity toward the connection likeliest to already hold this request's prompt-cache prefix (`open-sse/services/combo/promptCacheAffinity.ts`); disabled by default (#8008) |
| `resetWindowAffinity` | 0.00 | Bias toward connections whose quota reset window is favorable (disabled by default) |
-**Sum:** `0.20 + 0.15 + 0.15 + 0.12 + 0.08 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.00 = 1.0` (validated by `validateWeights()`).
+**Sum:** `0.20 + 0.15 + 0.15 + 0.12 + 0.08 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.05 + 0.00 + 0.00 = 1.0` (validated by `validateWeights()`).
## Mode Packs
@@ -237,7 +238,7 @@ resolved values feed the engine's existing `config.modePack` / `config.budgetCap
## All Routing Strategies
-OmniRoute's combo engine supports **18 routing strategies** (declared in `src/shared/constants/routingStrategies.ts` → `ROUTING_STRATEGY_VALUES`). The Auto Combo engine itself is exposed under the `auto` strategy; the others are available for persisted combos.
+OmniRoute's combo engine supports **19 routing strategies** (declared in `src/shared/constants/routingStrategies.ts` → `ROUTING_STRATEGY_VALUES`). The Auto Combo engine itself is exposed under the `auto` strategy; the others are available for persisted combos.
| Strategy | Description |
| :------------------ | :--------------------------------------------------------------------------------------------------------------------------- |
@@ -257,6 +258,7 @@ OmniRoute's combo engine supports **18 routing strategies** (declared in `src/sh
| `auto` | Use Auto Combo scoring (9-factor) — **recommended** |
| `lkgp` | Last-Known-Good Path (sticky route to last successful target) |
| `context-optimized` | Pick target with best fit for current context size |
+| `cache-optimized` | Reorder targets by prompt-cache affinity — the connection likeliest to already hold this request's cached prefix is tried first (`open-sse/services/combo/promptCacheAffinity.ts`, #8008) |
| `fusion` 🧬 | Fan out to a panel of models in parallel, then synthesize one answer via a judge (see below) |
| `pipeline` | Run targets sequentially, threading each step's output into the next step's input; only the final answer is returned (#6396) |
@@ -715,5 +717,5 @@ intentionally excluded from CI because they require live credentials and VPS acc
| `open-sse/services/autoCombo/autoPrefix.ts` | `auto/` prefix parser + 6 variants |
| `open-sse/services/autoCombo/virtualFactory.ts` | Builds in-memory `AutoComboConfig` from live connections |
| `open-sse/services/autoCombo/providerRegistryAccessor.ts` | Test hook for mocking provider registry |
-| `src/shared/constants/routingStrategies.ts` | `ROUTING_STRATEGY_VALUES` (18 strategies) |
+| `src/shared/constants/routingStrategies.ts` | `ROUTING_STRATEGY_VALUES` (19 strategies) |
| `src/sse/handlers/chat.ts` | Integration: auto-prefix short-circuit |
diff --git a/docs/security/GUARDRAILS.md b/docs/security/GUARDRAILS.md
index 9653207129..68252680c4 100644
--- a/docs/security/GUARDRAILS.md
+++ b/docs/security/GUARDRAILS.md
@@ -20,14 +20,15 @@ request. Blocking is an explicit decision (`block: true`), never an accident.
## Built-in Guardrails
-The registry auto-loads three guardrails in priority order on import
+The registry auto-loads four guardrails in priority order on import
(see `registry.ts` → `registerDefaultGuardrails()`):
-| Priority | Name | Stage(s) | File |
-| -------- | ------------------ | -------------- | -------------------- |
-| `5` | `vision-bridge` | `preCall` | `visionBridge.ts` |
-| `10` | `pii-masker` | `pre` + `post` | `piiMasker.ts` |
-| `20` | `prompt-injection` | `preCall` | `promptInjection.ts` |
+| Priority | Name | Stage(s) | File |
+| -------- | -------------------- | -------------- | --------------------- |
+| `5` | `vision-bridge` | `preCall` | `visionBridge.ts` |
+| `10` | `pii-masker` | `pre` + `post` | `piiMasker.ts` |
+| `20` | `prompt-injection` | `preCall` | `promptInjection.ts` |
+| `95` | `credential-masker` | `pre` + `post` | `credentialMasker.ts` |
Lower priority numbers run **first**.
@@ -118,6 +119,36 @@ the pattern loop. Injection directives sit near the top of an input, so this
caps regex CPU/GC on multi-hundred-KB payloads without weakening detection (cf.
#3932, #4041).
+### Credential Masker (`credentialMasker.ts`)
+
+Runs on **both** stages, last in the default chain (priority `95`). Redacts
+well-known API-key / secret-token patterns from the outbound payload (message
+content, tool-call arguments, tool results) **and** the provider response, so a
+credential pasted into a prompt (or echoed back by a tool result) is not leaked
+to the upstream provider or back to the client.
+
+- **Opt-in only**, same convention as PII redaction (Hard Rule #20-adjacent):
+ disabled unless `settings.credentialRedactionEnabled === true` **or**
+ `CREDENTIAL_REDACTION_ENABLED=true`. With it off, the guardrail is a no-op —
+ it never blocks and never rewrites.
+- `redactCredentials()` walks the full payload/response tree (`walkValue()`,
+ prototype-pollution-safe, cycle-safe via `WeakSet`) and replaces matches with
+ a `[REDACTED:]` placeholder, cloning only the branches that actually
+ changed.
+- `CREDENTIAL_PATTERNS` covers LLM provider keys (OpenAI, OpenAI-proj,
+ Anthropic, Google, Hugging Face, Replicate), VCS/SaaS tokens (GitHub, Slack,
+ Linear, Notion, npm, Postman, Discord), payment keys (Stripe, Square), cloud
+ keys (AWS access key, Twilio, SendGrid, Mailgun), private keys / JWTs,
+ credential-bearing connection strings (`mongodb://user:pass@...`, etc.), and
+ a generic `Authorization`/`x-api-key`/`api-key`/`apikey` header-value
+ pattern. Header-shaped keys (`authorization`, `x-api-key`, `api-key`,
+ `apikey`) are redacted structurally (value only, scheme prefix like
+ `Bearer `/`Basic ` preserved) rather than via the generic text regex.
+- The guardrail never blocks; it only rewrites (`modifiedPayload` /
+ `modifiedResponse`) and annotates (`meta.credentialsRedacted`, `meta.count`).
+
+Regression guard: `tests/unit/credential-masker-guardrail.test.ts`.
+
## Base Contract (`base.ts`)
```typescript