From fd487ce594f73f190a5c26bceb90363a5bc9571a Mon Sep 17 00:00:00 2001 From: Mihaly Bodo Date: Tue, 11 Aug 2026 13:00:00 +0200 Subject: [PATCH] fix(sse): apply Azure param rules on azure-ai and clamp gpt-4o-mini output tokens (#9787) * fix(sse): apply Azure request-param rules on the azure-ai wire path Azure rejects several stock Chat Completions params on its newer deployments and returns HTTP 400 rather than ignoring them: max_tokens -> 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead. reasoning_effort -> Function tools with reasoning_effort are not supported. Those rules lived inline in AzureOpenAIExecutor, so they only covered the azure-openai provider. azure-ai (Azure AI Foundry) had no executor entry and fell through to the bare DefaultExecutor, so the SAME Azure deployment succeeded on one connection and 400'd on the other. Every agentic client sends tools on every turn, so azure-ai failed on the first request. Extract the rules to open-sse/executors/azureParamRules.ts, add an AzureAiExecutor that inherits DefaultExecutor's azure-ai URL/header/apiType handling unchanged and applies the shared rules, and register it for azure-ai. Also widen the deployment pattern to cover gpt-chat-latest: it is a moving alias that resolves to a GPT-5-era model and rejects max_tokens, but carries no version number for the token-boundary pattern to key on. Verified against the base regex - gpt-chat-latest did not match, which is exactly the observed 400. Regression guard: tests/unit/azure-param-rules.test.ts, including an assertion that getExecutor("azure-ai") no longer resolves to a bare DefaultExecutor. * fix(sse): clamp Azure gpt-4o-mini completion tokens to its 16384 ceiling Azure gpt-4o-mini deployments accept at most 16384 completion tokens and 400 on anything larger: max_tokens is too large: 32000. This model supports at most 16384 completion tokens, whereas you provided 32000. The 32000 is OmniRoute's own doing: adjustMaxTokens raises any smaller max_tokens to DEFAULT_MIN_TOKENS (32000) whenever tools are present, to avoid truncated tool arguments. That floor has no upper bound, so an agentic client asking for far less still trips the model ceiling on its first turn. Add scoped maxOutputCap rules in paramSupport.ts for both Azure wire paths. PROVIDER_MAX_TOKENS is the wrong lever here - it is provider-wide, and the same Azure resource also serves GPT-5 deployments with a much higher ceiling. Regression guard: tests/unit/azure-max-output-clamp.test.ts, which also pins that the clamp does not leak to gpt-5.1 or to gpt-4o-mini on other providers.