mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-08-06 23:32:12 +03:00
Z.AI's glm-4.6v vision endpoint enforces a 32768 max_tokens ceiling server-side and 400s when a client sends a larger explicit max_tokens (e.g. a client defaulting to 65536). paramSupport.ts's STRIP_RULES already has a working clampToModelMaxOutput/maxOutputCap mechanism (used today for a VolcEngine Kimi rule) but had no entry for zai/glm + glm-4.6v. Added two rules: "zai" uses a fixed maxOutputCap (glm-4.6v is only reachable there as a custom model attached to the connection, so it is not in PROVIDER_MODELS["zai"] and clampToModelMaxOutput would find no catalog ceiling); "glm" uses clampToModelMaxOutput (glm-4.6v IS in the registry catalog there, GLM_SHARED_MODELS, maxOutputTokens: 32768). Also discovered and fixed a second, deeper bug the "glm" rule alone would not have caught: GlmExecutor.execute() drives its own fetch flow (executeTransport()/transformForTransport()) and never runs through DefaultExecutor.execute()'s stripUnsupportedParams() call site — so a STRIP_RULES clamp entry for provider "glm" was dead code until transformForTransport() now calls stripUnsupportedParams() directly. Regression tests: tests/unit/zai-glm-max-tokens-clamp-7364.test.ts (reused from the triage plan-file's RED probe, sanity assertion updated to lock the fix instead of the bug) and tests/unit/glm-executor-max-tokens-clamp-7364.test.ts (proves the real GlmExecutor.transformForTransport wiring, not just the STRIP_RULES entry in isolation). Gates run: check-file-size, check-complexity, check-cognitive-complexity, typecheck:core, eslint (suppressions), tests/unit/zai-glm-max-tokens-clamp-7364.test.ts, tests/unit/glm-executor-max-tokens-clamp-7364.test.ts, tests/unit/executors-strip-unsupported-params.test.ts, tests/unit/nvidia-minimax-thinking-strip.test.ts, tests/unit/glm-executor.test.ts — all green. Refs #7364