mirror of
https://github.com/diegosouzapw/OmniRoute.git
synced 2026-09-13 18:32:12 +03:00
* fix(sse): trust finish_reason:length/max_tokens over the reasoning-ratio heuristic in response quality validation A truncated response with empty content and reasoning_content present was only rejected by validateResponseQuality() when reasoning consumed >=90% of completion_tokens. A response truncated at a lower ratio (e.g. 63%) passed through as "valid" even though the caller received no usable content and finish_reason was explicitly "length" (or the alternate "max_tokens" naming some providers use) -- an unambiguous truncation signal the validator wasn't reading. Reproduced live against nvidia/nemotron-3-super-120b-a12b: content:null, finish_reason:length, reasoning_tokens 645/1024 (63%). Trust finish_reason directly when it's reported, falling back to the existing token-ratio heuristic only when it isn't. Does not affect the deliberate-tiny-probe case (e.g. max_tokens:1 connectivity pings) -- those never produce reasoning_content, so the branch this change is in doesn't run for them. * docs(changelog): add fragment for #12262 --------- Co-authored-by: brick30llc-ctrl <admin@brick30.com>