Files
sim/apps
Waleed ea4f2691e3 fix(providers): correct max-tokens param and add schema guidance for NVIDIA/Z.ai (#5569)
* fix(providers): correct max-tokens param and add schema guidance for NVIDIA/Z.ai

- NVIDIA NIM and Z.ai both document max_tokens for output-length control,
  not OpenAI's newer max_completion_tokens - the latter was silently
  ignored by both vLLM-served NIM models and Z.ai's GLM models
- Z.ai has no json_schema response_format mode (only text/json_object),
  so structured-output requests now also inject the expected schema into
  the system prompt as best-effort guidance, since the request param
  alone can't enforce field names/types

* style(providers): replace inline comments with TSDoc, per CLAUDE.md

Consolidated the scattered narrative // comments in nvidia/index.ts and
zai/index.ts into a single TSDoc block per provider documenting the
provider-specific API quirks; removed the rest where they only restated
what the code already shows. Also dropped an inline comment on zai's
modelPatterns field in models.ts.

* fix(providers): scope Z.ai schema guidance to the response_format call only

- schemaGuidance now falls back to the bare responseFormat object when
  .schema is absent, matching the schema-or-format fallback used
  elsewhere in the codebase (was silently injecting nothing for callers
  that pass a bare JSON schema)
- guidance is now only appended to the messages sent alongside an
  actual response_format (the immediate call, or whichever pass
  deferResponseFormat applies it to) instead of every turn of an
  active tool loop, where it wrongly told the model to return final
  JSON instead of continuing to call tools
2026-07-10 14:20:57 -07:00
..