Enable container recommendation in the custom-tool eval deps

make_deps now sets, on the inference_services.custom_tool block:

- container_recommendation_enabled: so the custom_tool eval exercises the full
  container-resolution path (dedicated container critic + quay.io recommender +
  deterministic override) rather than only the model's raw container choice.
- max_tokens (16384, mirroring the agent's DEFAULT_MAX_TOKENS): the producer emits
  a full tool definition (inputs/outputs/configfiles) that overruns the 4000-token
  default block and would otherwise fail with a token-limit error before any
  response is generated.

Config resolves per-key, so model/api still come from the default block. This adds
an extra model call plus outbound quay.io lookups to custom_tool cases.
This commit is contained in:
mvdbeek
2026-07-07 16:39:17 +02:00
parent 01b43065f5
commit 87f0cbc83e
+15 -1
View File
@@ -76,6 +76,13 @@ def make_deps(
Routes all agents through `inference_services.default`, so the router
handoff targets resolve via the registry without needing a live Galaxy.
The ``custom_tool`` block additionally turns on ``container_recommendation_enabled``
so the custom-tool eval exercises the full container-resolution path -- the
dedicated container critic, the quay.io recommender, and the deterministic
override. Config is resolved per-key (agent block then ``default``), so only the
flag is set here; model/api/temperature still come from ``default``. This adds an
extra model call plus outbound quay.io lookups to custom_tool cases.
"""
config = MagicMock()
config.ai_api_key = None
@@ -88,7 +95,14 @@ def make_deps(
"api_base_url": base_url,
"temperature": temperature,
"max_tokens": max_tokens,
}
},
"custom_tool": {
"container_recommendation_enabled": True,
# The producer emits a full tool definition (inputs/outputs/configfiles),
# which overruns the 4000-token default block -- mirror the agent's own
# DEFAULT_MAX_TOKENS so generation isn't truncated in the eval.
"max_tokens": 16384,
},
}
# MagicMock for trans/user so handoff targets that touch deps.trans.app
# (history, error_analysis, tools) don't crash before the model returns.