Promote the max_tokens fallback in BaseGalaxyAgent from a hardcoded
literal to a class attribute (DEFAULT_MAX_TOKENS = 8192) so individual
agents can declare their own ceiling without overriding the whole
_get_max_tokens method. The history, orchestrator, and custom_tool
agents now bump theirs to 16384 -- those are the ones that produce
long structured output (history listings, multi-step plans, full
UserToolSource JSON schemas) and were the most likely to silently
truncate. Router, tool recommendation, and error_analysis stay on the
8k default; their typical responses sit comfortably under that.
inference_services.<agent>.max_tokens config still wins over the
class default.
Also refreshed the example configs in doc/source/admin/ai_agents.md
so they no longer show 2000 as a representative override value, and
added a parenthetical to the param description noting the new
defaults.
The 2000-token default in _get_max_tokens() was causing the history
agent to fail with "Model token limit (2000) exceeded before any
response was generated" on histories with more than a handful of
datasets. The cap was an artifact of an earlier era -- every backend
we currently support (Maverick / Llama-3.3-70B at 131k context, Qwen3
at 32k, gpt-oss-120b at ~128k, plus the Anthropic and OpenAI
providers) handles 8k output comfortably with plenty of headroom for
the prompt. 8k is high enough that the common agents (history,
error_analysis, orchestrator) don't truncate mid-answer, but still
well under the smallest window we point at (Qwen3's 32k) so a runaway
response can't blow up the total context. Per-agent overrides via
inference_services.<agent>.max_tokens still work the same way; this
only changes the fallback when nothing is configured.
The browser's native EventSource auto-retry gives up once readyState flips to
CLOSED — typical for a 4xx/5xx response with no text/event-stream body — and
the previous onerror handler did not detect that, leaving the client stranded
on the polling fallback for the rest of the session.
Take ownership of the reconnect loop on the client: close the source on
onerror+CLOSED, schedule a reopen with full-jitter exponential backoff capped
at 30 s, and reset the counter on a successful onopen. The replay-on-open
viewer-subscription path is unchanged.
Tested with a Playwright-only integration_selenium test that fails the first
SSE request with a 503 via page.route() and asserts the client reconnects and
delivers a subsequently-pushed notification.
WorkflowIndexPayload and InvocationIndexPayload were defined in the
webapps service layer, which forced operations.py (in galaxy-app) to
reach across the package boundary into galaxy-webapps just to
construct them. They're plain Pydantic models that extend a base
already living in galaxy.schema.schema, so move them next to their
parents. Services keep the same names via re-export so external
consumers don't break, and operations.py can now import them at
module level alongside the rest of the galaxy.schema imports.
Per Marius' review on PR 22625 -- Galaxy only uses local imports when
there's a real reason (circular dep, optional dep). None of these had
one, so they belong at the top of the file alongside the rest.
test_mcp_run_user_tool was building its history and input dataset with
the default test interactor, while the MCP-side run authenticated as a
freshly-provisioned UDT user -- so the MCP call was reaching for
resources owned by a different user. Switch to
DatasetPopulator(self._get_interactor(api_key=api_key)) so populator-
side and MCP-side calls share an identity. Also assert on the run-tool
response shape (jobs[].tool_id, outputs[].output_name) in addition to
the end-to-end content check.
test_mcp_delete_user_tool now also tries to run the UDT after
deletion and asserts the call errors with "deactivated" -- regression
test for the deactivation guard in run_user_tool.
deactivate_unprivileged_tool deliberately only flips the per-user
UserDynamicToolAssociation.active flag, leaving DynamicTool.active
intact so other users with associations to the same DynamicTool aren't
affected (the model schema permits many-to-many, even though the
current create path is 1:1). That means a user who deactivates "their"
UDT can still resolve it by UUID through the toolbox -- and run it via
tools_service._create -- because get_unprivileged_tool_by_uuid doesn't
filter by association.active either.
Add a runtime preflight in run_user_tool that fails the call when
either the underlying tool or the calling user's association is
inactive. Also surfaces unauthenticated and unowned errors as clean
ValueErrors before reaching the deeper service layer.
Tightening the chokepoint (DynamicToolManager.get_unprivileged_tool_by_uuid)
to filter by association.active would close this across all entry points
but is a meaningful behavior change for the existing UnprivilegedToolsApi
endpoints; leaving that for a separate review.
Mirrors run_tool but passes tool_uuid in the payload, which
tools_service._create routes to the toolbox's unprivileged-tool
resolver. Closes UDT parity with the standalone galaxy-mcp server.
Validates the representation through DynamicUnprivilegedToolCreatePayload
and delegates to DynamicToolManager.create_unprivileged_tool, mirroring
the POST /api/unprivileged_tools endpoint. The MCP wrapper carries the
full GalaxyUserTool schema in its docstring (required fields, the
common container-as-string mistake, a worked example) so agents can
construct valid representations without round-tripping.
Wraps DynamicToolManager.list_unprivileged_tools so the in-process MCP
server can enumerate a user's user-defined tools, matching the parity
gap with the standalone galaxy-mcp server. Also wires up the lazy
DynamicToolManager property on AgentOperationsManager and a
_setup_udt_user helper on the smoke test class for the rest of the UDT
work.
GCard's responsive width rules use `@container cards-list (max-width: ...)`
to drop cards from 1/3 width to 1/2 / full width on narrow viewports.
RecentDownloads.vue had no `container: cards-list / inline-size`
declaration anywhere up the tree, so those queries never fired and
cards stayed at `calc(100% / 3)` regardless of viewport width — which,
combined with the recently fixed title squashing, was producing
per-syllable title wrapping in the export card on narrow screens.
Match the idiom used by WorkflowCardList and HistoryCardList.
Add responsive layout to GCard header that prevents badges from
squashing the title. Title section now has min-width of 50% and
badges wrap to next line when needed.
- Replace global _user_counter with uuid4().hex[:8] for xdist safety
- Drop test_determinism_identical_requests (duplicate of test_deterministic_ordering)
- Lift inline imports (sqlalchemy event, HistoryGraphBuilder) to top
- Lower scale defaults (500/100/10/50 -> 250/60/5/20)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- direction: Literal["backward","forward","both"] in service and manager
- HistoriesService.graph annotated -> HistoryGraphResponse
- Pin 404 for missing seed_scope and missing history (was loose 4xx)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>