These are available in a list(n) x list (m) -> list(nxm) version (a cross product that produces two flat lists) and a list(n) x list(m) -> list(n):list(m) version (a cross product that produces two nested lists).
After two lists have been run through one of these two tools - the result is two new lists that can be passed into another tool to perform all-against-all operations using Galaxy's normal collection mapping semantics.
The choice of which to use will depend on how you want to continue to process the all-against-all results after the next step in an analysis. My sense is the flat version is "easier" to think about and pick through manually and the nested version perserves more structure if additional collection operation tools will be used to filter or aggregate the results.
Some considerations:
Apply Rules?
I do not believe the Apply Rules tool semanatics would allow these operations but certainly the Apply Rules tool could be used to convert the result of the flat version to the nested version or vice versa - so no metadata is really lost per se between the two versions. I think it is still worth including both versions though - they both have utility (both for instance are baked into CWL's workflow semantics - https://docs.sevenbridges.com/docs/about-parallelizing-tool-executions#nested-cross-product) and avoiding requiring complex Apply Rules programs for simple workflows is probably ideal.
One Tool vs Two?
Marius and I agree that few simpler tools for these kinds of operations are better. The tool help can be more focused and avoiding the conditional and conditional outputs make the static analysis done for instance by the workflow editor simpler.
Updating regular tool data tables for frequently changing data is a
little tricky, and it was requested that we allow loading options
dynamically from a URL. The eventual goals of this work is to enable
exploring resources with great APIs (NCBI datasets is a good exmaple)
from within Galaxy tools.
This doesn't do any of that yet ... it's just a simple URL one can
query with a GET request, and the data there must be in the format
["name", "value", True|False].
Following steps:
- [x] allow templating (Cheetah ? JS?)
- [x] allow post-processing to extract name-value pairs
- jq-like syntax ?
- [x] JS expression ?
- add pagination (requires work on vue-multiselect)
- add request caching
Some potential issues:
- Slow server responses: use async library and cancel request after N
seconds ?
- Large respons ... stream response and cancel after configurable max
response size ?
Works for both files and collections. Workflow defaults override tool defaults.
TODO:
- Unit test case to ensure this only works for non-default, non-multi data parameters.
- Implement XSD once syntax is finalized.