Allow multiple collections to be fed to a multi data parameter in one reduction step. Fixes#750 and will really simplify certain classes of tools.
Rebased original with fixes for rerun of such reductions.
Manually tested workflow execution and everything seems fine. The workflow editor already thought this was possible, so that is another bug corrected by this enhancement.
To run the associated API test, execute the following command:
./run_tests.sh -with_framework_test_tools -api test/api/test_tools.py:ToolsTestCase.test_reduce_multiple_lists_on_multi_data
Conflicts:
static/maps/mvc/dataset/dataset-choice.js.map
static/maps/mvc/form/form-select-content.js.map
static/scripts/mvc/form/form-select-content.js
By default if single data parameters are mapped over by collections, the collections are matched up (e.g. for two collections the items are paired off and two collections of dimension (N) will result in running N jobs and producing an (N) dimensional output for each define output). In the parameter meta value wrapper (where batch mode can be set with the ``batch`` flag), the linked flag is now respected for collection operations. If ``linked`` is ``False``, then the cross product of the inputs will be used to map over the tool. In the above simplest case, this would cause an NxN (``list:list``) collection to be created for each tool output.
This operation was previously available for individual datasets, but when supplied a collection it would not create an implicit collection pulling together all the relevant datasets with the correct structure - it would just run the jobs and leave the datasets uncollected. Now a collection with the correct structure and identifiers is created.
Limitations:
-----------------
This does not enable tool form support but the API for collections matches that for doing cross product operations over sets of individual datasets, so once support is added to the tool form for that collection support should be trivial.
Testing:
-----------------
The following test case has been extended to now ensure that an implicit collection is created and that it has the right dimensionality/structure and contents.
./run_tests.sh -with_framework_test_tools -api test/api/test_tools.py:ToolsTestCase.test_map_over_two_collections_unlinked
- Add example tool demonstrating/testing specifing format via conditional output actions.
- Add API test testing mapping collections over tools without output action formatting.
- Add API test testing more complex actions using the Cut1 tool.
Tools may now use $input.element_identifier during tool evalution for input 'data' parameters with the following semantics:
- If the input was specified as a single dataset by the user - this just fallbacks to providing the $input.name.
- If the input was mapped over a collection (to produce many jobs) or if the input is a 'multiple="true"' input that was provided a collection - the $input.element_identifier will be the element identifier for the corresponding collection item (generally much more useful the dataset name - since if preserved throughout workflows).
'data_collection' parameters already can access this kind of information - but it is something of a best practice to use simple 'data' parameters since they are compatible with more traditional un-collected datasets.
This commit really needs more comments - but Philip Mabon has been patiently waiting for this functionality for a long time.
Normal selects seem to be prevented from execution with invalid parameter values, but not columns. Values are escaped properly so shell exploitation isn't the problem - but as a usability thing Galaxy should prevent execution and provide a warning message.
Models:
Track whether dataset collections have been populated yet.
Dataset collections are still effectively immutable once populated - but dynamic output collections require them to be sort of like `final` fields in Java (analogy courtesy of JJ) - allowing them to be declared before they are initialized or populated. This is tracked by the `populated_state` field.
Tools:
Output collections can now describe `discover_datasets` elements just like datasets - except in this case instead of dynamically populating new datasets in the history - they will comprise the collection. `designation` has been reused to serve as the element_identifier for the collection element corresponding to the dataset.
See Pull Request 356 for more information on the discover_datasets tag https://bitbucket.org/galaxy/galaxy-central/pull-request/356/enhancements-for-runtime-discovered.
Workflows:
Update workflow execution and recovery for dynamic output collections.
Galaxy workflow data flow before collections
* - * - * - * - * - *
Galaxy worfklow data flow after collections (iteration 1)
* - * - * \
* - * - *
* - * - * / \
* - * - *
* - * - * \ /
* - * - *
* - * - * /
Galaxy worfklow data flow after this commit
/ * - * \
* - * * - *
/ \ * - * / \
/ \
/ \
/ / * - * \ \
* - * -- * - * * - * -- * - *
\ \ * - * / /
\ /
\ /
\ / * - * \ /
* - * * - *
\ * - * /
By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)
There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.
Model:
The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.
Workflow:
Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.
Tool Testing:
This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.
Tests:
Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.
Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.
How to use:
1.) Place multiple tools with different IDs in your tool conf.
2.) ... ummm ... no step 2 - just use the tools.
Implementation:
The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.
To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.
Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
./run_tests.sh -with_framework_test_tools -api test/api/test_tools.py:ToolsTestCase.test_multidata_param
It was what was requested, but I am not sure I love this behavior - seems like for consistency that should maybe be a list of lists? I can see the other side of the argument though.
Like the reductions - was previously constrained by sequeezing these values into simple strings - now the tool form will target the API I think this expanded version is a little more straight-forward (though verbose). Adds consistency with rest of the tool form API changes.
Old tool form needed to encode every value as a string so I had done "__collection_reduction__|<hdca_id>" to distinguish that value from an "<hda_id>" - since hdca and hdas can have the same encoded ids. The new tool form API is going to use the API which allows for richer object representations - so {"src": "hda", "id": "<hda_id>"} versus {"src": "hdca", "id": "<hdca_id>"} should be enough to distinguish between passing an HDA and an HDCA to a multiple input data parameter.
Adding consistency allowing each parameter to be wrapped in a object describing the meta-properties of the submitting value - this was requested by Sam to make the new tool form easier to manage, it makes multi-running properties work for non-data parameters, and allows linked/unlinked specification of parameters.
Collections can be mapped over 'data' parameters and sufficiently nested collections can be mapped over 'data_collection' parameters (for instance a list of 5 pairs can be supplied to a tool taking in a pair and 5 jobs will be executed). I (perhaps poorly) term these concepts collection mapping and subcollection mapping.
Prior to this changeset - the API for doing collection 'mapping' and 'subcollection mapping' was somewhat more divergent and the tool execution code explicitly forbid doing both kinds of mappings in the same tool execution even if the effective collections could be matched (e.g. it could not map a 'data' parameter over a list of 5 datasets and a pair parameter a list of 5 pairs in the same execution).
This changeset should remedy this - as long as the effective collection mappings can match up such jobs should be possible. The workflow editor (and I think runner) already thought this was possible, so this changeset reduces the tool-workflow impedance mismatch - an existing problem exacerbated by recent dataset collections introduction.
This all needs much more testing - test workflows execute this way, functional test of a tool execution that combines collection mapping and subcollection mapping, etc....
Replace ad-hoc tools API test method for skipping tests with a more general purpose decorator. Use new decorator to specify required tools for workflow tests.