By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)
There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.
Model:
The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.
Workflow:
Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.
Tool Testing:
This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.
Tests:
Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.
Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.
How to use:
1.) Place multiple tools with different IDs in your tool conf.
2.) ... ummm ... no step 2 - just use the tools.
Implementation:
The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.
To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.
Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
Leave empty module in galaxy.tools.filters and modify config to ensure backward compatibility (filters in this old directory will continue to work for now).
Update config/galaxy.ini.sample with more discussion of ToolBox filter.
- Allow overriding the base module location for ToolBox filters (needed for OS packages, etc...).
- Allow scanning multiple base modules for filters.
- Unit tests for module loading functionality, custom tool, label, and section filters, default hidden and require_login filters.
Using ToolBox._xxx instead of ToolBox.__xxx because realistically ToolBox is still much to large to grok and so I imagine it will need to be broken up even more (base class focused on just the panel details perhaps - or mixins - etc...).
Create one public method load_item on ToolBox that should be used externally instead of load_tool_tag_set, load_section_tag_set, etc.... This reduces the duplication within the toolbox and between the toolbox and the Tool Shed's ToolPanelManager. Probably more importantly it also is another step down the road toward hiding implementation details such as toolbox.tool_panel and toolbox.intergrated_tool_panel from the ToolPanelManager.
Already had unit test coverage for ToolPanelManager stuff here - but added test coverage for loading labels, workflows, and tool directories.
Now with unit test. Part of broader effort to make ToolBox interface more explicit, tested, and hide implementation details from other parts of code (such as ToolPanelManager).
Also reworked the code to now build up and re-parse XML structures - just use a dictionary - thanks to 34b3e1c.
Hides some details of writing out integrated tool panel, reindexing, etc... from tool shed install code and reduces duplicatation in tool shed's tool_panel_manager.
Add unit tests.
Do not convert rst to a mako template until needed, this is a costly operation and has the potential to speed update Galaxy start time. Move logic related to parsing of simple help text blocks out of tool and into the new tool parser interface and add implementation for YAML-based tools as well as unit tests for both. More advanced, multi-page tool help is still possible, workflows with the old tool form, but is only available to XML-based tools.
Long term this could allow Galaxy to support - multiple tooling formats (Galaxy-like YAML, CWL http://bit.ly/cwltooldesc, etc...). But I think it is also important from a purely design perspective - this is a core logic class integrating different components - they should not also be doing XML parsing.
To verify the interface for parsing tools is expressive enough to allow multiple useful implementations, I built a test YAML tool description that implements many of the same features as Galaxy but smooths out rough edges (uses exit codes for job failure by default for instance). Loading these tools is disabled by default and it is not documented how to enable them because they are not intended to be part of Galaxy's public API.
Allow tool panel to contain a <tool_dir dir="foo" /> element. Tools in such a directory will be loaded at startup time.
Additionally, if watch_tools is set to True in config/galaxy.ini and watchdog (http://pythonhosted.org/watchdog/) is available to Galaxy - all tools will be dynamically reloaded as they are modified and new tools that are added to the tool_dir directories will be dynamically populated (no need to restart Galaxy).
With refactoring to reduce cyclomatic complexity. Also renaming 'input_ext' what it actually is 'random_input_ext'. We should fix that or at least issue a huge warning if we detect 'input' could have reasonable been different things.
... so internals of tool XML description are only utilized in the galaxy.tools module (and submodules). Add some unit tests for this new method for finding externally referenced files.
Fixes multirun inside of conditional, repeats if tool form state updated (e.g. because conditional param updated or repeat block added).
More tests - ugly tests - but tests.
Unit tests verifing the fixed, correct behavior included. Bug introduced in e8c84dd715782e7c1d709d8068e6033b835f7f39 as an unintended consequence of duplicating the first dataset in the dictionary of input datasets generated by the tool action code.
Fixes https://trello.com/c/LCEKxImR.
Previously LWR jut serialized tool requirements, by transitioning to this class LWR should be able to resolve tool shed installed packages as well - requiring the additional repository and tool dependency context.
When running normal tools with normal data inputs across dataset collections in parallel - a collection will be created for each output with a "structure" matching that of the input (for instance pairs will create pairs, lists will create lists with the same identifiers, etc...). A previous changeset added the ability to run the tool in parallel - this changeset extends that functionality to create analogous collections from these parallel runs.
For example, if one filters a pair of FASTQ files a pair is created out of the result. Likewise, if one filters a list of FASTQ files - a list collection with same cardinality is built from the results.
There is a lot left TODO here - for one a lot of this logic should be moved into the dataset_collection module. The matching needs to be exact right now - not a problem for pairs (every 'element' has name 'left' or 'right') but for lists with element names - these have to match exactly - but a 'list' like samp1_l, samp2_l, samp3_l should be able to match against samp1_r, samp2_r, samp3_r and create a new list samp1, samp2, and samp3. Even if there is no matching prefixes a new 'unlabeled' list should be able to be created.
Allow replacing data parameter inputs with collections - this will cause the tool to produce multiple jobs for the submission - one for each combination of input parameters after being matched up (linked). In addition to various unit tests, functional tests demonstrate the API usage in `test/functional/api/test_tools.py`.