This differs from a traditional tool in that its inputs don't need to be in an 'ok' state and instead of creating new datasets and duplicating data on disk, new HDAs are created from the existing datasets.
This special class of tools leverages the infrastructure for tool inputs, tool state tracking, tool module for workflows, tool API, etc... without actually producing command-line jobs. Instead these tools are provided the input model objects and are expected to produce output model objects directly. This provides an oppertunity to copy HDAs without copying the underlying datasets.
The first driving use case for these tools are also included - namely tools that allow zipping and unzipping paired collections. These tools can be mapped over lists (e.g. list:paired to (list, list) or the inverse) using much of the existing infrastructure for tools. Test cases included that validate these work with mapping operations and in workflows.
The most obvious advantage of these versus traditional tools that do the same thing is that the data isn't copied on disk - new HDAs are created directly from the source datasets.
Testing:
This PR includes various API test cases for functionality, these can be run with the following command:
```
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_unzip_collection
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_zip_inputs
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_zip_list_inputs
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_workflow_run_zip_collections
```
Refactor test verification into galaxy-lib-compat module.
- Refactor a bunch class methods in ``test.base.twilltestcase`` into module functions in ``galaxy.tools.verify``.
- Move ``test.base.test_data`` to ``galaxy.tools.verify.test_data``.
- Move ``test.base.asserts`` to ``galaxy.tools.verify.asserts``.
Remove duplication in execution of these methods between composite and normal test outputs. This also entailed reworking the parsing of the composite test outputs to bring them inline with normal outputs. In addition to simplify removing duplication, this means many more tests can be made over composite outputs - such as md5 checks and test assertions. I've added a new framework test tool to verify this.
- Fixes a bug in lib/galaxy/tools/linters/general.py found through galaxy-lib.
- Brings in docstring and import order linting fixes from the gxformat2 project (https://github.com/jmchilton/gxformat2).
I was unable to reproduce #1531 in testing, but if the problem is something to do with stale state the following sledge hammer should fix it.
Runt the new test case with:
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_delete_intermediate_datasets_pja_1
Long term this is just not a solution, but it is the strategy we use for mapped over output HDAs also.
Rebased and fixed based on comments from @nsoranzo indicating the previous attempt did absolutely nothing to fix the problem.
xref https://github.com/galaxyproject/tools-iuc/pull/412/files
Conflicts:
lib/galaxy/tools/actions/__init__.py
Running a workflow or showing a workflow can both restore the previous behavior by passing legacy=True as an API parameter. By changing these two endpoints in tandem I believe backward compatiblity for most existing code should be maintained unless:
- The external application saved these workflow IDs previously and re-runs workflows without refetching the workflow definition. I could imagine Refinery for instance might do this and will have to update indexed workflows or add legacy=True to workflow requests.
- The external application contacted the database directly after using this API endpoint to fetch more information about the step (seems unlikely).
See conversation:
- http://dev.list.galaxyproject.org/workflow-API-step-order-vs-step-id-in-bioblend-td4668367.html
Rebased with changes suggested by @nsoranzo.
Details:
- Add a new workflow module describing subworkflows.
- Add workflow list to editor side panel - with options to link in a subworkflow module or copy the target workflow into the workflow being editted node for node.
- Update workflow, workflow step, and workflow invocation models to track subworkflow connections and execution.
- Extend workflow outputs with concepts of labels (and UUIDs while I'm there) to match workflow inputs. This allow us to have something to label outputs with in the workflow editor and to reference in the format 2 workflow description language.
- Extend workflow editor UI to allow labeling workflow outputs (and enforce that these are unique across a workflow).
- Extend workflow invocation and progress tracking to allow invoking a subworkflow as part of another workflow invocation.
- Extend workflow import and export code to allow a nested representation of workflows.
- Update format 2 workflow description to allow testing nested workflows.
Most relevant new and modified test cases can be run using the following commands:
```
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_run_subworkflow_simple
./run_tests.sh -api test/api/test_workflows_from_yaml.py:WorkflowsFromYamlApiTestCase.test_subworkflow_simple
./run_tests.sh -api test/api/test_workflows_from_yaml.py:WorkflowsFromYamlApiTestCase.test_outputs
nosetests test/unit/test_galaxy_mapping.py
nosetests test/unit/workflows/test_workflow_progress.py
```
- Implement a input parameter module that mirrors data and collection input modules but has a type that can currently be one of text, integer, float, color, and boolean.
- Allow connections between these and tool step inputs.
- Extend model to support this.
- Add new input types for format 2 workflow definitions for various types that all map to this kind of step. Typed inputs such as this match well with CWL workflow inputs.
Someday I imagine these will be superior to just marking a tool input "Specify at Runtime" for all the same reasons input steps are superior to leaving inputs unattached.
Allow multiple collections to be fed to a multi data parameter in one reduction step. Fixes#750 and will really simplify certain classes of tools.
Rebased original with fixes for rerun of such reductions.
Manually tested workflow execution and everything seems fine. The workflow editor already thought this was possible, so that is another bug corrected by this enhancement.
To run the associated API test, execute the following command:
./run_tests.sh -with_framework_test_tools -api test/api/test_tools.py:ToolsTestCase.test_reduce_multiple_lists_on_multi_data
Conflicts:
static/maps/mvc/dataset/dataset-choice.js.map
static/maps/mvc/form/form-select-content.js.map
static/scripts/mvc/form/form-select-content.js
Copy CWL style of specifying a outputs specification block at the top-level of the workflow specification and within that use the "source" attribute to specify the output. Reuse the specification used by "$link"s to describe this output - namely <label_or_order_index>[#<output_name=output>].
In downstream work this is used for subworkflows and nested tools. Not sure which of these would potentially hit main line Galaxy first so setting this all up in its own commit.
Lots of other tests would fail if simple uploads aren't working, but it is better to be direct and very explicit about why this particular test is failing.
This strategy proved to work around certain race conditions in tool testing so hopefully it will solve the transiently failing job searching and filtering test cases.