This special class of tools leverages the infrastructure for tool inputs, tool state tracking, tool module for workflows, tool API, etc... without actually producing command-line jobs. Instead these tools are provided the input model objects and are expected to produce output model objects directly. This provides an oppertunity to copy HDAs without copying the underlying datasets.
The first driving use case for these tools are also included - namely tools that allow zipping and unzipping paired collections. These tools can be mapped over lists (e.g. list:paired to (list, list) or the inverse) using much of the existing infrastructure for tools. Test cases included that validate these work with mapping operations and in workflows.
The most obvious advantage of these versus traditional tools that do the same thing is that the data isn't copied on disk - new HDAs are created directly from the source datasets.
Testing:
This PR includes various API test cases for functionality, these can be run with the following command:
```
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_unzip_collection
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_zip_inputs
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_zip_list_inputs
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_workflow_run_zip_collections
```
Refactor test verification into galaxy-lib-compat module.
- Refactor a bunch class methods in ``test.base.twilltestcase`` into module functions in ``galaxy.tools.verify``.
- Move ``test.base.test_data`` to ``galaxy.tools.verify.test_data``.
- Move ``test.base.asserts`` to ``galaxy.tools.verify.asserts``.
Remove duplication in execution of these methods between composite and normal test outputs. This also entailed reworking the parsing of the composite test outputs to bring them inline with normal outputs. In addition to simplify removing duplication, this means many more tests can be made over composite outputs - such as md5 checks and test assertions. I've added a new framework test tool to verify this.
I was unable to reproduce #1531 in testing, but if the problem is something to do with stale state the following sledge hammer should fix it.
Runt the new test case with:
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_delete_intermediate_datasets_pja_1
Running a workflow or showing a workflow can both restore the previous behavior by passing legacy=True as an API parameter. By changing these two endpoints in tandem I believe backward compatiblity for most existing code should be maintained unless:
- The external application saved these workflow IDs previously and re-runs workflows without refetching the workflow definition. I could imagine Refinery for instance might do this and will have to update indexed workflows or add legacy=True to workflow requests.
- The external application contacted the database directly after using this API endpoint to fetch more information about the step (seems unlikely).
See conversation:
- http://dev.list.galaxyproject.org/workflow-API-step-order-vs-step-id-in-bioblend-td4668367.html
Rebased with changes suggested by @nsoranzo.
Details:
- Add a new workflow module describing subworkflows.
- Add workflow list to editor side panel - with options to link in a subworkflow module or copy the target workflow into the workflow being editted node for node.
- Update workflow, workflow step, and workflow invocation models to track subworkflow connections and execution.
- Extend workflow outputs with concepts of labels (and UUIDs while I'm there) to match workflow inputs. This allow us to have something to label outputs with in the workflow editor and to reference in the format 2 workflow description language.
- Extend workflow editor UI to allow labeling workflow outputs (and enforce that these are unique across a workflow).
- Extend workflow invocation and progress tracking to allow invoking a subworkflow as part of another workflow invocation.
- Extend workflow import and export code to allow a nested representation of workflows.
- Update format 2 workflow description to allow testing nested workflows.
Most relevant new and modified test cases can be run using the following commands:
```
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_run_subworkflow_simple
./run_tests.sh -api test/api/test_workflows_from_yaml.py:WorkflowsFromYamlApiTestCase.test_subworkflow_simple
./run_tests.sh -api test/api/test_workflows_from_yaml.py:WorkflowsFromYamlApiTestCase.test_outputs
nosetests test/unit/test_galaxy_mapping.py
nosetests test/unit/workflows/test_workflow_progress.py
```
- Implement a input parameter module that mirrors data and collection input modules but has a type that can currently be one of text, integer, float, color, and boolean.
- Allow connections between these and tool step inputs.
- Extend model to support this.
- Add new input types for format 2 workflow definitions for various types that all map to this kind of step. Typed inputs such as this match well with CWL workflow inputs.
Someday I imagine these will be superior to just marking a tool input "Specify at Runtime" for all the same reasons input steps are superior to leaving inputs unattached.
This is a slight tightening up of the ad-hoc, experimental workflow YAML definition used by the test framework (and by Kyle, but lets just admit Kyle is part of Galaxy's test framework).
This format is still defined entirely client side by transcoding the YAML or python object description into real (or format 1) Galaxy workflows - so these cannot realisticaly be declared part of Galaxy's public interface and can remain experimental and subject to change.
In addition to refactoring the code implementing these workflows for upstream modifications to explore new features, the format itself has been made slightly more stringent in two ways:
- 'steps' must now be explicit (previously converted auto-convert list to dict).
- Ensure the workflow declared a 'class' is defined and the class is 'GalaxyWorkflow'.
This will help with discovery of workflow objects with downstream tooling (planemo for workflows?) and brings the Galaxy definition slightly more inline with the CWL definition for workflows (very slightly).
Runtime post job actions are post job actions inserted when the workflow is invoked instead of being part of the workflow object in the database. The bug noticed by @kellrott was that these actions were being appended to the original workflow post job actions instead of being transient things just attached to the jobs themselves.
This fixes that problem and adds a test to try to prevent regressions.
Going to use bioblend for performance testing instead of galaxy_interactor. This refactoring allows all of those helpers to be reused with bioblend backing the helpers instead of galaxy_interactor.
It would fail when being run with the rest of the suite and not on its own - because it was using the same id for the workflow id and invocation id - which is obviously wrong unless it is a completely fresh database :).
Models:
Track whether dataset collections have been populated yet.
Dataset collections are still effectively immutable once populated - but dynamic output collections require them to be sort of like `final` fields in Java (analogy courtesy of JJ) - allowing them to be declared before they are initialized or populated. This is tracked by the `populated_state` field.
Tools:
Output collections can now describe `discover_datasets` elements just like datasets - except in this case instead of dynamically populating new datasets in the history - they will comprise the collection. `designation` has been reused to serve as the element_identifier for the collection element corresponding to the dataset.
See Pull Request 356 for more information on the discover_datasets tag https://bitbucket.org/galaxy/galaxy-central/pull-request/356/enhancements-for-runtime-discovered.
Workflows:
Update workflow execution and recovery for dynamic output collections.
Galaxy workflow data flow before collections
* - * - * - * - * - *
Galaxy worfklow data flow after collections (iteration 1)
* - * - * \
* - * - *
* - * - * / \
* - * - *
* - * - * \ /
* - * - *
* - * - * /
Galaxy worfklow data flow after this commit
/ * - * \
* - * * - *
/ \ * - * / \
/ \
/ \
/ / * - * \ \
* - * -- * - * * - * -- * - *
\ \ * - * / /
\ /
\ /
\ / * - * \ /
* - * * - *
\ * - * /
By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)
There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.
Model:
The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.
Workflow:
Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.
Tool Testing:
This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.
Tests:
Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.
Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.
Feature requested by Kyle. Implemented only in the API at this point - not sure it is a feature valuable to UI consumers.
Pass in PJAs along with step parameters map but keyed on __POST_JOB_ACTIONS__. JSON definition same as when defining PJA in the workflow definition JSON.
Includes test cases for normal use and for use after delayed workflow steps have been evaluated by the new workflow scheduling stuff.
How to use:
1.) Place multiple tools with different IDs in your tool conf.
2.) ... ummm ... no step 2 - just use the tools.
Implementation:
The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.
To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.
Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
Pass in the parameter 'inputs_by' as 'step_uuid' to the workflow run command to use this.
Specifing inputs by UUID has the nice advantage that it survives workflows saves - so if one sets up an API script or something to target a workflow - saving the workflow in the editor doesn't need to break the script as long as inputs were not added or deleted. The UUID (like the order_index) has the advantage of the step id that it is predeterminable - so one can set it up a workflow script against any Galaxy and the script doesn't need to be adapted to raw ids the steps get assigned in that instance.
Give every step a UUID that can be preserved across edit to the workflow. Likewise - allow every step to be given a label (a unique short name for that workflow) - that allows for a human consumable way to reference steps for use in tests and when driving workflows via the API. Workflow editor doesn't yet (and might never) display these attributes but it does preserve them across workflow saves.
Implement automated testing that uploading workflows, updating workflows, and exporting them handle UUID and labels. Manually tested workflow editor preserves labels and UUID across changes.
This means steps that are not connecting an output of one step to the input of another. This could potentially address all sorts of untraditional (in a Galaxy sense) workflows where some sort of data is managed externally. The most important use I think I have heard discussed is that of data managers - this can be used in cases where data managers depend on one another (grab the fasta files in one step, index them in another) or workflows where a downstream analysis depends on index data populated via data managers in earlier steps.
Not really sure how to represent these in the workflow editor - but this is a power user feature anyway so hopefully that is not super pressing. The YAML to workflow DSL supports the operation (see test cases) so these power users (a euphemism for Dan I guess) can just use that for now.
This logic for building up editor representation of the workflow.
Introduce concept of a workflow to dict style - with 'export' and 'editor' as first cracks.