Works for both files and collections. Workflow defaults override tool defaults.
TODO:
- Unit test case to ensure this only works for non-default, non-multi data parameters.
- Implement XSD once syntax is finalized.
Trying out https://github.com/dropbox/pyannotate here.
It's still a lot of work even with pyannotate, but I think this fixes
some minor oddities, like assigning non-mapped attributes to
WorkflowInvocation instances.
Upgrade syntax using `pyupgrade --py36-plus` .
Manually drop several `six` imports.
Also:
- Remove broken pr_cache in scripts/bootstrap_history.py
- Fix broken prefix removal in lib/galaxy/tool_util/deps/mulled/mulled_build.py
Implement markdown backend and frontend components as well generator plugin framework to allow customizable workflow invocation reports.
Architectural Choices on the Client
Old-style embedded objects vs components
In theory, this could make really nice use of excellent VueJS components for dataset display, dataset collection display, workflow display, etc.. These aren't available yet, we don't even really have Backbone stuff that exist very well outside history panels, so this re-uses "components" from Galaxy Pages for datasets and reuses the collection display for history panels as a stand-alone display for collections (I got this trick from DIsplayStructured.vue). This isn't ideal and I know that but at least reusing things this way will provide a path forward for migrating both pages and this new Markdown language (which could easily replace or co-exist with pages HTML someday) together seamlessly as the real modern components become available and doesn't increase the overall work needed to integrate newer style components. In fact, this might even be the impetus for creating and polishing more of these components.
markdown-it vs markdown-it-vue
I tried this with markdown-it-vue also but it added very little in terms of reducing client code, obscured entirely how to attach plugins to the markdown rendering process, and brought in many, many more extra packages that we don't need or want.
Given there is no representation of invocations in the GUI and not editor for reports config this is a bit challenging still. But here is goes:
- Source Galaxy's virtualenv.
- Login to a user and grab an API key, set in the following command:
- ``GALAXY_TEST_EXTERNAL=http://localhost:8080 GALAXY_TEST_USER_API_KEY=38175dfc69d9009992b69efcca8b209b pytest test/api/test_workflows.py::WorkflowsApiTestCase::test_workflow_invocation_report_custom``
- In the web browser go to http://localhost:8080/api/invocations and grab the latest invocation ID. Replace it in the follow query string.
- Navigate to http://localhost:8080/workflows/invocations/report?id=c887f1d0da42bdfe
We don't track workflow step inputs in any formal way in our model currently. This has resulted in some current hacks and prevents future enhancements. This commit splits WorkflowStepConnection into two models WorkflowStepInput and WorkflowStepConnection - normalizing the previous table workflow_step_connection on input step and input name.
In terms of current hacks forced on it by restricting all of tool state to be confined to a big JSON blob in the database - we have problems distinguishing keys and values when walking tool state. As we store more and more JSON blobs inside of the giant tool state blob - the worse this problem gets. Take for instance checking for runtime parameters or the rules parameter values - these both use JSON blobs that aren't simple values, so it is hard to tell looking at the tool state blob in the database or the workflow export to tell what is a key or what is a value. Tracking state as normalized inputs with default values and explicit attributes runtime values should allow much more percise state definition and construction.
This variant of the models would also potentially allow defining runtime values with non-tool default values (so default values defined for the workflow but still explicitly settable at runtime). The combinations of overriding defaults and defining runtime values were not representable before.
In terms of future enhancements, there is a lot we cannot track with the current models - such as map/reduce options for collection operations (https://github.com/galaxyproject/galaxy/issues/4623#issuecomment-389544980). This should enable a lot of that. Obviously there are a lot of attributes defined here that are not yet utilized, but I'm using most (all?) of them downstream in the CWL branch. I'd rather populate this table fully realized and fill in the implementation around it as work continues to stream in from the CWL branch - to keep things simple and avoid extra database migrations. But I understand if this feels like speculative complexity we want to avoid despite the implementation being readily available for inspection downstream.
Refactor get_data_inputs into get_all_inputs and use the resulting dictionaries to reason about if collection mapping should occur during invocation of tools. Using these dictionaries instead of explicit tool input objects should allow reuse within other module types since they may produce the same interface.
Workflow Invocations
--------------------
The workflow invocation outputs half of this is relatively straight forward. It is modelled somewhat on job outputs, output datasets and output dataset collections are now tracked for each workflow invocation and exposed via the workflow invocation API. This required adding new tables (linked to WorkflowInvocations and WorkflowOutputs) that track these output associations.
Previously one could imagine backtracking this information for simple tool steps via the WorkflowInvocationStep -> Job table, but for steps that have many jobs (i.e. mapping over a collection) or for non-tool steps such information was more difficult to recover (and simply couldn't be recovered from the API at all or even internally without significant knowledge of the underlying workflow).
Workflow Invocation Steps
-------------------------
Tracking the outputs of WorkflowInvocationSteps was not previously done at all, one would have to follow the Job table as well. A signficant downside to this is that one cannot map over empty collections in a workflow - since no such job would exist. Tracking job outputs for WorkflowInvocationSteps is not a simple matter of just attaching outputs to an existing table because we had no concept of a workflow step tracked - since there could be many WorklfowInvocationSteps corresponding to the same combination of WorkflowInvocation and WorkflowStep. That should feel wrong and that is because it is - when collections were added the possiblity of having many jobs for the same combination of WorkflowInvocation and WorkflowStep was added. I should have split WorkflowInvocationSteps into WorkflowInvocationSteps and WorkflowInvocationStepJobAssociations at that time but didn't. This commit now does it - effectively normalizing the ``workflow_invocation_step`` table by introducing the new ``workflow_invocation_step_job_association`` table.
Splitting up the WorkflowInvocationStep table this way allows recovering the mapped over output (e.g. the implicitly created collection from all the jobs) as well the outputs from the individual jobs (by walking WorkflowInvocationStep -> WorkflowInvocationStepJobAssociation -> Job -> JobToOutput*Association).
This split up involves failrly substantial changes to the workflow module interface. Any place a list of WorkflowInvocationSteps was assumed, I reworked it to just expect a single WorkflowInvocationStep. I vastly simplified recover_mapping to just use the persisted outputs (this was needed in order to also implment empty collection mapping in workflows). This also fixes a bug (or implements a missing feature) where Subworkflow moudles had no recover_mapping methods - so for instance if a tool that produces dynamic collections appeared anywhere in a workflow after a subworkflow step - that workflow would not complete scheduling properly.
Now that we have a way to reference the set of jobs corresponding to a workflow step within an invocation, we can start to track partial scheduling of such steps. This is outlined in https://github.com/galaxyproject/galaxy/issues/3883 and refactoring toward this goal is included here - including adding a state to WorkflowInvocationStep so Galaxy can determine if it has started scheduling this step and an index when scheduling jobs so it can tell how far into a scheduling things have gone as well as augmenting the tool executor to take a maximum number of jobs to execute and allow recovery of existing jobs for collection building purposes.
*Applications*
These changes will enable:
- A simple, consistent API for finding workflow outputs that can be consumed by Planemo for testing workflows.
- Mapping over empty collections in workflows.
- Re-scheduling workflow invocations that include subworkflow steps.
- Partial scheduling within steps requiring a large number of jobs when scheduling workflow invocations.
- Add flake8-import-order to flake8 Pipfile and remove py27-lint-imports
and py27-lint-imports-include-list tox envs
- Fix most E201 and E202 errors reported by flake8-import-order v0.15,
but pin flake8-import-order to v0.14.3 until
https://github.com/PyCQA/flake8-import-order/issues/123
is fixed
This let us drop 2 jobs on Travis per each job.
Details:
- Add a new workflow module describing subworkflows.
- Add workflow list to editor side panel - with options to link in a subworkflow module or copy the target workflow into the workflow being editted node for node.
- Update workflow, workflow step, and workflow invocation models to track subworkflow connections and execution.
- Extend workflow outputs with concepts of labels (and UUIDs while I'm there) to match workflow inputs. This allow us to have something to label outputs with in the workflow editor and to reference in the format 2 workflow description language.
- Extend workflow editor UI to allow labeling workflow outputs (and enforce that these are unique across a workflow).
- Extend workflow invocation and progress tracking to allow invoking a subworkflow as part of another workflow invocation.
- Extend workflow import and export code to allow a nested representation of workflows.
- Update format 2 workflow description to allow testing nested workflows.
Most relevant new and modified test cases can be run using the following commands:
```
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_run_subworkflow_simple
./run_tests.sh -api test/api/test_workflows_from_yaml.py:WorkflowsFromYamlApiTestCase.test_subworkflow_simple
./run_tests.sh -api test/api/test_workflows_from_yaml.py:WorkflowsFromYamlApiTestCase.test_outputs
nosetests test/unit/test_galaxy_mapping.py
nosetests test/unit/workflows/test_workflow_progress.py
```