Runtime post job actions are post job actions inserted when the workflow is invoked instead of being part of the workflow object in the database. The bug noticed by @kellrott was that these actions were being appended to the original workflow post job actions instead of being transient things just attached to the jobs themselves.
This fixes that problem and adds a test to try to prevent regressions.
Going to use bioblend for performance testing instead of galaxy_interactor. This refactoring allows all of those helpers to be reused with bioblend backing the helpers instead of galaxy_interactor.
It would fail when being run with the rest of the suite and not on its own - because it was using the same id for the workflow id and invocation id - which is obviously wrong unless it is a completely fresh database :).
Models:
Track whether dataset collections have been populated yet.
Dataset collections are still effectively immutable once populated - but dynamic output collections require them to be sort of like `final` fields in Java (analogy courtesy of JJ) - allowing them to be declared before they are initialized or populated. This is tracked by the `populated_state` field.
Tools:
Output collections can now describe `discover_datasets` elements just like datasets - except in this case instead of dynamically populating new datasets in the history - they will comprise the collection. `designation` has been reused to serve as the element_identifier for the collection element corresponding to the dataset.
See Pull Request 356 for more information on the discover_datasets tag https://bitbucket.org/galaxy/galaxy-central/pull-request/356/enhancements-for-runtime-discovered.
Workflows:
Update workflow execution and recovery for dynamic output collections.
Galaxy workflow data flow before collections
* - * - * - * - * - *
Galaxy worfklow data flow after collections (iteration 1)
* - * - * \
* - * - *
* - * - * / \
* - * - *
* - * - * \ /
* - * - *
* - * - * /
Galaxy worfklow data flow after this commit
/ * - * \
* - * * - *
/ \ * - * / \
/ \
/ \
/ / * - * \ \
* - * -- * - * * - * -- * - *
\ \ * - * / /
\ /
\ /
\ / * - * \ /
* - * * - *
\ * - * /
By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)
There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.
Model:
The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.
Workflow:
Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.
Tool Testing:
This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.
Tests:
Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.
Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.
Feature requested by Kyle. Implemented only in the API at this point - not sure it is a feature valuable to UI consumers.
Pass in PJAs along with step parameters map but keyed on __POST_JOB_ACTIONS__. JSON definition same as when defining PJA in the workflow definition JSON.
Includes test cases for normal use and for use after delayed workflow steps have been evaluated by the new workflow scheduling stuff.
How to use:
1.) Place multiple tools with different IDs in your tool conf.
2.) ... ummm ... no step 2 - just use the tools.
Implementation:
The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.
To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.
Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
Pass in the parameter 'inputs_by' as 'step_uuid' to the workflow run command to use this.
Specifing inputs by UUID has the nice advantage that it survives workflows saves - so if one sets up an API script or something to target a workflow - saving the workflow in the editor doesn't need to break the script as long as inputs were not added or deleted. The UUID (like the order_index) has the advantage of the step id that it is predeterminable - so one can set it up a workflow script against any Galaxy and the script doesn't need to be adapted to raw ids the steps get assigned in that instance.
Give every step a UUID that can be preserved across edit to the workflow. Likewise - allow every step to be given a label (a unique short name for that workflow) - that allows for a human consumable way to reference steps for use in tests and when driving workflows via the API. Workflow editor doesn't yet (and might never) display these attributes but it does preserve them across workflow saves.
Implement automated testing that uploading workflows, updating workflows, and exporting them handle UUID and labels. Manually tested workflow editor preserves labels and UUID across changes.
This means steps that are not connecting an output of one step to the input of another. This could potentially address all sorts of untraditional (in a Galaxy sense) workflows where some sort of data is managed externally. The most important use I think I have heard discussed is that of data managers - this can be used in cases where data managers depend on one another (grab the fasta files in one step, index them in another) or workflows where a downstream analysis depends on index data populated via data managers in earlier steps.
Not really sure how to represent these in the workflow editor - but this is a power user feature anyway so hopefully that is not super pressing. The YAML to workflow DSL supports the operation (see test cases) so these power users (a euphemism for Dan I guess) can just use that for now.
This logic for building up editor representation of the workflow.
Introduce concept of a workflow to dict style - with 'export' and 'editor' as first cracks.
Fill out the dataset collection parameter test and add a new test for a workflow that includes subcollection mapping. Move toward orchestrating jobs in this file via a high-level YAML description of steps and test data as well as some higher-level methods for testing various stuff about extracted workflows (step counts, input types, tools used, connected-ness, etc...).
Allow ordering jobs index API by create_time instead of update_time.
test_workflows.py was pretty unwieldy and the workflow extraction tests were very different than the other tests in that file (e.g. very different helper functions).
New module type that pauses a workflow and gives time for the runner to review it before proceeding with execution.
- Introduce concept of beta workflow modules - I guess we should just keep the pause module as beta until their is a UI to support it.
- Extracts base class ouptut of InputModule for modules that are "simple" - i.e. their configuration state is represented as a dictionary and configuration form is rendered via the generic template.
- Test Cases (for this changeset and a bunch of stuff that is now testable with the pause module in place from the scheduling framework commit).
Models:
Workflow invocations have been augmented with significantly more state - inputs, parameters, runtime step state, are all being tracked now. Workflow invocations have a state that can be changed over time, the UUIDs generated for workflow invocations in Pull Request #465 have to be persisted so they can be reused when scheduling new jobs for theworkflow invocation. Workflow invocation steps now have an action parameter for persisting state provided by users during the execution of the workflow (see forthcoming PauseModule for further details).
Some initial elements of these model changes were based on model changes in Kyle Ellrott's Galaxy farm work (https://bitbucket.org/kellrott/galaxy-farm/branch/workflow_migrate). I made heavy modifications to the model to enforce referential integrity on parameter to workflow step mappings and made some cosmetic changes various other details.
Scheduling Plugins:
Used the pattern setup with dependency resolvers and job metrics to build a dynamic plugin infrastructure for defining workflow schedulers. I hesistate calling anything with only one implementation a plugin infrastructure, but I am confident enough that the combination of persisted workflow request combined with scheduler tag could be used to build a galaxy-farm plugin that would wait for another Galaxy instance to become available and it would pull the workflow down and
This work piggy backs on Galaxy job handlers to have workflow scheduled in the background (i.e. during submission each workflow being scheduled in the background is assigned a unique job handler and only that job handler thread will process the workflow). It should be pretty easy to allow the definition of a new kind of handler - that is a workflow handler instead of a job handler if that is of interest.
I will probably move a bunch of stuff that is happening in workflow/scheduling_manager.py more into the scheduler itself so that it can be more configurable and closer to a true plugin.
API:
There are a number of new API points here for flushing out dealing with workflow invocations (called usages in existing parlance).
- POST /api/workflows/{encoded_workflow_id}/usage
Schedule a worklfow to be run in the background and return just the workflow invocation information.
RESTfully speaking this should be plural but the matching GET endpoint is likewise usage and not usages - so I am favoring consistency over RESTful correctness here. Also, likewise creating a 'usage' feel like odd - I would like to make all of the usage endpoints aliases to a more RESTfully correct invocations endpoints.
The existing workflow run API endpoints still work and still work the way they use usually - but the output now includes all of the workflow invocation to_dict stuff as well as the list of outputs it initially used. Once everything is scheduled this way - that list of outputs is going to have to disappear but hopefully people can start using the invocation stuff now to help the transition.
- DELETE /api/workflows/{workflow_id}/usage/{usage_id}
Cancel a scheduled workflow invocation.
- GET /api/workflows/{workflow_id}/usage/{usage_id}/steps/{step_id}
Get information about a workflow invocation step.
- PUT /api/workflows/{workflow_id}/usage/{usage_id}/steps/{step_id}
Update a workflow invocation step - for ones with modifiable state. Extension point added to workflow modules to support this but it is unused by all existing worklfow modules. A subsequent PauseModule will use this to either continue or cancel a workflow invocation at a particular step.
Modules:
Workflow modules can now define new methods for dealing with recovering state and interacting with user requests.
Testing:
One can issue a workflow request by running the following test.
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_workflow_request
Including same test with workflow parameter substition - gotta admit I thought that workflow test was going to fail - so this week is looking pretty good :).
There may now be multiple WorkflowInvocationSteps for each WorkflowStep for steps that are mapped over collections - so to_dict creating a dictionary of this information indexed on order step is problematic because only one WorkflowInvocationStep will be represented per step. Instead now just returning a big list of all of the invocations - which contains all of the same information. This is a backward incompatible API change for the workflow invocation API.
Also update the input mapping stuff with logic for dealing with data collection inputs.
Sharing a workflow with a user was previously sufficient to grant access to all invocations of that workflow. This isn't a huge problem since the information potentially leaking out was limitted to invocation counts, various encoded ids, and update times. Still I think no information about invocations should be avaialble as a result of sharing a workflow - and upcoming changes to Galaxy will result in much more information being made available via the workflow invocation API.
Convert workflow API to use new-style API decorator throughout.
Improved workflow exception handling - now more explicit status code setting and returning of raw error strings in the API. Most obvious exceptional paths out of the API are now coming through MessageExceptions of things even more specific. Add API tests for many of these exceptional conditions.
Use more helper methods to reduce method length, duplication.
Previous route was "POST /api/workflows/upload", this still works but should be condisdered deprecated in favor
of POSTing to /api/workflows with 'workflow' in the payload.
The old 'ds_map' parameter remains in place for backward compatibility and with the same behavior. A new parameter 'inputs' can now be specified instead however, and its keys corresponding to the steps 'order_index' instead of the raw unencoded database ids 'ds_map' uses. This variant has the nice property that this map can be constructed without prior knowledge of how Galaxy will assign ids during import - hence it is easier to use and more portable.
Additionally, 'inputs' is more flexiable and can revert to the old behavior by specifying a new parameter 'inputs_by' as 'step_id'. 'inputs_by' can also be 'name' - this is even more human friendly because it will assign the ids based on data input names (this is what I intend to use mostly, but I have not made it the default because not all workflows will have data inputs with distinct names).
Finally, the new name 'inputs' will be more appropriate once data collection inputs can be explicitly mapped via the API.
Probably should do something like this server side so workflow API can be used more deterministically (wouldn't need to import a workflow and then hit the API to know how to use it).