- Reduce duplication between initializing generic modules and tool modules.
- Switch tool_id to content_id as variable names throughout the client.
- Rename get_tool_id to get_content_id on workflow modules.
- Add some minimal documentation to the workflow module about get_content_id.
Downstream in the subworkflow commit I switch the over-the-wire communication to use content_id instead of tool_id also and use content_ids to refer to workflow ids in subworkflow moduls.
How to use:
1.) Place multiple tools with different IDs in your tool conf.
2.) ... ummm ... no step 2 - just use the tools.
Implementation:
The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.
To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.
Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
Models:
Workflow invocations have been augmented with significantly more state - inputs, parameters, runtime step state, are all being tracked now. Workflow invocations have a state that can be changed over time, the UUIDs generated for workflow invocations in Pull Request #465 have to be persisted so they can be reused when scheduling new jobs for theworkflow invocation. Workflow invocation steps now have an action parameter for persisting state provided by users during the execution of the workflow (see forthcoming PauseModule for further details).
Some initial elements of these model changes were based on model changes in Kyle Ellrott's Galaxy farm work (https://bitbucket.org/kellrott/galaxy-farm/branch/workflow_migrate). I made heavy modifications to the model to enforce referential integrity on parameter to workflow step mappings and made some cosmetic changes various other details.
Scheduling Plugins:
Used the pattern setup with dependency resolvers and job metrics to build a dynamic plugin infrastructure for defining workflow schedulers. I hesistate calling anything with only one implementation a plugin infrastructure, but I am confident enough that the combination of persisted workflow request combined with scheduler tag could be used to build a galaxy-farm plugin that would wait for another Galaxy instance to become available and it would pull the workflow down and
This work piggy backs on Galaxy job handlers to have workflow scheduled in the background (i.e. during submission each workflow being scheduled in the background is assigned a unique job handler and only that job handler thread will process the workflow). It should be pretty easy to allow the definition of a new kind of handler - that is a workflow handler instead of a job handler if that is of interest.
I will probably move a bunch of stuff that is happening in workflow/scheduling_manager.py more into the scheduler itself so that it can be more configurable and closer to a true plugin.
API:
There are a number of new API points here for flushing out dealing with workflow invocations (called usages in existing parlance).
- POST /api/workflows/{encoded_workflow_id}/usage
Schedule a worklfow to be run in the background and return just the workflow invocation information.
RESTfully speaking this should be plural but the matching GET endpoint is likewise usage and not usages - so I am favoring consistency over RESTful correctness here. Also, likewise creating a 'usage' feel like odd - I would like to make all of the usage endpoints aliases to a more RESTfully correct invocations endpoints.
The existing workflow run API endpoints still work and still work the way they use usually - but the output now includes all of the workflow invocation to_dict stuff as well as the list of outputs it initially used. Once everything is scheduled this way - that list of outputs is going to have to disappear but hopefully people can start using the invocation stuff now to help the transition.
- DELETE /api/workflows/{workflow_id}/usage/{usage_id}
Cancel a scheduled workflow invocation.
- GET /api/workflows/{workflow_id}/usage/{usage_id}/steps/{step_id}
Get information about a workflow invocation step.
- PUT /api/workflows/{workflow_id}/usage/{usage_id}/steps/{step_id}
Update a workflow invocation step - for ones with modifiable state. Extension point added to workflow modules to support this but it is unused by all existing worklfow modules. A subsequent PauseModule will use this to either continue or cancel a workflow invocation at a particular step.
Modules:
Workflow modules can now define new methods for dealing with recovering state and interacting with user requests.
Testing:
One can issue a workflow request by running the following test.
./run_tests.sh -api test/api/test_workflows.py:WorkflowsApiTestCase.test_workflow_request
(Lot easier now that I understand what all of the module methods are doing and have an example of a 4th module downstream.) Now with even more unit tests.
It was stubbing out connection stuff which worked fine when running test on its own - but after any test sets up the sql alchemy mappings this breaks down because events are trying to fire. This just uses actual model classes - which is just as easy anyway really.
... this will allow linking to actual steps (referential integrity) if/when we store workflow requests in the database and failing faster.
Also more unit tests.
In particular, this moves get_job_dict out of workflow controller. I am making big changes to this downstream in dataset collection work so I want unit tests - additionally get_job_dict isn't really a great name for what this becomes - so I am renaming it to summarize.