A temporary script in the job working directory will be created to
import and call it (trusting that `$PYTHONPATH` in a job script is
always set to `galaxy/lib`). This is so the auto-detect button can defer
command line generation until job preparation time for the case that
handlers and web processes run from different paths.
Related Trello card: https://trello.com/c/v2eCOYZi
(cherry picked from commit b043d43a35)
Move set_metadata script from galaxy_utils to galaxy, remove unused
imports.
(cherry picked from commit e093f58d63)
Restore old set_metadata files to prevent failure of jobs running at
upgrade time.
(cherry picked from commit 4443b64eae)
Fix unit test broken by set_metadata changes.
(cherry picked from commit d242286187)
Use the job working directory for creating MetadataFiles in external
set_metadata, rather than new_files_path.
(cherry picked from commit 8116d2c917)
It would fail when being run with the rest of the suite and not on its own - because it was using the same id for the workflow id and invocation id - which is obviously wrong unless it is a completely fresh database :).
These aren't ideal tests - but it is some indication that things are working that the API will import the workflow and produce a representation. Should follow up at some point and verify the representation is in fact the correct one.
... for tool tests. Longer term this functionality should be dropped (i.e. after all the tools are out of Galaxy) and the test-data always lives next to the tool - but for now it decreases the size of the repository ahead of a potential move to github.
Repeatedly one wants to do things like create symbolic links and then call a helper script - the 'interpreter' tag doesn't allow shell commands before calling a helper script so they could not be used in this fashion. The previous pattern for doing this was then to use the ToolShed only 'set_environment' requirement tag. These were onerous to setup and the resulting tools were no longer portable to non-ToolShed installed contexts - I believe this variant is more robust and elegant.
More information https://trello.com/c/0pgF5PBQ and here https://trello.com/c/XK5SqE1i.
Tools may now use $input.element_identifier during tool evalution for input 'data' parameters with the following semantics:
- If the input was specified as a single dataset by the user - this just fallbacks to providing the $input.name.
- If the input was mapped over a collection (to produce many jobs) or if the input is a 'multiple="true"' input that was provided a collection - the $input.element_identifier will be the element identifier for the corresponding collection item (generally much more useful the dataset name - since if preserved throughout workflows).
'data_collection' parameters already can access this kind of information - but it is something of a best practice to use simple 'data' parameters since they are compatible with more traditional un-collected datasets.
This commit really needs more comments - but Philip Mabon has been patiently waiting for this functionality for a long time.
Like done for Docker to isolate metadata commands from the environment modifications required to resolve tool dependencies. Should allow for Python 3 dependencies (originally also allowed samtools - but Nate other commit resolved that problem also).
Just meant as a config option for now - it will become the default once tested more thoroughly. For now enable it by setting enable_beta_tool_command_isolation to True in galaxy.ini.
Normal selects seem to be prevented from execution with invalid parameter values, but not columns. Values are escaped properly so shell exploitation isn't the problem - but as a usability thing Galaxy should prevent execution and provide a warning message.
Models:
Track whether dataset collections have been populated yet.
Dataset collections are still effectively immutable once populated - but dynamic output collections require them to be sort of like `final` fields in Java (analogy courtesy of JJ) - allowing them to be declared before they are initialized or populated. This is tracked by the `populated_state` field.
Tools:
Output collections can now describe `discover_datasets` elements just like datasets - except in this case instead of dynamically populating new datasets in the history - they will comprise the collection. `designation` has been reused to serve as the element_identifier for the collection element corresponding to the dataset.
See Pull Request 356 for more information on the discover_datasets tag https://bitbucket.org/galaxy/galaxy-central/pull-request/356/enhancements-for-runtime-discovered.
Workflows:
Update workflow execution and recovery for dynamic output collections.
Galaxy workflow data flow before collections
* - * - * - * - * - *
Galaxy worfklow data flow after collections (iteration 1)
* - * - * \
* - * - *
* - * - * / \
* - * - *
* - * - * \ /
* - * - *
* - * - * /
Galaxy worfklow data flow after this commit
/ * - * \
* - * * - *
/ \ * - * / \
/ \
/ \
/ / * - * \ \
* - * -- * - * * - * -- * - *
\ \ * - * / /
\ /
\ /
\ / * - * \ /
* - * * - *
\ * - * /
By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)
There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.
Model:
The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.
Workflow:
Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.
Tool Testing:
This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.
Tests:
Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.
Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.