Commit Graph
72 Commits
Author SHA1 Message Date
John Chilton cf8fd5c597 Add test case demonstrating data params don't need namespace...
... unlike other parameters. Is this a bug? Has been this way for at least three releases it seems.

https://github.com/galaxyproject/tools-devteam/commit/c01b354711d982f61f1fd086148692cb084aa8a5#commitcomment-9922496
2015-02-25 20:07:20 -05:00
John Chilton b56bd85b6a Expose $__tool_directory__ to tools.
Repeatedly one wants to do things like create symbolic links and then call a helper script - the 'interpreter' tag doesn't allow shell commands before calling a helper script so they could not be used in this fashion. The previous pattern for doing this was then to use the ToolShed only 'set_environment' requirement tag. These were onerous to setup and the resulting tools were no longer portable to non-ToolShed installed contexts - I believe this variant is more robust and elegant.

More information https://trello.com/c/0pgF5PBQ and here https://trello.com/c/XK5SqE1i.
2015-02-03 10:29:34 -05:00
John Chilton 2f15eb0d78 Expose improved sample tracking to tools for implicit map/reduce ops.
Tools may now use $input.element_identifier during tool evalution for input 'data' parameters with the following semantics:

 - If the input was specified as a single dataset by the user - this just fallbacks to providing the $input.name.
 - If the input was mapped over a collection (to produce many jobs) or if the input is a 'multiple="true"' input that was provided a collection - the $input.element_identifier will be the element identifier for the corresponding collection item (generally much more useful the dataset name - since if preserved throughout workflows).

'data_collection' parameters already can access this kind of information - but it is something of a best practice to use simple 'data' parameters since they are compatible with more traditional un-collected datasets.

This commit really needs more comments - but Philip Mabon has been patiently waiting for this functionality for a long time.
2015-02-02 16:29:42 -05:00
John Chilton 08ab0c6376 A simpler way to configure output pairs (exploiting static structure).
See example and comments in test/functional/tools/collection_creates_pair_from_type.xml.
2015-01-15 09:30:00 -05:00
John Chilton 2cb7c8d73e More configurable format and metadata handling for output collections.
Imporvements to testing code.
2015-01-15 09:30:00 -05:00
John Chilton 4c5c8a47db Allow tools to output collections with a dynamic number of datasets.
Models:

Track whether dataset collections have been populated yet.

Dataset collections are still effectively immutable once populated - but dynamic output collections require them to be sort of like `final` fields in Java (analogy courtesy of JJ) - allowing them to be declared before they are initialized or populated. This is tracked by the `populated_state` field.

Tools:

Output collections can now describe `discover_datasets` elements just like datasets - except in this case instead of dynamically populating new datasets in the history - they will comprise the collection. `designation` has been reused to serve as the element_identifier for the collection element corresponding to the dataset.

See Pull Request 356 for more information on the discover_datasets tag https://bitbucket.org/galaxy/galaxy-central/pull-request/356/enhancements-for-runtime-discovered.

Workflows:

Update workflow execution and recovery for dynamic output collections.

Galaxy workflow data flow before collections

* - * - * - * - * - *

Galaxy worfklow data flow after collections (iteration 1)

* - * - * \
           * - * - *
* - * - * /         \
                     * - * - *
* - * - * \         /
           * - * - *
* - * - * /

Galaxy worfklow data flow after this commit

              / * - * \
         * - *         * - *
        /     \ * - * /     \
       /                     \
      /                       \
     /        / * - * \        \
* - * -- * - *         * - * -- * - *
     \        \ * - * /        /
      \                       /
       \                     /
        \     / * - * \     /
         * - *         * - *
              \ * - * /
2015-01-15 09:30:00 -05:00
John Chilton 44f7317fa5 Allow tools to output collections with static or determinable structure.
By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)

There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.

Model:

The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.

Workflow:

Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.

Tool Testing:

This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.

Tests:

Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.

Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.
2015-01-15 09:30:00 -05:00
John Chilton 7e45ca2a72 Allow multiple tools with the same id in ToolBox.
How to use:

 1.) Place multiple tools with different IDs in your tool conf.
 2.) ... ummm ... no step 2 - just use the tools.

Implementation:

The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.

To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.

Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
2014-12-31 18:21:10 -05:00
John Chilton 8e060c77bc Implement expect_failure and expect_exit_code on test tag of tools. 2014-12-17 14:22:08 -05:00
John Chilton 77ee2ace67 Allow tool tests to assert properties about command, standard output, and standard error.
See test/functional/tools/job_properties.xml for example. Provides full access to assertion based XML tags - tabular and XML based assertions probably not so interesting so these include - has_text, has_line, not_has_text, has_text_matching, has_line_matching.
2014-12-17 12:34:39 -05:00
John Chilton 0a1a871fd0 Merge next-stable. 2014-12-17 09:41:26 -05:00
John Chilton c4d29ad1b6 Fix gz/zip upload of files to test framework.
With test case to ensure they don't break again. Problem was due to a small switch of an if to an elif in 7114d15a6d22 which had been sort of masking over an older bug.

Thanks to Peter Cock for reporting the issue.
2014-12-17 09:38:49 -05:00
John Chilton 82e117fadd Add code_file example tool from Bjoern. 2014-12-15 14:37:22 -05:00
John Chilton b0e43e567e Allow specification of multiple data manager configuration files...
... use this to create API functional tests for data managers.
2014-12-15 00:20:35 -05:00
John Chilton ec24396140 Refactoring: Start abstracting XML processing out of Tool and param classes.
Long term this could allow Galaxy to support - multiple tooling formats (Galaxy-like YAML, CWL http://bit.ly/cwltooldesc, etc...). But I think it is also important from a purely design perspective - this is a core logic class integrating different components - they should not also be doing XML parsing.

To verify the interface for parsing tools is expressive enough to allow multiple useful implementations, I built a test YAML tool description that implements many of the same features as Galaxy but smooths out rough edges (uses exit codes for job failure by default for instance). Loading these tools is disabled by default and it is not documented how to enable them because they are not intended to be part of Galaxy's public API.
2014-12-11 00:26:13 -05:00
John Chilton 251155887e Merge next-stable. 2014-12-05 20:36:48 -05:00
John Chilton 5b6c092aa0 Twill tool tests allowed selects specified by display text instead of value.
This didn't work for API tests - restore this functionality and add test case in test/functional/tools/multi_select.xml to verify this behavior is correct.
2014-12-05 20:35:22 -05:00
John Chilton 2a23da50bd Test case for passing null to optional tool input. 2014-12-04 12:28:23 -05:00
John Chilton 7ffccb7075 Implement tool test demonstrating pull request 569.
The new assign_primary_output attribute on discover_datasets tag.
2014-11-25 22:13:17 -05:00
John Chilton d57088409c Fix data column parameters pointed at multiple data parameters.
Before it would just die with an unhelpful server side exception - now it attempts to build, validate, and use a useful set of columns.
2014-09-26 10:49:43 -04:00
John Chilton 6ff787d769 More tool functional tests for validation stuff.
Test default sanitization in repeat. Basic test of simpler santizer and mapping.
2014-09-16 21:29:10 -04:00
Nate Coraor 8fadc6392b Rename universe_wsgi.ini -> galaxy.ini everywhere. 2014-09-15 16:43:51 -04:00
John Chilton 21f4d49a9e Synchronize validation of workflows between web and API controllers.
Reduces code duplication and does more correct checking of workflow step replacement parameters. More parameter checking functional tests.
2014-09-10 11:51:10 -04:00
John Chilton c5511cd747 Add very basic functional tests for tool param validation.
Including same test with workflow parameter substition - gotta admit I thought that workflow test was going to fail - so this week is looking pretty good :).
2014-09-08 10:33:54 -04:00
John Chilton abf4005749 Add another example tool - to illustrate three ways to collect files for concatenation. 2014-09-05 15:19:18 -04:00
John Chilton b21bb6f2ad More discovered dataset testing tweaks. 2014-09-05 15:19:18 -04:00
John Chilton c955e41867 Fill out tool test for configured discovered datasets.
Add new example based on Anton's bamtools split work and label all the outputs as visible - this should likely be the default but until it is might as well demonstrate them as visible since that is how they are most useful.
2014-09-03 15:17:49 -04:00
John Chilton c1ee4c9b31 Add functional test tool demonstrating output filters.
Add new 'expect_num_outputs' attribute to 'test' element in tool XML to verify the produced number - needed to test output filtering.
2014-09-03 08:43:25 -04:00
John Chilton fcce8b8218 Add functional test tool demonstrating setting output dataset formats.
Directly setting format attribute, setting to format to 'input' (ambigious, non-deterministic and should be deprecated IMO), using format_source, and using change_format actions.
2014-09-03 08:43:25 -04:00
John Chilton db07824a12 Add functional test tool demonstrating special parameters.
:(
2014-09-03 08:43:25 -04:00
John Chilton ef1877b468 Small tweaks to some functional test tools.
Add data labels so the tools works in the workflow editor and added a conditional switches to some with collection params and multiple input data parameters to test some state-y logic in workflow editor.
2014-09-03 08:43:25 -04:00
John Chilton 401fad5e25 Add some more basic test tools for building simple workflows for testing. 2014-08-27 16:32:26 -04:00
John Chilton 565580a564 Update readme for functional test tools directory. 2014-08-27 16:32:26 -04:00
John Chilton ec34167345 Fix multi-page tool help parsing.
Patch thanks to Hans-Philipp Brachvogel.
2014-08-12 08:19:04 -04:00
John Chilton 5979ac41c8 Allow using data collection steps via workflow API.
Implement API test for this and fixup test for previous commit related improved workflow run endpoint.
2014-07-28 19:18:40 -04:00
John Chilton ebe93b1673 Merged in jmchilton/galaxy-central-fork-1 (pull request #440)
Initial BibTeX/DOI citation support in tools and histories.
2014-07-28 13:15:32 -04:00
John Chilton cb9fc095fc Initial BibTeX/DOI citation support in tools and histories.
Allow tool authors to specify citations using a DOI or BibTeX. BibTeX can be specified either by pointing at a BibTeX file in the tool directory or by embedding bibtex entries right in tool citation blocks. If referencing a file parallel to the tool, the file should contain only a single BibTeX entry, this restriction can be easily lifted by adding a BibTeX parser as a Python dependency for Galaxy - but I do not have permission to do this.

These citations will appear at the bottom of the tool form in a formatted way but the user will have to option to select RAW BibTeX for copying and pasting. Likewise, the history menu now has an option allowing users to aggregatesuch citations across an analysis a comparable list of citations.

UI interactions are implemented using a Backbone model and view and data is fetched from the Galaxy server as BibTeX using the API. Two API entry points have been added - one to fetch the BibTeX entries for a tool and another for a history.

BibTeX entries for citations annotated with DOIs will be fetched from http://dx.doi.org/ and cached using Beaker.

Additional Limitations:

 - I am not super happy with a few different GUI elements of this. It is ugly and I didn't write the BibTeX parser but I did write the code that takes parsed BibTeX and converts it to a formatted entry. If merged, I will outline a Trello card to follow up and improve the UI and find some more standard way to build a formatted HTML citation from a parsed BibTeX entry.

 - BibTeX Limitations: LaTeX embedded in the BibTeX entries doesn't render properly when producing a "pretty" citation in the GUI (should still be exported to citation managers properly though). Cross references aren't supported at this time.

Alternative Implementations:

BibTex/DOI vs. PROV:

There was some discussion of PROV encoding citation information on the development mailing list. The citation tags on tools are typed so this could certainly be added - but it was discussed at the BOSC 2014 codefest and there was some conensus that tool authors are more likely to already have BibTeX or DOIs available for the tool's references and the major reference managers end users will likely plug these citations into while writing papers are more likely to be able to consume BibTeX than anything else.

The bench biologist using Galaxy is the consumer of this work I most concerned with - if we need to convert citations into other formats such as EndNote or Word's bibliography support there is a suite of tools we could optionally plug into Galaxy (http://sourceforge.net/p/bibutils/home/Bibutils/) to enable this down the road. PROV seems to have such an ecosystem to leverage - I could not even find a tool to convert it BibTeX.

Parse BibTeX Client vs Server:

As mentioned above it would be nice in some ways to be able to parse and reason about BibTeX on the backend - but it would require adding a new dependency to Python. Since we have to ship BibTeX to the browser anyway to allow users to copy and paste it - I decided it was easier to start with parsing and formatting BibTeX on the client side. I therefore added the following library BSD JavaScript dependency https://github.com/mayanklahiri/bib2json to enable this. I would be happy to revisit this decision and produce formatted entries server side - there seem to be more Python options for doing this than JavaScript.
2014-07-23 15:43:13 -05:00
John Chilton 3020dcd037 Allow discovered datasets to use input data format in 'ext' definition. 2014-06-29 13:31:33 -05:00
John Chilton 3d9da261d1 Allow tool conf 'tool_path' resolution relative to conf file.
Relative 'tool_path' directories are resolved relative to GALAXY_ROOT (or really working directory). This is probably the expected behavior - but I think it is advantageous to be able to find tools relative to the tool_conf file also. This can now be done using string.Template resolution of ${tool_conf_dir}.
2014-06-23 16:32:24 -05:00
John Chilton 1c5691964e Implement optional collection params.
Was already parsing optional attribute but I put exactly zero thought into the implementation so these didn't work at all I don't think. This fills out the implementation, adds a test tool, and some cheetah helpers to facilitate this: "#if $collect_param" will fail if input not supplied or collection is empty and "#if $collect_param.is_input_supplied" will fail is input not supplied (i.e. empty collections will pass this check).
2014-07-25 10:28:50 -05:00
John Chilton 9150a884aa Bugfix in test tool demoing two collection params.
Bit problematic functional tests passed despite this bug.
2014-07-24 15:30:51 -05:00
John Chilton 78fc3c9850 Add functional tool tests exercising multiple collection parameters at once. 2014-07-24 15:06:04 -05:00
John Chilton f2f21e2bed Allow tools and deployers to specify optional Docker-based dependency resolution.
Testing it out:
---------------

 - Install [docker](http://docker.io) (tough, but getting easier).
 - Copy `test/functional/tools/catDocker.xml` to somewhere in `tools/` and add to `tool_conf.xml`.
 - Add `<param id="docker_enabled">true</param>` to your favorite job destination.
 - Run the tool.

Description and Configuration:
------------------------------

Works with all stock job runners including remote jobs with the LWR.

Supports file system isolation allowing deployer to determine what paths are exposed to container and optionally allowing these to be read-only. They can be overridden or extended but defaults are provided that attempt to guess what should be read-only and what should be writable based on Galaxy's configuration and the job destination. Lots of details and discussion in job_conf.xml.sample_advanced.

`$GALAXY_SLOTS` (however it is configured for the given runner) is passed into the container at runtime and will be available.

Tools are allowed to explicitly annotate what container should be used to run the tool. I added in hooks to allow a more expansive approach where containers could be linked to requirements and resolved that way. To be clear, this is not implemented at all but the class ContainerRegistry is instantiated, passed the list of requirements, and given the chance to return a list of potential containers... if someone wants to implement this someday.

From a reproducibility stand-point it makes sense for tool author's to have control over which container is selected, but there is this security and isolation aspect to these enhancements as well. So there are some more advanced options that allow deployers (instead of tool authors) to decide which containers are selected for jobs. `docker_default_container_id` can be added to a destination to cause that container to be used for all un-mapped tools - which will result in every job on that destination being run in a docker container. If the deployer does not even trust those tools annotated with image ids - they can go a step further and set `docker_container_id_override` instead. This will likewise cause all jobs to run in a container - but the tool details themselves will be ignored and *EVERY* tool will use the specified container.

Additional advanced docker options are available to control memory, enable network access (disabled by default), where docker is found, if and how sudo is used, etc.... These are all documented in `job_conf.xml.sample_advanced`.

Implementation Details:
-----------------------

Metadata is set outside the container - so the container itself only needs to supply the underlying application and doesn't need to be configured with Galaxy for instance. Likewise - traditional `tool_dependency_dir` based dependency resolution is disabled when job is run in a container - for now it is assumed the container will supply these dependencies.

What's Next:
------------

If implementation is merged, much is left to be discussed and worked through - how to fetch and control what images are fetched (right now the code just assumes if you have docker enabled all referenced images are available), where to fetch images from, tool shed integration (host a tool shed docker repository?, services to build docker images preconfigured with tool shed depedencies?), etc.... This is meant as more of a foundation for the dependency resolution and job runner portions of this.
2014-06-07 01:45:05 -05:00
John Chilton e3fbe97116 Bugfix: More left/right -> forward/reverse fixes. 2014-05-27 12:52:41 -05:00
John Chilton 3e7b1b511a Bugfix: Hadn't implemented clean way to iterate of collection element names in tools.
Perhaps stretching the definition of bugfix here to get this in next-stable.
2014-05-27 12:52:41 -05:00
John Chilton d7b4b36756 Change paired collections terminology from left/right to forward/reverse.
If you have paired datasets in your database they will no longer work - send me an e-mail and I can give you an SQL update statement.
2014-05-15 13:01:36 -05:00
John Chilton 3688e02475 Add framework test tool that has both data and data_collection param - with test case.
Tool test is mildly useful to verify this works for single execution - but more useful as basis for future changesets testing mixed collection and subcollection mapping tool executions.
2014-05-07 15:39:03 -05:00
John Chilton 5410d476e8 Dataset collections - "reduce" with existing tools.
Allow users to select dataset collections in place of individual datasets for data tool parameters with multiple="true" enabled (if all elements of collection would be valid as input to this parameter).

Restrict collection reductions to flat collections. If a user wanted to reduce a nested collection they probably want to map of the subcollections reducing each and building a collection of the reductions. TODO: The sentence is probably unintelligiable, need to provide a concrete example.

A functional test demonstrating these reductions in included.
2014-05-06 08:54:30 -05:00
John Chilton f42d97fef3 Dataset collections - tool parameters - allow tool tests.
With example tool tests demonstrating some basic features available to cheetah for data_collection parameters.
2014-05-06 08:54:30 -05:00
John Chilton ddcad1582c Merged in jmchilton/galaxy-central-fork-1 (pull request #356)
Enhancements for Runtime Discovered (Collected Primary) Datasets
2014-05-06 08:13:29 -05:00