Commit Graph
124 Commits
Author SHA1 Message Date
John Chilton f823278efd Fix bug in workflows test for implicit connections between steps.
It would fail when being run with the rest of the suite and not on its own - because it was using the same id for the workflow id and invocation id - which is obviously wrong unless it is a completely fresh database :).
2015-03-03 09:37:30 -05:00
John Chilton fc0a6d8764 Fix intermittently failing test test_workflow_run_dynamic_output_collections.
Love it when the bugs correspond to actual TODOs I left in the code.
2015-03-03 09:30:55 -05:00
John Chilton 391abd5fd6 Remove assert False from test/api/test_workflows_from_yaml.py.
These aren't ideal tests - but it is some indication that things are working that the API will import the workflow and produce a representation. Should follow up at some point and verify the representation is in fact the correct one.
2015-03-02 22:39:58 -05:00
John Chilton f8bfe89c42 Clarify failing test test_tools.py. 2015-03-02 22:39:58 -05:00
John Chilton a20938fc48 Attempt to fix intermittently failing jobs API test.
Retry check several times in case there is some sort of timing problem where a history turns okay - before a job. Improve assertion error message.
2015-03-02 22:39:50 -05:00
John Chilton b1f3654208 Tool data API testing fixes.
- Add files I didn't commit previously in test/functional/tool-data/.
 - Comment out test for tool data table deletion since Dan fixed the problem of regular users running data managers through the API (https://github.com/galaxyproject/galaxy/commit/48f77dc742acf01ddbafafcc4634e69378f1f020) that I previously exploited to write the test.
2015-03-02 22:36:24 -05:00
Carl Eberhard 14383c0868 API, tests: add datasets API tests in test/API and casperjs sections; Browser tests: add datasets api, allow api to send extra params for history show & index 2015-02-03 16:11:10 -05:00
John Chilton 25e1517f37 Drop test that is failing due to improved HDA accessibility consistency. 2015-02-03 12:20:41 -05:00
John Chilton 2f15eb0d78 Expose improved sample tracking to tools for implicit map/reduce ops.
Tools may now use $input.element_identifier during tool evalution for input 'data' parameters with the following semantics:

 - If the input was specified as a single dataset by the user - this just fallbacks to providing the $input.name.
 - If the input was mapped over a collection (to produce many jobs) or if the input is a 'multiple="true"' input that was provided a collection - the $input.element_identifier will be the element identifier for the corresponding collection item (generally much more useful the dataset name - since if preserved throughout workflows).

'data_collection' parameters already can access this kind of information - but it is something of a best practice to use simple 'data' parameters since they are compatible with more traditional un-collected datasets.

This commit really needs more comments - but Philip Mabon has been patiently waiting for this functionality for a long time.
2015-02-02 16:29:42 -05:00
John Chilton 4939c45e3d Test cases for select validation handling.
Normal selects seem to be prevented from execution with invalid parameter values, but not columns. Values are escaped properly so shell exploitation isn't the problem - but as a usability thing Galaxy should prevent execution and provide a warning message.
2015-01-27 12:14:56 -05:00
John Chilton facc29e8a9 Update workflow extraction backend for output collections. 2015-01-15 09:30:00 -05:00
John Chilton 2cb7c8d73e More configurable format and metadata handling for output collections.
Imporvements to testing code.
2015-01-15 09:30:00 -05:00
John Chilton 4c5c8a47db Allow tools to output collections with a dynamic number of datasets.
Models:

Track whether dataset collections have been populated yet.

Dataset collections are still effectively immutable once populated - but dynamic output collections require them to be sort of like `final` fields in Java (analogy courtesy of JJ) - allowing them to be declared before they are initialized or populated. This is tracked by the `populated_state` field.

Tools:

Output collections can now describe `discover_datasets` elements just like datasets - except in this case instead of dynamically populating new datasets in the history - they will comprise the collection. `designation` has been reused to serve as the element_identifier for the collection element corresponding to the dataset.

See Pull Request 356 for more information on the discover_datasets tag https://bitbucket.org/galaxy/galaxy-central/pull-request/356/enhancements-for-runtime-discovered.

Workflows:

Update workflow execution and recovery for dynamic output collections.

Galaxy workflow data flow before collections

* - * - * - * - * - *

Galaxy worfklow data flow after collections (iteration 1)

* - * - * \
           * - * - *
* - * - * /         \
                     * - * - *
* - * - * \         /
           * - * - *
* - * - * /

Galaxy worfklow data flow after this commit

              / * - * \
         * - *         * - *
        /     \ * - * /     \
       /                     \
      /                       \
     /        / * - * \        \
* - * -- * - *         * - * -- * - *
     \        \ * - * /        /
      \                       /
       \                     /
        \     / * - * \     /
         * - *         * - *
              \ * - * /
2015-01-15 09:30:00 -05:00
John Chilton 44f7317fa5 Allow tools to output collections with static or determinable structure.
By "static" I mean tools such as a FASTQ de-interlacer that would produce a "paired" collection with two datasets everytime. By "determinable" I mean tools that perform N->N operations within the same job - such as a tool that needs to normalize a bunch of datasets all at once and not in separate jobs. (For N->N collection operations that should or can be done in N separate jobs tool authors should just write tools that operate over a dataset and produce a dataset and let the end-user 'map over' that operation.)

There are still large classes of operations where the structure of the output collection cannot be pre-determined - such as splitting files (e.g. bam files by read group) - that are not implemented in this commit.

Model:

The models have been updated to do a more thorough job of tracking collection outputs. Jobs just producing HistoryDatasetCollectionAssociations works fine for simple jobs producing collections - but you don't want to map a list over a tool that produces a pair and produce a bunch of pairs HDCAs and a list:pair HDCA- you just want a bunch of pieces and the one list:pair at that the top.

Workflow:

Workflows containing such operations can be executed - but the workflow editor has not been updated to handle this complexity (and it will require a significant overhaul) so such tools are not available in the workflow editor.

Tool Testing:

This commit also introduces a new tool XML syntax for describing tests on output collections. See files test/functional/tools/collection_creates_list.xml and test/functional/tools/collection_creates_pair.xml for examples.

Tests:

Includes two tools to test this - one that uses explicit pair output names and one that iterates over the structure of input list to produce an output list.

Includes several new tools API tests that test the tools described above via the API and implicit mapping over such tools. Includes two new workflow API tests - one that verifies a simple workflow with output collections works and one that verifies mapping over workflow steps in collections works.
2015-01-15 09:30:00 -05:00
John Chilton 630287500d Allow specification of post job actions at workflow invocation.
Feature requested by Kyle. Implemented only in the API at this point - not sure it is a feature valuable to UI consumers.

Pass in PJAs along with step parameters map but keyed on __POST_JOB_ACTIONS__. JSON definition same as when defining PJA in the workflow definition JSON.

Includes test cases for normal use and for use after delayed workflow steps have been evaluated by the new workflow scheduling stuff.
2015-01-05 15:38:27 -05:00
John Chilton 7e45ca2a72 Allow multiple tools with the same id in ToolBox.
How to use:

 1.) Place multiple tools with different IDs in your tool conf.
 2.) ... ummm ... no step 2 - just use the tools.

Implementation:

The Tool Shed allows tool lineages by assigning each tool version a GUID and tracking versions in a database. This
implementation works by simply allowing the ToolBox to contain multiple tools with the same ID and orders them by the version specified by the tool author.

To track enable this a second tool lineage has been introduced that just uses tool versions instead of a database (non-toolshed installed tools are not longer placed into the Tool Shed install database). The ToolBox has been updated to allow multiple versions per tool id (defaulting to the 'latest' version for all operations which do not specify a version). Both jobs and workflow steps would track tool versions but did not use that version when fetching tools from the Toolbox - these components have been updated to try to use the tool version.

Unit tests working through most of the ToolBox and tool panel have been added, as well as functional tests exercising the tools API and to ensure workflows now at least attempt to respect tool versions (still kind of silently switches versions in some cases). Manual tests against the new tool form seem to demonstrate the tool switching and tool re-running work with only minor changes to the tools API and the job handler.
2014-12-31 18:21:10 -05:00
John Chilton c70cc95128 Expose datatype converter information via datatype API.
With test case.
2014-12-25 12:24:23 -05:00
John Chilton f68a103ae0 Suite of API tests for exist datatypes endpoint. 2014-12-25 12:24:23 -05:00
John Chilton 285896dfe1 Increase consistency between job index and show.
Add tests for new date range and history filtering as well as to ensure only admins get to see external_id and command_line and that users cannot see each other's jobs.

Always allow admins to views all jobs on index (instead of only when user_details is specified) and show. Add user_email and external_id to show (for admins) to bring it inline with index and add command_line to index (for admins) to bring it in line with show.

Show still allow more details including job standard error and output as well as job metrics.
2014-12-18 21:50:15 -05:00
John Chilton 7c9096ae12 Allow workflow inputs to be specifiable by UUID without explicit inputs_by.
Since there is no chance of conflict with step.id or step.order_index, just allow inputs or ds_map respectively to be indexed by UUID.
2014-12-15 22:20:12 -05:00
John Chilton 623d8ce9c9 Allow specifing workflow inputs by step UUID.
Pass in the parameter 'inputs_by' as 'step_uuid' to the workflow run command to use this.

Specifing inputs by UUID has the nice advantage that it survives workflows saves - so if one sets up an API script or something to target a workflow - saving the workflow in the editor doesn't need to break the script as long as inputs were not added or deleted. The UUID (like the order_index) has the advantage of the step id that it is predeterminable - so one can set it up a workflow script against any Galaxy and the script doesn't need to be adapted to raw ids the steps get assigned in that instance.
2014-12-15 22:20:12 -05:00
John Chilton 4f9e0bab8d Allow specifing workflow parameter replacements by step UUID. 2014-12-15 22:20:12 -05:00
John Chilton f7493341fa Ensure unique step label and UUIDs across workflows during create/update. 2014-12-15 22:20:12 -05:00
John Chilton 7dfdf530cf Augment workflow step model for improved tracking.
Give every step a UUID that can be preserved across edit to the workflow. Likewise - allow every step to be given a label (a unique short name for that workflow) - that allows for a human consumable way to reference steps for use in tests and when driving workflows via the API. Workflow editor doesn't yet (and might never) display these attributes but it does preserve them across workflow saves.

Implement automated testing that uploading workflows, updating workflows, and exporting them handle UUID and labels. Manually tested workflow editor preserves labels and UUID across changes.
2014-12-15 22:20:12 -05:00
John Chilton e0a5e82bda Allow implicit connections between workflow steps (no editor GUI yet).
This means steps that are not connecting an output of one step to the input of another. This could potentially address all sorts of untraditional (in a Galaxy sense) workflows where some sort of data is managed externally. The most important use I think I have heard discussed is that of data managers - this can be used in cases where data managers depend on one another (grab the fasta files in one step, index them in another) or workflows where a downstream analysis depends on index data populated via data managers in earlier steps.

Not really sure how to represent these in the workflow editor - but this is a power user feature anyway so hopefully that is not super pressing. The YAML to workflow DSL supports the operation (see test cases) so these power users (a euphemism for Dan I guess) can just use that for now.
2014-12-15 22:20:11 -05:00
John Chilton 58ce775d14 Move more workflow logic out of controller into manager.
This logic for building up editor representation of the workflow.

Introduce concept of a workflow to dict style - with 'export' and 'editor' as first cracks.
2014-12-15 22:20:11 -05:00
John Chilton 2565f8b256 Refactor updating workflow contents out of mixin into manager.
Introduce an API endpoint for this operation with test case.
2014-12-15 22:20:11 -05:00
John Chilton c526ed87da Refactor helper out for detailed encoding of a workflow for API. 2014-12-15 22:20:11 -05:00
John Chilton d6a03acef9 Implement test case for Pull Request #577. 2014-12-15 00:20:35 -05:00
John Chilton b0e43e567e Allow specification of multiple data manager configuration files...
... use this to create API functional tests for data managers.
2014-12-15 00:20:35 -05:00
John Chilton 8536ca4315 Allow downloading index files via tool data API.
This provides direct access to the files to admins - probably still wise to provide some mechanism to download a copressed archive of these files.
2014-12-15 00:20:35 -05:00
John Chilton 4262ff7673 Implement a detailed break down of data table fields via API.
Originally this approach was laid out by Kyle Ellrott in this pull request (https://bitbucket.org/galaxy/galaxy-central/pull-request/531/add-downloads-to-tool-data-api/diff). The changes to galaxy.tools.data are entirely his contribution, I only reworked the API and endpoint slightly and did some stylistic fixes and refactoring.
2014-12-15 00:20:35 -05:00
John Chilton 33463efe13 Allow multiple tool data table config files...
... use this functionality to implement tests for tool data API.
2014-12-15 00:20:35 -05:00
John Chilton e5e3b82460 Minor de-duplication in test_workflows.py API tests. 2014-12-13 13:02:35 -05:00
John Chilton 78472d24c4 Extend workflow YAML syntax to work with pause steps. 2014-12-13 12:58:34 -05:00
John Chilton 5b3700a53e Fix for YAML workflow DSL respect multiple input connections.
With test case for this.
2014-12-13 12:58:34 -05:00
John Chilton f80c5f2ca7 Test for importing workflows with annotations. 2014-12-13 12:58:34 -05:00
John Chilton c515e18fc4 Improvements to to yaml_to_workflow for Kyle.
Add UUID to workflows (his contribution) - add shortcuts for rename an hide actions with tests (his request, my implementation).
2014-12-04 16:18:20 -05:00
John Chilton 2a23da50bd Test case for passing null to optional tool input. 2014-12-04 12:28:23 -05:00
John Chilton 79b18e1337 Easier debugging of API test problems. 2014-12-04 09:16:01 -05:00
John Chilton 5512f531d9 Improvements to test/api/test_workflow_extraction.py.
Fill out the dataset collection parameter test and add a new test for a workflow that includes subcollection mapping. Move toward orchestrating jobs in this file via a high-level YAML description of steps and test data as well as some higher-level methods for testing various stuff about extracted workflows (step counts, input types, tools used, connected-ness, etc...).

Allow ordering jobs index API by create_time instead of update_time.
2014-12-02 22:17:25 -05:00
John Chilton 93e2389484 Introduce another helper to simplify test_workflow_extraction.py. 2014-12-01 21:18:56 -05:00
John Chilton 942cdc48b2 Simplify interface to test helper _extract_and_download_workflow.
Makes tests in this file marginally more readable.
2014-12-01 21:18:56 -05:00
John Chilton 10f62a087f Simplify new test_workflow_extraction.py.
Now makes sense to assume all tests consume a history.
2014-12-01 21:18:56 -05:00
John Chilton 9cc3ce15a1 Split all workflow extraction testing out of test/api/test_workflows.py.
test_workflows.py was pretty unwieldy and the workflow extraction tests were very different than the other tests in that file (e.g. very different helper functions).
2014-12-01 21:18:56 -05:00
John Chilton 8c9cb97878 Script to build workflows from simpler YAML description. 2014-12-01 21:18:56 -05:00
John Chilton 9c4dc70f97 Add file missed with pull request 505. 2014-12-01 21:18:56 -05:00
John Chilton 12f8776908 Tool test - verify that if 'batch' meta-parameter is False - matching is not attempted. 2014-11-24 13:02:18 -05:00
John Chilton bf4be0e936 Tool API test for mixing batch and non-batch input datasets. 2014-11-24 12:56:05 -05:00
John Chilton cd0faf9e96 Tool test for broken behavior where nested parameter replacements are passed in for a workflow step. 2014-11-14 13:01:10 -05:00