Adding consistency allowing each parameter to be wrapped in a object describing the meta-properties of the submitting value - this was requested by Sam to make the new tool form easier to manage, it makes multi-running properties work for non-data parameters, and allows linked/unlinked specification of parameters.
(Lot easier now that I understand what all of the module methods are doing and have an example of a 4th module downstream.) Now with even more unit tests.
Including same test with workflow parameter substition - gotta admit I thought that workflow test was going to fail - so this week is looking pretty good :).
There may now be multiple WorkflowInvocationSteps for each WorkflowStep for steps that are mapped over collections - so to_dict creating a dictionary of this information indexed on order step is problematic because only one WorkflowInvocationStep will be represented per step. Instead now just returning a big list of all of the invocations - which contains all of the same information. This is a backward incompatible API change for the workflow invocation API.
Also update the input mapping stuff with logic for dealing with data collection inputs.
Sharing a workflow with a user was previously sufficient to grant access to all invocations of that workflow. This isn't a huge problem since the information potentially leaking out was limitted to invocation counts, various encoded ids, and update times. Still I think no information about invocations should be avaialble as a result of sharing a workflow - and upcoming changes to Galaxy will result in much more information being made available via the workflow invocation API.
Add new example based on Anton's bamtools split work and label all the outputs as visible - this should likely be the default but until it is might as well demonstrate them as visible since that is how they are most useful.
It was stubbing out connection stuff which worked fine when running test on its own - but after any test sets up the sql alchemy mappings this breaks down because events are trying to fire. This just uses actual model classes - which is just as easy anyway really.
With refactoring to reduce cyclomatic complexity. Also renaming 'input_ext' what it actually is 'random_input_ext'. We should fix that or at least issue a huge warning if we detect 'input' could have reasonable been different things.
Directly setting format attribute, setting to format to 'input' (ambigious, non-deterministic and should be deprecated IMO), using format_source, and using change_format actions.
Add data labels so the tools works in the workflow editor and added a conditional switches to some with collection params and multiple input data parameters to test some state-y logic in workflow editor.
Detail bug report from Michael Crusoe here : https://trello.com/c/0mdGCx4P.
This also fixes a test case added in 289e48b which was both attempting to assert something wrong and was incorrectly implemented. Augmenting the workflow editor test suite with some actually valid test cases that assert the correct behaviors.
For a longer explaination - workflow input terminals have two related concepts 'canAccept' and 'attachable'. An output terminal is 'attachable' if in the abstract it could be attached to the input regardless of whether the input is already filled or not. 'canAccept' is more stateful in that an output terminal is 'canAccept'able if the input terminal is not filled ('_inputFilled') and it is 'attachable'.
So - the problem was 'attachable' was not correctly defined for a multiple input data parameters. 'attachable' was asserting that any connected input terminal could not be 'attachable' by an collection output terminal - so on these asynchronous node state changes collections attached to multiple input data parameters were being wiped out. The more percise/correct distinction is that if a multiple input data parameter has single inputs connected to it - it cannot also have a collection connected to it (yet anyway). This fixes the mentioned bug.
The problematic test case was conflating 'attachable' and 'canAccept'able - I have fixed the test case to verify the correct 'attachable' logic and added newer, higher level test cases to test the canAccept logic and the actual behavior the end user would observe (of the connector being destroy).
Actually do some verification of imported history, eliminiate use of deprecated mixin, timeout operations that 'wait', break up big function and name test better.
In particular break out mixins defining context-based utilities for interacting with app, an user, and a history. In downstream work on workflow scheduling this proves nessecary and sufficient to create an context for 'execute'-ing tools (create jobs) outside of web threads in response to workflow requests.
Unit tests for some of this.
Add utility to RuleHelper to simplify bursting reasoning slightly, setup a stock rule for bursting between two static job destinations - mostly for demonstration purprose but it might be useful in some settings.
... if only in a limitted sort of way. Provide ability to hash jobs by history, workflow, user, etc... and choose among various destinations semi-randomly based on that to distribute work across destinations based on these factors.
Add a stock rule as in the form of a new dynamic destination type ("choose_one") that demonstrate using this to quickly proxy to some fixed static destinations. Add examples to job_conf.xml.sample_advanced.
Generate a workflow invocation UUID that is stored with each job in the workflow and allow dynamic job destinations to consume these. Should allow for grouping jobs from the same workflow together during resource allocation (a sort of first attempt at dealing with data locality in workflows).
In other words, send extra job destination parameters to dynamic rule functions as arguments (in addition to those dynamically populated by Galaxy itself). This enables greater parameterization of rule functions and should lead to cleaner separation of logic and data (i.e. sites can program rules that restrict access to users, but which users can be populated at a higher level in `job_conf.xml`).
For example the following dynamic job rule:
def cluster1(app, memory="4096", cores="1", hours="48"):
native_spec = "--time=%s:00:00 --nodes=1 --ntasks=%s --mem=%s" % ( hours, cores, memory )
return JobDestination( "cluster1", params=dict( native_specification=native_spec ) )
Could then be called with various parameters in job_conf.xml as follows:
<destination id="short_job" type="dyanmic">
<param id="function">cluster1</param>
<param id="hours">1</param>
</destination>
<destination id="big_job" type="dynamic">
<param id="function">cluster1</param>
<param id="cores">8</param>
<param id="memory">32768</param>
</destination>
If multiple destinations map to the same underlying resource (a very typical case), this could allow rule developer to reason about cluster as a whole - though perhaps clunkily - by supplying all destination ids mapping to that cluster.