Tests overridding various output properties (ext, dbkey, name, visible) with file name and galaxy.json, collecting from new_file_path versus job_working_directory, setting job output associations, and logic related to adding outputs to copied histories.
An example job_metrics_conf.xml.sample is included that describes which plugins are enabled and how they are configured. This will be updated for each new plugin added. By default not instrumentation or data collection occurs - but if a job_metrics.xml file is present it will serve as the default for all job destination. Additionally, individual job destinations may disable, load a different job metrics file, or define metrics directly in job_conf.xml in an embedded fashion. See comment at top of job_metrics_conf.xml for more information.
This commit include an initial plugin (named 'core') to demonstrate the framework and capture the highest priority data - namely the number of cores allocated to the job and the runtime of the job on the cluster. These two pieces of information alone should provide a much clearer picture of what Galaxy is actually allocating cluster compute cycles to.
Current limitations - This only works with job runners utilizing the job script module and the LWR (it utilizes the job script module on the remote server), hence it won't yet work with...
- Local job runner - I do have a downstream fork of Galaxy where I have reworked the local job runner to use the common job script template.
https://github.com/jmchilton/galaxy-central/commits/local_job_scripthttps://github.com/jmchilton/galaxy-central/commit/949db2cd14c7191cedf1febeb5f4a2f53123123e
- CLI runner - CLI runner needs to be reworked to use the job script module anyway so GALAXY_SLOTS works - the LWR version of the CLI runner uses the job script module - this work just needs to be back ported to Galaxy.
If a job_metrics_conf.xml is present and some jobs route to the above destinations - the jobs won't fail but annoying errors will appear in the logs. Simply attach a 'metrics="off"' those these specific job destinations to disable any attempt to use metrics for these jobs and disable these errors.
Workflow run template was causing these methods to be called wrong (and has been for a LONG time), this in turn caused me to misunderstand what the spec for how other_values in tool parameters should operate when writing the tests.
The real underlying bug should be fixed with 55cf8bb.
In particular, this moves get_job_dict out of workflow controller. I am making big changes to this downstream in dataset collection work so I want unit tests - additionally get_job_dict isn't really a great name for what this becomes - so I am renaming it to summarize.
This reduces code duplication related dataset_collectors now and abstracts out important functionality I reuse further to collect dataset collections downstream.
Was originally added in 28d43f4 as DatasetParamContext and backed out of right away. I have reworked it so that it no longer breaks implicit conversion, the relevant classes and methods have less generic names, and it has a healthy set of test cases.
In addition to basic tests on matching datasets to parameters, selections, and implicit conversions - these tests include testing of dataset security inconjuction with data_destination tools as well as filtering data parameters on other data parameters.
Test optional datasets can be used in tool evaluation (in test_evaluation.py).
Add test_data_parameters.py which test many random DataToolParameter behaviors. Test various paths to DataToolParameter.to_python - including recently enhanced ability to use optional dataset with 'multiple=True' data parameters. Test filtering on datatypes, implicit conversion options (both existing conversions and new ones) both when building HTML forms and picking intial values for workflows. Test special handling of hidden datasets. Tests picking intial datasets when optional and without repeats when used in subsequent calls.
Previous changes enabled targetted rewriting of specific kinds of paths - working directory, inputs, outputs, extra files, version path, etc.... This change allows rewriting remaining 'unstructured' paths - namely data indices.
Right now tool evaluation framework uses this capability only for SelectParameter values and fields - which is where these paths will be for data indices. Changeset lays out the recipe for doing this and the functionality could easily be extended for arbitrary parameters or other specific kinds of inputs.
The default ComputeEnvironment does not rewrite any paths obviously, but the abstract base class docstring lays out how to extend a ComputeEnvironment to do this:
def unstructured_path_rewriter( self ):
""" Return a function that takes in a value, determines if it is path
to be rewritten (will be passed non-path values as well - onus is on
this function to determine both if its input is a path and if it should
be rewritten.)
"""
The LwrComputeEnviroment has been updated to provide such a rewriter - it will rewrite such paths, and create a dict of paths that need to be transferred, etc.... The LWR server and client side infrastructure that enables this can be found in this changeset - https://bitbucket.org/jmchilton/lwr/commits/63981e79696337399edb42be5614bc7218cdf95f.
This changeset includes tests for changes to wrappers and the tool evaluation module to enable this.
Add version_path method to ComputeEnviornment interface - provide default implementation delegating to existing JobWrapper method as well as LwrComputeEnviornment piggy backing on recently added version command support.
Build full command to do this in JobWrapper.prepare so that ComputeEnviornment is available and potentially remote path can be used.
Pull code out of job wrapper and tool for building and evaluating against template environments and move them into a new ToolEvaluator class (in galaxy/tools/evalution.py). Introduce an abstraction (ComputeEnvironment) for various paths that get evaluated that may be different on a remote server (inputs, outputs, working directory, tools and config directory) and evaluate the template against an instance of this class. Created a default instance of this class (SharedComputeEnvironment). The idea will be that the LWR should be able to an LwrComputeEnvironment and send this to the JobWrapper when building up job inputs - nothing in this commit is LWR specific though so other runners should be able to remotely stage jobs using other mechanisms as well.
This commit adds extensive unit tests of this tool evaluation - testing many different branches through the code, with and without path rewriting, testing job hooks, config files, testing the cheetah evaluation of simple parameters, conditionals, repeats, and non-job stuff like $__app__ and $__root_dir__. As well as a new test case class for JobWrapper and TaskWrapper - though this just tests the relevant portions of that class - namely prepare and version handling.
For clarity - more work toward reduces tools/__init__.py to a more managable size. This also includes an initial suite of test cases for these wrappers - testing simple select wrapper, select wrapper with file options, select wrapper with drilldown widget, raw object wrapper, input value wrapper, and the dataset file name wrapper with and without false paths.
Want to refactor some stuff around in DefaultToolAction so can be reused when dealing with dataset collections downstream in https://github.com/jmchilton/galaxy-central/tree/collections_1 - so creating unit tests to ensure functionality is not changing.
This changeset also reworks test_execution.py moving more stuff to test/unit/tools_support.py to share between test files.
Slightly confusing that state params and raw incoming passed to that method, so pull out rerun_remap_job_id sooner and just pass that along (it was the only incoming was used for). Use the oppertunity to isolate potential errors with decoding rerun_remap_job_id and include more informative error message.
Add unit test to test invalid rerun_remap_job_ids.
Going to be doing some more work on tool state stuff so it will be good to have a way to test that. This also brings in test/unit/tools_support.py from Pull Request #287 (would be overkill for just these tests, but it is useful for future tests coming down the pipe.)
Add stronger test cases, though these are still insufficient to exercise the error. The only way I was able to get the history contents to come out misordered was to use postgres. Nonetheless, using postgres this now orders datasets properly.
... or would that be back-into 714f5b1?
714f5b1 - Move history contents filtering logic into model.
Filtering for deleted, visible, and ids with ORM should prevent loading unneeded objects into memory (a history with 1000 items and 5 visible will now only cause 5 items to be loaded instead of 1000 just to filter out 5). Simplifies history_contents API logic somewhat, allows easier 'unit' testing (included), and provides a clearer entry point for additional 'showing' additional contents types (read dataset collections).
Filtering for deleted, visible, and ids with ORM should prevent loading unneeded objects into memory (a history with 1000 items and 5 visible will now only cause 5 items to be loaded instead of 1000 just to filter out 5). Simplifies history_contents API logic somewhat, allows easier 'unit' testing (included), and provides a clearer entry point for additional 'showing' additional contents types (read dataset collections).
It can be 'none', 'local', or 'remote'. At this point this allows one to disable command dependency injection if set to 'none' or 'remote'. Actually implementing logic for 'remote' resolution will follow.
In subsequent changes - 'remote' will make backward-imcompatiable API calls to the LWR - this is why 'none' is also provided as an option - it will be backward compatiable way to disable dependency resolution for an LWR destination.
Old behavior can be reverted by setting up a dependency_resolvers_conf.xml with the following contents
```
<dependency_resolvers>
<tool_shed_packages />
<galaxy_packages />
</dependency_resolvers>
```
The new behavior corresponds to the following dependency_resolvers_conf.xml contents:
```
<dependency_resolvers>
<tool_shed_packages />
<galaxy_packages />
<galaxy_packages versionless="true" />
</dependency_resolvers>
```
Still think galaxy_packages should come before tool_shed_packages in resolution order so that deployers can fix broken tool shed installs with manual installs or deploy optimized versions of packages without having to mess with a dependency_resolvers_conf.xml file.
Add unit tests both for the old default parameters and the ability to replace these defaults.
This is all in case one wants to run the set metadata command on a remote server with different galaxy path, output paths, etc....