Properly order the installation of the repositories adding
prior_installation_required="True". Otherwise filter_1440 is installed first
and ends in INSTALLED status because of the logic of
Repository.create_tool_dependency_with_initialized_env_sh_file() method
contained in
lib/tool_shed/galaxy_install/tool_dependencies/recipe/tag_handler.py .
This strategy proved to work around certain race conditions in tool testing so hopefully it will solve the transiently failing job searching and filtering test cases.
Fetch HIDs for these datasets all at once and flush once for all assignments. Creating jobs from 4 threads each job producing 10 outputs - this resulted in reducing the runtime of this portion of the code by 90% and saving around 20 seconds per job.
Before:
galaxy.tools.actions INFO 2015-11-30 15:53:46,235 Add outputs to history (22494.829 ms)
galaxy.tools.actions INFO 2015-11-30 15:53:47,295 Add outputs to history (22239.062 ms)
galaxy.tools.actions INFO 2015-11-30 15:53:48,320 Add outputs to history (23781.079 ms)
galaxy.tools.actions INFO 2015-11-30 15:53:51,490 Add outputs to history (21820.567 ms)
galaxy.tools.actions INFO 2015-11-30 15:54:24,506 Add outputs to history (25353.837 ms)
After:
galaxy.tools.actions INFO 2015-11-30 16:08:47,640 Add outputs to history (2781.675 ms)
galaxy.tools.actions INFO 2015-11-30 16:08:47,860 Add outputs to history (3177.738 ms)
galaxy.tools.actions INFO 2015-11-30 16:08:48,776 Add outputs to history (2425.528 ms)
galaxy.tools.actions INFO 2015-11-30 16:09:02,942 Add outputs to history (2579.022 ms)
Should work when mapping over collections or for big muli-run tool submissions.
Because of database tension with sqlalchemy it is not strictly a linear increase, but the end user walltime experience for a 24 dataset collection being submitted using 4 threads instead 1 drops execution time from 68 seconds to 35.
Rebased with fixes thanks to @nsoranzo - https://github.com/jmchilton/galaxy/commit/7f6514a21222ea787a0676764e4a4fe825891b27#commitcomment-14712999.
- Don't refresh job when this is the only thread fetching it.
- Avoid a bunch of unnecessary flushes, just flush once essentially during process.
Brings this process from taking over 1 second per job on average on sqlite for cat1 on my laptop to around 400 ms on average.
Add proper idpDB sniffing logic (akin to MzSQlite) and test case
Add H5 sniffing based on magic number and test case
Add SQLite/IdpDB/MzSQlite/GeminiSQLite and H5 to default datatypes in registry.py
Continue to delay calculation of this but auto-compute it if requested and module injection code hasn't been called explicitly. Moves logic internal to class where it belongs also.