From 5539a083e21b601bfd2f87fa3417f23bf295a690 Mon Sep 17 00:00:00 2001 From: Nate Coraor Date: Wed, 18 Jun 2014 15:55:32 -0400 Subject: [PATCH] Add some new features to job execution: - Runner states - Allow runner plugins to provide job finish/failure conditions back to Galaxy so that actions can be taken on specific actions. Currently only the WALLTIME_REACHED state is defined. This can be set on the (Asynchronous)JobState's `runner_state` attribute. Only the slurm runner currently does this. - Runner state handlers - Pluggable interface for defining actions to take when runner state actions occur. Any (non _) python file in galaxy.jobs.runners.state_handlers will be loaded, but handler function names should match the step in the job lifecycle where they should be used. Only the 'failure' method is currently implemented, but adding more would be trivial. Clever parameterization a la the dynamic runner would be a nice improvement here. - Destination resubmission - Destinations in the job config can specify a new destination that jobs should be resubmitted to under certain conditions (currently the only condition implemented is walltime_reached on the original destination...) - Resubmit on walltime reached state handler plugin - The actual resubmission implementation. - RESUBMITTED Job state - Allows resubmitted jobs to bypass the normal ready to run checks and begin execution immediately. - RESUBMITTED DatasetInstance state - This was the best method Carl and I could come up with for persisting the resubmitted state so that it would be visible in the UI. It's not perfect but I didn't want to alter the schema and the only place it could go (job table) is not eagerloaded on history status updates. The resubmission code will not actully set this state yet (it is commented out) until the UI can cope with it. Bonus: once this is done we can pretty easily add a "job concurrency limit reached" to give users a visual cue on jobs waiting for that reason. This has been tested pretty extensively with job recovery, concurrency limits and multiprocess setups, which is to say that it will surely fail miserably in production. --- job_conf.xml.sample_advanced | 24 +++++++++ lib/galaxy/jobs/__init__.py | 33 +++++++++++- lib/galaxy/jobs/handler.py | 43 +++++++++++++-- lib/galaxy/jobs/mapper.py | 12 ++--- lib/galaxy/jobs/runners/__init__.py | 35 ++++++++++-- lib/galaxy/jobs/runners/slurm.py | 5 ++ .../jobs/runners/state_handler_factory.py | 54 +++++++++++++++++++ .../jobs/runners/state_handlers/__init__.py | 0 .../jobs/runners/state_handlers/resubmit.py | 39 ++++++++++++++ lib/galaxy/model/__init__.py | 5 +- 10 files changed, 234 insertions(+), 16 deletions(-) create mode 100644 lib/galaxy/jobs/runners/state_handler_factory.py create mode 100644 lib/galaxy/jobs/runners/state_handlers/__init__.py create mode 100644 lib/galaxy/jobs/runners/state_handlers/resubmit.py diff --git a/job_conf.xml.sample_advanced b/job_conf.xml.sample_advanced index 346b0f2461f..e4fb8975a13 100644 --- a/job_conf.xml.sample_advanced +++ b/job_conf.xml.sample_advanced @@ -323,6 +323,30 @@ --> 8 + + + + --time=00:05:00 --nodes=1 + + + + + -l h_rt=96:00:00 + +