From 7c6b184557dcfd052abf1424fc024f77838c4a70 Mon Sep 17 00:00:00 2001 From: Nate Coraor Date: Thu, 25 Oct 2018 15:10:08 -0400 Subject: [PATCH] Clarify that maximum job runner worker shutdown time is a multiple of `monitor_thread_join_timeout` based on the configured number of runner plugins. --- config/galaxy.yml.sample | 2 +- doc/source/admin/galaxy_options.rst | 3 ++- doc/source/admin/scaling.md | 15 ++++++++++----- lib/galaxy/webapps/galaxy/config_schema.yml | 3 ++- 4 files changed, 15 insertions(+), 8 deletions(-) diff --git a/config/galaxy.yml.sample b/config/galaxy.yml.sample index b87093cd5c5..e9d91554771 100644 --- a/config/galaxy.yml.sample +++ b/config/galaxy.yml.sample @@ -1025,7 +1025,7 @@ galaxy: # and which can cause job errors if not shut down cleanly. If using # supervisord, consider also increasing the value of `stopwaitsecs`. # If using job handler mules, consider also setting the `mule-reload- - # mercy` uWSGI option. + # mercy` uWSGI option. See the Galaxy Admin Documentation for more. #monitor_thread_join_timeout: 30 # Write thread status periodically to 'heartbeat.log', (careful, uses diff --git a/doc/source/admin/galaxy_options.rst b/doc/source/admin/galaxy_options.rst index 288b3490620..e68276436ad 100644 --- a/doc/source/admin/galaxy_options.rst +++ b/doc/source/admin/galaxy_options.rst @@ -2103,7 +2103,8 @@ jobs, and which can cause job errors if not shut down cleanly. If using supervisord, consider also increasing the value of `stopwaitsecs`. If using job handler mules, consider also setting - the `mule-reload-mercy` uWSGI option. + the `mule-reload-mercy` uWSGI option. See the Galaxy Admin + Documentation for more. :Default: ``30`` :Type: int diff --git a/doc/source/admin/scaling.md b/doc/source/admin/scaling.md index f7b5307b61f..5af4d2ae955 100644 --- a/doc/source/admin/scaling.md +++ b/doc/source/admin/scaling.md @@ -414,9 +414,13 @@ servicing web requests, but some parts of Galaxy's job preparation/submission an take quite a bit of time to complete and are not entirely reentrant: job errors or state inconsistencies can occur if interrupted (although every effort has been made to minimize such possibilities). By default, Galaxy will wait up to 30 seconds for the threads allocated for these operations to terminate after instructing them to shut down. You can change -this behavior by increasing the value of `monitor_thread_join_timeout` in the `galaxy` section of `galaxy.yml`. If you -increase this near to or above 60, you should set the appropriate uWSGI `*-restart-mercy` option to a higher value. If -using **uWSGI all-in-one**, set `worker-reload-mercy`, and if using **uWSGI + Mule job handling**, set +this behavior by increasing the value of `monitor_thread_join_timeout` in the `galaxy` section of `galaxy.yml`. The +maximum amount of time that Galaxy will take to shut down job runner workers is `monitor_thread_join_timeout * +runner_plugin_count` since each plugin is shut down sequentially (`runner_plugin_count` is the number of ``s in +your `job_conf.xml`). + +Thus you should set the appropriate uWSGI `*-restart-mercy` option to a value higher than the maximum job runner worker +shutdown time. If using **uWSGI all-in-one**, set `worker-reload-mercy`, and if using **uWSGI + Mule job handling**, set `mule-reload-mercy` (both in the `uwsgi` section of `galaxy.yml`). **Signals** @@ -546,8 +550,9 @@ substitution in the `command` and `process_name` fields. We've set `numprocs = 3 processes. Supervisord will loop over `0..numprocs` and launch `handler0`, `handler1`, and `handler2` processes automatically for us, templating out the command string so each handler receives a different log file and name. -The value of `stopwaitsecs` should be at least as large as the `monitor_thread_join_timeout` Galaxy option, which -defaults to `30`. +The value of `stopwaitsecs` should be at least as large as `monitor_thread_join_timeout * runner_plugin_count`, which +is `30` in the default configuration (`monitor_thread_join_timeout` is a Galaxy configuration option and +`runner_plugin_count` is the number of ``s in your `job_conf.xml`). Lastly, collect the tasks defined above into a single group. If you are not using webless handlers this is as simple as: diff --git a/lib/galaxy/webapps/galaxy/config_schema.yml b/lib/galaxy/webapps/galaxy/config_schema.yml index db3efec896d..cd60a7c8fe9 100644 --- a/lib/galaxy/webapps/galaxy/config_schema.yml +++ b/lib/galaxy/webapps/galaxy/config_schema.yml @@ -1572,7 +1572,8 @@ mapping: workers, which are responsible for preparing/submitting and collecting/finishing jobs, and which can cause job errors if not shut down cleanly. If using supervisord, consider also increasing the value of `stopwaitsecs`. If using job - handler mules, consider also setting the `mule-reload-mercy` uWSGI option. + handler mules, consider also setting the `mule-reload-mercy` uWSGI option. See + the Galaxy Admin Documentation for more. use_heartbeat: type: bool