Reason for the Slurm warning
----------------------------
The Linux kernel memory controller (responsible for the memory cgroup) may
run out of cgroup subsystem state (CSS) IDs. When a cgroup is created it is
assigned a CSS ID to manage its state. Upon removal of the cgroup the
corresponding state information (e.g. cache entries) may still exist and
therefore the CSS ID is still held. When multiple frequent short-lived jobs
are run on a cluster node, the number of available CSS IDs becomes exhausted
and creating a new memory cgroup results in a ENOSPC ("No space left on
device"), which is reported by SLURM with the message "unable to add
task[pid=<PID>] to memory cg '(null)'" (even if jobs run fine).
There is a bugfix for the kernel that releases the CSS ID upon cgroup
destruction, but it appears to be only available after Linux 4.4. We use
CentOS 7, which is based on Linux 3.10. The only temporary solutions is to
reboot the affected cluster node.
Thanks to @tuxtobin for the detailed analysis above.
Reason for this patch
---------------------
Many tools rely on an empty stderr to determine if the job was successful,
either because they were never updated to use `<stdio>`/`detect_errors`, or
because the underlying tool returns a non-zero exit code when successful,
e.g. `tranalign` from
https://toolshed.g2.bx.psu.edu/view/devteam/emboss_5/832c20329690 .
Even using a `<regex>` inside `<stdio>` would not work because the Slurm
warning contains the word `error`.
This is necessary when mapping over a collection over an input
that is referenced in the output section, like so:
```
...
<when value="paired_collection">
<action type="format">
<option type="from_param" name="library.input_1" param_attribute="reverse.ext" />
</action>
</when>
...
```
1) galaxy.yml
2) galaxy.ini
3) universe_wsgi.ini
Also:
- properly quote variables which may contain spaces
- no need to pass the config file to `./scripts/manage_tool_dependencies.py`
- simplify `scripts/maintenance.sh` by moving the `cd` up and reusing
functions from `scripts/common_startup_functions.sh`
Also don't fail if no config file is found, just print a message on
stderr.
This allows the removal of the code to set `GALAXY_CONFIG_FILE` from
`scripts/common_startup.sh` .
to store --pid-file and --log-file options of `./scripts/paster.py` .
https://github.com/galaxyproject/galaxy/pull/6239 fixed 3 bugs:
- `GALAXY_RUN_ALL=1 ./run.sh` was reusing the same pid and log files for
all Galaxy processes;
- `./run.sh restart` was always using `paster.pid` and `paster.log` instead
of the configured files
- `./run.sh status` was always using `paster.pid`
but also introduced a regression, i.e. the plain `./run.sh` started
writing to the log file instead of the standard output (console).
This fixes the regression by using a separate variable for these options
instead of conflating them in `paster_args` or `server_args`.