Reason for the Slurm warning
----------------------------
The Linux kernel memory controller (responsible for the memory cgroup) may
run out of cgroup subsystem state (CSS) IDs. When a cgroup is created it is
assigned a CSS ID to manage its state. Upon removal of the cgroup the
corresponding state information (e.g. cache entries) may still exist and
therefore the CSS ID is still held. When multiple frequent short-lived jobs
are run on a cluster node, the number of available CSS IDs becomes exhausted
and creating a new memory cgroup results in a ENOSPC ("No space left on
device"), which is reported by SLURM with the message "unable to add
task[pid=<PID>] to memory cg '(null)'" (even if jobs run fine).
There is a bugfix for the kernel that releases the CSS ID upon cgroup
destruction, but it appears to be only available after Linux 4.4. We use
CentOS 7, which is based on Linux 3.10. The only temporary solutions is to
reboot the affected cluster node.
Thanks to @tuxtobin for the detailed analysis above.
Reason for this patch
---------------------
Many tools rely on an empty stderr to determine if the job was successful,
either because they were never updated to use `<stdio>`/`detect_errors`, or
because the underlying tool returns a non-zero exit code when successful,
e.g. `tranalign` from
https://toolshed.g2.bx.psu.edu/view/devteam/emboss_5/832c20329690 .
Even using a `<regex>` inside `<stdio>` would not work because the Slurm
warning contains the word `error`.
This is necessary when mapping over a collection over an input
that is referenced in the output section, like so:
```
...
<when value="paired_collection">
<action type="format">
<option type="from_param" name="library.input_1" param_attribute="reverse.ext" />
</action>
</when>
...
```
Also don't fail if no config file is found, just print a message on
stderr.
This allows the removal of the code to set `GALAXY_CONFIG_FILE` from
`scripts/common_startup.sh` .
The tool wouldn't load after 8af243c176 and that caused the relevant tests to "SKIP" instead of failing outright.
./run_tests.sh -api test/api/test_tools.py:ToolsTestCase.test_map_over_collection_type_source
This reverts commit 91e97b6096.
The commit to revert causes
https://github.com/galaxyproject/galaxy/issues/6219, where the
`optional="True"` flag is not repected anymore. This is because we raise
an exception before checking that `value` may be legitimately `None`.
Fix handling comparison of older hashes in Python 2 and fix the hanlding of both for Python 3.
The previous iteration gave me some errors in Python 3, I brought in and adapted @mvdbeek's Python 3 fixes to fix these:
```
======================================================================
ERROR: test_security_passwords.test_hash_and_check
----------------------------------------------------------------------
Traceback (most recent call last):
File "/Users/john/workspace/galaxy-lib/.tox/py34/lib/python3.4/site-packages/nose/case.py", line 198, in runTest
self.test(*self.arg)
File "/Users/john/workspace/galaxy-lib/tests/test_security_passwords.py", line 12, in test_hash_and_check
simple_pass_hash = passwords.hash_password(simple_pass)
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 21, in hash_password
return hash_password_PBKDF2(password)
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 44, in hash_password_PBKDF2
hashed = pbkdf2_bin(smart_str(password), salt, COST_FACTOR, KEY_LENGTH, getattr(hashlib, HASH_FUNCTION))
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 78, in pbkdf2_bin
rv = u = _pseudorandom(salt + _pack_int(block))
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 75, in _pseudorandom
return [ord(_) for _ in h.digest()]
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 75, in <listcomp>
return [ord(_) for _ in h.digest()]
TypeError: ord() expected string of length 1, but int found
======================================================================
ERROR: test_security_passwords.test_hash_consistent
----------------------------------------------------------------------
Traceback (most recent call last):
File "/Users/john/workspace/galaxy-lib/.tox/py34/lib/python3.4/site-packages/nose/case.py", line 198, in runTest
self.test(*self.arg)
File "/Users/john/workspace/galaxy-lib/tests/test_security_passwords.py", line 38, in test_hash_consistent
assert passwords.check_password(simple_pass, simple_pass_hash)
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 30, in check_password
if check_password_PBKDF2(guess, hashed):
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 53, in check_password_PBKDF2
hashed_guess = pbkdf2_bin(smart_str(guess), salt, int(cost_factor), KEY_LENGTH, getattr(hashlib, hash_function))
File "/Users/john/workspace/galaxy-lib/galaxy/security/passwords.py", line 78, in pbkdf2_bin
rv = u = _pseudorandom(salt + _pack_int(block))
TypeError: Can't convert 'bytes' object to str implicitly
```
Added tests that all now pass in Python 3 - including generating some passwords in Python 3 and verifying they pass when checked in Python 2 and vice versa.
Fix the following traceback:
```
galaxy.jobs.runners.slurm ERROR 2018-05-22 13:01:07,050 (1024816/14094939) Failure in SLURM _complete_terminal_job(), job final state will be: failed
Traceback (most recent call last):
File "/tgac/services/galaxy/prod/galaxy/lib/galaxy/jobs/runners/slurm.py", line 86, in _complete_terminal_job
slurm_state = _get_slurm_state()
File "/tgac/services/galaxy/prod/galaxy/lib/galaxy/jobs/runners/slurm.py", line 81, in _get_slurm_state
job_info_dict = dict([out_param.split('=', 1) for out_param in stdout.split()])
ValueError: dictionary update sequence element #52 has length 1; 2 is required
```
Update tool XSD to encourage using fully qualified inputs.
Reorganized the existing tests for consistency with other test tools I think.
(Rebased with fixes, including XSD fixes from @nsoranzo.)
This should work around issues like
```
galaxy.jobs.runners ERROR 2018-05-16 10:15:55,522 [p:13768,w:0,m:1] [ShellRunner.work_thread-3] (34130) Unhandled exception calling queue_job
Traceback (most recent call last):
File "lib/galaxy/jobs/runners/__init__.py", line 112, in run_next
method(arg)
File "lib/galaxy/jobs/runners/cli.py", line 94, in queue_job
chmod_out = shell.execute("chmod +x %s" % tool_script)
File "lib/galaxy/jobs/runners/util/cli/shell/rsh.py", line 77, in execute
self.connect()
File "lib/galaxy/jobs/runners/util/cli/shell/rsh.py", line 69, in connect
timeout=self.timeout)
File "/bioinfo/guests/mvandenb/galaxy/.venv/local/lib/python2.7/site-packages/paramiko/client.py", line 424, in connect
passphrase,
File "/bioinfo/guests/mvandenb/galaxy/.venv/local/lib/python2.7/site-packages/paramiko/client.py", line 714, in _auth
raise saved_exception
SSHException: not a valid OPENSSH private key file
```
that can intermittently occur with rapid submission of many small jobs.