This does undo an attempt that Nate made to auto sniff fastq.gz files when they can't be converted anyway, but that was kind of broken in its own way.
From Gitter:
jmchilton: natefoo "# Link mode can't decompress anyway, so enable sniffing for keep-compressed datatypes even when auto_decompress is enabled" Did you test this hack? I get the thought process but it seems like it results in uploaded fastq.gz files defaulting to fastqcssanger.gz? Maybe it didn't at some point though?
Nate Coraor: ummmmmmmmmmmm maybe probably not
John Chilton: I have four fixes for compressed datatypes that all fall apart because of this hack... can I just use the configured sniff order or do you want me to try to preserve this and write a test case
Nate Coraor: nah go ahead and de-hack
John Chilton: There is a good argument to be made for making the behavior more consistent anyway right? Cool thanks
Previously sniffing would happen on the original file (before carriage returns and tabular spaces were converted) if in_place was false and on the converted file if it was true.
Allows describing hierarchical data in JSON or inferring structure from archives or directories.
Datasets or archive sources can be specified via uploads, URLs, paths (if admin && allow_path_paste), library_import_dir/user_library_import_dir, and/or FTP imports. Unlike existing API endpoints, a mix of these on a per file basis is allowed and they work seemlessly between libraries and histories.
Supported "archives" include gzip, zip, bagit directories, bagit achives (with fetching and validations of downloads).
The existing upload API endpoint is quite rough to work with both in terms of adding parameters (e.g. the file type and dbkey hanlding in 4563 was difficult to implement, terribly hacky, and should seemingly have been trivial) and in terms of building requests (one needs to build a tool form - not describe sensible inputs in JSON). This API is built to be intelligable from an API standpoint instead of being constrained to the older style tool form. Additionally it built with hierarchical data in mind in a way that would not be easy at all enhancing the tool form components we don't even render.
This implements 5159 though much simpler YAML descriptions of data libraries should be possible basically as the API descriptions. We can replace the data library script in Ephemeris https://github.com/galaxyproject/ephemeris/blob/master/ephemeris/setup_data_libraries.py with one that converts a simple YAML file into an API call and allows many new options for free.
In future PRs I'll add filtering options to this and it will serve as the backend to 4733.
There seems to be no UnvalidatedValue class anymore. So I removed these
checks which seems to make the tool functional again.
Note that there is one more instance of such a check in the Galaxy
sources (in tools/parameters/basic.py). I guess this can also be
removed?
I like this better for three reasons:
- Since usually it is scripts producing this JSON - we have the most control at that point for determining the failure and we don't have to deal with an artificial dependency between the tool's stdio and the output.
- At some point we could potentially allow some datasets to be ok now even though the job fails.
- It is a cleaner interface at the Python level between job finish and output collection IMO (no need for isinstance checking).
I think setting each of these variables once and simplifing the context they are used in (in the case of purge_upload) makes it more clear what each variable is and how it is set. I also think one, more detailed comment for each variable helps.
Note: This will break run-as-user uploads started prior to the upgrade to 18.XX and executed after the upgrade. It is a small switch to restore the old behavior but I'm not sure it is worth the complexity it adds to the file.
```
run_as_real_user = dataset.get('run_as_real_user', False) or dataset_get.('in_place', True)
```
Remove the need to call `Binary.register_unsniffable_binary_ext()` for
each unsniffable binary datatype.
Fix https://github.com/galaxyproject/galaxy/issues/3441 , where the upload
of files of a datatype defined in datatypes_conf.xml as subclass of an
unsniffable binary datatype ended up with "The uploaded binary file
contains inappropriate content" because it was not possible to register the
subclassed datatype as unsniffable.
Also remove unused `stop_err()` function in upload.py .
Also, when sniffing binary files, sniff images together with the other
formats and respect sniff order.
Also remove `is_multi_byte` from:
- `stream_to_open_named_file()`
- `stream_to_file()`
Remove the now unused `get_image_ext()` and `Binary.is_sniffable_binary()`
and all the calls to `Binary.register_sniffable_binary_format()`.
Also, merge its 3 parameters `gzip_only`, `bz2_only`, `zip_only` into
`compressed_formats` (a list of allowed formats).
As a consequence of the changes in `get_fileobj()`, update:
- `files_diff()`
- `get_file_peek()`, which now determines that a file is binary when a
`UnicodeDecodeError` exception is raised and doesn't need
`is_multi_byte` any more
- `iter_headers()` and `get_headers()`, which now return Unicode and don't
need `is_multi_byte` parameter any more
As a consequence of the changes in `get_file_peek()`, update:
- `set_peek()`, which now doesn't need `is_multi_byte` any more
As a consequence of the changes in `get_headers()`, update:
- `guess_ext` and `is_column_based()`, which now determine that a file is
binary when a `UnicodeDecodeError` exception is raised and don't need
`is_multi_byte` any more
As a consequence of the changes to `guess_ext`, update:
- `handle_uploaded_dataset_file() doesn't need `is_multi_byte` any more
Also, remove duplicated calls to `get_file_peek()` in
lib/galaxy/datatypes/molecules.py and lib/galaxy/datatypes/msa.py
The `is_multi_byte` was not removed from the signature of `get_file_peek()`
and `set_peek()` in order to preserve compatibility for ToolShed datatypes,
thanks @jmchilton for the review.
shutil.move tries to move files by renaming them. If that fails with
OSError (due to permission or cross-filesystem rename) it falls back to
copying files followed by removing them
(https://github.com/python/cpython/blob/2.7/Lib/shutil.py#L279). By
using shutil.move and catching permission problems we avoid an
unnecessary copy if source and destination are on the same filesystem.
Also avoids shutil.move if the upload tool is run as real-user which
should fix https://github.com/galaxyproject/galaxy/issues/4300.
When importing files from the FTP folder the admin can choose to prevent
purging of imported files. In this case the source file should not be
modified in-place by galaxy.
This also modifies relevant bare `except:` statements and modifies
the meaning of `in_place` in upload.py from no external chown script to
do not edit files in place if we keep the source file.
This fixes https://github.com/galaxyproject/galaxy/issues/4527.