Commit Graph
2477 Commits
Author SHA1 Message Date
John Chilton cdbae706e6 Respect auto_decompress in fetch API, synchronize upload APIs. 2018-05-15 07:39:14 -04:00
John Chilton b4de30b5fc Small fixes for upload flags (auto_decompress, check_content, space_to_tab). 2018-05-15 06:28:54 -04:00
John Chilton 8a58711e39 Remove validate_mode prop for sniffing, no longer needed with previous fixes.
This does undo an attempt that Nate made to auto sniff fastq.gz files when they can't be converted anyway, but that was kind of broken in its own way.

From Gitter:

jmchilton: natefoo "# Link mode can't decompress anyway, so enable sniffing for keep-compressed datatypes even when auto_decompress is enabled" Did you test this hack? I get the thought process but it seems like it results in uploaded fastq.gz files defaulting to fastqcssanger.gz? Maybe it didn't at some point though?

Nate Coraor: ummmmmmmmmmmm maybe probably not

John Chilton: I have four fixes for compressed datatypes that all fall apart because of this hack... can I just use the configured sniff order or do you want me to try to preserve this and write a test case

Nate Coraor: nah go ahead and de-hack

John Chilton: There is a good argument to be made for making the behavior more consistent anyway right? Cool thanks
2018-05-10 18:35:31 -04:00
Nicola Soranzo 2515267ca5 Add version attribute to tool requirements
instead of relying on `lib/galaxy/tools/deps/resolvers/default_conda_mapping.yml` .
Follow-up on https://github.com/galaxyproject/galaxy/pull/5544 .

See https://github.com/galaxyproject/galaxy/pull/5544/files#r183909746 for
an explanation why that's preferrable for future-proof reproducibility.

Also, small fixes to `tools/evolution/codingSnps.xml` .
2018-04-25 15:50:35 +01:00
Martin Cech 34016542d9 Merge pull request #5544 from natefoo/split-ucsc-requirements
Update all tools/converters using UCSC binaries to depend only on the binary they are using
2018-04-24 16:23:13 -04:00
John Chilton 7712f69b69 Rev tool versions for tools with updated ucsc tool requirements. 2018-04-24 14:40:28 -04:00
John Chilton 09fcf13aaf Merge pull request #5872 from abretaud/filter_sanitize
Filter tool: fix too strict sanitizing
2018-04-20 07:17:55 -04:00
Anthony Bretaudeau dadb4fcda7 change back version 2018-04-20 09:45:06 +02:00
Anthony Bretaudeau dc4902353f add test case for fixed bug 2018-04-20 09:44:27 +02:00
John Chilton d11f76b031 More flake8 fixes for sorter.py. 2018-04-19 12:58:28 -04:00
John Chilton 145903eae3 Merge remote-tracking branch 'jmchilton/dev' into sort_header 2018-04-19 12:57:18 -04:00
Anthony Bretaudeau 0b191abaf1 pass the condition by a json file instead of arg 2018-04-19 15:24:29 +02:00
mvdbeek 2e18ed2c97 Set ext explicitly to dataset.file_type
since it is being used again later. Thanks @nsoranzo!
2018-04-19 10:40:17 +02:00
mvdbeek 6c020a6dbc Fix uploads of link-only datasets to data library
`ext` was never being set in the case of link only datasets, where I think the
right thing is to fall back on `dataset.file_type`.
Fixes https://github.com/galaxyproject/galaxy/issues/5915.
2018-04-19 10:40:17 +02:00
Nicola Soranzo 89ac586fff flake8 fixes 2018-04-18 22:11:32 +01:00
Anthony Bretaudeau bb9963dddc fix quote escaping 2018-04-10 14:20:53 +02:00
Anthony Bretaudeau 16f9e58af5 quotes 2018-04-10 12:02:18 +02:00
Anthony Bretaudeau de5ac42fa9 more tolerant sanitizing 2018-04-10 11:36:08 +02:00
Nicola Soranzo 6e36af5c0e Fix linking option for library dataset uploads
Broken in commit a45cbfeb8d .

Found by BioBlend tests:
https://travis-ci.org/galaxyproject/bioblend/jobs/356761960
2018-03-22 10:36:58 +00:00
Nate Coraor 11917108f3 Improve extension stripping 2018-03-16 16:33:39 -04:00
Nate Coraor 0163a57f0e Additional fixes for failing tests 2018-03-15 16:56:41 -04:00
Nate Coraor ef6f7f62ab Memory usage bug fixes and code cleanup/simplification for upload tool 2018-03-15 14:32:02 -04:00
John Chilton 495d1258fd Consistent sniffing regardless of in_place.
Previously sniffing would happen on the original file (before carriage returns and tabular spaces were converted) if in_place was false and on the converted file if it was true.
2018-03-08 10:33:42 -05:00
John Chilton 1c6cc02415 Hierarchical upload API optimized for folders & collections.
Allows describing hierarchical data in JSON or inferring structure from archives or directories.

Datasets or archive sources can be specified via uploads, URLs, paths (if admin && allow_path_paste), library_import_dir/user_library_import_dir, and/or FTP imports. Unlike existing API endpoints, a mix of these on a per file basis is allowed and they work seemlessly between libraries and histories.

Supported "archives" include gzip, zip, bagit directories, bagit achives (with fetching and validations of downloads).

The existing upload API endpoint is quite rough to work with both in terms of adding parameters (e.g. the file type and dbkey hanlding in 4563 was difficult to implement, terribly hacky, and should seemingly have been trivial) and in terms of building requests (one needs to build a tool form - not describe sensible inputs in JSON). This API is built to be intelligable from an API standpoint instead of being constrained to the older style tool form. Additionally it built with hierarchical data in mind in a way that would not be easy at all enhancing the tool form components we don't even render.

This implements 5159 though much simpler YAML descriptions of data libraries should be possible basically as the API descriptions. We can replace the data library script in Ephemeris https://github.com/galaxyproject/ephemeris/blob/master/ephemeris/setup_data_libraries.py with one that converts a simple YAML file into an API call and allows many new options for free.

In future PRs I'll add filtering options to this and it will serve as the backend to 4733.
2018-03-08 10:33:36 -05:00
John Chilton a45cbfeb8d Stronger assertion about link_data_only option in upload.py. 2018-03-08 10:32:19 -05:00
John Chilton ebac0fd9e5 Upload simplification - just base this check on link_only. 2018-03-08 10:32:19 -05:00
Martin Cech e5d0c1c0ee Merge branch 'merge-in-18.01' into dev 2018-03-07 13:35:00 -05:00
Martin Cech 10f8c3bc82 Merge branch 'release_17.09' into release_18.01
dropped all static content from merge except of styles
2018-03-07 12:47:59 -05:00
Marius van den Beek f8d44c2b69 Merge pull request #5631 from martenson/genomespace
[17.09] Changed GenomeSpace token handling to use manual OpenID association only
2018-03-06 10:13:40 +01:00
lecorguille 376aa79825 subprocess 2018-03-05 17:13:58 +01:00
Matthias Bernt a6593d4caf python3 fix
strings can not be written to byte mode file handlers.
writing resulted in:
TypeError: a bytes-like object is required, not 'str'

Alternative would be to make string byte.
2018-03-05 12:43:43 +00:00
Nuwan Goonasekera 6b782ad16d Changed GenomeSpace token handling to use manual OpenID association only 2018-03-01 15:03:12 -05:00
John Chilton d9d8f5d327 Eliminate precreated datasets concept in upload API.
Was used by the async controller in the past I think (https://github.com/galaxyproject/galaxy/blob/v13.01/lib/galaxy/webapps/galaxy/controllers/tool_runner.py#L256), but doesn't seem to be used now.
2018-02-26 11:00:05 -05:00
Nicola Soranzo 75b0618705 Do not remove external path files during library uploads
xref. https://github.com/galaxyproject/galaxy/pull/5264#issuecomment-367172820
2018-02-21 16:02:36 +00:00
Nate Coraor 4b09027ebb Update all tools/converters using UCSC binaries to depend only on the
binary they are using.
2018-02-15 10:54:04 -05:00
lecorguille d4faabae95 add requirements 2018-01-23 09:27:39 +01:00
Nicola Soranzo b59fb06fd9 Remove unused import 2018-01-17 16:26:56 +00:00
Matthias Bernt 4b2bf122b6 fix for microbial import tool
There seems to be no UnvalidatedValue class anymore. So I removed these
checks which seems to make the tool functional again.

Note that there is one more instance of such a check in the Galaxy
sources (in tools/parameters/basic.py). I guess this can also be
removed?
2018-01-17 15:55:10 +00:00
John Chilton 3f75a2d3a6 Re-work upload clarification from #5206.
See post-merge discussion on that issue.
2018-01-04 09:30:16 -05:00
Dannon Baker 6bf5d663b4 Merge pull request #5229 from jmchilton/upload_refactor
Refactor upload.py toward reuse
2017-12-18 11:15:37 -05:00
John Chilton 7e1bff7d69 Upload refactor - change upload.py to use exceptions.
Make decomposing and reuse of this easier downstream and feels cleaner to me.
2017-12-15 13:24:44 -05:00
John Chilton cafac19f65 Upload refactor - make link_data_only a bool.
Since it is a bool.
2017-12-15 13:24:44 -05:00
John Chilton 09f51f59af Upload optimization - eliminate second call to check_binary in upload.py. 2017-12-15 13:24:44 -05:00
Nicola Soranzo 6d3eadedbd Merge pull request #5227 from jmchilton/merge_1709
Merge 17.09.
2017-12-15 18:14:34 +00:00
John Chilton 90ba35a9a4 Merge remote-tracking branch 'jmchilton/release_17.09' into merge_1709 2017-12-15 12:02:07 -05:00
John Chilton 121285b40b Let ToolProvidedMetadata interface more directly decide if it has failed outputs.
I like this better for three reasons:

- Since usually it is scripts producing this JSON - we have the most control at that point for determining the failure and we don't have to deal with an artificial dependency between the tool's stdio and the output.
- At some point we could potentially allow some datasets to be ok now even though the job fails.
- It is a cleaner interface at the Python level between job finish and output collection IMO (no need for isinstance checking).
2017-12-15 08:34:35 -05:00
Dannon Baker 3b7f40198e Merge pull request #5215 from nsoranzo/python3
Python3: finish first pass on whole codebase
2017-12-14 14:38:17 -05:00
Nicola Soranzo 14259b11d9 Python3: finish first pass on whole codebase 2017-12-14 12:01:01 +00:00
John Chilton 099c1562fe Re-organize edge case upload options for my own clarity.
I think setting each of these variables once and simplifing the context they are used in (in the case of purge_upload) makes it more clear what each variable is and how it is set. I also think one, more detailed comment for each variable helps.

Note: This will break run-as-user uploads started prior to the upgrade to 18.XX and executed after the upgrade. It is a small switch to restore the old behavior but I'm not sure it is worth the complexity it adds to the file.

```
run_as_real_user = dataset.get('run_as_real_user', False) or dataset_get.('in_place', True)
```
2017-12-13 12:54:03 -05:00
Nicola Soranzo b3dc7923db Python3: use collections.Mapping instead of removed DictMixin 2017-12-12 15:30:53 +00:00