mirror of
https://github.com/galaxyproject/galaxy.git
synced 2026-09-24 16:30:27 +08:00
*Overview*
This work defines an interface for interacting with "filesystem"-like entities during "upload". In addition to being a pluggable framework adding important new capabilities to Galaxy, this is a generalization and formalization of existing file sources (e.g. the directories described by `library_import_dir`, `user_library_dir`, and `ftp_upload_dir`).
*Plugin Infrastructure*
This introduces a new plugin `FilesSource` to represent sources of directories and files during "upload". A `FilesSource` plugin should be able to index directories and download (called 'realize' to be generic) files to local posix directories. Indexing is used by the remote_files API to provide the client with hierarchies to navigate and to build URIs for the files. The 'realize' operation is used by the 'upload1' and '__DATA_FETCH__' and tools during upload to bring the files into Galaxy as datasets.
An instance of the `ConfiguredFileSources` class is responsible for managing individual instances of `FilesSource` plugins. It has methods to map URIs to the appropriate plugin instance.
The `ConfiguredFileSources` class tracks the loaded plugins and reuses the go to `galaxy.util.plugin_config` module for loading YAML (or XML) definitions of plugins (the same dependency resolvers, job metrics, auth backends, etc. do). A `ConfiguredFileSources` object can serialize itself to a file and re-materialize it during job execution to allow using this abstraction during uploads.
When operating within the Galaxy app, the `ConfiguredFileSources` uses an adapter pattern to parse user-level information from Galaxy's `trans` object. During serialization, the `ConfiguredFileSources` object is expected to encode all the required information about the user that is needed into the output JSON description of the file sources. This is because the web transaction won't be available remotely during the upload job. These objects working in such different ways between the Galaxy process and in the remote job is mildly jarring - so unit tests have been written to ensure this all functions properly.
*Plugin Implementations*
The `FilesSource` interface has a helper implementation base class `BaseFilesSource` that provides some assistance for plugin development. Additionally, the base class `PyFilesystem2FilesSource` extends `BaseFilesSource` but assumes a PyFilesystem2 implementation exists to target the file source of interest - so the plugin author need only provide a PyFilesystem `FS` object describing the target. This commit includes three concrete implementations - posix, webdav, and dropbox. `posix` extends `BaseFilesSource` while the others are light-weight extensions of `PyFilesystem2FilesSource`.
**posix**
While one could imagine a very lightweight implementation based on `PyFilesystem2FilesSource` this fully worked through plugin is implemented directly to ensure we respect Galaxy's strong security checks on paths containing symlinks and preserve the semantics `user_library_import_symlink_allowlist`.
**webdav**
Galaxy tools for integrating OwnCloud exist - see https://github.com/shiltemann/Galaxy-Owncloud-Integration, part of the driver for this work was extending that idea to provide more integrated UX for uploading that data. So this work includes a WebDav plugin (and associated test cases) that could potentially target OwnCloud.
This plugin was a good exercise in flushing and testing the PyFilesystem2 interface but the PyFilesystem2 WebDAV implementation seems a bit fragile... we might want to replace it with more direct APIs but we can take a wait and see approach.
The config YAML for a webdav plugin that lets user's target their own OwnCloud servers configured via user preferences might look something like:
```
- type: webdav
id: owncloud1
label: OwnCloud
doc: User-configured OwnCloud files
url: ${user.preferences['owncloud|url']}
login: ${user.preferences['owncloud|username']}
password: ${user.preferences['webdav|password']}
```
The configuration would provide a user's OwnCloud files at `gxfiles://owncloud1/`.
If instead, a big centralized WebDav server is made available with public data for all users (mirroring use cases of `library_import_dir`) - a simpler configuration not requiring user preferences might be something like:
```
- type: webdav
id: lab
label: Lab WebDAV server
doc: Our lab's research files managed at ourlab.org.
url: http://ourlab.org:7083
login: ${environ.get('WEBDAV_LOGIN')}
password: ${environ.get('WEBDAV_PASSWORD')}
```
The configuration would provide a these WebDAV files at `gxfiles://lab/`.
These two examples demonstrate basic templating is allowed inside the YAML configuration. These are Cheetah templates exposing very specific views of the 'user', 'config', and the whole 'environ' available to the Galaxy server.
**dropbox**
The Dropbox PyFilesystem2 plugin is even easier to configure, all that is needed is a Dropbox access token (this can be configured from the settings menu and may be isolated to a specific app specific folder for added security on the user's part).
An example of such a plugin might be:
```
- type: dropbox
id: dropbox1
label: Dropbox Files
doc: Your Dropbox files - configure an access token via the user preferences
accessToken: ${user.preferences['dropbox|access_token']}
```
The configuration would provide a user's Dropbox files at `gxfiles://dropbox1/`.
**gxftp**
This is an automatically populated plugin (if `ftp_upload_dir` is configured in Galaxy) that provides the user's FTP files at `gxftp://`.
**gximport**
This is an automatically populated plugin (if `library_import_dir` is configured in Galaxy) that provides Galaxy's library import files at `gximport://`.
**gxuserimport**
This is an automatically populated plugin (if `user_library_import_dir` is configured in Galaxy) that provides the requesting user's Galaxy's user library import files at `gximportfiles://`.
*Why not a tool?*
One could imagine a tool - but the upload dialog has many advanced options for selecting how to ingest files (convert tabs and newlines, select format vs. detect, select dbkey, organize into collections, organize via rules, etc...). It would be next to impossible to provide all these same options via a normal tool and the user experience would be very different than using the upload components in Galaxy - which have been optimized and designed for this task.
That said - one future direction I would like to take this is to be able to mark plugins as writable and implement a new tool form input type "export_directory" or something like that. This could then be used to write data export tools. This could be used to write generalizations of the the cloud send tool.
*`ObjectStore` vs `FilesSource`*
ObjectStores provide datasets not files, the files are organized logically in a very flat way around a dataset. `FilesSource` s instead provide files and directories, not datasets. A `FilesSource` is meant to be browsed in hierarchical fashion - and also has no concept of extra files, etc..
*Future Work*
- This is hopefully going to serve as the basis of a first pass at Terra integration with Galaxy using the FISS lib. Having an implementation based on `PyFilesytem2` means we could potentially integrate support for S3, Basespace, Google Drive, OneDrive, etc..
- Tool form support for selecting files for import and directories for export.
- Allow writing collection archives, history export, etc.. to the `FilesSource` - this would really enhance the UI around getting big stuff out of Galaxy potentially I think.
Rebase into galaxy.files...
234 lines
10 KiB
XML
234 lines
10 KiB
XML
<?xml version="1.0"?>
|
|
<tool name="Upload File" id="upload1" version="1.1.7" workflow_compatible="false" profile="16.04">
|
|
<description>from your computer</description>
|
|
<action module="galaxy.tools.actions.upload" class="UploadToolAction"/>
|
|
<configfiles>
|
|
<file_sources filename="file_sources.json" />
|
|
</configfiles>
|
|
<command>
|
|
python '$__tool_directory__/upload.py' '$GALAXY_ROOT_DIR' '$GALAXY_DATATYPES_CONF_FILE' '$paramfile'
|
|
#set $outnum = 0
|
|
#while $varExists('output%i' % $outnum):
|
|
#set $output = $getVar('output%i' % $outnum)
|
|
#set $outnum += 1
|
|
#set $file_name = $output.file_name
|
|
## FIXME: This is not future-proof for other uses of external_filename (other than for use by the library upload's "link data" feature)
|
|
#if $output.dataset.dataset.external_filename:
|
|
#set $file_name = "None"
|
|
#end if
|
|
'${output.dataset.dataset.id}:${output.files_path}:${file_name}'
|
|
#end while
|
|
</command>
|
|
<inputs nginx_upload="true">
|
|
<param name="file_type" type="select" label="File Format" help="Which format? See help below">
|
|
<options from_parameter="tool.app.datatypes_registry.upload_file_formats" transform_lines="[ "%s%s%s" % ( line, self.separator, line ) for line in obj ]">
|
|
<column name="value" index="1"/>
|
|
<column name="name" index="0"/>
|
|
<filter type="sort_by" column="0"/>
|
|
<filter type="add_value" name="Auto-detect" value="auto" index="0"/>
|
|
</options>
|
|
</param>
|
|
<param name="file_count" type="hidden" value="auto" />
|
|
<upload_dataset name="files" title="Specify Files for Dataset" file_type_name="file_type" metadata_ref="files_metadata">
|
|
<param name="file_data" type="file" label="File" ajax-upload="true" help="TIP: Due to browser limitations, uploading files larger than 2GB is guaranteed to fail. To upload large files, use the URL method (below) or FTP (if enabled by the site administrator).">
|
|
</param>
|
|
<param name="url_paste" type="text" area="true" label="URL/Text" help="Here you may specify a list of URLs (one per line) or paste the contents of a file."/>
|
|
<param name="ftp_files" type="ftpfile" label="Files uploaded via FTP"/>
|
|
<!-- Swap the following parameter for the select one that follows to
|
|
enable the to_posix_lines option in the Web GUI. See Bitbucket
|
|
Pull Request 171 for more information. -->
|
|
<param name="uuid" type="hidden" required="False" />
|
|
<param name="to_posix_lines" type="hidden" value="Yes" />
|
|
<param name="auto_decompress" type="hidden" value="Yes" />
|
|
<!-- allow per-file override of dbkey -->
|
|
<param name="file_type" type="hidden" value="" />
|
|
<param name="dbkey" type="hidden" value="" />
|
|
|
|
<!--
|
|
<param name="to_posix_lines" type="select" display="checkboxes" multiple="True" label="Convert universal line endings to Posix line endings" help="Turn this option off if you upload a gzip, bz2 or zip archive which contains a binary file." value="Yes">
|
|
<option value="Yes" selected="true">Yes</option>
|
|
</param>
|
|
-->
|
|
<param name="space_to_tab" type="select" display="checkboxes" multiple="True" label="Convert spaces to tabs" help="Use this option if you are entering intervals by hand.">
|
|
<option value="Yes">Yes</option>
|
|
</param>
|
|
<param name="NAME" type="hidden" help="Name for dataset in upload"></param>
|
|
</upload_dataset>
|
|
<param name="force_composite" type="hidden" value="false" />
|
|
<param name="dbkey" type="genomebuild" label="Genome" />
|
|
<conditional name="files_metadata" value_from="self:app.datatypes_registry.get_upload_metadata_params" value_ref="file_type" value_ref_in_group="False" />
|
|
<!-- <param name="other_dbkey" type="text" label="Or user-defined Genome" /> -->
|
|
</inputs>
|
|
<help>
|
|
**Auto-detect**
|
|
|
|
The system will attempt to detect Axt, Fasta, Fastqsolexa, Gff, Gff3, Html, Lav, Maf, Tabular, Wiggle, Bed and Interval (Bed with headers) formats. If your file is not detected properly as one of the known formats, it most likely means that it has some format problems (e.g., different number of columns on different rows). You can still coerce the system to set your data to the format you think it should be. You can also upload compressed files, which will automatically be decompressed.
|
|
|
|
-----
|
|
|
|
**Ab1**
|
|
|
|
A binary sequence file in 'ab1' format with a '.ab1' file extension. You must manually select this 'File Format' when uploading the file.
|
|
|
|
-----
|
|
|
|
**Axt**
|
|
|
|
blastz pairwise alignment format. Each alignment block in an axt file contains three lines: a summary line and 2 sequence lines. Blocks are separated from one another by blank lines. The summary line contains chromosomal position and size information about the alignment. It consists of 9 required fields.
|
|
|
|
-----
|
|
|
|
**Bam**
|
|
|
|
A binary file compressed in the BGZF format with a '.bam' file extension.
|
|
|
|
-----
|
|
|
|
**Bed**
|
|
|
|
* Tab delimited format (tabular)
|
|
* Does not require header line
|
|
* Contains 3 required fields:
|
|
|
|
- chrom - The name of the chromosome (e.g. chr3, chrY, chr2_random) or contig (e.g. ctgY1).
|
|
- chromStart - The starting position of the feature in the chromosome or contig. The first base in a chromosome is numbered 0.
|
|
- chromEnd - The ending position of the feature in the chromosome or contig. The chromEnd base is not included in the display of the feature. For example, the first 100 bases of a chromosome are defined as chromStart=0, chromEnd=100, and span the bases numbered 0-99.
|
|
|
|
* May contain 9 additional optional BED fields:
|
|
|
|
- name - Defines the name of the BED line. This label is displayed to the left of the BED line in the Genome Browser window when the track is open to full display mode or directly to the left of the item in pack mode.
|
|
- score - A score between 0 and 1000. If the track line useScore attribute is set to 1 for this annotation data set, the score value will determine the level of gray in which this feature is displayed (higher numbers = darker gray).
|
|
- strand - Defines the strand - either '+' or '-'.
|
|
- thickStart - The starting position at which the feature is drawn thickly (for example, the start codon in gene displays).
|
|
- thickEnd - The ending position at which the feature is drawn thickly (for example, the stop codon in gene displays).
|
|
- itemRgb - An RGB value of the form R,G,B (e.g. 255,0,0). If the track line itemRgb attribute is set to "On", this RBG value will determine the display color of the data contained in this BED line. NOTE: It is recommended that a simple color scheme (eight colors or less) be used with this attribute to avoid overwhelming the color resources of the Genome Browser and your Internet browser.
|
|
- blockCount - The number of blocks (exons) in the BED line.
|
|
- blockSizes - A comma-separated list of the block sizes. The number of items in this list should correspond to blockCount.
|
|
- blockStarts - A comma-separated list of block starts. All of the blockStart positions should be calculated relative to chromStart. The number of items in this list should correspond to blockCount.
|
|
|
|
* Example::
|
|
|
|
chr22 1000 5000 cloneA 960 + 1000 5000 0 2 567,488, 0,3512
|
|
chr22 2000 6000 cloneB 900 - 2000 6000 0 2 433,399, 0,3601
|
|
|
|
-----
|
|
|
|
**Fasta**
|
|
|
|
A sequence in FASTA format consists of a single-line description, followed by lines of sequence data. The first character of the description line is a greater-than (">") symbol in the first column. All lines should be shorter than 80 characters::
|
|
|
|
>sequence1
|
|
atgcgtttgcgtgc
|
|
gtcggtttcgttgc
|
|
>sequence2
|
|
tttcgtgcgtatag
|
|
tggcgcggtga
|
|
|
|
-----
|
|
|
|
**FastqSolexa**
|
|
|
|
FastqSolexa is the Illumina (Solexa) variant of the Fastq format, which stores sequences and quality scores in a single file::
|
|
|
|
@seq1
|
|
GACAGCTTGGTTTTTAGTGAGTTGTTCCTTTCTTT
|
|
+seq1
|
|
hhhhhhhhhhhhhhhhhhhhhhhhhhPW@hhhhhh
|
|
@seq2
|
|
GCAATGACGGCAGCAATAAACTCAACAGGTGCTGG
|
|
+seq2
|
|
hhhhhhhhhhhhhhYhhahhhhWhAhFhSIJGChO
|
|
|
|
Or::
|
|
|
|
@seq1
|
|
GAATTGATCAGGACATAGGACAACTGTAGGCACCAT
|
|
+seq1
|
|
40 40 40 40 35 40 40 40 25 40 40 26 40 9 33 11 40 35 17 40 40 33 40 7 9 15 3 22 15 30 11 17 9 4 9 4
|
|
@seq2
|
|
GAGTTCTCGTCGCCTGTAGGCACCATCAATCGTATG
|
|
+seq2
|
|
40 15 40 17 6 36 40 40 40 25 40 9 35 33 40 14 14 18 15 17 19 28 31 4 24 18 27 14 15 18 2 8 12 8 11 9
|
|
|
|
-----
|
|
|
|
**Gff**
|
|
|
|
GFF lines have nine required fields that must be tab-separated.
|
|
|
|
-----
|
|
|
|
**Gff3**
|
|
|
|
The GFF3 format addresses the most common extensions to GFF, while preserving backward compatibility with previous formats.
|
|
|
|
-----
|
|
|
|
**Interval (Genomic Intervals)**
|
|
|
|
- Tab delimited format (tabular)
|
|
- File must start with definition line in the following format (columns may be in any order).::
|
|
|
|
#CHROM START END STRAND
|
|
|
|
- CHROM - The name of the chromosome (e.g. chr3, chrY, chr2_random) or contig (e.g. ctgY1).
|
|
- START - The starting position of the feature in the chromosome or contig. The first base in a chromosome is numbered 0.
|
|
- END - The ending position of the feature in the chromosome or contig. The chromEnd base is not included in the display of the feature. For example, the first 100 bases of a chromosome are defined as chromStart=0, chromEnd=100, and span the bases numbered 0-99.
|
|
- STRAND - Defines the strand - either '+' or '-'.
|
|
|
|
- Example::
|
|
|
|
#CHROM START END STRAND NAME COMMENT
|
|
chr1 10 100 + exon myExon
|
|
chrX 1000 10050 - gene myGene
|
|
|
|
-----
|
|
|
|
**Lav**
|
|
|
|
Lav is the primary output format for BLASTZ. The first line of a .lav file begins with #:lav..
|
|
|
|
-----
|
|
|
|
**MAF**
|
|
|
|
TBA and multiz multiple alignment format. The first line of a .maf file begins with ##maf. This word is followed by white-space-separated "variable=value" pairs. There should be no white space surrounding the "=".
|
|
|
|
-----
|
|
|
|
**Scf**
|
|
|
|
A binary sequence file in 'scf' format with a '.scf' file extension. You must manually select this 'File Format' when uploading the file.
|
|
|
|
-----
|
|
|
|
**Sff**
|
|
|
|
A binary file in 'Standard Flowgram Format' with a '.sff' file extension.
|
|
|
|
-----
|
|
|
|
**Tabular (tab delimited)**
|
|
|
|
Any data in tab delimited format (tabular)
|
|
|
|
-----
|
|
|
|
**Table (delimiter-separated)**
|
|
|
|
Any delimiter-separated tabular data (CSV or TSV).
|
|
|
|
-----
|
|
|
|
**Wig**
|
|
|
|
The wiggle format is line-oriented. Wiggle data is preceded by a track definition line, which adds a number of options for controlling the default display of this track.
|
|
|
|
-----
|
|
|
|
**Other text type**
|
|
|
|
Any text file
|
|
</help>
|
|
</tool>
|