mirror of
https://github.com/galaxyproject/galaxy.git
synced 2026-09-24 16:30:27 +08:00
More interface work for converters and feature extractors
This commit is contained in:
@@ -5,11 +5,11 @@
|
||||
</command>
|
||||
<inputs>
|
||||
<page>
|
||||
<param format="gff" name="input1" type="data" label="Select a GFF file"/>
|
||||
<param format="gff" name="input1" type="data" label="Select GFF data"/>
|
||||
</page>
|
||||
<page>
|
||||
<conditional name="column_choice">
|
||||
<param name="col" type="select" label="Select a field">
|
||||
<param name="col" type="select" label="Select field">
|
||||
<option value="0" selected="true">Column 1 / Sequence name</option>
|
||||
<option value="1">Column 2 / Source</option>
|
||||
<option value="2">Column 3 / Feature</option>
|
||||
@@ -17,19 +17,19 @@
|
||||
<option value="7">Column 8 / Frame</option>
|
||||
</param>
|
||||
<when value="0">
|
||||
<param name="feature" label="Select Features" type="select" multiple="true" dynamic_options="get_features( input1, 0 )"/>
|
||||
<param name="feature" label="Select features" type="select" multiple="true" dynamic_options="get_features( input1, 0 )"/>
|
||||
</when>
|
||||
<when value="1">
|
||||
<param name="feature" label="Select Features" type="select" multiple="true" dynamic_options="get_features( input1, 1 )"/>
|
||||
<param name="feature" label="Select features" type="select" multiple="true" dynamic_options="get_features( input1, 1 )"/>
|
||||
</when>
|
||||
<when value="2">
|
||||
<param name="feature" label="Select Features" type="select" multiple="true" dynamic_options="get_features( input1, 2 )"/>
|
||||
<param name="feature" label="Select features" type="select" multiple="true" dynamic_options="get_features( input1, 2 )"/>
|
||||
</when>
|
||||
<when value="6">
|
||||
<param name="feature" label="Select Features" type="select" multiple="true" dynamic_options="get_features( input1, 6 )"/>
|
||||
<param name="feature" label="Select features" type="select" multiple="true" dynamic_options="get_features( input1, 6 )"/>
|
||||
</when>
|
||||
<when value="7">
|
||||
<param name="feature" label="Select Features" type="select" multiple="true" dynamic_options="get_features( input1, 7 )"/>
|
||||
<param name="feature" label="Select features" type="select" multiple="true" dynamic_options="get_features( input1, 7 )"/>
|
||||
</when>
|
||||
</conditional>
|
||||
<!--
|
||||
@@ -55,22 +55,45 @@
|
||||
-->
|
||||
<help>
|
||||
|
||||
.. class:: infomark
|
||||
**What it does**
|
||||
|
||||
**Info:** This tool extracts selected features from the input GFF file into an output GFF file. To convert the output GFF file to BED format, please use *GFF-to-BED converter* under *Convert Formats* section.
|
||||
This tool extracts selected features from a history item in GFF format. To convert the output GFF file to BED format, please use **GFF-to-BED converter** under **Convert Formats** section.
|
||||
|
||||
The interface for this tool contains two pages (steps):
|
||||
|
||||
* **Step 1 of 2**. Choose GFF data from history to extract features from.
|
||||
* **Step 2 of 2**. Choose field and features to be extracted. Once you select the field the **Select features** widget will refresh and display non-redundant set of values present in that column.
|
||||
|
||||
----
|
||||
|
||||
.. class:: warningmark
|
||||
|
||||
If the chosen column is absent in the input GFF file, the tool displays a message in the "Select Features" drop-down menu. Running the tool in this case will return an empty output.
|
||||
Also, running the tool on columns with ONLY one feature will return an output same as the input.
|
||||
If the chosen column is absent in the input GFF file, the tool displays a message in the "Select Features" drop-down menu. Running the tool in this case will return an empty output. Also, running the tool on columns with ONLY one feature will return an output same as the input.
|
||||
|
||||
|
||||
-----
|
||||
|
||||
**Example**
|
||||
|
||||
Selecting **promoter** from the following GFF-formatted data::
|
||||
|
||||
chr22 GeneA enhancer 10000000 10001000 500 + . TGA
|
||||
chr22 GeneA promoter 10010000 10010100 900 + . TGA
|
||||
chr22 GeneB promoter 10020000 10025000 400 - . TGB
|
||||
chr22 GeneB CCDS2220 10030000 10065000 800 - . TGB
|
||||
|
||||
will produce the following output::
|
||||
|
||||
chr22 GeneA promoter 10010000 10010100 900 + . TGA
|
||||
chr22 GeneB promoter 10020000 10025000 400 - . TGB
|
||||
|
||||
----
|
||||
|
||||
**Syntax**
|
||||
.. class:: infomark
|
||||
|
||||
- **GFF format** General Feature Format is a format for describing genes and other features associated with DNA, RNA and Protein sequences. GFF lines have nine tab-separated fields::
|
||||
**About formats**
|
||||
|
||||
**GFF format** General Feature Format is a format for describing genes and other features associated with DNA, RNA and Protein sequences. GFF lines have nine tab-separated fields::
|
||||
|
||||
1. seqname - Must be a chromosome or scaffold.
|
||||
2. source - The program that generated this feature.
|
||||
@@ -82,21 +105,6 @@ Also, running the tool on columns with ONLY one feature will return an output sa
|
||||
8. frame - If the feature is a coding exon, frame should be a number between 0-2 that represents the reading frame of the first base. If the feature is not a coding exon, the value should be '.'.
|
||||
9. group - All lines with the same group are linked together into a single item.
|
||||
|
||||
-----
|
||||
|
||||
**Example**
|
||||
|
||||
- Input GFF file::
|
||||
|
||||
chr22 TeleGene enhancer 10000000 10001000 500 + . TG1
|
||||
chr22 TeleGene promoter 10010000 10010100 900 + . TG1
|
||||
chr22 TeleGene promoter 10020000 10025000 400 - . TG2
|
||||
chr22 TeleGene CCDS2220 10030000 10065000 800 - . TG2
|
||||
|
||||
- Running this tool on the above input by choosing field as *Column 3 / Feature* and feature as *promoter* will produce the following output::
|
||||
|
||||
chr22 TeleGene promoter 10010000 10010100 900 + . TG1
|
||||
chr22 TeleGene promoter 10020000 10025000 400 - . TG2
|
||||
|
||||
</help>
|
||||
<code file="extract_GFF_Features_code.py"/>
|
||||
|
||||
+31
-24
@@ -15,18 +15,43 @@
|
||||
</tests>
|
||||
<help>
|
||||
|
||||
**Syntax**
|
||||
**What it does**
|
||||
|
||||
This tool converts a BED format file to GFF format.
|
||||
This tool converts data from BED format to GFF format (scroll down for format description).
|
||||
|
||||
- **BED format** Browser Extensible Data format was designed at UCSC for displaying data tracks in the Genome Browser. It has three required fields and twelve additional optional ones::
|
||||
--------
|
||||
|
||||
**Example**
|
||||
|
||||
The following data in BED format::
|
||||
|
||||
chr3 214671 265280 Hs.517745 300 +
|
||||
chrX 156881 157496 Hs.530320 300 +
|
||||
|
||||
Will be converted to GFF (**note** that the start coordinate is incremented by 1)::
|
||||
|
||||
## gff-version 2
|
||||
## bed2gff.pl $Rev: 601 $
|
||||
|
||||
chr3 bed2gff Hs.517745 214672 265280 . + . score "300";
|
||||
chrX bed2gff Hs.530320 156882 157496 . + . score "300";
|
||||
|
||||
------
|
||||
|
||||
.. class:: informark
|
||||
|
||||
**About formats**
|
||||
|
||||
**BED format** Browser Extensible Data format was designed at UCSC for displaying data tracks in the Genome Browser. It has three required fields and several additional optional ones:
|
||||
|
||||
The first three BED fields (required) are::
|
||||
|
||||
The first three BED fields (required) are:
|
||||
1. chrom - The name of the chromosome (e.g. chr1, chrY_random).
|
||||
2. chromStart - The starting position in the chromosome. (The first base in a chromosome is numbered 0.)
|
||||
3. chromEnd - The ending position in the chromosome, plus 1 (i.e., a half-open interval).
|
||||
|
||||
The twelve additional BED fields (optional) are:
|
||||
The additional BED fields (optional) are::
|
||||
|
||||
4. name - The name of the BED line.
|
||||
5. score - A score between 0 and 1000.
|
||||
6. strand - Defines the strand - either '+' or '-'.
|
||||
@@ -40,7 +65,7 @@ This tool converts a BED format file to GFF format.
|
||||
14. expIds - A comma-separated list of experiment ids. The number of items in this list should correspond to expCount.
|
||||
15. expScores - A comma-separated list of experiment scores. All of the expScores should be relative to expIds. The number of items in this list should correspond to expCount.
|
||||
|
||||
- **GFF format** General Feature Format is a format for describing genes and other features associated with DNA, RNA and Protein sequences. GFF lines have nine tab-separated fields::
|
||||
**GFF format** General Feature Format is a format for describing genes and other features associated with DNA, RNA and Protein sequences. GFF lines have nine tab-separated fields::
|
||||
|
||||
1. seqname - Must be a chromosome or scaffold.
|
||||
2. source - The program that generated this feature.
|
||||
@@ -52,23 +77,5 @@ This tool converts a BED format file to GFF format.
|
||||
8. frame - If the feature is a coding exon, frame should be a number between 0-2 that represents the reading frame of the first base. If the feature is not a coding exon, the value should be '.'.
|
||||
9. group - All lines with the same group are linked together into a single item.
|
||||
|
||||
------
|
||||
|
||||
**Example**
|
||||
|
||||
- BED format::
|
||||
|
||||
#chrom chromStart chromEnd name score strand thickStart thickEnd reserved blockCount blockSizes blockStarts
|
||||
chr3 214671 265280 Hs.517745 300 + 214671 265280 0 3 104,80,2030, 0,46624,48579,
|
||||
chrX 156881 157496 Hs.530320 300 + 156881 157496 0 2 231,384, 0,231,
|
||||
|
||||
- Convert the above file to GFF format::
|
||||
|
||||
## gff-version 2
|
||||
## bed2gff.pl $Rev: 601 $
|
||||
|
||||
chr3 bed2gff Hs.517745 214672 265280 . + . score "300";
|
||||
chrX bed2gff Hs.530320 156882 157496 . + . score "300";
|
||||
|
||||
</help>
|
||||
</tool>
|
||||
|
||||
+37
-31
@@ -15,30 +15,40 @@
|
||||
</tests>
|
||||
<help>
|
||||
|
||||
**Syntax**
|
||||
**What it does**
|
||||
|
||||
This tool converts a GFF format file to BED format.
|
||||
This tool converts data from GFF format to BED format (scroll down for format description).
|
||||
|
||||
- **GFF format** General Feature Format is a format for describing genes and other features associated with DNA, RNA and Protein sequences. GFF lines have nine tab-separated fields::
|
||||
--------
|
||||
|
||||
1. seqname - Must be a chromosome or scaffold.
|
||||
2. source - The program that generated this feature.
|
||||
3. feature - The name of this type of feature. Some examples of standard feature types are "CDS", "start_codon", "stop_codon", and "exon".
|
||||
4. start - The starting position of the feature in the sequence. The first base is numbered 1.
|
||||
5. end - The ending position of the feature (inclusive).
|
||||
6. score - A score between 0 and 1000. If there is no score value, enter ".".
|
||||
7. strand - Valid entries include '+', '-', or '.' (for don't know/care).
|
||||
8. frame - If the feature is a coding exon, frame should be a number between 0-2 that represents the reading frame of the first base. If the feature is not a coding exon, the value should be '.'.
|
||||
9. group - All lines with the same group are linked together into a single item.
|
||||
**Example**
|
||||
|
||||
- **BED format** Browser Extensible Data format was designed at UCSC for displaying data tracks in the Genome Browser. It has three required fields and twelve additional optional ones::
|
||||
The following data in GFF format::
|
||||
|
||||
chr22 GeneA enhancer 10000000 10001000 500 + . TGA
|
||||
chr22 GeneA promoter 10010000 10010100 900 + . TGA
|
||||
|
||||
Will be converted to BED (**note** that 1 is subtracted from the start coordinate)::
|
||||
|
||||
chr22 9999999 10001000 enhancer 0 +
|
||||
chr22 10009999 10010100 promoter 0 +
|
||||
|
||||
------
|
||||
|
||||
.. class:: infomark
|
||||
|
||||
**About formats**
|
||||
|
||||
**BED format** Browser Extensible Data format was designed at UCSC for displaying data tracks in the Genome Browser. It has three required fields and several additional optional ones:
|
||||
|
||||
The first three BED fields (required) are::
|
||||
|
||||
The first three BED fields (required) are:
|
||||
1. chrom - The name of the chromosome (e.g. chr1, chrY_random).
|
||||
2. chromStart - The starting position in the chromosome. (The first base in a chromosome is numbered 0.)
|
||||
3. chromEnd - The ending position in the chromosome, plus 1 (i.e., a half-open interval).
|
||||
3. chromEnd - The ending position in the chromosome, plus 1 (i.e., a half-open interval).
|
||||
|
||||
The additional BED fields (optional) are::
|
||||
|
||||
The twelve additional BED fields (optional) are:
|
||||
4. name - The name of the BED line.
|
||||
5. score - A score between 0 and 1000.
|
||||
6. strand - Defines the strand - either '+' or '-'.
|
||||
@@ -50,23 +60,19 @@ This tool converts a GFF format file to BED format.
|
||||
12. blockStarts - A comma-separated list of block starts. All of the blockStart positions should be calculated relative to chromStart. The number of items in this list should correspond to blockCount.
|
||||
13. expCount - The number of experiments.
|
||||
14. expIds - A comma-separated list of experiment ids. The number of items in this list should correspond to expCount.
|
||||
15. expScores - A comma-separated list of experiment scores. All of the expScores should be relative to expIds. The number of items in this list should correspond to expCount.
|
||||
15. expScores - A comma-separated list of experiment scores. All of the expScores should be relative to expIds. The number of items in this list should correspond to expCount.
|
||||
|
||||
------
|
||||
**GFF format** General Feature Format is a format for describing genes and other features associated with DNA, RNA and Protein sequences. GFF lines have nine tab-separated fields::
|
||||
|
||||
**Example**
|
||||
|
||||
- GFF format::
|
||||
|
||||
chr22 TeleGene enhancer 10000000 10001000 500 + . TG1
|
||||
chr22 TeleGene promoter 10010000 10010100 900 + . TG1
|
||||
chr22 TeleGene promoter 10020000 10025000 800 - . TG2
|
||||
|
||||
- Convert the above file to BED format::
|
||||
|
||||
chr22 9999999 10001000 enhancer 0 +
|
||||
chr22 10009999 10010100 promoter 0 +
|
||||
chr22 10019999 10025000 promoter 0 -
|
||||
1. seqname - Must be a chromosome or scaffold.
|
||||
2. source - The program that generated this feature.
|
||||
3. feature - The name of this type of feature. Some examples of standard feature types are "CDS", "start_codon", "stop_codon", and "exon".
|
||||
4. start - The starting position of the feature in the sequence. The first base is numbered 1.
|
||||
5. end - The ending position of the feature (inclusive).
|
||||
6. score - A score between 0 and 1000. If there is no score value, enter ".".
|
||||
7. strand - Valid entries include '+', '-', or '.' (for don't know/care).
|
||||
8. frame - If the feature is a coding exon, frame should be a number between 0-2 that represents the reading frame of the first base. If the feature is not a coding exon, the value should be '.'.
|
||||
9. group - All lines with the same group are linked together into a single item.
|
||||
|
||||
</help>
|
||||
</tool>
|
||||
|
||||
@@ -2,17 +2,15 @@
|
||||
<description>expander</description>
|
||||
<command interpreter="python">ucsc_gene_bed_to_exon_bed.py --input=$input1 --output=$out_file1 --region=$region "--exons"</command>
|
||||
<inputs>
|
||||
<param name="input1" type="data" format="bed" label="UCSC Gene Table"/>
|
||||
<param name="region" type="select">
|
||||
<label>Feature Type</label>
|
||||
<option value="transcribed">Coding + UTR</option>
|
||||
<option value="coding">Coding</option>
|
||||
<option value="utr3">3' UTR</option>
|
||||
<option value="utr5">5' UTR</option>
|
||||
<label>Extract</label>
|
||||
<option value="transcribed">Coding Exons + UTR Exons</option>
|
||||
<option value="coding">Coding Exons only</option>
|
||||
<option value="utr5">5'-UTR Exons</option>
|
||||
<option value="utr3">3'-UTR Exons</option>
|
||||
<option value="intron">Introns</option>
|
||||
</param>
|
||||
|
||||
|
||||
<param name="input1" type="data" format="bed" label="from" help="this history item must contain a 12 field BED (see below)"/>
|
||||
</inputs>
|
||||
<outputs>
|
||||
<data name="out_file1" format="bed"/>
|
||||
@@ -26,18 +24,44 @@
|
||||
</tests>
|
||||
<help>
|
||||
|
||||
**Syntax**
|
||||
.. class:: warningmark
|
||||
|
||||
This tool converts a UCSC gene bed format file to a list of bed format lines corresponding to requested features of each gene.
|
||||
This tool works only on a BED file that contains at least 12 fields (see **Example** and **About formats** below). The output will be empty if applied to a BED file with 3 or 6 fields.
|
||||
|
||||
- **BED format** Browser Extensible Data format was designed at UCSC for displaying data tracks in the Genome Browser. It has three required fields and twelve additional optional ones::
|
||||
------
|
||||
|
||||
**What it does**
|
||||
|
||||
BED format can be used to represent a single gene in just one line, which contains the information about exons, coding sequence location (CDS), and positions of untranslated regions (UTRs). This tool *unpacks* this information by converting a single line describing a gene into a collection of lines representing individual exons, introns, UTRs, etc.
|
||||
|
||||
-------
|
||||
|
||||
**Example**
|
||||
|
||||
Extracting **Coding Exons + UTR Exons** from the following two BED lines::
|
||||
|
||||
chr7 127475281 127491632 NM_000230 0 + 127486022 127488767 0 3 29,172,3225, 0,10713,13126
|
||||
chr7 127486011 127488900 D49487 0 + 127486022 127488767 0 2 155,490, 0,2399
|
||||
|
||||
will return::
|
||||
|
||||
chr7 127475281 127475310 NM_000230 0 +
|
||||
chr7 127485994 127486166 NM_000230 0 +
|
||||
chr7 127488407 127491632 NM_000230 0 +
|
||||
chr7 127486011 127486166 D49487 0 +
|
||||
chr7 127488410 127488900 D49487 0 +
|
||||
|
||||
------
|
||||
|
||||
.. class:: infomark
|
||||
|
||||
**About formats**
|
||||
|
||||
**BED format** Browser Extensible Data format was designed at UCSC for displaying data tracks in the Genome Browser. It has three required fields and additional optional ones. In the specific case of this tool the following fields must be present::
|
||||
|
||||
The first three BED fields (required) are:
|
||||
1. chrom - The name of the chromosome (e.g. chr1, chrY_random).
|
||||
2. chromStart - The starting position in the chromosome. (The first base in a chromosome is numbered 0.)
|
||||
3. chromEnd - The ending position in the chromosome, plus 1 (i.e., a half-open interval).
|
||||
|
||||
The twelve additional BED fields (optional) are:
|
||||
4. name - The name of the BED line.
|
||||
5. score - A score between 0 and 1000.
|
||||
6. strand - Defines the strand - either '+' or '-'.
|
||||
@@ -47,26 +71,7 @@ This tool converts a UCSC gene bed format file to a list of bed format lines cor
|
||||
10. blockCount - The number of blocks (exons) in the BED line.
|
||||
11. blockSizes - A comma-separated list of the block sizes. The number of items in this list should correspond to blockCount.
|
||||
12. blockStarts - A comma-separated list of block starts. All of the blockStart positions should be calculated relative to chromStart. The number of items in this list should correspond to blockCount.
|
||||
13. expCount - The number of experiments.
|
||||
14. expIds - A comma-separated list of experiment ids. The number of items in this list should correspond to expCount.
|
||||
15. expScores - A comma-separated list of experiment scores. All of the expScores should be relative to expIds. The number of items in this list should correspond to expCount.
|
||||
|
||||
-----
|
||||
|
||||
**Example**
|
||||
|
||||
- A UCSC gene bed format file::
|
||||
|
||||
chr7 127475281 127491632 NM_000230 0 + 127486022 127488767 0 3 29,172,3225, 0,10713,13126
|
||||
chr7 127486011 127488900 D49487 0 + 127486022 127488767 0 2 155,490, 0,2399
|
||||
|
||||
- Converts the above file to a list of bed lines, which has the transcribed regions and overlap with exons. (if user selects **Coding + UTR**)::
|
||||
|
||||
chr7 127475281 127475310 NM_000230 0 +
|
||||
chr7 127485994 127486166 NM_000230 0 +
|
||||
chr7 127488407 127491632 NM_000230 0 +
|
||||
chr7 127486011 127486166 D49487 0 +
|
||||
chr7 127488410 127488900 D49487 0 +
|
||||
|
||||
</help>
|
||||
</tool>
|
||||
|
||||
Reference in New Issue
Block a user