More changes plus modification requested by Madelaine Gogol

This commit is contained in:
Anton Nekrutenko
2008-04-28 16:08:30 +00:00
parent 529301036c
commit e2ca36ea59
4 changed files with 42 additions and 35 deletions
+1 -1
View File
@@ -2,7 +2,7 @@
<description>of a file</description>
<command interpreter="perl">remove_beginning.pl $input $num_lines $out_file1</command>
<inputs>
<param name="num_lines" size="5" type="integer" value="5" label="Remove first" help="lines"/>
<param name="num_lines" size="5" type="integer" value="1" label="Remove first" help="lines"/>
<param format="txt" name="input" type="data" label="from"/>
</inputs>
<outputs>
+2 -2
View File
@@ -5,12 +5,12 @@
<param format="tabular" name="input" type="data" label="Sort Query" />
<param name="column" label="on column" type="data_column" data_ref="input" accept_default="true" />
<param name="order" type="select" label="in">
<option value="ASC">Ascending order</option>
<option value="DESC">Descending order</option>
<option value="ASC">Ascending order</option>
</param>
<param name="style" type="select" label="Flavor">
<option value="alpha">Alphabetical sort</option>
<option value="num">Numerical sort</option>
<option value="alpha">Alphabetical sort</option>
</param>
</inputs>
<outputs>
+6 -8
View File
@@ -2,16 +2,15 @@
<description>for Metagenomics Projects</description>
<command interpreter="python">megablast_wrapper.py $source_select $input_query $output1 $word_size $iden_cutoff $disc_word $disc_type $filter_query ${GALAXY_DATA_INDEX_DIR}</command>
<inputs>
<param name="source_select" type="select" display="radio" label="Target database">
<param name="source_select" type="select" display="radio" label="Choose target database">
<options from_file="blastdb.loc" name_col="0" value_col="0"/>
</param>
<param name="input_query" type="data" format="fasta" label="Sequence file"/>
<param name="word_size" type="select" label="Word size (-W)" help="Size of best perfect match">
<option value="28">28</option>
<option value="16">16</option>
<option value="12">12</option>
</param>
<param name="iden_cutoff" type="float" size="15" value="0.0" label="Identity percentage cut-off (-p)" help="no cutoff if 0" />
<param name="iden_cutoff" type="float" size="15" value="90.0" label="Identity percentage cut-off (-p)" help="no cutoff if 0" />
<param name="disc_word" type="integer" size="15" value="0" label="Length of a discontiguous word template (-t)" help="contiguous word if 0. Only support 16, 18, or 21."/>
<param name="disc_type" type="select" label="Type of a discontiguous word template (-N)">
<option value="0">0 - coding</option>
@@ -43,16 +42,15 @@
</tests>
<help>
.. class:: infomark
.. class:: warningmark
**TIP**. Searching might require few hours before returning the result.
**Note**. Database searches may take substantial amount of time. For large input datasets it is advisable to allow overnight processing.
-----
**What it does**
This tool runs **megablast** (for information about megablast, please see the reference below) and search against a nucleotide database.
The query sequence could be any nucleotide sequences in fasta format.
This tool runs **megablast** (for information about megablast, please see the reference below) a high performance nucleotide local aligner developed by Webb Miller and colleagues.
-----
@@ -81,7 +79,7 @@ The query sequence could be any nucleotide sequences in fasta format.
**Reference**
**Megablast**: Zhang et al. A Greedy Algorithm for Aligning DNA Sequences. 2000. JCB: 203-214.
Zhang et al. A Greedy Algorithm for Aligning DNA Sequences. 2000. JCB: 203-214.
</help>
</tool>
+33 -24
View File
@@ -1,5 +1,5 @@
<tool id="megablast_xml_parser" name="Parse megablast xml output">
<description> </description>
<tool id="megablast_xml_parser" name="Parse">
<description>XML output from blast programs</description>
<command interpreter="python">megablast_xml_parser.py $input1 $output1</command>
<inputs>
<param name="input1" type="data" format="txt" label="Megablast XML output" />
@@ -20,40 +20,49 @@
.. class:: warningmark
**IMPORTANT**. Please **zip** your xml file and upload the zipped file. This will help speeding up the process. (hint: under shell command, try: gzip -c your_xml_file > your_xml_file.gz)
.. class:: infomark
**TIP**. To get xml output from megablast, add option **-m 7** in the command line.
Blast XML output **must** be uploaded to Galaxy in zipped form.
-----
**What it does**
This tool is for users to upload their own megablast xml output and converts to tabular format, showing all information for retrieving alignment blocks.
This tool will process XML output of any NCBI blast tool (if you run your own blast jobs, the XML output can be generated with **-m 7** option).
-----
**Example**
**Output fields**
- Query sequence::
&gt;seq1
CGGACAGCGCCGCCACCAACAAAGCCACCA
- Parsed output::
+------+------+-----------------+----------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
| Qid | Qlen | Sid | Slen | HSP |
+------+------+-----------------+----------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
| seq1 | 30 | gnl|BL_ORD_ID|0 | 5528445 | 59.96 | 8.38112e-11 | 1 | 30 | 5436010 | 5436039 | 1 | 1 | 30 | 30 | CGGACAGCGCCGCCACCAACAAAGCCACCA | CGGACAGCGCCGCCACCAACAAAGCCACCA | |||||||||||||||||||||||||||||| |
+------+------+-----------------+----------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
This tools returns tab-delimted output with the following fields::
-----
Description Example
----------------------------------------- -----------------
**Reference**
1. Name of the query sequence Seq1
2. Length of the query sequence 30
3. Name of target sequence gnl|BL_ORD_ID|0
4. Length of target sequence 5528445
5. Alignment bit score 59.96
6. E-value 8.38112e-11
7. Start of alignment within query 1
8. End of alignment within query 30
9. Start of alignment within target 5436010
10. End of alignment within target 5436039
11. Query frame 1
12. Target frame 1
13. Number of identical bases within 29
the alignment
14. Alignment length 30
15. Aligned portion (sequence) of query CGGACAGCGCCGCCACCAACAAAGCCACCA
16. Aligned portion (sequence) of target CGGACAGCGCCGCCACCAACAAAGCCATCA
17. Midline indicating positions of ||||||||||||||||||||||||||| ||
matches within the alignment
------
.. class:: infomark
Note that this form of output does not contain alignment identify value. However, it can be computed by dividing the number of identical bases wityh the alignment (Field 13) by the alignment length (Field 14) using *Text Manipulation->Compute* tool
**Megablast**: Zhang et al. A Greedy Algorithm for Aligning DNA Sequences. 2000. JCB: 203-214.
</help>