mirror of
https://github.com/galaxyproject/galaxy.git
synced 2026-09-24 16:30:27 +08:00
Lca commit. The tool itself is written by guru with minor modification made by me.
This commit is contained in:
@@ -156,6 +156,7 @@
|
||||
<tool file="taxonomy/t2t_report.xml" />
|
||||
<tool file="taxonomy/t2ps_wrapper.xml" />
|
||||
<tool file="taxonomy/find_diag_hits.xml" />
|
||||
<tool file="taxonomy/lca.xml" />
|
||||
<tool file="taxonomy/poisson2test.xml" />
|
||||
</section>
|
||||
<section name="Solexa tools" id="solexa_tools">
|
||||
|
||||
+13
-6
@@ -99,7 +99,10 @@ def main():
|
||||
out_list[0] = str(prev_item)
|
||||
out_list[1] = str(prev_vals[0][0])
|
||||
out_list[2] = str(prev_vals[1][0])
|
||||
out_list[24] = str(prev_vals[23][0])
|
||||
try:
|
||||
out_list[24] = str(prev_vals[23][0])
|
||||
except:
|
||||
pass
|
||||
for k, col in enumerate(cols):
|
||||
if col >= 3 and col < 24:
|
||||
if len(set(prev_vals[k])) == 1:
|
||||
@@ -111,12 +114,12 @@ def main():
|
||||
k += 1
|
||||
|
||||
if rank_bound == 0:
|
||||
print >>fout, '\t'.join(out_list)
|
||||
print >>fout, '\t'.join(out_list).strip()
|
||||
#print 'n'*( 24 - rank_bound )
|
||||
else:
|
||||
#print '\t'.join(out_list[rank_bound:24])
|
||||
if ''.join(out_list[rank_bound:24]) != 'n'*( 24 - rank_bound ):
|
||||
print >>fout, '\t'.join(out_list)
|
||||
print >>fout, '\t'.join(out_list).strip()
|
||||
|
||||
block_valid = True
|
||||
prev_item = item
|
||||
@@ -145,7 +148,11 @@ def main():
|
||||
out_list[0] = str(prev_item)
|
||||
out_list[1] = str(prev_vals[0][0])
|
||||
out_list[2] = str(prev_vals[1][0])
|
||||
out_list[24] = str(prev_vals[23][0])
|
||||
try:
|
||||
out_list[24] = str(prev_vals[23][0])
|
||||
except:
|
||||
pass
|
||||
|
||||
for k, col in enumerate(cols):
|
||||
if col >= 3 and col < 24:
|
||||
if len(set(prev_vals[k])) == 1:
|
||||
@@ -157,12 +164,12 @@ def main():
|
||||
k += 1
|
||||
|
||||
if rank_bound == 0:
|
||||
print >>fout, '\t'.join(out_list)
|
||||
print >>fout, '\t'.join(out_list).strip()
|
||||
else:
|
||||
#print ''.join(out_list[rank_bound:24])
|
||||
#print 'n'*( 24 - rank_bound )
|
||||
if ''.join(out_list[rank_bound:24]) != 'n'*( 24 - rank_bound ):
|
||||
print >>fout, '\t'.join(out_list)
|
||||
print >>fout, '\t'.join(out_list).strip()
|
||||
|
||||
if skipped_lines > 0:
|
||||
print "Skipped %d invalid lines." % ( skipped_lines )
|
||||
|
||||
+48
-7
@@ -1,12 +1,12 @@
|
||||
<tool id="lca1" name="Least Common Ancestor" version="1.0.0">
|
||||
<tool id="lca1" name="Find lowest diagnostic rank" version="1.0.0">
|
||||
<description></description>
|
||||
<command interpreter="python">
|
||||
lca.py $input1 $out_file1 $rank_bound
|
||||
</command>
|
||||
<inputs>
|
||||
<param format="taxonomy" name="input1" type="data" label="Select taxonomy dataset"/>
|
||||
<param name="rank_bound" label="rank bound" type="select" help="Choose smallest diagnostic level">
|
||||
<option value="0">Everything</option>
|
||||
<param format="taxonomy" name="input1" type="data" label="for taxonomy dataset"/>
|
||||
<param name="rank_bound" label="require the lowest rank to be at least" type="select">
|
||||
<option value="0">No restriction</option>
|
||||
<option value="3">Superkingdom</option>
|
||||
<option value="4">Kingdom</option>
|
||||
<option value="5">Subkingdom</option>
|
||||
@@ -32,13 +32,54 @@
|
||||
</inputs>
|
||||
<outputs>
|
||||
<data format="taxonomy" name="out_file1" metadata_source="input1" />
|
||||
</outputs>
|
||||
</outputs>
|
||||
<tests>
|
||||
<test>
|
||||
<param name="input1" value="lca_input.taxonomy" ftype="taxonomy"/>
|
||||
<param name="rank_bound" value="0" />
|
||||
<output name="out_file1" file="lca_output.taxonomy" ftype="taxonomy"/>
|
||||
</test>
|
||||
</tests>
|
||||
|
||||
<help>
|
||||
<help>
|
||||
|
||||
**What it does**
|
||||
|
||||
When performing metagenomic analyses it is often necessary to identify sequence reads corresponding to a particular taxonomic group, or, in other words, diagnostic of a particular taxonomic rank. This utility performs this analysis. It takes data generated by *Taxonomy manipulation->Fetch Taxonomic Ranks* as input and outputs either a list of sequence reads unique to a particular taxonomic rank, or a list of taxonomic ranks and the count of unique reads corresponding to each rank.
|
||||
This tool identifies the lowest taxonomic rank for which a mategenomic sequencing read is diagnostic. It takes datasets produced by *Fetch Taxonomic Ranks* tool (aka Taxonomy format) as the input.
|
||||
|
||||
-------
|
||||
|
||||
**Example**
|
||||
|
||||
Suppose you have two reads, **read_1** and **read_2**, with the following taxonomic profiles (scroll sideways to see the entire dataset)::
|
||||
|
||||
read_1 1 root superkingdom1 kingdom1 subkingdom1 superphylum1 phylum1 subphylum1 superclass1 class1 subclass1 superorder1 order1 suborder1 superfamily1 family1 subfamily1 tribe1 subtribe1 genus1 subgenus1 species1 subspecies1
|
||||
read_1 2 root superkingdom1 kingdom1 subkingdom1 superphylum1 phylum1 subphylum1 superclass1 class1 subclass1 superorder1 order1 suborder1 superfamily1 family1 subfamily1 tribe1 subtribe1 genus2 subgenus2 species2 subspecies2
|
||||
read_2 3 root superkingdom1 kingdom1 subkingdom1 superphylum1 phylum3 subphylum3 superclass3 class3 subclass3 superorder3 order3 suborder3 superfamily3 family3 subfamily3 tribe3 subtribe3 genus3 subgenus3 species3 subspecies3
|
||||
read_2 4 root superkingdom1 kingdom1 subkingdom1 superphylum1 phylum4 subphylum4 superclass4 class4 subclass4 superorder4 order4 suborder4 superfamily4 family4 subfamily4 tribe4 subtribe4 genus4 subgenus4 species4 subspecies4
|
||||
|
||||
For **read_1** taxonomic labels are consistent until the genus level, where the taxonomy splits into two branches, one ending with *subspecies1* and the other with *subspecies2*. This implies **that the lowest taxomomic rank read_1 can identify is SUBTRIBE**. Similarly, read_2 is diagnostic up until the **superphylum** level. As a results the output of this tool will be::
|
||||
|
||||
read_1 2 root superkingdom1 kingdom1 subkingdom1 superphylum1 phylum1 subphylum1 superclass1 class1 subclass1 superorder1 order1 suborder1 superfamily1 family1 subfamily1 tribe1 subtribe1 n n n n
|
||||
read_2 3 root superkingdom1 kingdom1 subkingdom1 superphylum1 n n n n n n n n n n n n n n n n n
|
||||
|
||||
where, **n** means *EMPTY*.
|
||||
|
||||
--------
|
||||
|
||||
**What's up with the drop down?**
|
||||
|
||||
Why do we need the *require the lowest rank to be at least* dropdown? Let's look at the above example again. Suppose you need to find only those reads that are diagnostic on at least phylum level. To do this you need to set the *require the lowest rank to be at least* to **phylum**. As a result your output will look like this::
|
||||
|
||||
read_1 2 root superkingdom1 kingdom1 subkingdom1 superphylum1 phylum1 subphylum1 superclass1 class1 subclass1 superorder1 order1 suborder1 superfamily1 family1 subfamily1 tribe1 subtribe1 n n n n
|
||||
|
||||
.. class:: infomark
|
||||
|
||||
Note, that **read_2** is now omitted as it matches two phyla (**phylum3** and **phylum4**) and therefore is not diagnostic (but rather cosmopolitan) on *phylum* level.
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
</help>
|
||||
</tool>
|
||||
Reference in New Issue
Block a user