Integration of Assaf Gordon's Solexa toolkit into Galaxy - (all tools except 3 are working fine).

This commit is contained in:
Guruprasad Anada
2009-02-20 14:17:46 -05:00
parent 028a662bad
commit 6fd54350eb
26 changed files with 949 additions and 0 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.5 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 5.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 12 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 6.9 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 13 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 13 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 10 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 10 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 12 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 44 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 9.8 KiB

+19
View File
@@ -288,6 +288,25 @@
<tool file="fasta_tools/fasta_to_tabular.xml" />
<tool file="fasta_tools/tabular_to_fasta.xml" />
</section>
<section name="FASTA/Q Information" id="cshl_library_information">
<tool file="fastx_toolkit/fastq_qual_stat.xml" />
<tool file="fastx_toolkit/fastq_quality_boxplot.xml" />
<tool file="fastx_toolkit/fastq_nucleotides_distribution.xml" />
<!-- <tool file="fastx_toolkit/fasta_clipping_histogram.xml" /> -->
</section>
<section name="FASTA/Q Preprocessing" id="cshl_fastx_manipulation">
<tool file="fastx_toolkit/fastq_to_fasta.xml" />
<tool file="fastx_toolkit/fastq_qual_conv.xml" />
<!-- <tool file="fastx_toolkit/fastx_clipper.xml" /> -->
<tool file="fastx_toolkit/fastx_trimmer.xml" />
<tool file="fastx_toolkit/fastx_reverse_complement.xml" />
<tool file="fastx_toolkit/fastx_artifacts_filter.xml" />
<tool file="fastx_toolkit/fastq_quality_filter.xml" />
<!-- <tool file="fastx_toolkit/fasta_collapser.xml" /> -->
<!-- <tool file="fastx_toolkit/fastx_barcode_splitter.xml" /> -->
</section>
<section name="Short Read QC and Manipulation" id="short_read_analysis">
<tool file="metag_tools/short_reads_figure_score.xml" />
<tool file="metag_tools/short_reads_figure_high_quality_length.xml" />
@@ -0,0 +1,38 @@
<tool id="cshl_fasta_clipping_histogram" name="Length Distribution">
<description>chart</description>
<command>fasta_clipping_histogram.pl $input $outfile</command>
<inputs>
<param format="fasta" name="input" type="data" label="Library to analyze" />
</inputs>
<outputs>
<data format="png" name="outfile" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool creates a histogram image of sequence lengths distribution in a given fasta data set file.
**TIP:** Use this tool after clipping your library (with **FASTX Clipper tool**), to visualize the clipping results.
-----
**Output Examples**
In the following library, most sequences are 24-mers to 27-mers.
This could indicate an abundance of endo-siRNAs (depending of course of what you've tried to sequence in the first place).
.. image:: ../static/fastx_icons/fasta_clipping_histogram_1.png
In the following library, most sequences are 19,22 or 23-mers.
This could indicate an abundance of miRNAs (depending of course of what you've tried to sequence in the first place).
.. image:: ../static/fastx_icons/fasta_clipping_histogram_2.png
</help>
</tool>
<!-- FASTA-Clipping-Histogram is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
+75
View File
@@ -0,0 +1,75 @@
<tool id="cshl_fasta_collapser" name="Collapse">
<description>sequences</description>
<command>fasta_collapser.pl $input $output</command>
<inputs>
<param format="fasta" name="input" type="data" label="Library to collapse" />
</inputs>
<tests>
<test>
<param name="input" value="fasta_collapser1.fasta" />
<output name="output" file="fasta_collapser1.out" />
</test>
</tests>
<outputs>
<data format="fasta" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool collapses identical sequences in a FASTA file into a single sequence.
--------
**Example**
Example Input File (Sequence "ATAT" appears multiple times)::
>CSHL_2_FC0042AGLLOO_1_1_605_414
TGCG
>CSHL_2_FC0042AGLLOO_1_1_537_759
ATAT
>CSHL_2_FC0042AGLLOO_1_1_774_520
TGGC
>CSHL_2_FC0042AGLLOO_1_1_742_502
ATAT
>CSHL_2_FC0042AGLLOO_1_1_781_514
TGAG
>CSHL_2_FC0042AGLLOO_1_1_757_487
TTCA
>CSHL_2_FC0042AGLLOO_1_1_903_769
ATAT
>CSHL_2_FC0042AGLLOO_1_1_724_499
ATAT
Example Output file::
>1-1
TGCG
>2-4
ATAT
>3-1
TGGC
>4-1
TGAG
>5-1
TTCA
.. class:: infomark
Original Sequence Names / Lane descriptions (e.g. "CSHL_2_FC0042AGLLOO_1_1_742_502") are discarded.
The output seqeunce name is composed of two numbers: the first is the sequence's number, the second is the multiplicity value.
The following output::
>2-4
ATAT
means that the sequence "ATAT" is the second sequence in the file, and it appeared 4 times in the input FASTA file.
</help>
</tool>
@@ -0,0 +1,66 @@
<tool id="cshl_fastq_nucleotides_distribution" name="Nucleotides Distribution">
<description>chart</description>
<command>fastq_nucleotide_distribution_graph.sh -t '$input.name' -i $input -o $output</command>
<inputs>
<param format="txt" name="input" type="data" label="Statistics Text File (output of 'FASTQ Statistics' tool)" />
</inputs>
<outputs>
<data format="png" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
Creates a stacked-histogram graph for the nucleotide distribution in the Solexa library.
.. class:: infomark
**TIP:** Use the **FASTQ Statistics** tool to generate the report file needed for this tool.
-----
**Output Examples**
The following chart clearly shows the barcode used at the 5'-end of the library: **GATCT**
.. image:: ../static/fastx_icons/fastq_nucleotides_distribution_1.png
In the following chart, one can almost 'read' the most abundant sequence by looking at the dominant values: **TGATA TCGTA TTGAT GACTG AA...**
.. image:: ../static/fastx_icons/fastq_nucleotides_distribution_2.png
The following chart shows a growing number of unknown (N) nucleotides towards later cycles (which might indicate a sequencing problem):
.. image:: ../static/fastx_icons/fastq_nucleotides_distribution_3.png
But most of the time, the chart will look rather random:
.. image:: ../static/fastx_icons/fastq_nucleotides_distribution_4.png
</help>
</tool>
<!-- FASTQ-Nucleotides-Distribution is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
+82
View File
@@ -0,0 +1,82 @@
<tool id="cshl_fastq_qual_conv" name="Quality format converter">
<description>(ASCII-Numeric)</description>
<command>zcat -f $input | fastq_quality_converter $QUAL_FORMAT -o $output</command>
<inputs>
<param format="fastqsolexa" name="input" type="data" label="Library to convert" />
<param name="QUAL_FORMAT" type="select" label="Desired output format">
<option value="-a">ASCII (letters) quality scores</option>
<option value="-n">Numeric quality scores</option>
</param>
</inputs>
<tests>
<test>
<!-- ASCII to NUMERIC -->
<param name="input" value="fastq_qual_conv1.fastq" />
<param name="QUAL_FORMAT" value="Numeric quality scores" />
<output name="output" file="fastq_qual_conv1.out" />
</test>
<test>
<!-- ASCII to ASCII (basically, a no-op, but it should still produce a valid output -->
<param name="input" value="fastq_qual_conv1.fastq" />
<param name="QUAL_FORMAT" value="ASCII (letters) quality scores" />
<output name="output" file="fastq_qual_conv1a.out" />
</test>
<test>
<!-- NUMERIC to ASCII -->
<param name="input" value="fastq_qual_conv2.fastq" />
<param name="QUAL_FORMAT" value="ASCII (letters) quality scores" />
<output name="output" file="fastq_qual_conv2.out" />
</test>
<test>
<!-- NUMERIC to NUMERIC (basically, a no-op, but it should still produce a valid output -->
<param name="input" value="fastq_qual_conv2.fastq" />
<param name="QUAL_FORMAT" value="Numeric quality scores" />
<output name="output" file="fastq_qual_conv2n.out" />
</test>
</tests>
<outputs>
<data format="fastqsolexa" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
Converts a solexa FASTQ file to/from numeric or ASCII quality format.
.. class:: warningmark
Re-scaling is **not** performed. (e.g. conversion from Phred scale to Solexa scale).
-----
FASTQ with Numeric quality scores::
@CSHL__2_FC042AGWWWXX:8:1:120:202
ACGATAGATCGGAAGAGCTAGTATGCCGTTTTCTGC
+CSHL__2_FC042AGWWWXX:8:1:120:202
40 40 40 40 20 40 40 40 40 6 40 40 28 40 40 25 40 20 40 -1 30 40 14 27 40 8 1 3 7 -1 11 10 -1 21 10 8
@CSHL__2_FC042AGWWWXX:8:1:103:1185
ATCACGATAGATCGGCAGAGCTCGTTTACCGTCTTC
+CSHL__2_FC042AGWWWXX:8:1:103:1185
40 40 40 40 40 35 33 31 40 40 40 32 30 22 40 -0 9 22 17 14 8 36 15 34 22 12 23 3 10 -0 8 2 4 25 30 2
FASTQ with ASCII quality scores::
@CSHL__2_FC042AGWWWXX:8:1:120:202
ACGATAGATCGGAAGAGCTAGTATGCCGTTTTCTGC
+CSHL__2_FC042AGWWWXX:8:1:120:202
hhhhThhhhFhh\hhYhTh?^hN[hHACG?KJ?UJH
@CSHL__2_FC042AGWWWXX:8:1:103:1185
ATCACGATAGATCGGCAGAGCTCGTTTACCGTCTTC
+CSHL__2_FC042AGWWWXX:8:1:103:1185
hhhhhca_hhh`^Vh@IVQNHdObVLWCJ@HBDY^B
</help>
</tool>
<!-- FASTQ-Quality-Converter is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
+100
View File
@@ -0,0 +1,100 @@
<tool id="cshl_fastq_qual_stat" name="Quality Statistics">
<description></description>
<command>zcat -f $input | fastq_quality_stats -o $output</command>
<inputs>
<param format="fastqsolexa" name="input" type="data" label="Library to analyse" />
</inputs>
<tests>
<test>
<param name="input" value="fastq_stats1.fastq" />
<output name="output" file="fastq_stats1.out" />
</test>
</tests>
<outputs>
<data format="txt" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
Creates quality statistics report for the given Solexa/FASTQ library.
.. class:: infomark
**TIP:** This statistics report can be used as input for **Quality Score** and **Nucleotides Distribution** tools.
-----
**The output file will contain the following fields:**
* column = column number (1 to 36 for a 36-cycles read solexa file)
* count = number of bases found in this column.
* min = Lowest quality score value found in this column.
* max = Highest quality score value found in this column.
* sum = Sum of quality score values for this column.
* mean = Mean quality score value for this column.
* Q1 = 1st quartile quality score.
* med = Median quality score.
* Q3 = 3rd quartile quality score.
* IQR = Inter-Quartile range (Q3-Q1).
* lW = 'Left-Whisker' value (for boxplotting).
* rW = 'Right-Whisker' value (for boxplotting).
* A_Count = Count of 'A' nucleotides found in this column.
* C_Count = Count of 'C' nucleotides found in this column.
* G_Count = Count of 'G' nucleotides found in this column.
* T_Count = Count of 'T' nucleotides found in this column.
* N_Count = Count of 'N' nucleotides found in this column.
**Output Example**::
column count min max sum mean Q1 med Q3 IQR lW rW A_Count C_Count G_Count T_Count N_Count
1 6362991 -4 40 250734117 39.41 40 40 40 0 40 40 1396976 1329101 678730 2958184 0
2 6362991 -5 40 250531036 39.37 40 40 40 0 40 40 1786786 1055766 1738025 1782414 0
3 6362991 -5 40 248722469 39.09 40 40 40 0 40 40 2296384 984875 1443989 1637743 0
4 6362991 -5 40 247654797 38.92 40 40 40 0 40 40 1683197 1410855 1722633 1546306 0
5 6362991 -4 40 248214827 39.01 40 40 40 0 40 40 2536861 1167423 1248968 1409739 0
6 6362991 -5 40 248499903 39.05 40 40 40 0 40 40 1598956 1236081 1568608 1959346 0
7 6362991 -4 40 247719760 38.93 40 40 40 0 40 40 1692667 1822140 1496741 1351443 0
8 6362991 -5 40 245745205 38.62 40 40 40 0 40 40 2230936 1343260 1529928 1258867 0
9 6362991 -5 40 245766735 38.62 40 40 40 0 40 40 1702064 1306257 1336511 2018159 0
10 6362991 -5 40 245089706 38.52 40 40 40 0 40 40 1519917 1446370 1450995 1945709 0
11 6362991 -5 40 242641359 38.13 40 40 40 0 40 40 1717434 1282975 1387804 1974778 0
12 6362991 -5 40 242026113 38.04 40 40 40 0 40 40 1662872 1202041 1519721 1978357 0
13 6362991 -5 40 238704245 37.51 40 40 40 0 40 40 1549965 1271411 1973291 1566681 1643
14 6362991 -5 40 235622401 37.03 40 40 40 0 40 40 2101301 1141451 1603990 1515774 475
15 6362991 -5 40 230766669 36.27 40 40 40 0 40 40 2344003 1058571 1440466 1519865 86
16 6362991 -5 40 224466237 35.28 38 40 40 2 35 40 2203515 1026017 1474060 1651582 7817
17 6362991 -5 40 219990002 34.57 34 40 40 6 25 40 1522515 1125455 2159183 1555765 73
18 6362991 -5 40 214104778 33.65 30 40 40 10 15 40 1479795 2068113 1558400 1249337 7346
19 6362991 -5 40 212934712 33.46 30 40 40 10 15 40 1432749 1231352 1769799 1920093 8998
20 6362991 -5 40 212787944 33.44 29 40 40 11 13 40 1311657 1411663 2126316 1513282 73
21 6362991 -5 40 211369187 33.22 28 40 40 12 10 40 1887985 1846300 1300326 1318380 10000
22 6362991 -5 40 213371720 33.53 30 40 40 10 15 40 542299 3446249 516615 1848190 9638
23 6362991 -5 40 221975899 34.89 36 40 40 4 30 40 347679 1233267 926621 3855355 69
24 6362991 -5 40 194378421 30.55 21 40 40 19 -5 40 433560 674358 3262764 1992242 67
25 6362991 -5 40 199773985 31.40 23 40 40 17 -2 40 944760 325595 1322800 3769641 195
26 6362991 -5 40 179404759 28.20 17 34 40 23 -5 40 3457922 156013 1494664 1254293 99
27 6362991 -5 40 163386668 25.68 13 28 40 27 -5 40 1392177 281250 3867895 821491 178
28 6362991 -5 40 156230534 24.55 12 25 40 28 -5 40 907189 981249 4174945 299437 171
29 6362991 -5 40 163236046 25.65 13 28 40 27 -5 40 1097171 3418678 1567013 280008 121
30 6362991 -5 40 151309826 23.78 12 23 40 28 -5 40 3514775 2036194 566277 245613 132
31 6362991 -5 40 141392520 22.22 10 21 40 30 -5 40 1569000 4571357 124732 97721 181
32 6362991 -5 40 143436943 22.54 10 21 40 30 -5 40 1453607 4519441 38176 351107 660
33 6362991 -5 40 114269843 17.96 6 14 30 24 -5 40 3311001 2161254 155505 734297 934
34 6362991 -5 40 140638447 22.10 10 20 40 30 -5 40 1501615 1637357 18113 3205237 669
35 6362991 -5 40 138910532 21.83 10 20 40 30 -5 40 1532519 3495057 23229 1311834 352
36 6362991 -5 40 117158566 18.41 7 15 30 23 -5 40 4074444 1402980 63287 822035 245
</help>
</tool>
<!-- FASTQ-Statistics is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
@@ -0,0 +1,47 @@
<tool id="cshl_fastq_quality_boxplot" name="Quality Score">
<description>chart</description>
<command>fastq_quality_boxplot_graph.sh -t '$input.name' -i $input -o $output</command>
<inputs>
<param format="txt" name="input" type="data" label="Statistics report file (output of 'FASTQ Statistics' tool)" />
</inputs>
<outputs>
<data format="png" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
Creates a boxplot graph for the quality scores in the library.
.. class:: infomark
**TIP:** Use the **FASTQ Statistics** tool to generate the report file needed for this tool.
-----
**Output Examples**
* Black horizontal lines are medians
* Rectangular red boxes show the Inter-quartile Range (IQR) (top value is Q3, bottom value is Q1)
* Whiskers show outlier at max. 1.5*IQR
An excellent quality library (median quality is 40 for almost all 36 cycles):
.. image:: ../static/fastx_icons/fastq_quality_boxplot_1.png
A relatively good quality library (median quality degrades towards later cycles):
.. image:: ../static/fastx_icons/fastq_quality_boxplot_2.png
A low quality library (median drops quickly):
.. image:: ../static/fastx_icons/fastq_quality_boxplot_3.png
</help>
</tool>
<!-- FASTQ-Quality-Boxplot is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
@@ -0,0 +1,73 @@
<tool id="cshl_fastq_quality_filter" name="Quality Filter">
<description></description>
<command>zcat -f '$input' | fastq_quality_filter -q $quality -p $percent -v -o $output</command>
<inputs>
<param format="fastqsolexa" name="input" type="data" label="Library to filter" />
<param name="quality" size="4" type="integer" value="20">
<label>Quality cut-off value</label>
</param>
<param name="percent" size="4" type="integer" value="90">
<label>Percent of bases in sequence that must have quality equal to / higher than cut-off value</label>
</param>
</inputs>
<tests>
<test>
<!-- Test1: 100% of bases with quality 33 or higher (pretty steep requirement...) -->
<param name="input" value="fastq_qual_filter1.fastq" />
<param name="quality" value="33"/>
<param name="percent" value="100"/>
<output name="output" file="fastq_qual_filter1a.out" />
</test>
<test>
<!-- Test2: 80% of bases with quality 20 or higher -->
<param name="input" value="fastq_qual_filter1.fastq" />
<param name="quality" value="20"/>
<param name="percent" value="80"/>
<output name="output" file="fastq_qual_filter1b.out" />
</test>
</tests>
<outputs>
<data format="input" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool filters reads based on quality scores.
.. class:: infomark
Using **percent = 100** requires all cycles of all reads to be at least the quality cut-off value.
.. class:: infomark
Using **percent = 50** requires the median quality of the cycles (in each read) to be at least the quality cut-off value.
--------
Quality score distribution (of all cycles) is calculated for each read. If it is lower than the quality cut-off value - the read is discarded.
**Example**::
@CSHL_4_FC042AGOOII:1:2:214:584
GACAATAAAC
+CSHL_4_FC042AGOOII:1:2:214:584
30 30 30 30 30 30 30 30 20 10
Using **percent = 50** and **cut-off = 30** - This read will not be discarded (the median quality is higher than 30).
Using **percent = 90** and **cut-off = 30** - This read will be discarded (90% of the cycles do no have quality equal to / higher than 30).
Using **percent = 100** and **cut-off = 20** - This read will be discarded (not all cycles have quality equal to / higher than 20).
</help>
</tool>
<!-- FASTQ-Quality-Filter is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
+70
View File
@@ -0,0 +1,70 @@
<tool id="cshl_fastq_to_fasta" name="FASTQ to FASTA">
<description>converter</description>
<command>gunzip -cf $input | fastq_to_fasta $SKIPN $RENAMESEQ -o $output -v </command>
<inputs>
<param format="fastqsolexa" name="input" type="data" label="FASTQ Library to convert" />
<param name="SKIPN" type="select" label="Discard sequences with unknown (N) bases ">
<option value="">yes</option>
<option value="-n">no</option>
</param>
<param name="RENAMESEQ" type="select" label="Rename sequence names in output file (reduces file size)">
<option value="-r">yes</option>
<option value="">no</option>
</param>
</inputs>
<tests>
<test>
<!-- FASTQ-To-FASTA, keep N, don't rename -->
<param name="input" value="fastq_to_fasta1.fastq" />
<param name="SKIPN" value=""/>
<param name="RENAMESEQ" value=""/>
<output name="output" file="fastq_to_fasta1a.out" />
</test>
<test>
<!-- FASTQ-To-FASTA, discard N, rename -->
<param name="input" value="fastq_to_fasta1.fastq" />
<param name="SKIPN" value="no"/>
<param name="RENAMESEQ" value="yes"/>
<output name="output" file="fastq_to_fasta1b.out" />
</test>
</tests>
<outputs>
<data format="fasta" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool converts data from Solexa format to FASTA format (scroll down for format description).
--------
**Example**
The following data in Solexa-FASTQ format::
@CSHL_4_FC042GAMMII_2_1_517_596
GGTCAATGATGAGTTGGCACTGTAGGCACCATCAAT
+CSHL_4_FC042GAMMII_2_1_517_596
40 40 40 40 40 40 40 40 40 40 38 40 40 40 40 40 14 40 40 40 40 40 36 40 13 14 24 24 9 24 9 40 10 10 15 40
Will be converted to FASTA (with 'rename sequence names' = NO)::
>CSHL_4_FC042GAMMII_2_1_517_596
GGTCAATGATGAGTTGGCACTGTAGGCACCATCAAT
Will be converted to FASTA (with 'rename sequence names' = YES)::
>1
GGTCAATGATGAGTTGGCACTGTAGGCACCATCAAT
</help>
</tool>
<!-- FASTQ-to-FASTA is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
@@ -0,0 +1,80 @@
<tool id="cshl_fastx_artifacts_filter" name="Artifacts Filter">
<description></description>
<command>zcat -f '$input' | fastx_artifacts_filter -v -o "$output"</command>
<inputs>
<param format="fasta,fastqsolexa" name="input" type="data" label="Library to filter" />
</inputs>
<tests>
<test>
<!-- Filter FASTA file -->
<param name="input" value="fastx_artifacts1.fasta" />
<output name="output" file="fastx_artifacts1.out" />
</test>
<test>
<!-- Filter FASTQ file -->
<param name="input" value="fastx_artifacts2.fastq" />
<output name="output" file="fastx_artifacts2.out" />
</test>
</tests>
<outputs>
<data format="input" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool filters sequencing artifacts (reads with all but 3 identical bases).
--------
**The following is an example of sequences which will be filtered out**::
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAACACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAACACAAAAAAAAAAAAAAAAAAAAAAAAAAAAACACAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
CCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCC
AAAAACACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAACACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAACACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAA
AAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAA
AAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAA
AAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAA
</help>
</tool>
<!-- FASTX-Artifacts-filter is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
@@ -0,0 +1,68 @@
<tool id="cshl_fastx_barcode_splitter" name="Barcode Splitter">
<description></description>
<command>fastx_barcode_splitter_galaxy_wrapper.sh $BARCODE $input "$input.name" --mismatches $mismatches --partial $partial $EOL > $output </command>
<inputs>
<param format="txt" name="BARCODE" type="data" label="Barcodes to use" />
<param format="fasta,fastqsolexa" name="input" type="data" label="Library to split" />
<param name="EOL" type="select" label="Barcodes found at">
<option value="--bol">Start of sequence (5' end)</option>
<option value="--eol">End of sequence (3' end)</option>
</param>
<param name="mismatches" type="integer" size="3" value="2" label="Number of allowed mismatches" />
<param name="partial" type="integer" size="3" value="0" label="Number of allowed barcodes nucleotide deletions" />
</inputs>
<tests>
<test>
<!-- Split a FASTQ file -->
<param name="BARCODE" value="fastx_barcode_splitter1.txt" />
<param name="input" value="fastx_barcode_splitter1.fastq" />
<param name="EOL" value="Start of sequence (5' end)" />
<param name="mismatches" value="2" />
<param name="partial" value="0" />
<output name="output" file="fastx_barcode_splitter1.out" />
</test>
</tests>
<outputs>
<data format="html" name="output" />
</outputs>
<help>
**What it does**
This tool splits a solexa library (FASTQ file) or a regular FASTA file to several files, using barcodes as the split criteria.
--------
**Barcode file Format**
Barcode files are simple text files.
Each line should contain an identifier (descriptive name for the barcode), and the barcode itself (A/C/G/T), separated by a TAB character.
Example::
#This line is a comment (starts with a 'number' sign)
BC1 GATCT
BC2 ATCGT
BC3 GTGAT
BC4 TGTCT
For each barcode, a new FASTQ file will be created (with the barcode's identifier as part of the file name).
Sequences matching the barcode will be stored in the appropriate file.
One additional FASTQ file will be created (the 'unmatched' file), where sequences not matching any barcode will be stored.
The output of this tool is an HTML file, displaying the split counts and the file locations.
**Output Example**
.. image:: ../static/fastx_icons/barcode_splitter_output_example.png
</help>
</tool>
<!-- FASTX-barcode-splitter is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->
+109
View File
@@ -0,0 +1,109 @@
<tool id="cshl_fastx_clipper" name="Clip" version="1.0.1" >
<description>adapter sequences</description>
<command>
zcat -f $input | fastx_clipper -s $maxmismatches -l $minlength -a $clip_source.clip_sequence -d $keepdelta -o $output -v $KEEP_N $DISCARD_OPTIONS
</command>
<inputs>
<param format="fasta,fastqsolexa" name="input" type="data" label="Library to clip" />
<param name="maxmismatches" size="4" type="integer" value="2">
<label>Maximum number of mismatches allowed (when matching the adapter sequence)</label>
</param>
<param name="minlength" size="4" type="integer" value="15">
<label>Minimum sequence length (after clipping, sequences shorter than this length will be discarded)</label>
</param>
<conditional name="clip_source">
<param name="clip_source_list" type="select" label="Source">
<option value="prebuilt" selected="true">Standard (select from the list below)</option>
<option value="user">Enter custom sequence</option>
</param>
<when value="user">
<param name="clip_sequence" size="30" label="Enter custom clipping sequence" type="text" value="AATTGGCC" />
</when>
<when value="prebuilt">
<param name="clip_sequence" type="select" label="Choose Adapter">
<options from_file="fastx_clipper_sequences.txt">
<column name="name" index="1"/>
<column name="value" index="0"/>
</options>
</param>
</when>
</conditional>
<param name="keepdelta" size="2" type="integer" value="0">
<label>enter non-zero value to keep the adapter sequence and x bases that follow it</label>
<help>use this for hairpin barcoding. keep at 0 unless you know what you're doing.</help>
</param>
<param name="KEEP_N" type="select" label="Discard sequences with unknown (N) bases">
<option value="">Yes</option>
<option value="-n">No</option>
</param>
<param name="DISCARD_OPTIONS" type="select" label="Output options">
<option value="-c">Output only clipped seqeunces (i.e. sequences which contained the adapter)</option>
<option value="-C">Output only non-clipped seqeunces (i.e. sequences which did not contained the adapter)</option>
<option value="">Output both clipped and non-clipped sequences</option>
</param>
</inputs>
<tests>
<test>
<!-- Clip a FASTQ file -->
<param name="input" value="fastx_clipper1.fastq" />
<param name="maxmismatches" value="2" />
<param name="minlength" value="15" />
<param name="clip_source.clip_source_list" value="user" />
<param name="clip_source.clip_sequence" value="CAATTGGTTAATCCCCCTATATA" />
<param name="keepdelta" value="0" />
<param name="KEEP_N" value="-n" />
<param name="DISCARD_OPTIONS" value="-c" />
<output name="output" file="fastx_clipper1a.out" />
</test>
</tests>
<outputs>
<data format="input" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool clips adapters from the 3'-end of the sequences in a FASTA/FASTQ file.
--------
**Clipping Illustration:**
.. image:: ../static/fastx_icons/fastx_clipper_illustration.png
**Clipping Example:**
.. image:: ../static/fastx_icons/fastx_clipper_example.png
**In the above example:**
* Sequence no. 1 was discarded since it wasn't clipped (i.e. didn't contain the adapter sequence). (**Output** parameter).
* Sequence no. 5 was discarded --- it's length (after clipping) was shorter than 15 nt (**Minimum Sequence Length** parameter).
</help>
</tool>
@@ -0,0 +1,52 @@
<tool id="cshl_fastx_reverse_complement" name="Reverse-Complement">
<description>sequences</description>
<command>zcat -f '$input' | fastx_reverse_complement -v -o $output</command>
<inputs>
<param format="fasta,fastqsolexa" name="input" type="data" label="Library to reverse-complement" />
</inputs>
<tests>
<test>
<!-- Reverse-complement a FASTA file -->
<param name="input" value="fastx_rev_comp1.fasta" />
<output name="output" file="fastx_reverse_complement1.out" />
</test>
<test>
<!-- Reverse-complement a FASTQ file -->
<param name="input" value="fastx_rev_comp2.fastq" />
<output name="output" file="fastx_reverse_complement2.out" />
</test>
</tests>
<outputs>
<data format="input" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool reverse-complements each sequence in a library.
If the library is a FASTQ, the quality-scores are also reversed.
--------
**Example**
Input FASTQ file::
@CSHL_1_FC42AGWWWXX:8:1:3:740
TGTCTGTAGCCTCNTCCTTGTAATTCAAAGNNGGTA
+CSHL_1_FC42AGWWWXX:8:1:3:740
33 33 33 34 33 33 33 33 33 33 33 33 27 5 27 33 33 33 33 33 33 27 21 27 33 32 31 29 26 24 5 5 15 17 27 26
Output FASTQ file::
@CSHL_1_FC42AGWWWXX:8:1:3:740
TACCNNCTTTGAATTACAAGGANGAGGCTACAGACA
+CSHL_1_FC42AGWWWXX:8:1:3:740
26 27 17 15 5 5 24 26 29 31 32 33 27 21 27 33 33 33 33 33 33 27 5 27 33 33 33 33 33 33 33 33 34 33 33 33
</help>
</tool>
+70
View File
@@ -0,0 +1,70 @@
<tool id="cshl_fastx_trimmer" name="Trim">
<description>sequences</description>
<command>zcat -f '$input' | fastx_trimmer -v -f $first -l $last -o $output</command>
<inputs>
<param format="fasta,fastqsolexa" name="input" type="data" label="Library to clip" />
<param name="first" size="4" type="integer" value="1">
<label>First base to keep</label>
</param>
<param name="last" size="4" type="integer" value="21">
<label>Last base to keep</label>
</param>
</inputs>
<tests>
<test>
<!-- Trim a FASTA file - remove first four bases (e.g. a barcode) -->
<param name="input" value="fastx_trimmer1.fasta" />
<param name="first" value="5"/>
<param name="last" value="36"/>
<output name="output" file="fastx_trimmer1.out" />
</test>
<test>
<!-- Trim a FASTQ file - remove last 9 bases (e.g. keep only miRNA length sequences) -->
<param name="input" value="fastx_trimmer2.fastq" />
<param name="first" value="1"/>
<param name="last" value="27"/>
<output name="output" file="fastx_trimmer2.out" />
</test>
</tests>
<outputs>
<data format="input" name="output" metadata_source="input" />
</outputs>
<help>
**What it does**
This tool trims (cut bases from) sequences in a FASTA/Q file.
--------
**Example**
Input Fasta file (with 36 bases in each sequences)::
>1-1
TATGGTCAGAAACCATATGCAGAGCCTGTAGGCACC
>2-1
CAGCGAGGCTTTAATGCCATTTGGCTGTAGGCACCA
Trimming with First=1 and Last=21, we get a FASTA file with 21 bases in each sequences (starting from the first base)::
>1-1
TATGGTCAGAAACCATATGCA
>2-1
CAGCGAGGCTTTAATGCCATT
Trimming with First=6 and Last=10, will generate a FASTA file with 5 bases (bases 6,7,8,9,10) in each sequences::
>1-1
TCAGA
>2-1
AGGCT
</help>
</tool>
<!-- FASTX-Trimmer is part of the FASTX-toolkit, by A.Gordon (gordon@cshl.edu) -->