mirror of
https://github.com/galaxyproject/galaxy.git
synced 2026-09-24 16:30:27 +08:00
Migrate cuffcompare, cuffdiff, cufflinks, and cuffmerge from the distribution to the tool shed. Update the 'Review migration stages' layout.
This commit is contained in:
@@ -6,45 +6,49 @@ The following tools have been eliminated from the distribution:
|
||||
3: Compute Motif Frequencies For All Motifs motif by motif
|
||||
4: Compute Motif Frequencies in indel flanking regions
|
||||
5: CTD analysis of chemicals, diseases, or genes
|
||||
6: Delete Overlapping Indels from a chromosome indels file
|
||||
7: Separate pgSnp alleles into columns
|
||||
8: Draw Stacked Bar Plots for different categories and different criteria
|
||||
9: Length Distribution chart
|
||||
10: FASTA Width formatter
|
||||
11: RNA/DNA converter
|
||||
12: Draw quality score boxplot
|
||||
13: Quality format converter (ASCII-Numeric)
|
||||
14: Filter by quality
|
||||
15: FASTQ to FASTA converter
|
||||
16: Remove sequencing artifacts
|
||||
17: Barcode Splitter
|
||||
18: Clip adapter sequences
|
||||
19: Collapse sequences
|
||||
20: Draw nucleotides distribution chart
|
||||
21: Compute quality statistics
|
||||
22: Rename sequences
|
||||
23: Reverse- Complement
|
||||
24: Trim sequences
|
||||
25: FunDO human genes associated with disease terms
|
||||
26: HVIS visualization of genomic data with the Hilbert curve
|
||||
27: Fetch Indels from 3-way alignments
|
||||
28: Identify microsatellite births and deaths
|
||||
29: Extract orthologous microsatellites for multiple (>2) species alignments
|
||||
30: Mutate Codons with SNPs
|
||||
31: Pileup-to-Interval condenses pileup format into ranges of bases
|
||||
32: Filter pileup on coverage and SNPs
|
||||
33: Filter SAM on bitwise flag values
|
||||
34: Merge BAM Files merges BAM files together
|
||||
35: Generate pileup from BAM dataset
|
||||
36: SAM-to-BAM converts SAM format to BAM format
|
||||
37: Convert SAM to interval
|
||||
38: flagstat provides simple stats on BAM files
|
||||
39: MPileup SNP and indel caller
|
||||
40: rmdup remove PCR duplicates
|
||||
41: Slice BAM by provided regions
|
||||
42: Split paired end reads
|
||||
43: T Test for Two Samples
|
||||
44: Plotting tool for multiple series and graph types.
|
||||
6: Cuffcompare
|
||||
7: Cuffdiff
|
||||
8: Cufflinks
|
||||
9: Cuffmerge
|
||||
10: Delete Overlapping Indels from a chromosome indels file
|
||||
11: Separate pgSnp alleles into columns
|
||||
12: Draw Stacked Bar Plots for different categories and different criteria
|
||||
13: Length Distribution chart
|
||||
14: FASTA Width formatter
|
||||
15: RNA/DNA converter
|
||||
16: Draw quality score boxplot
|
||||
17: Quality format converter (ASCII-Numeric)
|
||||
18: Filter by quality
|
||||
19: FASTQ to FASTA converter
|
||||
20: Remove sequencing artifacts
|
||||
21: Barcode Splitter
|
||||
22: Clip adapter sequences
|
||||
23: Collapse sequences
|
||||
24: Draw nucleotides distribution chart
|
||||
25: Compute quality statistics
|
||||
26: Rename sequences
|
||||
27: Reverse- Complement
|
||||
28: Trim sequences
|
||||
29: FunDO human genes associated with disease terms
|
||||
30: HVIS visualization of genomic data with the Hilbert curve
|
||||
31: Fetch Indels from 3-way alignments
|
||||
32: Identify microsatellite births and deaths
|
||||
33: Extract orthologous microsatellites for multiple (>2) species alignments
|
||||
34: Mutate Codons with SNPs
|
||||
35: Pileup-to-Interval condenses pileup format into ranges of bases
|
||||
36: Filter pileup on coverage and SNPs
|
||||
37: Filter SAM on bitwise flag values
|
||||
38: Merge BAM Files merges BAM files together
|
||||
39: Generate pileup from BAM dataset
|
||||
40: SAM-to-BAM converts SAM format to BAM format
|
||||
41: Convert SAM to interval
|
||||
42: flagstat provides simple stats on BAM files
|
||||
43: MPileup SNP and indel caller
|
||||
44: rmdup remove PCR duplicates
|
||||
45: Slice BAM by provided regions
|
||||
46: Split paired end reads
|
||||
47: T Test for Two Samples
|
||||
48: Plotting tool for multiple series and graph types.
|
||||
|
||||
The tools are now available in the repositories respectively:
|
||||
|
||||
@@ -53,45 +57,49 @@ The tools are now available in the repositories respectively:
|
||||
3: compute_motif_frequencies_for_all_motifs
|
||||
4: compute_motifs_frequency
|
||||
5: ctd_batch
|
||||
6: delete_overlapping_indels
|
||||
7: divide_pg_snp
|
||||
8: draw_stacked_barplots
|
||||
9: fasta_clipping_histogram
|
||||
10: fasta_formatter
|
||||
11: fasta_nucleotide_changer
|
||||
12: fastq_quality_boxplot
|
||||
13: fastq_quality_converter
|
||||
14: fastq_quality_filter
|
||||
15: fastq_to_fasta
|
||||
16: fastx_artifacts_filter
|
||||
17: fastx_barcode_splitter
|
||||
18: fastx_clipper
|
||||
19: fastx_collapser
|
||||
20: fastx_nucleotides_distribution
|
||||
21: fastx_quality_statistics
|
||||
22: fastx_renamer
|
||||
23: fastx_reverse_complement
|
||||
24: fastx_trimmer
|
||||
25: hgv_fundo
|
||||
26: hgv_hilbertvis
|
||||
27: indels_3way
|
||||
28: microsatellite_birthdeath
|
||||
29: multispecies_orthologous_microsats
|
||||
30: mutate_snp_codon
|
||||
31: pileup_interval
|
||||
32: pileup_parser
|
||||
33: sam_bitwise_flag_filter
|
||||
34: sam_merge
|
||||
35: sam_pileup
|
||||
36: sam_to_bam
|
||||
37: sam2interval
|
||||
38: samtools_flagstat
|
||||
39: samtools_mpileup
|
||||
40: samtools_rmdup
|
||||
41: samtools_slice_bam
|
||||
42: split_paired_reads
|
||||
43: t_test_two_samples
|
||||
44: xy_plot
|
||||
6: cuffcompare
|
||||
7: cuffdiff
|
||||
8: cufflinks
|
||||
9: cuffmerge
|
||||
10: delete_overlapping_indels
|
||||
11: divide_pg_snp
|
||||
12: draw_stacked_barplots
|
||||
13: fasta_clipping_histogram
|
||||
14: fasta_formatter
|
||||
15: fasta_nucleotide_changer
|
||||
16: fastq_quality_boxplot
|
||||
17: fastq_quality_converter
|
||||
18: fastq_quality_filter
|
||||
19: fastq_to_fasta
|
||||
20: fastx_artifacts_filter
|
||||
21: fastx_barcode_splitter
|
||||
22: fastx_clipper
|
||||
23: fastx_collapser
|
||||
24: fastx_nucleotides_distribution
|
||||
25: fastx_quality_statistics
|
||||
26: fastx_renamer
|
||||
27: fastx_reverse_complement
|
||||
28: fastx_trimmer
|
||||
29: hgv_fundo
|
||||
30: hgv_hilbertvis
|
||||
31: indels_3way
|
||||
32: microsatellite_birthdeath
|
||||
33: multispecies_orthologous_microsats
|
||||
34: mutate_snp_codon
|
||||
35: pileup_interval
|
||||
36: pileup_parser
|
||||
37: sam_bitwise_flag_filter
|
||||
38: sam_merge
|
||||
39: sam_pileup
|
||||
40: sam_to_bam
|
||||
41: sam2interval
|
||||
42: samtools_flagstat
|
||||
43: samtools_mpileup
|
||||
44: samtools_rmdup
|
||||
45: samtools_slice_bam
|
||||
46: split_paired_reads
|
||||
47: t_test_two_samples
|
||||
48: xy_plot
|
||||
|
||||
from the main Galaxy tool shed at http://toolshed.g2.bx.psu.edu
|
||||
and will be installed into your local Galaxy instance at the
|
||||
|
||||
@@ -15,6 +15,18 @@
|
||||
<repository owner="devteam" changeset_revision="4a32700dcaa2" name="ctd_batch" description="Galaxy wrappers for the tool CTD: analysis of chemicals, diseases, or genes">
|
||||
<tool id="ctdBatch_1" version="1.0.0" file="ctd.xml" />
|
||||
</repository>
|
||||
<repository owner="devteam" changeset_revision="9d35cf35634e" name="cuffcompare" description="Galaxy wrappers for the Cuffcompare tool.">
|
||||
<tool id="cuffcompare" version="0.0.5" file="cuffcompare_wrapper.xml" />
|
||||
</repository>
|
||||
<repository owner="devteam" changeset_revision="0dabb2ed6eb1" name="cuffdiff" description="Galaxy wrappers for the Cuffdiff tool.">
|
||||
<tool id="cuffdiff" version="0.0.6" file="cuffdiff_wrapper.xml" />
|
||||
</repository>
|
||||
<repository owner="devteam" changeset_revision="b50aacc8ae49" name="cufflinks" description="Galaxy wrappers for the Cufflinks tool.">
|
||||
<tool id="cufflinks" version="0.0.6" file="cufflinks_wrapper.xml" />
|
||||
</repository>
|
||||
<repository owner="devteam" changeset_revision="dbbd37e013aa" name="cuffmerge" description="Galaxy wrappers for the Cuffmerge tool.">
|
||||
<tool id="cuffmerge" version="0.0.5" file="cuffmerge_wrapper.xml" />
|
||||
</repository>
|
||||
<repository owner="devteam" changeset_revision="f16000dc644b" name="delete_overlapping_indels" description="Galaxy wrappers for the tool Delete Overlapping Indels: from a chromosome indels file">
|
||||
<tool id="delete_overlapping_indels" version="1.0.0" file="delete_overlapping_indels.xml" />
|
||||
</repository>
|
||||
|
||||
@@ -35,6 +35,7 @@
|
||||
</p>
|
||||
</div>
|
||||
<table class="grid">
|
||||
<% from tool_shed.util.shed_util_common import to_html_string %>
|
||||
%for stage in migration_stages_dict.keys():
|
||||
<%
|
||||
migration_command = 'sh ./scripts/migrate_tools/%04d_tools.sh' % stage
|
||||
@@ -54,7 +55,7 @@
|
||||
<tr>
|
||||
<td bgcolor="#FFFFCC">
|
||||
<div class="form-row">
|
||||
<p>${migration_info} <b>Run commands from the Galaxy installation directory!</b></p>
|
||||
<p>${to_html_string(migration_info)} <b>Run commands from the Galaxy installation directory!</b></p>
|
||||
<p>
|
||||
%if tool_dependencies:
|
||||
This migration stage includes tools that have tool dependencies that can be automatically installed. To install them, run:<br/>
|
||||
|
||||
+119
-71
@@ -1,12 +1,12 @@
|
||||
<?xml version='1.0' encoding='utf-8'?>
|
||||
<?xml version="1.0"?>
|
||||
<toolbox>
|
||||
<section id="getext" name="Get Data">
|
||||
<tool file="data_source/upload.xml" />
|
||||
<section name="Get Data" id="getext">
|
||||
<tool file="data_source/upload.xml"/>
|
||||
<tool file="data_source/ucsc_tablebrowser.xml" />
|
||||
<tool file="data_source/ucsc_tablebrowser_test.xml" />
|
||||
<tool file="data_source/ucsc_tablebrowser_archaea.xml" />
|
||||
<tool file="data_source/bx_browser.xml" />
|
||||
<tool file="data_source/ebi_sra.xml" />
|
||||
<tool file="data_source/ebi_sra.xml"/>
|
||||
<tool file="data_source/microbial_import.xml" />
|
||||
<tool file="data_source/biomart.xml" />
|
||||
<tool file="data_source/biomart_test.xml" />
|
||||
@@ -32,19 +32,19 @@
|
||||
<tool file="genomespace/genomespace_importer.xml" />
|
||||
<tool file="validation/fix_errors.xml" />
|
||||
</section>
|
||||
<section id="send" name="Send Data">
|
||||
<section name="Send Data" id="send">
|
||||
<tool file="data_destination/epigraph.xml" />
|
||||
<tool file="data_destination/epigraph_test.xml" />
|
||||
<tool file="genomespace/genomespace_exporter.xml" />
|
||||
</section>
|
||||
<section id="EncodeTools" name="ENCODE Tools">
|
||||
<section name="ENCODE Tools" id="EncodeTools">
|
||||
<tool file="encode/gencode_partition.xml" />
|
||||
<tool file="encode/random_intervals.xml" />
|
||||
</section>
|
||||
<section id="liftOver" name="Lift-Over">
|
||||
<section name="Lift-Over" id="liftOver">
|
||||
<tool file="extract/liftOver_wrapper.xml" />
|
||||
</section>
|
||||
<section id="textutil" name="Text Manipulation">
|
||||
<section name="Text Manipulation" id="textutil">
|
||||
<tool file="filters/fixedValueColumn.xml" />
|
||||
<tool file="stats/column_maker.xml" />
|
||||
<tool file="filters/catWrapper.xml" />
|
||||
@@ -65,25 +65,25 @@
|
||||
<tool file="stats/dna_filtering.xml" />
|
||||
<tool file="new_operations/tables_arithmetic_operations.xml" />
|
||||
</section>
|
||||
<section id="filter" name="Filter and Sort">
|
||||
<section name="Filter and Sort" id="filter">
|
||||
<tool file="stats/filtering.xml" />
|
||||
<tool file="filters/sorter.xml" />
|
||||
<tool file="filters/grep.xml" />
|
||||
|
||||
<label id="gff" text="GFF" />
|
||||
<label text="GFF" id="gff" />
|
||||
<tool file="filters/gff/extract_GFF_Features.xml" />
|
||||
<tool file="filters/gff/gff_filter_by_attribute.xml" />
|
||||
<tool file="filters/gff/gff_filter_by_feature_count.xml" />
|
||||
<tool file="filters/gff/gtf_filter_by_attribute_values_list.xml" />
|
||||
</section>
|
||||
<section id="group" name="Join, Subtract and Group">
|
||||
<section name="Join, Subtract and Group" id="group">
|
||||
<tool file="filters/joiner.xml" />
|
||||
<tool file="filters/compare.xml" />
|
||||
<tool file="new_operations/subtract_query.xml" />
|
||||
<tool file="filters/compare.xml"/>
|
||||
<tool file="new_operations/subtract_query.xml"/>
|
||||
<tool file="stats/grouping.xml" />
|
||||
<tool file="new_operations/column_join.xml" />
|
||||
</section>
|
||||
<section id="convert" name="Convert Formats">
|
||||
<section name="Convert Formats" id="convert">
|
||||
<tool file="filters/axt_to_concat_fasta.xml" />
|
||||
<tool file="filters/axt_to_fasta.xml" />
|
||||
<tool file="filters/axt_to_lav.xml" />
|
||||
@@ -95,38 +95,39 @@
|
||||
<tool file="maf/maf_to_interval.xml" />
|
||||
<tool file="maf/maf_to_fasta.xml" />
|
||||
<tool file="fasta_tools/tabular_to_fasta.xml" />
|
||||
<tool file="fastq/fastq_to_fasta.xml" />
|
||||
<tool file="filters/wiggle_to_simple.xml" />
|
||||
<tool file="filters/sff_extractor.xml" />
|
||||
<tool file="filters/gtf2bedgraph.xml" />
|
||||
<tool file="filters/wig_to_bigwig.xml" />
|
||||
<tool file="filters/bed_to_bigbed.xml" />
|
||||
</section>
|
||||
<section id="features" name="Extract Features">
|
||||
<section name="Extract Features" id="features">
|
||||
<tool file="filters/ucsc_gene_bed_to_exon_bed.xml" />
|
||||
</section>
|
||||
<section id="fetchSeq" name="Fetch Sequences">
|
||||
<section name="Fetch Sequences" id="fetchSeq">
|
||||
<tool file="extract/extract_genomic_dna.xml" />
|
||||
</section>
|
||||
<section id="fetchAlign" name="Fetch Alignments">
|
||||
<section name="Fetch Alignments" id="fetchAlign">
|
||||
<tool file="maf/interval2maf_pairwise.xml" />
|
||||
<tool file="maf/interval2maf.xml" />
|
||||
<tool file="maf/maf_split_by_species.xml" />
|
||||
<tool file="maf/maf_split_by_species.xml"/>
|
||||
<tool file="maf/interval_maf_to_merged_fasta.xml" />
|
||||
<tool file="maf/genebed_maf_to_fasta.xml" />
|
||||
<tool file="maf/maf_stats.xml" />
|
||||
<tool file="maf/maf_thread_for_species.xml" />
|
||||
<tool file="maf/maf_limit_to_species.xml" />
|
||||
<tool file="maf/maf_limit_size.xml" />
|
||||
<tool file="maf/maf_by_block_number.xml" />
|
||||
<tool file="maf/maf_reverse_complement.xml" />
|
||||
<tool file="maf/maf_filter.xml" />
|
||||
<tool file="maf/genebed_maf_to_fasta.xml"/>
|
||||
<tool file="maf/maf_stats.xml"/>
|
||||
<tool file="maf/maf_thread_for_species.xml"/>
|
||||
<tool file="maf/maf_limit_to_species.xml"/>
|
||||
<tool file="maf/maf_limit_size.xml"/>
|
||||
<tool file="maf/maf_by_block_number.xml"/>
|
||||
<tool file="maf/maf_reverse_complement.xml"/>
|
||||
<tool file="maf/maf_filter.xml"/>
|
||||
</section>
|
||||
<section id="scores" name="Get Genomic Scores">
|
||||
<section name="Get Genomic Scores" id="scores">
|
||||
<tool file="stats/wiggle_to_simple.xml" />
|
||||
<tool file="stats/aggregate_binned_scores_in_intervals.xml" />
|
||||
<tool file="extract/phastOdds/phastOdds_tool.xml" />
|
||||
</section>
|
||||
<section id="bxops" name="Operate on Genomic Intervals">
|
||||
<section name="Operate on Genomic Intervals" id="bxops">
|
||||
<tool file="new_operations/intersect.xml" />
|
||||
<tool file="new_operations/subtract.xml" />
|
||||
<tool file="new_operations/merge.xml" />
|
||||
@@ -140,19 +141,21 @@
|
||||
<tool file="new_operations/flanking_features.xml" />
|
||||
<tool file="annotation_profiler/annotation_profiler.xml" />
|
||||
</section>
|
||||
<section id="stats" name="Statistics">
|
||||
<section name="Statistics" id="stats">
|
||||
<tool file="stats/gsummary.xml" />
|
||||
<tool file="filters/uniq.xml" />
|
||||
<tool file="stats/cor.xml" />
|
||||
<tool file="stats/generate_matrix_for_pca_lda.xml" />
|
||||
<tool file="stats/lda_analy.xml" />
|
||||
<tool file="stats/plot_from_lda.xml" />
|
||||
<tool file="regVariation/t_test_two_samples.xml" />
|
||||
<tool file="regVariation/compute_q_values.xml" />
|
||||
<tool file="stats/MINE.xml" />
|
||||
|
||||
<label id="gff" text="GFF" />
|
||||
<label text="GFF" id="gff" />
|
||||
<tool file="stats/count_gff_features.xml" />
|
||||
</section>
|
||||
<section id="dwt" name="Wavelet Analysis">
|
||||
<section name="Wavelet Analysis" id="dwt">
|
||||
<tool file="discreteWavelet/execute_dwt_var_perFeature.xml" />
|
||||
<!--
|
||||
Keep this section/tools commented until all of the tools have functional tests
|
||||
@@ -162,10 +165,11 @@
|
||||
<tool file="discreteWavelet/execute_dwt_var_perClass.xml" />
|
||||
-->
|
||||
</section>
|
||||
<section id="plots" name="Graph/Display Data">
|
||||
<section name="Graph/Display Data" id="plots">
|
||||
<tool file="plotting/histogram2.xml" />
|
||||
<tool file="plotting/scatterplot.xml" />
|
||||
<tool file="plotting/bar_chart.xml" />
|
||||
<tool file="plotting/xy_plot.xml" />
|
||||
<tool file="plotting/boxplot.xml" />
|
||||
<tool file="visualization/GMAJ.xml" />
|
||||
<tool file="visualization/LAJ.xml" />
|
||||
@@ -173,48 +177,57 @@
|
||||
<tool file="maf/vcf_to_maf_customtrack.xml" />
|
||||
<tool file="mutation/visualize.xml" />
|
||||
</section>
|
||||
<section id="regVar" name="Regional Variation">
|
||||
<section name="Regional Variation" id="regVar">
|
||||
<tool file="regVariation/windowSplitter.xml" />
|
||||
<tool file="regVariation/featureCounter.xml" />
|
||||
<tool file="regVariation/WeightedAverage.xml" />
|
||||
<tool file="regVariation/quality_filter.xml" />
|
||||
<tool file="regVariation/maf_cpg_filter.xml" />
|
||||
<tool file="regVariation/getIndels_2way.xml" />
|
||||
<tool file="regVariation/getIndels_3way.xml" />
|
||||
<tool file="regVariation/getIndelRates_3way.xml" />
|
||||
<tool file="regVariation/substitutions.xml" />
|
||||
<tool file="regVariation/substitution_rates.xml" />
|
||||
<tool file="regVariation/microsats_alignment_level.xml" />
|
||||
<tool file="regVariation/microsats_mutability.xml" />
|
||||
<tool file="regVariation/delete_overlapping_indels.xml" />
|
||||
<tool file="regVariation/compute_motifs_frequency.xml" />
|
||||
<tool file="regVariation/compute_motif_frequencies_for_all_motifs.xml" />
|
||||
<tool file="regVariation/categorize_elements_satisfying_criteria.xml" />s
|
||||
<tool file="regVariation/draw_stacked_barplots.xml" />
|
||||
<tool file="regVariation/multispecies_MicrosatDataGenerator_interrupted_GALAXY.xml" />
|
||||
<tool file="regVariation/microsatellite_birthdeath.xml" />
|
||||
</section>
|
||||
<section id="multReg" name="Multiple regression">
|
||||
<section name="Multiple regression" id="multReg">
|
||||
<tool file="regVariation/linear_regression.xml" />
|
||||
<tool file="regVariation/logistic_regression_vif.xml" />
|
||||
<tool file="regVariation/best_regression_subsets.xml" />
|
||||
<tool file="regVariation/rcve.xml" />
|
||||
<tool file="regVariation/partialR_square.xml" />
|
||||
</section>
|
||||
<section id="multVar" name="Multivariate Analysis">
|
||||
<section name="Multivariate Analysis" id="multVar">
|
||||
<tool file="multivariate_stats/pca.xml" />
|
||||
<tool file="multivariate_stats/cca.xml" />
|
||||
<tool file="multivariate_stats/kpca.xml" />
|
||||
<tool file="multivariate_stats/kcca.xml" />
|
||||
</section>
|
||||
<section id="hyphy" name="Evolution">
|
||||
<section name="Evolution" id="hyphy">
|
||||
<tool file="hyphy/hyphy_branch_lengths_wrapper.xml" />
|
||||
<tool file="hyphy/hyphy_nj_tree_wrapper.xml" />
|
||||
<tool file="hyphy/hyphy_dnds_wrapper.xml" />
|
||||
<tool file="evolution/mutate_snp_codon.xml" />
|
||||
<tool file="evolution/codingSnps.xml" />
|
||||
<tool file="evolution/add_scores.xml" />
|
||||
</section>
|
||||
<section id="motifs" name="Motif Tools">
|
||||
<tool file="meme/meme.xml" />
|
||||
<tool file="meme/fimo.xml" />
|
||||
<section name="Motif Tools" id="motifs">
|
||||
<tool file="meme/meme.xml"/>
|
||||
<tool file="meme/fimo.xml"/>
|
||||
<tool file="rgenetics/rgWebLogo3.xml" />
|
||||
</section>
|
||||
<section id="clustal" name="Multiple Alignments">
|
||||
<section name="Multiple Alignments" id="clustal">
|
||||
<tool file="rgenetics/rgClustalw.xml" />
|
||||
</section>
|
||||
<section id="tax_manipulation" name="Metagenomic analyses">
|
||||
<section name="Metagenomic analyses" id="tax_manipulation">
|
||||
<tool file="taxonomy/gi2taxonomy.xml" />
|
||||
<tool file="taxonomy/t2t_report.xml" />
|
||||
<tool file="taxonomy/t2ps_wrapper.xml" />
|
||||
@@ -222,35 +235,38 @@
|
||||
<tool file="taxonomy/lca.xml" />
|
||||
<tool file="taxonomy/poisson2test.xml" />
|
||||
</section>
|
||||
<section id="fasta_manipulation" name="FASTA manipulation">
|
||||
<section name="FASTA manipulation" id="fasta_manipulation">
|
||||
<tool file="fasta_tools/fasta_compute_length.xml" />
|
||||
<tool file="fasta_tools/fasta_filter_by_length.xml" />
|
||||
<tool file="fasta_tools/fasta_concatenate_by_species.xml" />
|
||||
<tool file="fasta_tools/fasta_to_tabular.xml" />
|
||||
<tool file="fasta_tools/tabular_to_fasta.xml" />
|
||||
<tool file="fastx_toolkit/fasta_formatter.xml" />
|
||||
<tool file="fastx_toolkit/fasta_nucleotide_changer.xml" />
|
||||
<tool file="fastx_toolkit/fastx_collapser.xml" />
|
||||
</section>
|
||||
<section id="NGS_QC" name="NGS: QC and manipulation">
|
||||
<section name="NGS: QC and manipulation" id="NGS_QC">
|
||||
|
||||
<label id="fastqcsambam" text="FastQC: fastq/sam/bam" />
|
||||
<label text="FastQC: fastq/sam/bam" id="fastqcsambam" />
|
||||
<tool file="rgenetics/rgFastQC.xml" />
|
||||
|
||||
<label id="illumina" text="Illumina fastq" />
|
||||
<label text="Illumina fastq" id="illumina" />
|
||||
<tool file="fastq/fastq_groomer.xml" />
|
||||
<tool file="fastq/fastq_paired_end_splitter.xml" />
|
||||
<tool file="fastq/fastq_paired_end_joiner.xml" />
|
||||
<tool file="fastq/fastq_stats.xml" />
|
||||
|
||||
<label id="454" text="Roche-454 data" />
|
||||
<label text="Roche-454 data" id="454" />
|
||||
<tool file="metag_tools/short_reads_figure_score.xml" />
|
||||
<tool file="metag_tools/short_reads_trim_seq.xml" />
|
||||
<tool file="fastq/fastq_combiner.xml" />
|
||||
|
||||
<label id="solid" text="AB-SOLiD data" />
|
||||
<label text="AB-SOLiD data" id="solid" />
|
||||
<tool file="next_gen_conversion/solid2fastq.xml" />
|
||||
<tool file="solid_tools/solid_qual_stats.xml" />
|
||||
<tool file="solid_tools/solid_qual_boxplot.xml" />
|
||||
|
||||
<label id="generic_fastq" text="Generic FASTQ manipulation" />
|
||||
<label text="Generic FASTQ manipulation" id="generic_fastq" />
|
||||
<tool file="fastq/fastq_filter.xml" />
|
||||
<tool file="fastq/fastq_trimmer.xml" />
|
||||
<tool file="fastq/fastq_trimmer_by_quality.xml" />
|
||||
@@ -258,10 +274,25 @@
|
||||
<tool file="fastq/fastq_paired_end_interlacer.xml" />
|
||||
<tool file="fastq/fastq_paired_end_deinterlacer.xml" />
|
||||
<tool file="fastq/fastq_manipulation.xml" />
|
||||
<tool file="fastq/fastq_to_fasta.xml" />
|
||||
<tool file="fastq/fastq_to_tabular.xml" />
|
||||
<tool file="fastq/tabular_to_fastq.xml" />
|
||||
|
||||
<label id="fastx_toolkit" text="FASTX-Toolkit for FASTQ data" />
|
||||
<label text="FASTX-Toolkit for FASTQ data" id="fastx_toolkit" />
|
||||
<tool file="fastx_toolkit/fastq_quality_converter.xml" />
|
||||
<tool file="fastx_toolkit/fastx_quality_statistics.xml" />
|
||||
<tool file="fastx_toolkit/fastq_quality_boxplot.xml" />
|
||||
<tool file="fastx_toolkit/fastx_nucleotides_distribution.xml" />
|
||||
<tool file="fastx_toolkit/fastq_to_fasta.xml" />
|
||||
<tool file="fastx_toolkit/fastq_quality_filter.xml" />
|
||||
<tool file="fastx_toolkit/fastq_to_fasta.xml" />
|
||||
<tool file="fastx_toolkit/fastx_artifacts_filter.xml" />
|
||||
<tool file="fastx_toolkit/fastx_barcode_splitter.xml" />
|
||||
<tool file="fastx_toolkit/fastx_clipper.xml" />
|
||||
<tool file="fastx_toolkit/fastx_collapser.xml" />
|
||||
<tool file="fastx_toolkit/fastx_renamer.xml" />
|
||||
<tool file="fastx_toolkit/fastx_reverse_complement.xml" />
|
||||
<tool file="fastx_toolkit/fastx_trimmer.xml" />
|
||||
</section>
|
||||
<!--
|
||||
Keep this section commented until it includes tools that
|
||||
@@ -274,23 +305,24 @@
|
||||
<tool file="sr_assembly/velveth.xml" />
|
||||
</section>
|
||||
-->
|
||||
<section id="solexa_tools" name="NGS: Mapping">
|
||||
<section name="NGS: Mapping" id="solexa_tools">
|
||||
<tool file="sr_mapping/bowtie2_wrapper.xml" />
|
||||
<tool file="sr_mapping/bfast_wrapper.xml" />
|
||||
<tool file="metag_tools/megablast_wrapper.xml" />
|
||||
<tool file="metag_tools/megablast_xml_parser.xml" />
|
||||
<tool file="sr_mapping/PerM.xml" />
|
||||
<tool file="sr_mapping/srma_wrapper.xml" />
|
||||
<tool file="sr_mapping/mosaik.xml" />
|
||||
<tool file="sr_mapping/mosaik.xml"/>
|
||||
</section>
|
||||
<section id="indel_analysis" name="NGS: Indel Analysis">
|
||||
<section name="NGS: Indel Analysis" id="indel_analysis">
|
||||
<tool file="indels/sam_indel_filter.xml" />
|
||||
<tool file="indels/indel_sam2interval.xml" />
|
||||
<tool file="indels/indel_table.xml" />
|
||||
<tool file="indels/indel_analysis.xml" />
|
||||
</section>
|
||||
<section id="ngs-rna-tools" name="NGS: RNA Analysis">
|
||||
<section name="NGS: RNA Analysis" id="ngs-rna-tools">
|
||||
|
||||
<label id="rna_seq" text="RNA-seq" />
|
||||
<label text="RNA-seq" id="rna_seq" />
|
||||
<tool file="ngs_rna/tophat_wrapper.xml" />
|
||||
<tool file="ngs_rna/tophat2_wrapper.xml" />
|
||||
<tool file="ngs_rna/tophat_color_wrapper.xml" />
|
||||
@@ -305,74 +337,90 @@
|
||||
<tool file="ngs_rna/trinity_all.xml" />
|
||||
-->
|
||||
|
||||
<label id="filtering" text="Filtering" />
|
||||
<label text="Filtering" id="filtering" />
|
||||
<tool file="ngs_rna/filter_transcripts_via_tracking.xml" />
|
||||
</section>
|
||||
<section id="samtools" name="NGS: SAM Tools">
|
||||
<section name="NGS: SAM Tools" id="samtools">
|
||||
<tool file="samtools/sam_bitwise_flag_filter.xml" />
|
||||
<tool file="samtools/sam2interval.xml" />
|
||||
<tool file="samtools/sam_to_bam.xml" />
|
||||
<tool file="samtools/bam_to_sam.xml" />
|
||||
<tool file="samtools/sam_merge.xml" />
|
||||
<tool file="samtools/samtools_mpileup.xml" />
|
||||
<tool file="samtools/sam_pileup.xml" />
|
||||
<tool file="samtools/pileup_parser.xml" />
|
||||
<tool file="samtools/pileup_interval.xml" />
|
||||
<tool file="samtools/samtools_flagstat.xml" />
|
||||
<tool file="samtools/samtools_rmdup.xml" />
|
||||
<tool file="samtools/samtools_slice_bam.xml" />
|
||||
</section>
|
||||
<section id="gatk" name="NGS: GATK Tools (beta)">
|
||||
<label id="gatk_bam_utilities" text="Alignment Utilities" />
|
||||
<section name="NGS: GATK Tools (beta)" id="gatk">
|
||||
<label text="Alignment Utilities" id="gatk_bam_utilities"/>
|
||||
<tool file="gatk/depth_of_coverage.xml" />
|
||||
<tool file="gatk/print_reads.xml" />
|
||||
|
||||
<label id="gatk_realignment" text="Realignment" />
|
||||
<label text="Realignment" id="gatk_realignment" />
|
||||
<tool file="gatk/realigner_target_creator.xml" />
|
||||
<tool file="gatk/indel_realigner.xml" />
|
||||
|
||||
<label id="gatk_recalibration" text="Base Recalibration" />
|
||||
<label text="Base Recalibration" id="gatk_recalibration" />
|
||||
<tool file="gatk/count_covariates.xml" />
|
||||
<tool file="gatk/table_recalibration.xml" />
|
||||
<tool file="gatk/analyze_covariates.xml" />
|
||||
|
||||
<label id="gatk_genotyping" text="Genotyping" />
|
||||
<label text="Genotyping" id="gatk_genotyping" />
|
||||
<tool file="gatk/unified_genotyper.xml" />
|
||||
|
||||
<label id="gatk_annotation" text="Annotation" />
|
||||
<label text="Annotation" id="gatk_annotation" />
|
||||
<tool file="gatk/variant_annotator.xml" />
|
||||
|
||||
<label id="gatk_filtration" text="Filtration" />
|
||||
<label text="Filtration" id="gatk_filtration" />
|
||||
<tool file="gatk/variant_filtration.xml" />
|
||||
<tool file="gatk/variant_select.xml" />
|
||||
|
||||
<label id="gatk_variant_quality_score_recalibration" text="Variant Quality Score Recalibration" />
|
||||
<label text="Variant Quality Score Recalibration" id="gatk_variant_quality_score_recalibration" />
|
||||
<tool file="gatk/variant_recalibrator.xml" />
|
||||
<tool file="gatk/variant_apply_recalibration.xml" />
|
||||
|
||||
<label id="gatk_variant_utilities" text="Variant Utilities" />
|
||||
<label text="Variant Utilities" id="gatk_variant_utilities"/>
|
||||
<tool file="gatk/variants_validate.xml" />
|
||||
<tool file="gatk/variant_eval.xml" />
|
||||
<tool file="gatk/variant_combine.xml" />
|
||||
</section>
|
||||
<section id="peak_calling" name="NGS: Peak Calling">
|
||||
<section name="NGS: Peak Calling" id="peak_calling">
|
||||
<tool file="peak_calling/macs_wrapper.xml" />
|
||||
<tool file="peak_calling/sicer_wrapper.xml" />
|
||||
<tool file="peak_calling/ccat_wrapper.xml" />
|
||||
<tool file="genetrack/genetrack_indexer.xml" />
|
||||
<tool file="genetrack/genetrack_peak_prediction.xml" />
|
||||
</section>
|
||||
<section id="ngs-simulation" name="NGS: Simulation">
|
||||
<section name="NGS: Simulation" id="ngs-simulation">
|
||||
<tool file="ngs_simulation/ngs_simulation.xml" />
|
||||
</section>
|
||||
<section id="hgv" name="Phenotype Association">
|
||||
<section name="Phenotype Association" id="hgv">
|
||||
<tool file="evolution/codingSnps.xml" />
|
||||
<tool file="evolution/add_scores.xml" />
|
||||
<tool file="phenotype_association/sift.xml" />
|
||||
<tool file="phenotype_association/linkToGProfile.xml" />
|
||||
<tool file="phenotype_association/linkToDavid.xml" />
|
||||
<tool file="phenotype_association/linkToDavid.xml"/>
|
||||
<tool file="phenotype_association/ctd.xml" />
|
||||
<tool file="phenotype_association/funDo.xml" />
|
||||
<tool file="phenotype_association/snpFreq.xml" />
|
||||
<tool file="phenotype_association/ldtools.xml" />
|
||||
<tool file="phenotype_association/pass.xml" />
|
||||
<tool file="phenotype_association/gpass.xml" />
|
||||
<tool file="phenotype_association/beam.xml" />
|
||||
<tool file="phenotype_association/lps.xml" />
|
||||
<tool file="phenotype_association/hilbertvis.xml" />
|
||||
<tool file="phenotype_association/freebayes.xml" />
|
||||
<tool file="phenotype_association/master2pg.xml" />
|
||||
<tool file="phenotype_association/vcf2pgSnp.xml" />
|
||||
<tool file="phenotype_association/dividePgSnpAlleles.xml" />
|
||||
</section>
|
||||
<section id="vcf_tools" name="VCF Tools">
|
||||
<section name="VCF Tools" id="vcf_tools">
|
||||
<tool file="vcf_tools/intersect.xml" />
|
||||
<tool file="vcf_tools/annotate.xml" />
|
||||
<tool file="vcf_tools/filter.xml" />
|
||||
<tool file="vcf_tools/extract.xml" />
|
||||
</section>
|
||||
</toolbox>
|
||||
</toolbox>
|
||||
|
||||
@@ -1,138 +0,0 @@
|
||||
#!/usr/bin/env python
|
||||
|
||||
# Supports Cuffcompare versions v1.3.0 and newer.
|
||||
|
||||
import optparse, os, shutil, subprocess, sys, tempfile
|
||||
|
||||
def stop_err( msg ):
|
||||
sys.stderr.write( '%s\n' % msg )
|
||||
sys.exit()
|
||||
|
||||
# Copied from sam_to_bam.py:
|
||||
def check_seq_file( dbkey, cached_seqs_pointer_file ):
|
||||
seq_path = ''
|
||||
for line in open( cached_seqs_pointer_file ):
|
||||
line = line.rstrip( '\r\n' )
|
||||
if line and not line.startswith( '#' ) and line.startswith( 'index' ):
|
||||
fields = line.split( '\t' )
|
||||
if len( fields ) < 3:
|
||||
continue
|
||||
if fields[1] == dbkey:
|
||||
seq_path = fields[2].strip()
|
||||
break
|
||||
return seq_path
|
||||
|
||||
def __main__():
|
||||
#Parse Command Line
|
||||
parser = optparse.OptionParser()
|
||||
parser.add_option( '-r', dest='ref_annotation', help='An optional "reference" annotation GTF. Each sample is matched against this file, and sample isoforms are tagged as overlapping, matching, or novel where appropriate. See the refmap and tmap output file descriptions below.' )
|
||||
parser.add_option( '-R', action="store_true", dest='ignore_nonoverlap', help='If -r was specified, this option causes cuffcompare to ignore reference transcripts that are not overlapped by any transcript in one of cuff1.gtf,...,cuffN.gtf. Useful for ignoring annotated transcripts that are not present in your RNA-Seq samples and thus adjusting the "sensitivity" calculation in the accuracy report written in the transcripts accuracy file' )
|
||||
parser.add_option( '-s', dest='use_seq_data', action="store_true", help='Causes cuffcompare to look into for fasta files with the underlying genomic sequences (one file per contig) against which your reads were aligned for some optional classification functions. For example, Cufflinks transcripts consisting mostly of lower-case bases are classified as repeats. Note that <seq_dir> must contain one fasta file per reference chromosome, and each file must be named after the chromosome, and have a .fa or .fasta extension.')
|
||||
|
||||
# Wrapper / Galaxy options.
|
||||
parser.add_option( '', '--dbkey', dest='dbkey', help='The build of the reference dataset' )
|
||||
parser.add_option( '', '--index_dir', dest='index_dir', help='GALAXY_DATA_INDEX_DIR' )
|
||||
parser.add_option( '', '--ref_file', dest='ref_file', help='The reference dataset from the history' )
|
||||
|
||||
# Outputs.
|
||||
parser.add_option( '', '--combined-transcripts', dest='combined_transcripts' )
|
||||
|
||||
(options, args) = parser.parse_args()
|
||||
|
||||
# output version # of tool
|
||||
try:
|
||||
tmp = tempfile.NamedTemporaryFile().name
|
||||
tmp_stdout = open( tmp, 'wb' )
|
||||
proc = subprocess.Popen( args='cuffcompare 2>&1', shell=True, stdout=tmp_stdout )
|
||||
tmp_stdout.close()
|
||||
returncode = proc.wait()
|
||||
stdout = None
|
||||
for line in open( tmp_stdout.name, 'rb' ):
|
||||
if line.lower().find( 'cuffcompare v' ) >= 0:
|
||||
stdout = line.strip()
|
||||
break
|
||||
if stdout:
|
||||
sys.stdout.write( '%s\n' % stdout )
|
||||
else:
|
||||
raise Exception
|
||||
except:
|
||||
sys.stdout.write( 'Could not determine Cuffcompare version\n' )
|
||||
|
||||
# Set/link to sequence file.
|
||||
if options.use_seq_data:
|
||||
if options.ref_file != 'None':
|
||||
# Sequence data from history.
|
||||
# Create symbolic link to ref_file so that index will be created in working directory.
|
||||
seq_path = "ref.fa"
|
||||
os.symlink( options.ref_file, seq_path )
|
||||
else:
|
||||
# Sequence data from loc file.
|
||||
cached_seqs_pointer_file = os.path.join( options.index_dir, 'sam_fa_indices.loc' )
|
||||
if not os.path.exists( cached_seqs_pointer_file ):
|
||||
stop_err( 'The required file (%s) does not exist.' % cached_seqs_pointer_file )
|
||||
# If found for the dbkey, seq_path will look something like /galaxy/data/equCab2/sam_index/equCab2.fa,
|
||||
# and the equCab2.fa file will contain fasta sequences.
|
||||
seq_path = check_seq_file( options.dbkey, cached_seqs_pointer_file )
|
||||
if seq_path == '':
|
||||
stop_err( 'No sequence data found for dbkey %s, so sequence data cannot be used.' % options.dbkey )
|
||||
|
||||
# Build command.
|
||||
|
||||
# Base.
|
||||
cmd = "cuffcompare -o cc_output "
|
||||
|
||||
# Add options.
|
||||
if options.ref_annotation:
|
||||
cmd += " -r %s " % options.ref_annotation
|
||||
if options.ignore_nonoverlap:
|
||||
cmd += " -R "
|
||||
if options.use_seq_data:
|
||||
cmd += " -s %s " % seq_path
|
||||
|
||||
# Add input files.
|
||||
|
||||
# Need to symlink inputs so that output files are written to temp directory.
|
||||
for i, arg in enumerate( args ):
|
||||
input_file_name = "./input%i" % ( i+1 )
|
||||
os.symlink( arg, input_file_name )
|
||||
cmd += "%s " % input_file_name
|
||||
|
||||
# Debugging.
|
||||
print cmd
|
||||
|
||||
# Run command.
|
||||
try:
|
||||
tmp_name = tempfile.NamedTemporaryFile( dir="." ).name
|
||||
tmp_stderr = open( tmp_name, 'wb' )
|
||||
proc = subprocess.Popen( args=cmd, shell=True, stderr=tmp_stderr.fileno() )
|
||||
returncode = proc.wait()
|
||||
tmp_stderr.close()
|
||||
|
||||
# Get stderr, allowing for case where it's very large.
|
||||
tmp_stderr = open( tmp_name, 'rb' )
|
||||
stderr = ''
|
||||
buffsize = 1048576
|
||||
try:
|
||||
while True:
|
||||
stderr += tmp_stderr.read( buffsize )
|
||||
if not stderr or len( stderr ) % buffsize != 0:
|
||||
break
|
||||
except OverflowError:
|
||||
pass
|
||||
tmp_stderr.close()
|
||||
|
||||
# Error checking.
|
||||
if returncode != 0:
|
||||
raise Exception, stderr
|
||||
|
||||
# Copy outputs.
|
||||
shutil.copyfile( "cc_output.combined.gtf" , options.combined_transcripts )
|
||||
|
||||
# check that there are results in the output file
|
||||
cc_output_fname = "cc_output.stats"
|
||||
if len( open( cc_output_fname, 'rb' ).read().strip() ) == 0:
|
||||
raise Exception, 'The main output file is empty, there may be an error with your input file or settings.'
|
||||
except Exception, e:
|
||||
stop_err( 'Error running cuffcompare. ' + str( e ) )
|
||||
|
||||
if __name__=="__main__": __main__()
|
||||
@@ -1,223 +0,0 @@
|
||||
<tool id="cuffcompare" name="Cuffcompare" version="0.0.5">
|
||||
<!-- Wrapper supports Cuffcompare versions v1.3.0 and newer -->
|
||||
<description>compare assembled transcripts to a reference annotation and track Cufflinks transcripts across multiple experiments</description>
|
||||
<requirements>
|
||||
<requirement type="package">cufflinks</requirement>
|
||||
</requirements>
|
||||
<version_command>cuffcompare 2>&1 | head -n 1</version_command>
|
||||
<command interpreter="python">
|
||||
cuffcompare_wrapper.py
|
||||
|
||||
## Use annotation reference?
|
||||
#if $annotation.use_ref_annotation == "Yes":
|
||||
-r $annotation.reference_annotation
|
||||
#if $annotation.ignore_nonoverlapping_reference:
|
||||
-R
|
||||
#end if
|
||||
#end if
|
||||
|
||||
## Use sequence data?
|
||||
#if $seq_data.use_seq_data == "Yes":
|
||||
-s
|
||||
#if $seq_data.seq_source.index_source == "history":
|
||||
--ref_file=$seq_data.seq_source.ref_file
|
||||
#else:
|
||||
--ref_file="None"
|
||||
#end if
|
||||
--dbkey=${first_input.metadata.dbkey}
|
||||
--index_dir=${GALAXY_DATA_INDEX_DIR}
|
||||
#end if
|
||||
|
||||
## Outputs.
|
||||
--combined-transcripts=${transcripts_combined}
|
||||
|
||||
## Inputs.
|
||||
${first_input}
|
||||
#for $input_file in $input_files:
|
||||
${input_file.additional_input}
|
||||
#end for
|
||||
|
||||
</command>
|
||||
<inputs>
|
||||
<param format="gtf" name="first_input" type="data" label="GTF file produced by Cufflinks" help=""/>
|
||||
<repeat name="input_files" title="Additional GTF Input Files">
|
||||
<param format="gtf" name="additional_input" type="data" label="GTF file produced by Cufflinks" help=""/>
|
||||
</repeat>
|
||||
<conditional name="annotation">
|
||||
<param name="use_ref_annotation" type="select" label="Use Reference Annotation">
|
||||
<option value="No">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<when value="Yes">
|
||||
<param format="gff3,gtf" name="reference_annotation" type="data" label="Reference Annotation" help="Requires an annotation file in GFF3 or GTF format."/>
|
||||
<param name="ignore_nonoverlapping_reference" type="boolean" label="Ignore reference transcripts that are not overlapped by any transcript in input files"/>
|
||||
</when>
|
||||
<when value="No">
|
||||
</when>
|
||||
</conditional>
|
||||
<conditional name="seq_data">
|
||||
<param name="use_seq_data" type="select" label="Use Sequence Data" help="Use sequence data for some optional classification functions, including the addition of the p_id attribute required by Cuffdiff.">
|
||||
<option value="Yes">Yes</option>
|
||||
<option value="No">No</option>
|
||||
</param>
|
||||
<when value="No"></when>
|
||||
<when value="Yes">
|
||||
<conditional name="seq_source">
|
||||
<param name="index_source" type="select" label="Choose the source for the reference list">
|
||||
<option value="cached">Locally cached</option>
|
||||
<option value="history">History</option>
|
||||
</param>
|
||||
<when value="cached"></when>
|
||||
<when value="history">
|
||||
<param name="ref_file" type="data" format="fasta" label="Using reference file" />
|
||||
</when>
|
||||
</conditional>
|
||||
</when>
|
||||
</conditional>
|
||||
</inputs>
|
||||
|
||||
<outputs>
|
||||
<data format="txt" name="transcripts_accuracy" label="${tool.name} on ${on_string}: transcript accuracy"
|
||||
from_work_dir="cc_output.stats" />
|
||||
<data format="tabular" name="input1_tmap" label="${tool.name} on ${on_string}: data ${first_input.hid} tmap file"
|
||||
from_work_dir="cc_output.input1.tmap" />
|
||||
<data format="tabular" name="input1_refmap"
|
||||
label="${tool.name} on ${on_string}: data ${first_input.hid} refmap file"
|
||||
from_work_dir="cc_output.input1.refmap">
|
||||
<filter>annotation['use_ref_annotation'] == 'Yes'</filter>
|
||||
</data>
|
||||
<data format="tabular" name="input2_tmap" label="${tool.name} on ${on_string}: data ${input_files[0]['additional_input'].hid} tmap file" from_work_dir="cc_output.input2.tmap">
|
||||
<filter>len( input_files ) >= 1</filter>
|
||||
</data>
|
||||
<data format="tabular" name="input2_refmap"
|
||||
label="${tool.name} on ${on_string}: data ${input_files[0]['additional_input'].hid} refmap file"
|
||||
from_work_dir="cc_output.input2.refmap">
|
||||
<filter>annotation['use_ref_annotation'] == 'Yes' and len( input_files ) >= 1</filter>
|
||||
</data>
|
||||
<data format="tabular" name="transcripts_tracking" label="${tool.name} on ${on_string}: transcript tracking" from_work_dir="cc_output.tracking">
|
||||
<filter>len( input_files ) > 0</filter>
|
||||
</data>
|
||||
<data format="gtf" name="transcripts_combined" label="${tool.name} on ${on_string}: combined transcripts"/>
|
||||
</outputs>
|
||||
|
||||
<tests>
|
||||
<!--
|
||||
cuffcompare -r cuffcompare_in3.gtf -R cuffcompare_in1.gtf cuffcompare_in2.gtf
|
||||
-->
|
||||
<test>
|
||||
<param name="first_input" value="cuffcompare_in1.gtf" ftype="gtf"/>
|
||||
<param name="additional_input" value="cuffcompare_in2.gtf" ftype="gtf"/>
|
||||
<param name="use_ref_annotation" value="Yes"/>
|
||||
<param name="reference_annotation" value="cuffcompare_in3.gtf" ftype="gtf"/>
|
||||
<param name="ignore_nonoverlapping_reference" value="Yes"/>
|
||||
<param name="use_seq_data" value="No"/>
|
||||
<!-- Line diffs are the result of different locations for input files; this cannot be fixed as cuffcompare outputs
|
||||
full input path for each input. -->
|
||||
<output name="transcripts_accuracy" file="cuffcompare_out7.txt" lines_diff="16"/>
|
||||
<output name="input1_tmap" file="cuffcompare_out1.tmap"/>
|
||||
<output name="input1_refmap" file="cuffcompare_out2.refmap"/>
|
||||
<output name="input2_tmap" file="cuffcompare_out3.tmap"/>
|
||||
<output name="input2_refmap" file="cuffcompare_out4.refmap"/>
|
||||
<output name="transcripts_tracking" file="cuffcompare_out6.tracking"/>
|
||||
<output name="transcripts_combined" file="cuffcompare_out5.gtf"/>
|
||||
</test>
|
||||
</tests>
|
||||
|
||||
<help>
|
||||
**Cuffcompare Overview**
|
||||
|
||||
Cuffcompare is part of Cufflinks_. Cuffcompare helps you: (a) compare your assembled transcripts to a reference annotation and (b) track Cufflinks transcripts across multiple experiments (e.g. across a time course). Please cite: Trapnell C, Williams BA, Pertea G, Mortazavi AM, Kwan G, van Baren MJ, Salzberg SL, Wold B, Pachter L. Transcript assembly and abundance estimation from RNA-Seq reveals thousands of new transcripts and switching among isoforms. Nature Biotechnology doi:10.1038/nbt.1621
|
||||
|
||||
.. _Cufflinks: http://cufflinks.cbcb.umd.edu/
|
||||
|
||||
------
|
||||
|
||||
**Know what you are doing**
|
||||
|
||||
.. class:: warningmark
|
||||
|
||||
There is no such thing (yet) as an automated gearshift in expression analysis. It is all like stick-shift driving in San Francisco. In other words, running this tool with default parameters will probably not give you meaningful results. A way to deal with this is to **understand** the parameters by carefully reading the `documentation`__ and experimenting. Fortunately, Galaxy makes experimenting easy.
|
||||
|
||||
.. __: http://cufflinks.cbcb.umd.edu/manual.html#cuffcompare
|
||||
|
||||
------
|
||||
|
||||
**Input format**
|
||||
|
||||
Cuffcompare takes Cufflinks' GTF output as input, and optionally can take a "reference" annotation (such as from Ensembl_)
|
||||
|
||||
.. _Ensembl: http://www.ensembl.org
|
||||
|
||||
------
|
||||
|
||||
**Outputs**
|
||||
|
||||
Cuffcompare produces the following output files:
|
||||
|
||||
Transcripts Accuracy File:
|
||||
|
||||
Cuffcompare reports various statistics related to the "accuracy" of the transcripts in each sample when compared to the reference annotation data. The typical gene finding measures of "sensitivity" and "specificity" (as defined in Burset, M., Guigó, R. : Evaluation of gene structure prediction programs (1996) Genomics, 34 (3), pp. 353-367. doi: 10.1006/geno.1996.0298) are calculated at various levels (nucleotide, exon, intron, transcript, gene) for each input file and reported in this file. The Sn and Sp columns show specificity and sensitivity values at each level, while the fSn and fSp columns are "fuzzy" variants of these same accuracy calculations, allowing for a very small variation in exon boundaries to still be counted as a "match".
|
||||
|
||||
Transcripts Combined File:
|
||||
|
||||
Cuffcompare reports a GTF file containing the "union" of all transfrags in each sample. If a transfrag is present in both samples, it is thus reported once in the combined gtf.
|
||||
|
||||
Transcripts Tracking File:
|
||||
|
||||
This file matches transcripts up between samples. Each row contains a transcript structure that is present in one or more input GTF files. Because the transcripts will generally have different IDs (unless you assembled your RNA-Seq reads against a reference transcriptome), cuffcompare examines the structure of each the transcripts, matching transcripts that agree on the coordinates and order of all of their introns, as well as strand. Matching transcripts are allowed to differ on the length of the first and last exons, since these lengths will naturally vary from sample to sample due to the random nature of sequencing.
|
||||
If you ran cuffcompare with the -r option, the first and second columns contain the closest matching reference transcript to the one described by each row.
|
||||
|
||||
Here's an example of a line from the tracking file::
|
||||
|
||||
TCONS_00000045 XLOC_000023 Tcea|uc007afj.1 j \
|
||||
q1:exp.115|exp.115.0|100|3.061355|0.350242|0.350207 \
|
||||
q2:60hr.292|60hr.292.0|100|4.094084|0.000000|0.000000
|
||||
|
||||
In this example, a transcript present in the two input files, called exp.115.0 in the first and 60hr.292.0 in the second, doesn't match any reference transcript exactly, but shares exons with uc007afj.1, an isoform of the gene Tcea, as indicated by the class code j. The first three columns are as follows::
|
||||
|
||||
Column number Column name Example Description
|
||||
-----------------------------------------------------------------------
|
||||
1 Cufflinks transfrag id TCONS_00000045 A unique internal id for the transfrag
|
||||
2 Cufflinks locus id XLOC_000023 A unique internal id for the locus
|
||||
3 Reference gene id Tcea The gene_name attribute of the reference GTF record for this transcript, or '-' if no reference transcript overlaps this Cufflinks transcript
|
||||
4 Reference transcript id uc007afj.1 The transcript_id attribute of the reference GTF record for this transcript, or '-' if no reference transcript overlaps this Cufflinks transcript
|
||||
5 Class code c The type of match between the Cufflinks transcripts in column 6 and the reference transcript. See class codes
|
||||
|
||||
Each of the columns after the fifth have the following format:
|
||||
qJ:gene_id|transcript_id|FMI|FPKM|conf_lo|conf_hi
|
||||
|
||||
A transcript need be present in all samples to be reported in the tracking file. A sample not containing a transcript will have a "-" in its entry in the row for that transcript.
|
||||
|
||||
Class Codes
|
||||
|
||||
If you ran cuffcompare with the -r option, tracking rows will contain the following values. If you did not use -r, the rows will all contain "-" in their class code column::
|
||||
|
||||
Priority Code Description
|
||||
---------------------------------
|
||||
1 = Match
|
||||
2 c Contained
|
||||
3 j New isoform
|
||||
4 e A single exon transcript overlapping a reference exon and at least 10 bp of a reference intron, indicating a possible pre-mRNA fragment.
|
||||
5 i A single exon transcript falling entirely with a reference intron
|
||||
6 r Repeat. Currently determined by looking at the reference sequence and applied to transcripts where at least 50% of the bases are lower case
|
||||
7 p Possible polymerase run-on fragment
|
||||
8 u Unknown, intergenic transcript
|
||||
9 o Unknown, generic overlap with reference
|
||||
10 . (.tracking file only, indicates multiple classifications)
|
||||
|
||||
-------
|
||||
|
||||
**Settings**
|
||||
|
||||
All of the options have a default value. You can change any of them. Most of the options in Cuffcompare have been implemented here.
|
||||
|
||||
------
|
||||
|
||||
**Cuffcompare parameter list**
|
||||
|
||||
This is a list of implemented Cuffcompare options::
|
||||
|
||||
-r An optional "reference" annotation GTF. Each sample is matched against this file, and sample isoforms are tagged as overlapping, matching, or novel where appropriate. See the refmap and tmap output file descriptions below.
|
||||
-R If -r was specified, this option causes cuffcompare to ignore reference transcripts that are not overlapped by any transcript in one of cuff1.gtf,...,cuffN.gtf. Useful for ignoring annotated transcripts that are not present in your RNA-Seq samples and thus adjusting the "sensitivity" calculation in the accuracy report written in the transcripts_accuracy file
|
||||
</help>
|
||||
</tool>
|
||||
@@ -1,246 +0,0 @@
|
||||
<tool id="cuffdiff" name="Cuffdiff" version="0.0.6">
|
||||
<!-- Wrapper supports Cuffdiff versions 2.1.0-2.1.1 -->
|
||||
<description>find significant changes in transcript expression, splicing, and promoter use</description>
|
||||
<requirements>
|
||||
<requirement type="package">cufflinks</requirement>
|
||||
</requirements>
|
||||
<version_command>cuffdiff 2>&1 | head -n 1</version_command>
|
||||
<command>
|
||||
cuffdiff
|
||||
--FDR=$fdr
|
||||
--num-threads="4"
|
||||
--min-alignment-count=$min_alignment_count
|
||||
--library-norm-method=$library_norm_method
|
||||
--dispersion-method=$dispersion_method
|
||||
|
||||
## Set advanced data parameters?
|
||||
#if $additional.sAdditional == "Yes":
|
||||
-m $additional.frag_mean_len
|
||||
-s $additional.frag_len_std_dev
|
||||
#end if
|
||||
|
||||
## Multi-read correct?
|
||||
#if str($multiread_correct) == "Yes":
|
||||
-u
|
||||
#end if
|
||||
|
||||
## Bias correction?
|
||||
#if $bias_correction.do_bias_correction == "Yes":
|
||||
-b
|
||||
#if $bias_correction.seq_source.index_source == "history":
|
||||
--ref_file=$bias_correction.seq_source.ref_file
|
||||
#else:
|
||||
--ref_file="None"
|
||||
#end if
|
||||
--dbkey=${gtf_input.metadata.dbkey}
|
||||
--index_dir=${GALAXY_DATA_INDEX_DIR}
|
||||
#end if
|
||||
|
||||
#set labels = ','.join( [ str( $condition.name ) for $condition in $conditions ] )
|
||||
--labels $labels
|
||||
|
||||
## Inputs.
|
||||
$gtf_input
|
||||
#for $condition in $conditions:
|
||||
#set samples = ','.join( [ str( $sample.sample ) for $sample in $condition.samples ] )
|
||||
$samples
|
||||
#end for
|
||||
</command>
|
||||
<inputs>
|
||||
<param format="gtf,gff3" name="gtf_input" type="data" label="Transcripts" help="A transcript GFF3 or GTF file produced by cufflinks, cuffcompare, or other source."/>
|
||||
|
||||
<repeat name="conditions" title="Condition" min="2">
|
||||
<param name="name" title="Condition name" type="text" label="Name"/>
|
||||
<repeat name="samples" title="Replicate" min="1">
|
||||
<param name="sample" label="Add replicate" type="data" format="sam,bam"/>
|
||||
</repeat>
|
||||
</repeat>
|
||||
|
||||
<param name="library_norm_method" type="select" label="Library normalization method">
|
||||
<option value="geometric" selected="True">geometric</option>
|
||||
<option value="classic-fpkm">classic-fpkm</option>
|
||||
<option value="quartile">quartile</option>
|
||||
</param>
|
||||
|
||||
<param name="dispersion_method" type="select" label="Dispersion estimation method" help="If using only one sample per condition, you must use 'blind.'">
|
||||
<option value="pooled" selected="True">pooled</option>
|
||||
<option value="per-condition">per-condition</option>
|
||||
<option value="blind">blind</option>
|
||||
</param>
|
||||
|
||||
<param name="fdr" type="float" value="0.05" label="False Discovery Rate" help="The allowed false discovery rate."/>
|
||||
|
||||
<param name="min_alignment_count" type="integer" value="10" label="Min Alignment Count" help="The minimum number of alignments in a locus for needed to conduct significance testing on changes in that locus observed between samples."/>
|
||||
|
||||
<param name="multiread_correct" type="select" label="Use multi-read correct" help="Tells Cufflinks to do an initial estimation procedure to more accurately weight reads mapping to multiple locations in the genome.">
|
||||
<option value="No" selected="true">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
|
||||
<conditional name="bias_correction">
|
||||
<param name="do_bias_correction" type="select" label="Perform Bias Correction" help="Bias detection and correction can significantly improve accuracy of transcript abundance estimates.">
|
||||
<option value="No">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<when value="Yes">
|
||||
<conditional name="seq_source">
|
||||
<param name="index_source" type="select" label="Reference sequence data">
|
||||
<option value="cached">Locally cached</option>
|
||||
<option value="history">History</option>
|
||||
</param>
|
||||
<when value="cached"></when>
|
||||
<when value="history">
|
||||
<param name="ref_file" type="data" format="fasta" label="Using reference file" />
|
||||
</when>
|
||||
</conditional>
|
||||
</when>
|
||||
<when value="No"></when>
|
||||
</conditional>
|
||||
|
||||
<param name="include_read_group_files" type="select" label="Include Read Group Datasets" help="Read group datasets provide information on replicates.">
|
||||
<option value="No" selected="true">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
|
||||
<conditional name="additional">
|
||||
<param name="sAdditional" type="select" label="Set Additional Parameters? (not recommended for paired-end reads)">
|
||||
<option value="No">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<when value="No"></when>
|
||||
<when value="Yes">
|
||||
<param name="frag_mean_len" type="integer" value="200" label="Average Fragment Length"/>
|
||||
<param name="frag_len_std_dev" type="integer" value="80" label="Fragment Length Standard Deviation"/>
|
||||
</when>
|
||||
</conditional>
|
||||
</inputs>
|
||||
|
||||
<stdio>
|
||||
<regex match=".*" source="both" level="log" description="tool progress"/>
|
||||
</stdio>
|
||||
|
||||
<outputs>
|
||||
<!-- Optional read group datasets. -->
|
||||
<data format="tabular" name="isoforms_read_group" label="${tool.name} on ${on_string}: isoforms read group tracking" from_work_dir="isoforms.read_group_tracking" >
|
||||
<filter>(params['include_read_group_files'] == 'Yes'</filter>
|
||||
</data>
|
||||
<data format="tabular" name="genes_read_group" label="${tool.name} on ${on_string}: genes read group tracking" from_work_dir="genes.read_group_tracking" >
|
||||
<filter>(params['include_read_group_files'] == 'Yes'</filter>
|
||||
</data>
|
||||
<data format="tabular" name="cds_read_group" label="${tool.name} on ${on_string}: CDs read group tracking" from_work_dir="cds.read_group_tracking" >
|
||||
<filter>(params['include_read_group_files'] == 'Yes'</filter>
|
||||
</data>
|
||||
<data format="tabular" name="tss_groups_read_group" label="${tool.name} on ${on_string}: TSS groups read group tracking" from_work_dir="tss_groups.read_group_tracking" >
|
||||
<filter>(params['include_read_group_files'] == 'Yes'</filter>
|
||||
</data>
|
||||
|
||||
<!-- Standard datasets. -->
|
||||
<data format="tabular" name="splicing_diff" label="${tool.name} on ${on_string}: splicing differential expression testing" from_work_dir="splicing.diff" />
|
||||
<data format="tabular" name="promoters_diff" label="${tool.name} on ${on_string}: promoters differential expression testing" from_work_dir="promoters.diff" />
|
||||
<data format="tabular" name="cds_diff" label="${tool.name} on ${on_string}: CDS overloading diffential expression testing" from_work_dir="cds.diff" />
|
||||
<data format="tabular" name="cds_exp_fpkm_tracking" label="${tool.name} on ${on_string}: CDS FPKM differential expression testing" from_work_dir="cds_exp.diff" />
|
||||
<data format="tabular" name="cds_fpkm_tracking" label="${tool.name} on ${on_string}: CDS FPKM tracking" from_work_dir="cds.fpkm_tracking" />
|
||||
<data format="tabular" name="tss_groups_exp" label="${tool.name} on ${on_string}: TSS groups differential expression testing" from_work_dir="tss_group_exp.diff" />
|
||||
<data format="tabular" name="tss_groups_fpkm_tracking" label="${tool.name} on ${on_string}: TSS groups FPKM tracking" from_work_dir="tss_groups.fpkm_tracking" />
|
||||
<data format="tabular" name="genes_exp" label="${tool.name} on ${on_string}: gene differential expression testing" from_work_dir="gene_exp.diff" />
|
||||
<data format="tabular" name="genes_fpkm_tracking" label="${tool.name} on ${on_string}: gene FPKM tracking" from_work_dir="genes.fpkm_tracking" />
|
||||
<data format="tabular" name="isoforms_exp" label="${tool.name} on ${on_string}: transcript differential expression testing" from_work_dir="isoform_exp.diff" />
|
||||
<data format="tabular" name="isoforms_fpkm_tracking" label="${tool.name} on ${on_string}: transcript FPKM tracking" from_work_dir="isoforms.fpkm_tracking" />
|
||||
</outputs>
|
||||
|
||||
<tests>
|
||||
<test>
|
||||
<!--
|
||||
cuffdiff cuffcompare_out5.gtf cuffdiff_in1.sam cuffdiff_in2.sam
|
||||
-->
|
||||
<!--
|
||||
NOTE: as of version 0.0.6 of the wrapper, tests cannot be run because multiple inputs to a repeat
|
||||
element are not supported.
|
||||
<param name="gtf_input" value="cuffcompare_out5.gtf" ftype="gtf" />
|
||||
<param name="do_groups" value="No" />
|
||||
<param name="aligned_reads1" value="cuffdiff_in1.sam" ftype="sam" />
|
||||
<param name="aligned_reads2" value="cuffdiff_in2.sam" ftype="sam" />
|
||||
<param name="fdr" value="0.05" />
|
||||
<param name="min_alignment_count" value="0" />
|
||||
<param name="do_bias_correction" value="No" />
|
||||
<param name="do_normalization" value="No" />
|
||||
<param name="multiread_correct" value="No"/>
|
||||
<param name="sAdditional" value="No"/>
|
||||
<output name="splicing_diff" file="cuffdiff_out9.txt"/>
|
||||
<output name="promoters_diff" file="cuffdiff_out10.txt"/>
|
||||
<output name="cds_diff" file="cuffdiff_out11.txt"/>
|
||||
<output name="cds_exp_fpkm_tracking" file="cuffdiff_out4.txt"/>
|
||||
<output name="cds_fpkm_tracking" file="cuffdiff_out8.txt"/>
|
||||
<output name="tss_groups_exp" file="cuffdiff_out3.txt" lines_diff="200"/>
|
||||
<output name="tss_groups_fpkm_tracking" file="cuffdiff_out7.txt"/>
|
||||
<output name="genes_exp" file="cuffdiff_out2.txt" lines_diff="200"/>
|
||||
<output name="genes_fpkm_tracking" file="cuffdiff_out6.txt" lines_diff="200"/>
|
||||
<output name="isoforms_exp" file="cuffdiff_out1.txt" lines_diff="200"/>
|
||||
<output name="isoforms_fpkm_tracking" file="cuffdiff_out5.txt" lines_diff="200"/>
|
||||
-->
|
||||
</test>
|
||||
</tests>
|
||||
|
||||
<help>
|
||||
**Cuffdiff Overview**
|
||||
|
||||
Cuffdiff is part of Cufflinks_. Cuffdiff find significant changes in transcript expression, splicing, and promoter use. Please cite: Trapnell C, Williams BA, Pertea G, Mortazavi AM, Kwan G, van Baren MJ, Salzberg SL, Wold B, Pachter L. Transcript assembly and abundance estimation from RNA-Seq reveals thousands of new transcripts and switching among isoforms. Nature Biotechnology doi:10.1038/nbt.1621
|
||||
|
||||
.. _Cufflinks: http://cufflinks.cbcb.umd.edu/
|
||||
|
||||
------
|
||||
|
||||
**Know what you are doing**
|
||||
|
||||
.. class:: warningmark
|
||||
|
||||
There is no such thing (yet) as an automated gearshift in expression analysis. It is all like stick-shift driving in San Francisco. In other words, running this tool with default parameters will probably not give you meaningful results. A way to deal with this is to **understand** the parameters by carefully reading the `documentation`__ and experimenting. Fortunately, Galaxy makes experimenting easy.
|
||||
|
||||
.. __: http://cufflinks.cbcb.umd.edu/manual.html#cuffdiff
|
||||
|
||||
------
|
||||
|
||||
**Input format**
|
||||
|
||||
Cuffdiff takes Cufflinks or Cuffcompare GTF files as input along with two SAM files containing the fragment alignments for two or more samples.
|
||||
|
||||
------
|
||||
|
||||
**Outputs**
|
||||
|
||||
Cuffdiff produces many output files:
|
||||
|
||||
1. Transcript FPKM expression tracking.
|
||||
2. Gene FPKM expression tracking; tracks the summed FPKM of transcripts sharing each gene_id
|
||||
3. Primary transcript FPKM tracking; tracks the summed FPKM of transcripts sharing each tss_id
|
||||
4. Coding sequence FPKM tracking; tracks the summed FPKM of transcripts sharing each p_id, independent of tss_id
|
||||
5. Transcript differential FPKM.
|
||||
6. Gene differential FPKM. Tests difference sin the summed FPKM of transcripts sharing each gene_id
|
||||
7. Primary transcript differential FPKM. Tests difference sin the summed FPKM of transcripts sharing each tss_id
|
||||
8. Coding sequence differential FPKM. Tests difference sin the summed FPKM of transcripts sharing each p_id independent of tss_id
|
||||
9. Differential splicing tests: this tab delimited file lists, for each primary transcript, the amount of overloading detected among its isoforms, i.e. how much differential splicing exists between isoforms processed from a single primary transcript. Only primary transcripts from which two or more isoforms are spliced are listed in this file.
|
||||
10. Differential promoter tests: this tab delimited file lists, for each gene, the amount of overloading detected among its primary transcripts, i.e. how much differential promoter use exists between samples. Only genes producing two or more distinct primary transcripts (i.e. multi-promoter genes) are listed here.
|
||||
11. Differential CDS tests: this tab delimited file lists, for each gene, the amount of overloading detected among its coding sequences, i.e. how much differential CDS output exists between samples. Only genes producing two or more distinct CDS (i.e. multi-protein genes) are listed here.
|
||||
|
||||
-------
|
||||
|
||||
**Settings**
|
||||
|
||||
All of the options have a default value. You can change any of them. Most of the options in Cuffdiff have been implemented here.
|
||||
|
||||
------
|
||||
|
||||
**Cuffdiff parameter list**
|
||||
|
||||
This is a list of implemented Cuffdiff options::
|
||||
|
||||
-m INT Average fragement length; default 200
|
||||
-s INT Fragment legnth standard deviation; default 80
|
||||
-c INT The minimum number of alignments in a locus for needed to conduct significance testing on changes in that locus observed between samples. If no testing is performed, changes in the locus are deemed not significant, and the locus' observed changes don't contribute to correction for multiple testing. The default is 1,000 fragment alignments (up to 2,000 paired reads).
|
||||
--FDR FLOAT The allowed false discovery rate. The default is 0.05.
|
||||
--num-importance-samples INT Sets the number of importance samples generated for each locus during abundance estimation. Default: 1000
|
||||
--max-mle-iterations INT Sets the number of iterations allowed during maximum likelihood estimation of abundances. Default: 5000
|
||||
-N With this option, Cufflinks excludes the contribution of the top 25 percent most highly expressed genes from the number of mapped fragments used in the FPKM denominator. This can improve robustness of differential expression calls for less abundant genes and transcripts.
|
||||
|
||||
</help>
|
||||
</tool>
|
||||
@@ -1,227 +0,0 @@
|
||||
#!/usr/bin/env python
|
||||
|
||||
# Supports Cufflinks versions 1.3 and newer.
|
||||
|
||||
import optparse, os, shutil, subprocess, sys, tempfile
|
||||
from galaxy import eggs
|
||||
from galaxy.datatypes.util.gff_util import parse_gff_attributes, gff_attributes_to_str
|
||||
|
||||
def stop_err( msg ):
|
||||
sys.stderr.write( "%s\n" % msg )
|
||||
sys.exit()
|
||||
|
||||
# Copied from sam_to_bam.py:
|
||||
def check_seq_file( dbkey, cached_seqs_pointer_file ):
|
||||
seq_path = ''
|
||||
for line in open( cached_seqs_pointer_file ):
|
||||
line = line.rstrip( '\r\n' )
|
||||
if line and not line.startswith( '#' ) and line.startswith( 'index' ):
|
||||
fields = line.split( '\t' )
|
||||
if len( fields ) < 3:
|
||||
continue
|
||||
if fields[1] == dbkey:
|
||||
seq_path = fields[2].strip()
|
||||
break
|
||||
return seq_path
|
||||
|
||||
def __main__():
|
||||
#Parse Command Line
|
||||
parser = optparse.OptionParser()
|
||||
parser.add_option( '-1', '--input', dest='input', help=' file of RNA-Seq read alignments in the SAM format. SAM is a standard short read alignment, that allows aligners to attach custom tags to individual alignments, and Cufflinks requires that the alignments you supply have some of these tags. Please see Input formats for more details.' )
|
||||
parser.add_option( '-s', '--inner-dist-std-dev', dest='inner_dist_std_dev', help='The standard deviation for the distribution on inner distances between mate pairs. The default is 20bp.' )
|
||||
parser.add_option( '-I', '--max-intron-length', dest='max_intron_len', help='The minimum intron length. Cufflinks will not report transcripts with introns longer than this, and will ignore SAM alignments with REF_SKIP CIGAR operations longer than this. The default is 300,000.' )
|
||||
parser.add_option( '-F', '--min-isoform-fraction', dest='min_isoform_fraction', help='After calculating isoform abundance for a gene, Cufflinks filters out transcripts that it believes are very low abundance, because isoforms expressed at extremely low levels often cannot reliably be assembled, and may even be artifacts of incompletely spliced precursors of processed transcripts. This parameter is also used to filter out introns that have far fewer spliced alignments supporting them. The default is 0.05, or 5% of the most abundant isoform (the major isoform) of the gene.' )
|
||||
parser.add_option( '-j', '--pre-mrna-fraction', dest='pre_mrna_fraction', help='Some RNA-Seq protocols produce a significant amount of reads that originate from incompletely spliced transcripts, and these reads can confound the assembly of fully spliced mRNAs. Cufflinks uses this parameter to filter out alignments that lie within the intronic intervals implied by the spliced alignments. The minimum depth of coverage in the intronic region covered by the alignment is divided by the number of spliced reads, and if the result is lower than this parameter value, the intronic alignments are ignored. The default is 5%.' )
|
||||
parser.add_option( '-p', '--num-threads', dest='num_threads', help='Use this many threads to align reads. The default is 1.' )
|
||||
parser.add_option( '-m', '--inner-mean-dist', dest='inner_mean_dist', help='This is the expected (mean) inner distance between mate pairs. \
|
||||
For, example, for paired end runs with fragments selected at 300bp, \
|
||||
where each end is 50bp, you should set -r to be 200. The default is 45bp.')
|
||||
parser.add_option( '-G', '--GTF', dest='GTF', help='Tells Cufflinks to use the supplied reference annotation to estimate isoform expression. It will not assemble novel transcripts, and the program will ignore alignments not structurally compatible with any reference transcript.' )
|
||||
parser.add_option( '-g', '--GTF-guide', dest='GTFguide', help='use reference transcript annotation to guide assembly' )
|
||||
parser.add_option( '-u', '--multi-read-correct', dest='multi_read_correct', action="store_true", help='Tells Cufflinks to do an initial estimation procedure to more accurately weight reads mapping to multiple locations in the genome')
|
||||
|
||||
# Normalization options.
|
||||
parser.add_option( "-N", "--quartile-normalization", dest="do_normalization", action="store_true" )
|
||||
parser.add_option( "--no-effective-length-correction", dest="no_effective_length_correction", action="store_true" )
|
||||
|
||||
# Wrapper / Galaxy options.
|
||||
parser.add_option( '-A', '--assembled-isoforms-output', dest='assembled_isoforms_output_file', help='Assembled isoforms output file; formate is GTF.' )
|
||||
|
||||
# Advanced Options:
|
||||
parser.add_option( '--num-importance-samples', dest='num_importance_samples', help='Sets the number of importance samples generated for each locus during abundance estimation. Default: 1000' )
|
||||
parser.add_option( '--max-mle-iterations', dest='max_mle_iterations', help='Sets the number of iterations allowed during maximum likelihood estimation of abundances. Default: 5000' )
|
||||
|
||||
# Bias correction options.
|
||||
parser.add_option( '-b', dest='do_bias_correction', action="store_true", help='Providing Cufflinks with a multifasta file via this option instructs it to run our new bias detection and correction algorithm which can significantly improve accuracy of transcript abundance estimates.')
|
||||
parser.add_option( '', '--dbkey', dest='dbkey', help='The build of the reference dataset' )
|
||||
parser.add_option( '', '--index_dir', dest='index_dir', help='GALAXY_DATA_INDEX_DIR' )
|
||||
parser.add_option( '', '--ref_file', dest='ref_file', help='The reference dataset from the history' )
|
||||
|
||||
# Global model.
|
||||
parser.add_option( '', '--global_model', dest='global_model_file', help='Global model used for computing on local data' )
|
||||
|
||||
(options, args) = parser.parse_args()
|
||||
|
||||
# output version # of tool
|
||||
try:
|
||||
tmp = tempfile.NamedTemporaryFile().name
|
||||
tmp_stdout = open( tmp, 'wb' )
|
||||
proc = subprocess.Popen( args='cufflinks --no-update-check 2>&1', shell=True, stdout=tmp_stdout )
|
||||
tmp_stdout.close()
|
||||
returncode = proc.wait()
|
||||
stdout = None
|
||||
for line in open( tmp_stdout.name, 'rb' ):
|
||||
if line.lower().find( 'cufflinks v' ) >= 0:
|
||||
stdout = line.strip()
|
||||
break
|
||||
if stdout:
|
||||
sys.stdout.write( '%s\n' % stdout )
|
||||
else:
|
||||
raise Exception
|
||||
except:
|
||||
sys.stdout.write( 'Could not determine Cufflinks version\n' )
|
||||
|
||||
# If doing bias correction, set/link to sequence file.
|
||||
if options.do_bias_correction:
|
||||
if options.ref_file != 'None':
|
||||
# Sequence data from history.
|
||||
# Create symbolic link to ref_file so that index will be created in working directory.
|
||||
seq_path = "ref.fa"
|
||||
os.symlink( options.ref_file, seq_path )
|
||||
else:
|
||||
# Sequence data from loc file.
|
||||
cached_seqs_pointer_file = os.path.join( options.index_dir, 'sam_fa_indices.loc' )
|
||||
if not os.path.exists( cached_seqs_pointer_file ):
|
||||
stop_err( 'The required file (%s) does not exist.' % cached_seqs_pointer_file )
|
||||
# If found for the dbkey, seq_path will look something like /galaxy/data/equCab2/sam_index/equCab2.fa,
|
||||
# and the equCab2.fa file will contain fasta sequences.
|
||||
seq_path = check_seq_file( options.dbkey, cached_seqs_pointer_file )
|
||||
if seq_path == '':
|
||||
stop_err( 'No sequence data found for dbkey %s, so bias correction cannot be used.' % options.dbkey )
|
||||
|
||||
# Build command.
|
||||
|
||||
# Base; always use quiet mode to avoid problems with storing log output.
|
||||
cmd = "cufflinks -q --no-update-check"
|
||||
|
||||
# Add options.
|
||||
if options.inner_dist_std_dev:
|
||||
cmd += ( " -s %i" % int ( options.inner_dist_std_dev ) )
|
||||
if options.max_intron_len:
|
||||
cmd += ( " -I %i" % int ( options.max_intron_len ) )
|
||||
if options.min_isoform_fraction:
|
||||
cmd += ( " -F %f" % float ( options.min_isoform_fraction ) )
|
||||
if options.pre_mrna_fraction:
|
||||
cmd += ( " -j %f" % float ( options.pre_mrna_fraction ) )
|
||||
if options.num_threads:
|
||||
cmd += ( " -p %i" % int ( options.num_threads ) )
|
||||
if options.inner_mean_dist:
|
||||
cmd += ( " -m %i" % int ( options.inner_mean_dist ) )
|
||||
if options.GTF:
|
||||
cmd += ( " -G %s" % options.GTF )
|
||||
if options.GTFguide:
|
||||
cmd += ( " -g %s" % options.GTFguide )
|
||||
if options.multi_read_correct:
|
||||
cmd += ( " -u" )
|
||||
if options.num_importance_samples:
|
||||
cmd += ( " --num-importance-samples %i" % int ( options.num_importance_samples ) )
|
||||
if options.max_mle_iterations:
|
||||
cmd += ( " --max-mle-iterations %i" % int ( options.max_mle_iterations ) )
|
||||
if options.do_normalization:
|
||||
cmd += ( " -N" )
|
||||
if options.do_bias_correction:
|
||||
cmd += ( " -b %s" % seq_path )
|
||||
if options.no_effective_length_correction:
|
||||
cmd += ( " --no-effective-length-correction" )
|
||||
|
||||
# Debugging.
|
||||
print cmd
|
||||
|
||||
# Add input files.
|
||||
cmd += " " + options.input
|
||||
|
||||
#
|
||||
# Run command and handle output.
|
||||
#
|
||||
try:
|
||||
#
|
||||
# Run command.
|
||||
#
|
||||
tmp_name = tempfile.NamedTemporaryFile( dir="." ).name
|
||||
tmp_stderr = open( tmp_name, 'wb' )
|
||||
proc = subprocess.Popen( args=cmd, shell=True, stderr=tmp_stderr.fileno() )
|
||||
returncode = proc.wait()
|
||||
tmp_stderr.close()
|
||||
|
||||
# Error checking.
|
||||
if returncode != 0:
|
||||
raise Exception, "return code = %i" % returncode
|
||||
|
||||
#
|
||||
# Handle output.
|
||||
#
|
||||
|
||||
# Read standard error to get total map/upper quartile mass.
|
||||
total_map_mass = -1
|
||||
tmp_stderr = open( tmp_name, 'r' )
|
||||
for line in tmp_stderr:
|
||||
if line.lower().find( "map mass" ) >= 0 or line.lower().find( "upper quartile" ) >= 0:
|
||||
total_map_mass = float( line.split(":")[1].strip() )
|
||||
break
|
||||
tmp_stderr.close()
|
||||
|
||||
#
|
||||
# If there's a global model provided, use model's total map mass
|
||||
# to adjust FPKM + confidence intervals.
|
||||
#
|
||||
if options.global_model_file:
|
||||
# Global model is simply total map mass from original run.
|
||||
global_model_file = open( options.global_model_file, 'r' )
|
||||
global_model_total_map_mass = float( global_model_file.readline() )
|
||||
global_model_file.close()
|
||||
|
||||
# Ratio of global model's total map mass to original run's map mass is
|
||||
# factor used to adjust FPKM.
|
||||
fpkm_map_mass_ratio = total_map_mass / global_model_total_map_mass
|
||||
|
||||
# Update FPKM values in transcripts.gtf file.
|
||||
transcripts_file = open( "transcripts.gtf", 'r' )
|
||||
tmp_transcripts = tempfile.NamedTemporaryFile( dir="." ).name
|
||||
new_transcripts_file = open( tmp_transcripts, 'w' )
|
||||
for line in transcripts_file:
|
||||
fields = line.split( '\t' )
|
||||
attrs = parse_gff_attributes( fields[8] )
|
||||
attrs[ "FPKM" ] = str( float( attrs[ "FPKM" ] ) * fpkm_map_mass_ratio )
|
||||
attrs[ "conf_lo" ] = str( float( attrs[ "conf_lo" ] ) * fpkm_map_mass_ratio )
|
||||
attrs[ "conf_hi" ] = str( float( attrs[ "conf_hi" ] ) * fpkm_map_mass_ratio )
|
||||
fields[8] = gff_attributes_to_str( attrs, "GTF" )
|
||||
new_transcripts_file.write( "%s\n" % '\t'.join( fields ) )
|
||||
transcripts_file.close()
|
||||
new_transcripts_file.close()
|
||||
shutil.copyfile( tmp_transcripts, "transcripts.gtf" )
|
||||
|
||||
# TODO: update expression files as well.
|
||||
|
||||
# Set outputs. Transcript and gene expression handled by wrapper directives.
|
||||
shutil.copyfile( "transcripts.gtf" , options.assembled_isoforms_output_file )
|
||||
if total_map_mass > -1:
|
||||
f = open( "global_model.txt", 'w' )
|
||||
f.write( "%f\n" % total_map_mass )
|
||||
f.close()
|
||||
except Exception, e:
|
||||
# Read stderr so that it can be reported:
|
||||
tmp_stderr = open( tmp_name, 'rb' )
|
||||
stderr = ''
|
||||
buffsize = 1048576
|
||||
try:
|
||||
while True:
|
||||
stderr += tmp_stderr.read( buffsize )
|
||||
if not stderr or len( stderr ) % buffsize != 0:
|
||||
break
|
||||
except OverflowError:
|
||||
pass
|
||||
tmp_stderr.close()
|
||||
|
||||
stop_err( 'Error running cufflinks.\n%s\n%s' % ( str( e ), stderr ) )
|
||||
|
||||
if __name__=="__main__": __main__()
|
||||
@@ -1,235 +0,0 @@
|
||||
<tool id="cufflinks" name="Cufflinks" version="0.0.6">
|
||||
<!-- Wrapper supports Cufflinks versions v1.3.0 and newer -->
|
||||
<description>transcript assembly and FPKM (RPKM) estimates for RNA-Seq data</description>
|
||||
<requirements>
|
||||
<requirement type="package">cufflinks</requirement>
|
||||
</requirements>
|
||||
<version_command>cufflinks 2>&1 | head -n 1</version_command>
|
||||
<command interpreter="python">
|
||||
cufflinks_wrapper.py
|
||||
--input=$input
|
||||
--assembled-isoforms-output=$assembled_isoforms
|
||||
--num-threads="4"
|
||||
-I $max_intron_len
|
||||
-F $min_isoform_fraction
|
||||
-j $pre_mrna_fraction
|
||||
$effective_length_correction
|
||||
|
||||
## Include reference annotation?
|
||||
#if $reference_annotation.use_ref == "Use reference annotation":
|
||||
-G $reference_annotation.reference_annotation_file
|
||||
#end if
|
||||
#if $reference_annotation.use_ref == "Use reference annotation guide":
|
||||
-g $reference_annotation.reference_annotation_guide_file
|
||||
#end if
|
||||
|
||||
## Normalization?
|
||||
#if str($do_normalization) == "Yes":
|
||||
-N
|
||||
#end if
|
||||
|
||||
## Bias correction?
|
||||
#if $bias_correction.do_bias_correction == "Yes":
|
||||
-b
|
||||
#if $bias_correction.seq_source.index_source == "history":
|
||||
--ref_file=$bias_correction.seq_source.ref_file
|
||||
#else:
|
||||
--ref_file="None"
|
||||
#end if
|
||||
--dbkey=${input.metadata.dbkey}
|
||||
--index_dir=${GALAXY_DATA_INDEX_DIR}
|
||||
#end if
|
||||
|
||||
## Multi-read correct?
|
||||
#if str($multiread_correct) == "Yes":
|
||||
-u
|
||||
#end if
|
||||
|
||||
## Include global model if available.
|
||||
#if $global_model:
|
||||
--global_model=$global_model
|
||||
#end if
|
||||
</command>
|
||||
<inputs>
|
||||
<param format="sam,bam" name="input" type="data" label="SAM or BAM file of aligned RNA-Seq reads" help=""/>
|
||||
<param name="max_intron_len" type="integer" value="300000" min="1" max="600000" label="Max Intron Length" help=""/>
|
||||
<param name="min_isoform_fraction" type="float" value="0.10" min="0" max="1" label="Min Isoform Fraction" help=""/>
|
||||
<param name="pre_mrna_fraction" type="float" value="0.15" min="0" max="1" label="Pre MRNA Fraction" help=""/>
|
||||
<param name="do_normalization" type="select" label="Perform quartile normalization" help="Removes top 25% of genes from FPKM denominator to improve accuracy of differential expression calls for low abundance transcripts.">
|
||||
<option value="No" selected="true">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<conditional name="reference_annotation">
|
||||
<param name="use_ref" type="select" label="Use Reference Annotation">
|
||||
<option value="No" selected="true">No</option>
|
||||
<option value="Use reference annotation">Use reference annotation</option>
|
||||
<option value="Use reference annotation guide">Use reference annotation as guide</option>
|
||||
</param>
|
||||
<when value="No"></when>
|
||||
<when value="Use reference annotation">
|
||||
<param format="gff3,gtf" name="reference_annotation_file" type="data" label="Reference Annotation" help="Gene annotation dataset in GTF or GFF3 format."/>
|
||||
</when>
|
||||
<when value="Use reference annotation guide">
|
||||
<param format="gff3,gtf" name="reference_annotation_guide_file" type="data" label="Reference Annotation" help="Gene annotation dataset in GTF or GFF3 format."/>
|
||||
</when>
|
||||
</conditional>
|
||||
<conditional name="bias_correction">
|
||||
<param name="do_bias_correction" type="select" label="Perform Bias Correction" help="Bias detection and correction can significantly improve accuracy of transcript abundance estimates.">
|
||||
<option value="No" selected="true">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<when value="Yes">
|
||||
<conditional name="seq_source">
|
||||
<param name="index_source" type="select" label="Reference sequence data">
|
||||
<option value="cached" selected="true">Locally cached</option>
|
||||
<option value="history">History</option>
|
||||
</param>
|
||||
<when value="cached"></when>
|
||||
<when value="history">
|
||||
<param name="ref_file" type="data" format="fasta" label="Using reference file" />
|
||||
</when>
|
||||
</conditional>
|
||||
</when>
|
||||
<when value="No"></when>
|
||||
</conditional>
|
||||
|
||||
<param name="multiread_correct" type="select" label="Use multi-read correct" help="Tells Cufflinks to do an initial estimation procedure to more accurately weight reads mapping to multiple locations in the genome.">
|
||||
<option value="No" selected="true">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
|
||||
<param name="effective_length_correction" type="select" label="Use effective length correction" help="Cufflinks will not employ its 'effective' length normalization to transcript FPKM.">
|
||||
<option value="" selected="true">Yes</option>
|
||||
<option value="--no-effective-length-correction">No</option>
|
||||
</param>
|
||||
|
||||
<param name="global_model" type="hidden_data" label="Global model (for use in Trackster)" optional="True"/>
|
||||
</inputs>
|
||||
|
||||
<outputs>
|
||||
<data format="tabular" name="genes_expression" label="${tool.name} on ${on_string}: gene expression" from_work_dir="genes.fpkm_tracking"/>
|
||||
<data format="tabular" name="transcripts_expression" label="${tool.name} on ${on_string}: transcript expression" from_work_dir="isoforms.fpkm_tracking"/>
|
||||
<data format="gtf" name="assembled_isoforms" label="${tool.name} on ${on_string}: assembled transcripts"/>
|
||||
<data format="txt" name="total_map_mass" label="${tool.name} on ${on_string}: total map mass" hidden="true" from_work_dir="global_model.txt"/>
|
||||
</outputs>
|
||||
|
||||
<trackster_conf>
|
||||
<action type="set_param" name="global_model" output_name="total_map_mass"/>
|
||||
</trackster_conf>
|
||||
|
||||
<tests>
|
||||
<!--
|
||||
Simple test that uses test data included with cufflinks.
|
||||
-->
|
||||
<test>
|
||||
<param name="input" value="cufflinks_in.bam"/>
|
||||
<param name="max_intron_len" value="300000"/>
|
||||
<param name="min_isoform_fraction" value="0.05"/>
|
||||
<param name="pre_mrna_fraction" value="0.05"/>
|
||||
<param name="use_ref" value="No"/>
|
||||
<param name="do_normalization" value="No" />
|
||||
<param name="do_bias_correction" value="No"/>
|
||||
<param name="multiread_correct" value="No"/>
|
||||
<param name="effective_length_correction" value="Yes"/>
|
||||
<output name="genes_expression" format="tabular" lines_diff="2" file="cufflinks_out3.fpkm_tracking"/>
|
||||
<output name="transcripts_expression" format="tabular" lines_diff="2" file="cufflinks_out2.fpkm_tracking"/>
|
||||
<output name="assembled_isoforms" file="cufflinks_out1.gtf"/>
|
||||
<output name="global_model" file="cufflinks_out4.txt"/>
|
||||
</test>
|
||||
</tests>
|
||||
|
||||
<help>
|
||||
**Cufflinks Overview**
|
||||
|
||||
Cufflinks_ assembles transcripts, estimates their abundances, and tests for differential expression and regulation in RNA-Seq samples. It accepts aligned RNA-Seq reads and assembles the alignments into a parsimonious set of transcripts. Cufflinks then estimates the relative abundances of these transcripts based on how many reads support each one. Please cite: Trapnell C, Williams BA, Pertea G, Mortazavi AM, Kwan G, van Baren MJ, Salzberg SL, Wold B, Pachter L. Transcript assembly and abundance estimation from RNA-Seq reveals thousands of new transcripts and switching among isoforms. Nature Biotechnology doi:10.1038/nbt.1621
|
||||
|
||||
.. _Cufflinks: http://cufflinks.cbcb.umd.edu/
|
||||
|
||||
------
|
||||
|
||||
**Know what you are doing**
|
||||
|
||||
.. class:: warningmark
|
||||
|
||||
There is no such thing (yet) as an automated gearshift in expression analysis. It is all like stick-shift driving in San Francisco. In other words, running this tool with default parameters will probably not give you meaningful results. A way to deal with this is to **understand** the parameters by carefully reading the `documentation`__ and experimenting. Fortunately, Galaxy makes experimenting easy.
|
||||
|
||||
.. __: http://cufflinks.cbcb.umd.edu/manual.html
|
||||
|
||||
------
|
||||
|
||||
**Input formats**
|
||||
|
||||
Cufflinks takes a text file of SAM alignments as input. The RNA-Seq read mapper TopHat produces output in this format, and is recommended for use with Cufflinks. However Cufflinks will accept SAM alignments generated by any read mapper. Here's an example of an alignment Cufflinks will accept::
|
||||
|
||||
s6.25mer.txt-913508 16 chr1 4482736 255 14M431N11M * 0 0 \
|
||||
CAAGATGCTAGGCAAGTCTTGGAAG IIIIIIIIIIIIIIIIIIIIIIIII NM:i:0 XS:A:-
|
||||
|
||||
Note the use of the custom tag XS. This attribute, which must have a value of "+" or "-", indicates which strand the RNA that produced this read came from. While this tag can be applied to any alignment, including unspliced ones, it must be present for all spliced alignment records (those with a 'N' operation in the CIGAR string).
|
||||
The SAM file supplied to Cufflinks must be sorted by reference position. If you aligned your reads with TopHat, your alignments will be properly sorted already. If you used another tool, you may want to make sure they are properly sorted as follows::
|
||||
|
||||
sort -k 3,3 -k 4,4n hits.sam > hits.sam.sorted
|
||||
|
||||
NOTE: Cufflinks currently only supports SAM alignments with the CIGAR match ('M') and reference skip ('N') operations. Support for the other operations, such as insertions, deletions, and clipping, will be added in the future.
|
||||
|
||||
------
|
||||
|
||||
**Outputs**
|
||||
|
||||
Cufflinks produces three output files:
|
||||
|
||||
Transcripts and Genes:
|
||||
|
||||
This GTF file contains Cufflinks' assembled isoforms. The first 7 columns are standard GTF, and the last column contains attributes, some of which are also standardized (e.g. gene_id, transcript_id). There one GTF record per row, and each record represents either a transcript or an exon within a transcript. The columns are defined as follows::
|
||||
|
||||
Column number Column name Example Description
|
||||
-----------------------------------------------------
|
||||
1 seqname chrX Chromosome or contig name
|
||||
2 source Cufflinks The name of the program that generated this file (always 'Cufflinks')
|
||||
3 feature exon The type of record (always either "transcript" or "exon").
|
||||
4 start 77696957 The leftmost coordinate of this record (where 0 is the leftmost possible coordinate)
|
||||
5 end 77712009 The rightmost coordinate of this record, inclusive.
|
||||
6 score 77712009 The most abundant isoform for each gene is assigned a score of 1000. Minor isoforms are scored by the ratio (minor FPKM/major FPKM)
|
||||
7 strand + Cufflinks' guess for which strand the isoform came from. Always one of '+', '-' '.'
|
||||
7 frame . Cufflinks does not predict where the start and stop codons (if any) are located within each transcript, so this field is not used.
|
||||
8 attributes See below
|
||||
|
||||
Each GTF record is decorated with the following attributes::
|
||||
|
||||
Attribute Example Description
|
||||
-----------------------------------------
|
||||
gene_id CUFF.1 Cufflinks gene id
|
||||
transcript_id CUFF.1.1 Cufflinks transcript id
|
||||
FPKM 101.267 Isoform-level relative abundance in Reads Per Kilobase of exon model per Million mapped reads
|
||||
frac 0.7647 Reserved. Please ignore, as this attribute may be deprecated in the future
|
||||
conf_lo 0.07 Lower bound of the 95% confidence interval of the abundance of this isoform, as a fraction of the isoform abundance. That is, lower bound = FPKM * (1.0 - conf_lo)
|
||||
conf_hi 0.1102 Upper bound of the 95% confidence interval of the abundance of this isoform, as a fraction of the isoform abundance. That is, upper bound = FPKM * (1.0 + conf_lo)
|
||||
cov 100.765 Estimate for the absolute depth of read coverage across the whole transcript
|
||||
|
||||
|
||||
Transcripts only:
|
||||
This file is simply a tab delimited file containing one row per transcript and with columns containing the attributes above. There are a few additional attributes not in the table above, but these are reserved for debugging, and may change or disappear in the future.
|
||||
|
||||
Genes only:
|
||||
This file contains gene-level coordinates and expression values.
|
||||
|
||||
-------
|
||||
|
||||
**Cufflinks settings**
|
||||
|
||||
All of the options have a default value. You can change any of them. Most of the options in Cufflinks have been implemented here.
|
||||
|
||||
------
|
||||
|
||||
**Cufflinks parameter list**
|
||||
|
||||
This is a list of implemented Cufflinks options::
|
||||
|
||||
-m INT This is the expected (mean) inner distance between mate pairs. For, example, for paired end runs with fragments selected at 300bp, where each end is 50bp, you should set -r to be 200. The default is 45bp.
|
||||
-s INT The standard deviation for the distribution on inner distances between mate pairs. The default is 20bp.
|
||||
-I INT The minimum intron length. Cufflinks will not report transcripts with introns longer than this, and will ignore SAM alignments with REF_SKIP CIGAR operations longer than this. The default is 300,000.
|
||||
-F After calculating isoform abundance for a gene, Cufflinks filters out transcripts that it believes are very low abundance, because isoforms expressed at extremely low levels often cannot reliably be assembled, and may even be artifacts of incompletely spliced precursors of processed transcripts. This parameter is also used to filter out introns that have far fewer spliced alignments supporting them. The default is 0.05, or 5% of the most abundant isoform (the major isoform) of the gene.
|
||||
-j Some RNA-Seq protocols produce a significant amount of reads that originate from incompletely spliced transcripts, and these reads can confound the assembly of fully spliced mRNAs. Cufflinks uses this parameter to filter out alignments that lie within the intronic intervals implied by the spliced alignments. The minimum depth of coverage in the intronic region covered by the alignment is divided by the number of spliced reads, and if the result is lower than this parameter value, the intronic alignments are ignored. The default is 5%.
|
||||
-G Tells Cufflinks to use the supplied reference annotation to estimate isoform expression. It will not assemble novel transcripts, and the program will ignore alignments not structurally compatible with any reference transcript.
|
||||
-N With this option, Cufflinks excludes the contribution of the top 25 percent most highly expressed genes from the number of mapped fragments used in the FPKM denominator. This can improve robustness of differential expression calls for less abundant genes and transcripts.
|
||||
</help>
|
||||
</tool>
|
||||
@@ -1,138 +0,0 @@
|
||||
#!/usr/bin/env python
|
||||
|
||||
# Supports Cuffmerge versions 1.3 and newer.
|
||||
|
||||
import optparse, os, shutil, subprocess, sys, tempfile
|
||||
|
||||
def stop_err( msg ):
|
||||
sys.stderr.write( '%s\n' % msg )
|
||||
sys.exit()
|
||||
|
||||
# Copied from sam_to_bam.py:
|
||||
def check_seq_file( dbkey, cached_seqs_pointer_file ):
|
||||
seq_path = ''
|
||||
for line in open( cached_seqs_pointer_file ):
|
||||
line = line.rstrip( '\r\n' )
|
||||
if line and not line.startswith( '#' ) and line.startswith( 'index' ):
|
||||
fields = line.split( '\t' )
|
||||
if len( fields ) < 3:
|
||||
continue
|
||||
if fields[1] == dbkey:
|
||||
seq_path = fields[2].strip()
|
||||
break
|
||||
return seq_path
|
||||
|
||||
def __main__():
|
||||
#Parse Command Line
|
||||
parser = optparse.OptionParser()
|
||||
parser.add_option( '-g', dest='ref_annotation', help='An optional "reference" annotation GTF. Each sample is matched against this file, and sample isoforms are tagged as overlapping, matching, or novel where appropriate. See the refmap and tmap output file descriptions below.' )
|
||||
parser.add_option( '-s', dest='use_seq_data', action="store_true", help='Causes cuffmerge to look into for fasta files with the underlying genomic sequences (one file per contig) against which your reads were aligned for some optional classification functions. For example, Cufflinks transcripts consisting mostly of lower-case bases are classified as repeats. Note that <seq_dir> must contain one fasta file per reference chromosome, and each file must be named after the chromosome, and have a .fa or .fasta extension.')
|
||||
parser.add_option( '-p', '--num-threads', dest='num_threads', help='Use this many threads to align reads. The default is 1.' )
|
||||
|
||||
|
||||
# Wrapper / Galaxy options.
|
||||
parser.add_option( '', '--dbkey', dest='dbkey', help='The build of the reference dataset' )
|
||||
parser.add_option( '', '--index_dir', dest='index_dir', help='GALAXY_DATA_INDEX_DIR' )
|
||||
parser.add_option( '', '--ref_file', dest='ref_file', help='The reference dataset from the history' )
|
||||
|
||||
# Outputs.
|
||||
parser.add_option( '', '--merged-transcripts', dest='merged_transcripts' )
|
||||
|
||||
(options, args) = parser.parse_args()
|
||||
|
||||
# output version # of tool
|
||||
try:
|
||||
tmp = tempfile.NamedTemporaryFile().name
|
||||
tmp_stdout = open( tmp, 'wb' )
|
||||
proc = subprocess.Popen( args='cuffmerge -v 2>&1', shell=True, stdout=tmp_stdout )
|
||||
tmp_stdout.close()
|
||||
returncode = proc.wait()
|
||||
stdout = None
|
||||
for line in open( tmp_stdout.name, 'rb' ):
|
||||
if line.lower().find( 'merge_cuff_asms v' ) >= 0:
|
||||
stdout = line.strip()
|
||||
break
|
||||
if stdout:
|
||||
sys.stdout.write( '%s\n' % stdout )
|
||||
else:
|
||||
raise Exception
|
||||
except:
|
||||
sys.stdout.write( 'Could not determine Cuffmerge version\n' )
|
||||
|
||||
# Set/link to sequence file.
|
||||
if options.use_seq_data:
|
||||
if options.ref_file != 'None':
|
||||
# Sequence data from history.
|
||||
# Create symbolic link to ref_file so that index will be created in working directory.
|
||||
seq_path = "ref.fa"
|
||||
os.symlink( options.ref_file, seq_path )
|
||||
else:
|
||||
# Sequence data from loc file.
|
||||
cached_seqs_pointer_file = os.path.join( options.index_dir, 'sam_fa_indices.loc' )
|
||||
if not os.path.exists( cached_seqs_pointer_file ):
|
||||
stop_err( 'The required file (%s) does not exist.' % cached_seqs_pointer_file )
|
||||
# If found for the dbkey, seq_path will look something like /galaxy/data/equCab2/sam_index/equCab2.fa,
|
||||
# and the equCab2.fa file will contain fasta sequences.
|
||||
seq_path = check_seq_file( options.dbkey, cached_seqs_pointer_file )
|
||||
if seq_path == '':
|
||||
stop_err( 'No sequence data found for dbkey %s, so sequence data cannot be used.' % options.dbkey )
|
||||
|
||||
# Build command.
|
||||
|
||||
# Base.
|
||||
cmd = "cuffmerge -o cm_output "
|
||||
|
||||
# Add options.
|
||||
if options.num_threads:
|
||||
cmd += ( " -p %i " % int ( options.num_threads ) )
|
||||
if options.ref_annotation:
|
||||
cmd += " -g %s " % options.ref_annotation
|
||||
if options.use_seq_data:
|
||||
cmd += " -s %s " % seq_path
|
||||
|
||||
# Add input files to a file.
|
||||
inputs_file_name = tempfile.NamedTemporaryFile( dir="." ).name
|
||||
inputs_file = open( inputs_file_name, 'w' )
|
||||
for arg in args:
|
||||
inputs_file.write( arg + "\n" )
|
||||
inputs_file.close()
|
||||
cmd += inputs_file_name
|
||||
|
||||
# Debugging.
|
||||
print cmd
|
||||
|
||||
# Run command.
|
||||
try:
|
||||
tmp_name = tempfile.NamedTemporaryFile( dir="." ).name
|
||||
tmp_stderr = open( tmp_name, 'wb' )
|
||||
proc = subprocess.Popen( args=cmd, shell=True, stderr=tmp_stderr.fileno() )
|
||||
returncode = proc.wait()
|
||||
tmp_stderr.close()
|
||||
|
||||
# Get stderr, allowing for case where it's very large.
|
||||
tmp_stderr = open( tmp_name, 'rb' )
|
||||
stderr = ''
|
||||
buffsize = 1048576
|
||||
try:
|
||||
while True:
|
||||
stderr += tmp_stderr.read( buffsize )
|
||||
if not stderr or len( stderr ) % buffsize != 0:
|
||||
break
|
||||
except OverflowError:
|
||||
pass
|
||||
tmp_stderr.close()
|
||||
|
||||
# Error checking.
|
||||
if returncode != 0:
|
||||
raise Exception, stderr
|
||||
|
||||
if len( open( "cm_output/merged.gtf", 'rb' ).read().strip() ) == 0:
|
||||
raise Exception, 'The output file is empty, there may be an error with your input file or settings.'
|
||||
|
||||
# Copy outputs.
|
||||
shutil.copyfile( "cm_output/merged.gtf" , options.merged_transcripts )
|
||||
|
||||
except Exception, e:
|
||||
stop_err( 'Error running cuffmerge. ' + str( e ) )
|
||||
|
||||
if __name__=="__main__": __main__()
|
||||
@@ -1,129 +0,0 @@
|
||||
<tool id="cuffmerge" name="Cuffmerge" version="0.0.5">
|
||||
<!-- Wrapper supports Cuffmerge versions 1.3 and newer -->
|
||||
<description>merge together several Cufflinks assemblies</description>
|
||||
<requirements>
|
||||
<requirement type="package">cufflinks</requirement>
|
||||
</requirements>
|
||||
<command interpreter="python">
|
||||
cuffmerge_wrapper.py
|
||||
|
||||
--num-threads="4"
|
||||
|
||||
## Use annotation reference?
|
||||
#if $annotation.use_ref_annotation == "Yes":
|
||||
-g $annotation.reference_annotation
|
||||
#end if
|
||||
|
||||
## Use sequence data?
|
||||
#if $seq_data.use_seq_data == "Yes":
|
||||
-s
|
||||
#if $seq_data.seq_source.index_source == "history":
|
||||
--ref_file=$seq_data.seq_source.ref_file
|
||||
#else:
|
||||
--ref_file="None"
|
||||
#end if
|
||||
--dbkey=${first_input.metadata.dbkey}
|
||||
--index_dir=${GALAXY_DATA_INDEX_DIR}
|
||||
#end if
|
||||
|
||||
## Outputs.
|
||||
--merged-transcripts=${merged_transcripts}
|
||||
|
||||
## Inputs.
|
||||
${first_input}
|
||||
#for $input_file in $input_files:
|
||||
${input_file.additional_input}
|
||||
#end for
|
||||
|
||||
</command>
|
||||
<inputs>
|
||||
<param format="gtf" name="first_input" type="data" label="GTF file produced by Cufflinks" help=""/>
|
||||
<repeat name="input_files" title="Additional GTF Input Files">
|
||||
<param format="gtf" name="additional_input" type="data" label="GTF file produced by Cufflinks" help=""/>
|
||||
</repeat>
|
||||
<conditional name="annotation">
|
||||
<param name="use_ref_annotation" type="select" label="Use Reference Annotation">
|
||||
<option value="No">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<when value="Yes">
|
||||
<param format="gff3,gtf" name="reference_annotation" type="data" label="Reference Annotation" help="Requires an annotation file in GFF3 or GTF format."/>
|
||||
</when>
|
||||
<when value="No">
|
||||
</when>
|
||||
</conditional>
|
||||
<conditional name="seq_data">
|
||||
<param name="use_seq_data" type="select" label="Use Sequence Data" help="Use sequence data for some optional classification functions, including the addition of the p_id attribute required by Cuffdiff.">
|
||||
<option value="No">No</option>
|
||||
<option value="Yes">Yes</option>
|
||||
</param>
|
||||
<when value="No"></when>
|
||||
<when value="Yes">
|
||||
<conditional name="seq_source">
|
||||
<param name="index_source" type="select" label="Choose the source for the reference list">
|
||||
<option value="cached">Locally cached</option>
|
||||
<option value="history">History</option>
|
||||
</param>
|
||||
<when value="cached"></when>
|
||||
<when value="history">
|
||||
<param name="ref_file" type="data" format="fasta" label="Using reference file" />
|
||||
</when>
|
||||
</conditional>
|
||||
</when>
|
||||
</conditional>
|
||||
</inputs>
|
||||
|
||||
<outputs>
|
||||
<data format="gtf" name="merged_transcripts" label="${tool.name} on ${on_string}: merged transcripts"/>
|
||||
</outputs>
|
||||
|
||||
<tests>
|
||||
<!--
|
||||
cuffmerge -g cuffcompare_in3.gtf cuffcompare_in1.gtf cuffcompare_in2.gtf
|
||||
-->
|
||||
<test>
|
||||
<param name="first_input" value="cuffcompare_in1.gtf" ftype="gtf"/>
|
||||
<param name="additional_input" value="cuffcompare_in2.gtf" ftype="gtf"/>
|
||||
<param name="use_ref_annotation" value="Yes"/>
|
||||
<param name="reference_annotation" value="cuffcompare_in3.gtf" ftype="gtf"/>
|
||||
<param name="use_seq_data" value="No"/>
|
||||
<!-- oId assignment differ/are non-deterministic -->
|
||||
<output name="merged_transcripts" file="cuffmerge_out1.gtf" lines_diff="50"/>
|
||||
</test>
|
||||
</tests>
|
||||
|
||||
<help>
|
||||
**Cuffmerge Overview**
|
||||
|
||||
Cuffmerge is part of Cufflinks_. Please cite: Trapnell C, Williams BA, Pertea G, Mortazavi AM, Kwan G, van Baren MJ, Salzberg SL, Wold B, Pachter L. Transcript assembly and abundance estimation from RNA-Seq reveals thousands of new transcripts and switching among isoforms. Nature Biotechnology doi:10.1038/nbt.1621
|
||||
|
||||
.. _Cufflinks: http://cufflinks.cbcb.umd.edu/
|
||||
|
||||
------
|
||||
|
||||
**Know what you are doing**
|
||||
|
||||
.. class:: warningmark
|
||||
|
||||
There is no such thing (yet) as an automated gearshift in expression analysis. It is all like stick-shift driving in San Francisco. In other words, running this tool with default parameters will probably not give you meaningful results. A way to deal with this is to **understand** the parameters by carefully reading the `documentation`__ and experimenting. Fortunately, Galaxy makes experimenting easy.
|
||||
|
||||
.. __: http://cufflinks.cbcb.umd.edu/manual.html#cuffmerge
|
||||
|
||||
------
|
||||
|
||||
**Input format**
|
||||
|
||||
Cuffmerge takes Cufflinks' GTF output as input, and optionally can take a "reference" annotation (such as from Ensembl_)
|
||||
|
||||
.. _Ensembl: http://www.ensembl.org
|
||||
|
||||
------
|
||||
|
||||
**Outputs**
|
||||
|
||||
Cuffmerge produces the following output files:
|
||||
|
||||
Merged transcripts file:
|
||||
|
||||
Cuffmerge produces a GTF file that contains an assembly that merges together the input assemblies. </help>
|
||||
</tool>
|
||||
Reference in New Issue
Block a user