From 0cd1db87a262240b3c48f47ea67b076248d49b8d Mon Sep 17 00:00:00 2001 From: Richard Burhans Date: Tue, 28 Sep 2010 12:38:52 -0400 Subject: [PATCH] Updates to help text for Human Genome Variation tools and formatHelp.html --- static/formatHelp.html | 342 +++++++++++------- tools/evolution/add_scores.xml | 56 +-- tools/evolution/codingSnps.xml | 78 ++-- tools/human_genome_variation/beam.xml | 65 ++-- tools/human_genome_variation/ctd.xml | 24 +- tools/human_genome_variation/funDo.xml | 39 +- tools/human_genome_variation/gpass.xml | 25 +- tools/human_genome_variation/hilbertvis.xml | 7 +- tools/human_genome_variation/ldtools.xml | 78 ++-- tools/human_genome_variation/linkToDavid.xml | 12 +- .../human_genome_variation/linkToGProfile.xml | 9 +- tools/human_genome_variation/lps.xml | 53 ++- tools/human_genome_variation/pass.xml | 85 +++-- tools/human_genome_variation/sift.xml | 57 +-- tools/human_genome_variation/snpFreq.xml | 43 +-- 15 files changed, 568 insertions(+), 405 deletions(-) diff --git a/static/formatHelp.html b/static/formatHelp.html index 8f028538987..36bf5e098e7 100644 --- a/static/formatHelp.html +++ b/static/formatHelp.html @@ -1,51 +1,68 @@ -Galaxy data formats +Galaxy Data Formats -

Galaxy data formats

+

Galaxy Data Formats

-For problems with missing queries in the tool selection boxes the most common -reason is the tool only lists history items with data formats compatible with the -tool. Some formats are subsets of others and Galaxy should also list those -with compatible subformats as well. If the query is not showing up still and you -believe it is in the correct format you can click on the pencil icon and -manually change the format. This will not edit the file just change the -metadata for the file. Some cases you will need to actually change the -file format. For example, if the file is space delimited and a -tabular file is required; then the "Convert delimiters to TAB" tool under "Text -Manipulation" can be used to reformat the file. +
+ +

Dataset missing?

+

+If you have a dataset in your history that is not appearing in the +drop-down selector for a tool, the most common reason is that it has +the wrong format. Each Galaxy dataset has an associated file format +recorded in its metadata, and tools will only list datasets from your +history that have a format compatible with that particular tool. Of +course some of these datasets might not actually contain relevant +data, or even the correct columns needed by the tool, but filtering +by format at least makes the list to select from a bit shorter. +

+Some of the formats are defined hierarchically, going from very +general ones like tabular (which includes any text +file with tab-separated columns), to more restrictive sub-formats +like interval (where three of the columns +must be the chromosome, start position, and end position), and on +to even more specific ones such as BED or +GFF that have additional requirements. So for +example if a tool's required input format is tabular, then all of +your history items whose format is recorded as tabular will be +listed, along with those in all sub-formats that also qualify as +tabular (interval, BED, GFF, etc.). +

+There are two usual methods for changing a dataset's format in +Galaxy: if the file contents are already in the required format but +the metadata is wrong (perhaps because the Auto-detect feature of the +Upload File tool guessed it incorrectly), you can fix the metadata +manually by clicking on the pencil icon beside that dataset in your +history. Or, if the file contents really are in a different format, +Galaxy provides a number of format conversion tools (e.g. in the +Text Manipulation and Convert Formats categories). For instance, +if the tool you want to run requires tabular but your columns are +delimited by spaces or commas, you can use the "Convert delimiters +to TAB" tool under Text Manipulation to reformat your data. However +if your files are in a completely unsupported format, then you need +to convert them yourself before uploading.

-Some of the most commonly used formats are very similar. Start with the -basic tabular file. It has few requirements other than 1 or more columns of -data separated by tabs. Next is intervals which are tabular but they have the -added requirement that 3 of the columns must be the chromosome, start point, -and end point. There is optionally a strand and header labelling the -columns. Next is BED or GFF, which are also tabular and intervals, but with more -restrictions. BED can vary between 3 and 12 columns, with each being -precisely defined. Here -the order of the columns also matters, and only the end columns can be skipped. -Some groups of the columns have to be all there or all left off. GFF is -similar in setup but with all 9 columns required and different definitions. -See more detailed descriptions below.


-

Formats

+ +

Format Descriptions

-
+


+ Ab1

-A binary sequence file in 'ab1' format with a '.ab1' file extension. You must manually select this 'File Format' when uploading the file. - +A binary sequence file in 'ab1' format with a '.ab1' file extension. +You must manually select this file format when uploading the file.


AXT

-blastz pairwise alignment format. Each alignment block in an axt file contains three lines: a summary line and 2 sequence lines. Blocks are separated from one another by blank lines. The summary line contains chromosomal position and size information about the alignment. It consists of 9 required fields. -Click here for more information about axt format. +Used for pairwise alignment output from BLASTZ, after post-processing. +Each alignment block contains three lines: a summary line and two +sequence lines. Blocks are separated from one another by blank lines. +The summary line contains chromosomal position and size information +about the alignment, and consists of nine required fields. +More information

Can be converted to:
  • FASTA
    @@ -80,12 +103,13 @@ Convert Formats→AXT to LAV

-Bam +BAM

-A binary file compressed in the BGZF format with a '.bam' file extension. -SAM -format is the human readable text version of these files. +A binary file compressed in the BGZF format with a '.bam' file +extension. +SAM format +is the human readable text version of these files.

Can be converted to:
  • pileup
    @@ -96,23 +120,21 @@ NGS: SAM Tools→Pileup-to-Interval

-Binseq.zip - -

-A zipped archive consisting of binary sequence files in either 'ab1' or 'scf' format. All files in this archive must have the same file extension which is one of '.ab1' or '.scf'. You must manually select this 'File Format' when uploading the file. - -


- BED

-This describes a genomic interval, but has strict field specifications for use -in browsers. -
Click here for field specifications.
+This tab-separated format describes a genomic interval, but has +strict field specifications for use in genome browsers. BED files +can have from 3 to 12 columns, but the order of the columns matters, +and only the end ones can be omitted. Some groups of columns must +be all present or all absent. +Field specifications +

Example:

 chr22 1000 5000 cloneA 960 + 1000 5000 0 2 567,488, 0,3512
@@ -129,22 +151,35 @@ Convert Formats→BED-to-GFF
 
 

    -
  • also tabular -
  • also interval -
  • also BED +
  • also qualifies as tabular +
  • also qualifies as interval +
  • also qualifies as BED
-
BedGraph -is a BED file with the name column being a float value that is displayed -as a Wiggle in tracks. Unlike wiggles this score can be retrieved in its -exact value after being loaded as a track. +BedGraph is a BED file with the name column being a float value +that is displayed as a wiggle score in tracks. Unlike in Wiggle +format, the exact value of this score can be retrieved after being +loaded as a track.
-Fasta +Binseq.zip + +

+A zipped archive consisting of binary sequence files in either +'ab1' or 'scf' format. All files in this archive must have the same +file extension which is one of '.ab1' or '.scf'. You must manually +select this file format when uploading the file. +


+ +FASTA

A sequence in -FASTA format -consists of a single-line description, followed by lines of sequence data. The first character of the description line is a greater-than (">") symbol in the first column. All lines should be shorter than 80 characters:: +FASTA +format consists of a single-line description, followed by lines of +sequence data. The first character of the description line is a +greater-than (">") symbol. All lines should be shorter than 80 +characters.

 >sequence1
 atgcgtttgcgtgc
@@ -164,7 +199,8 @@ Convert Formats→FASTA-to-Tabular
 
 

FastqSolexa -is the Illumina (Solexa) variant of the Fastq format, which stores sequences and quality scores in a single file +is the Illumina (Solexa) variant of the Fastq format, which stores +sequences and quality scores in a single file.

 @seq1  
 GACAGCTTGGTTTTTAGTGAGTTGTTCCTTTCTTT  
@@ -174,9 +210,9 @@ hhhhhhhhhhhhhhhhhhhhhhhhhhPW@hhhhhh
 GCAATGACGGCAGCAATAAACTCAACAGGTGCTGG  
 +seq2  
 hhhhhhhhhhhhhhYhhahhhhWhAhFhSIJGChO
-
+
Or - +
 @seq1
 GAATTGATCAGGACATAGGACAACTGTAGGCACCAT
 +seq1
@@ -196,19 +232,22 @@ Convert Formats→FASTQ to FASTA
 fped
 
 

-Also known as the FBAT format, for use in the FBAT program. -It consists of a pedigree file and an phenotype file. +Also known as the FBAT format, for use with the +FBAT program. +It consists of a pedigree file and a phenotype file.


-Gff +GFF

    -
  • also tabular -
  • also interval +
  • also qualifies as tabular +
  • also qualifies as interval
-GFF lines have -
nine required fields that must be tab-separated. +GFF is a tab-separated format somewhat similar to BED, but it has +different columns and is more flexible. There are +nine required fields.
Can be converted to: