diff --git a/static/formatHelp.html b/static/formatHelp.html
index 8f028538987..36bf5e098e7 100644
--- a/static/formatHelp.html
+++ b/static/formatHelp.html
@@ -1,51 +1,68 @@
-
Galaxy data formats
+Galaxy Data Formats
-Galaxy data formats
+Galaxy Data Formats
-For problems with missing queries in the tool selection boxes the most common
-reason is the tool only lists history items with data formats compatible with the
-tool. Some formats are subsets of others and Galaxy should also list those
-with compatible subformats as well. If the query is not showing up still and you
-believe it is in the correct format you can click on the pencil icon and
-manually change the format. This will not edit the file just change the
-metadata for the file. Some cases you will need to actually change the
-file format. For example, if the file is space delimited and a
-tabular file is required; then the "Convert delimiters to TAB" tool under "Text
-Manipulation" can be used to reformat the file.
+
+
+
Dataset missing?
+
+If you have a dataset in your history that is not appearing in the
+drop-down selector for a tool, the most common reason is that it has
+the wrong format. Each Galaxy dataset has an associated file format
+recorded in its metadata, and tools will only list datasets from your
+history that have a format compatible with that particular tool. Of
+course some of these datasets might not actually contain relevant
+data, or even the correct columns needed by the tool, but filtering
+by format at least makes the list to select from a bit shorter.
+
+Some of the formats are defined hierarchically, going from very
+general ones like tabular (which includes any text
+file with tab-separated columns), to more restrictive sub-formats
+like interval (where three of the columns
+must be the chromosome, start position, and end position), and on
+to even more specific ones such as BED or
+GFF that have additional requirements. So for
+example if a tool's required input format is tabular, then all of
+your history items whose format is recorded as tabular will be
+listed, along with those in all sub-formats that also qualify as
+tabular (interval, BED, GFF, etc.).
+
+There are two usual methods for changing a dataset's format in
+Galaxy: if the file contents are already in the required format but
+the metadata is wrong (perhaps because the Auto-detect feature of the
+Upload File tool guessed it incorrectly), you can fix the metadata
+manually by clicking on the pencil icon beside that dataset in your
+history. Or, if the file contents really are in a different format,
+Galaxy provides a number of format conversion tools (e.g. in the
+Text Manipulation and Convert Formats categories). For instance,
+if the tool you want to run requires tabular but your columns are
+delimited by spaces or commas, you can use the "Convert delimiters
+to TAB" tool under Text Manipulation to reformat your data. However
+if your files are in a completely unsupported format, then you need
+to convert them yourself before uploading.
-Some of the most commonly used formats are very similar. Start with the
-basic tabular file. It has few requirements other than 1 or more columns of
-data separated by tabs. Next is intervals which are tabular but they have the
-added requirement that 3 of the columns must be the chromosome, start point,
-and end point. There is optionally a strand and header labelling the
-columns. Next is BED or GFF, which are also tabular and intervals, but with more
-restrictions. BED can vary between 3 and 12 columns, with each being
-precisely defined. Here
-the order of the columns also matters, and only the end columns can be skipped.
-Some groups of the columns have to be all there or all left off. GFF is
-similar in setup but with all 9 columns required and different definitions.
-See more detailed descriptions below.
-Formats
+
+Format Descriptions
-
+
+
Ab1
-A binary sequence file in 'ab1' format with a '.ab1' file extension. You must manually select this 'File Format' when uploading the file.
-
+A binary sequence file in 'ab1' format with a '.ab1' file extension.
+You must manually select this file format when uploading the file.
AXT
-blastz pairwise alignment format. Each alignment block in an axt file contains three lines: a summary line and 2 sequence lines. Blocks are separated from one another by blank lines. The summary line contains chromosomal position and size information about the alignment. It consists of 9 required fields.
-Click here for more information about axt format.
+Used for pairwise alignment output from BLASTZ, after post-processing.
+Each alignment block contains three lines: a summary line and two
+sequence lines. Blocks are separated from one another by blank lines.
+The summary line contains chromosomal position and size information
+about the alignment, and consists of nine required fields.
+More information
- Can be converted to:
- FASTA
@@ -80,12 +103,13 @@ Convert Formats→AXT to LAV
-Bam
+BAM
-A binary file compressed in the BGZF format with a '.bam' file extension.
-SAM
-format is the human readable text version of these files.
+A binary file compressed in the BGZF format with a '.bam' file
+extension.
+SAM format
+is the human readable text version of these files.
- Can be converted to:
- pileup
@@ -96,23 +120,21 @@ NGS: SAM Tools→Pileup-to-Interval
-Binseq.zip
-
-
-A zipped archive consisting of binary sequence files in either 'ab1' or 'scf' format. All files in this archive must have the same file extension which is one of '.ab1' or '.scf'. You must manually select this 'File Format' when uploading the file.
-
-
-
BED
-- also tabular
-
- also interval
+
- also qualifies as tabular
+
- also qualifies as interval
-This describes a genomic interval, but has strict field specifications for use
-in browsers.
-Click here for field specifications.
+This tab-separated format describes a genomic interval, but has
+strict field specifications for use in genome browsers. BED files
+can have from 3 to 12 columns, but the order of the columns matters,
+and only the end ones can be omitted. Some groups of columns must
+be all present or all absent.
+Field specifications
+
Example:
chr22 1000 5000 cloneA 960 + 1000 5000 0 2 567,488, 0,3512
@@ -129,22 +151,35 @@ Convert Formats→BED-to-GFF
-- also tabular
-
- also interval
-
- also BED
+
- also qualifies as tabular
+
- also qualifies as interval
+
- also qualifies as BED
-BedGraph
-is a BED file with the name column being a float value that is displayed
-as a Wiggle in tracks. Unlike wiggles this score can be retrieved in its
-exact value after being loaded as a track.
+BedGraph is a BED file with the name column being a float value
+that is displayed as a wiggle score in tracks. Unlike in Wiggle
+format, the exact value of this score can be retrieved after being
+loaded as a track.
-Fasta
+Binseq.zip
+
+
+A zipped archive consisting of binary sequence files in either
+'ab1' or 'scf' format. All files in this archive must have the same
+file extension which is one of '.ab1' or '.scf'. You must manually
+select this file format when uploading the file.
+
+
+FASTA
A sequence in
-FASTA format
-consists of a single-line description, followed by lines of sequence data. The first character of the description line is a greater-than (">") symbol in the first column. All lines should be shorter than 80 characters::
+FASTA
+format consists of a single-line description, followed by lines of
+sequence data. The first character of the description line is a
+greater-than (">") symbol. All lines should be shorter than 80
+characters.
>sequence1
atgcgtttgcgtgc
@@ -164,7 +199,8 @@ Convert Formats→FASTA-to-Tabular
FastqSolexa
-is the Illumina (Solexa) variant of the Fastq format, which stores sequences and quality scores in a single file
+is the Illumina (Solexa) variant of the Fastq format, which stores
+sequences and quality scores in a single file.
@seq1
GACAGCTTGGTTTTTAGTGAGTTGTTCCTTTCTTT
@@ -174,9 +210,9 @@ hhhhhhhhhhhhhhhhhhhhhhhhhhPW@hhhhhh
GCAATGACGGCAGCAATAAACTCAACAGGTGCTGG
+seq2
hhhhhhhhhhhhhhYhhahhhhWhAhFhSIJGChO
-
+
Or
-
+
@seq1
GAATTGATCAGGACATAGGACAACTGTAGGCACCAT
+seq1
@@ -196,19 +232,22 @@ Convert Formats→FASTQ to FASTA
fped
-Also known as the FBAT format, for use in the FBAT program.
-It consists of a pedigree file and an phenotype file.
+Also known as the FBAT format, for use with the
+FBAT program.
+It consists of a pedigree file and a phenotype file.
-Gff
+GFF
-- also tabular
-
- also interval
+
- also qualifies as tabular
+
- also qualifies as interval
-GFF lines have
-nine required fields that must be tab-separated.
+GFF is a tab-separated format somewhat similar to BED, but it has
+different columns and is more flexible. There are
+nine required fields.
- Can be converted to:
- BED
@@ -220,66 +259,80 @@ Convert Formats→GFF-to-BED
-- also tabular
-
- also interval
+
- also qualifies as tabular
+
- also qualifies as interval
-The
-GFF3
-format addresses the most common extensions to GFF, while preserving backward compatibility with previous formats.
+The GFF3
+format addresses the most common extensions to GFF, while preserving
+backward compatibility with previous formats.
GTF
-- also tabular
-
- also interval
+
- also qualifies as tabular
+
- also qualifies as interval
-GTF
-is a format for describing genes and other features associated with DNA, RNA and Protein sequences.
+GTF is a format for describing genes and other features
+associated with DNA, RNA, and protein sequences.
- Can be converted to:
-- BED graph
+ - BedGraph
Convert Formats→GTF-to-BEDGraph
-Html
+HTML
-This format is a html web page. Click the eye icon to view the dataset in
-your browser.
+This format is an HTML web page. Click the eye icon next to the
+dataset to view it in your browser.
-Interval (Genomic Intervals)
+Interval
-- also tabular
+
- also qualifies as tabular
-Required fields
+This Galaxy format represents genomic intervals. It is tab-separated,
+but has the added requirement that three of the columns must be the
+chromosome name, start position, and end position. An optional
+strand column can also be specified, and an initial header row can
+be used to label the columns, which do not have to be in any special
+order. Arbitrary additional columns can also be present.
+
+Required fields:
-- CHROM - The name of the chromosome (e.g. chr3, chrY, chr2_random) or contig (e.g. ctgY1).
-
- START - The starting position of the feature in the chromosome or contig. The first base in a chromosome is numbered 0.
-
- END - The ending position of the feature in the chromosome or contig. The chromEnd base is not included in the display of the feature. For example, the first 100 bases of a chromosome are defined as chromStart=0, chromEnd=100, and span the bases numbered 0-99.
+
- CHROM - The name of the chromosome (e.g. chr3, chrY, chr2_random)
+ or contig (e.g. ctgY1).
+
- START - The starting position of the feature in the chromosome or
+ contig. The first base in a chromosome is numbered 0.
+
- END - The ending position of the feature in the chromosome or
+ contig. This base is not included in the feature. For example,
+ the first 100 bases of a chromosome are described as START=0,
+ END=100, and span the bases numbered 0-99.
-Optional
+Optional:
-- STRAND - Defines the strand - either '+' or '-'.
-
- Headers
+
- STRAND - Defines the strand, either '+' or '-'.
+
- Header row
Example:
- #CHROM START END STRAND NAME COMMENT
- chr1 10 100 + exon myExon
- chrX 1000 10050 - gene myGene
+ #CHROM START END STRAND NAME COMMENT
+ chr1 10 100 + exon myExon
+ chrX 1000 10050 - gene myGene
- Can be converted to:
- BED
-The exact changes needed and tools to run can vary with what fields are in
-the interval file and what size BED you are converting to. In general you
-will likely use Text Manipulation→Compute, Cut or Merge Columns.
+The exact changes needed and tools to run will vary with what fields
+are in the interval file and what type of BED you are converting to.
+In general you will likely use Text Manipulation→Compute, Cut,
+or Merge Columns.
@@ -287,7 +340,8 @@ will likely use Text Manipulation→Compute, Cut or Merge Columns.
LAV
-is the primary output format for BLASTZ. The first line of a .lav file begins with #:lav..
+is the raw pairwise alignment format that is output by BLASTZ. The
+first line begins with #:lav.
- Can be converted to:
- BED
@@ -295,18 +349,20 @@ Convert Formats→LAV to BED
-Lped
+lped
-This is the linkage pedigree format (separate map and ped files).
-These files together describe SNPs, the map file has the position and an
-identifier for the SNP and the pedigree file has the alleles.
-To upload this format into Galaxy do not use auto-detect for the file
-format, instead select lped. You will then be given two sections for uploading
-files, one for the pedigree file and one for the map file.
-For more information see linkage pedigree
-or map
-or ped.
+This is the linkage pedigree format, which consists of separate
+map and ped files. Together these files
+describe SNPs; the map file contains the position and an identifier
+for the SNP, while the pedigree file has the alleles.
+To upload this format into Galaxy, do not use auto-detect for the
+file format; instead select lped. You will then be
+given two sections for uploading files, one for the pedigree file
+and one for the map file. For more information, see
+linkage pedigree,
+map,
+and/or ped.
- Can be converted to:
- pbed
Automatic
@@ -317,16 +373,20 @@ or ped.
MAF
-TBA and multiz multiple alignment format. The first line of a .maf file begins with ##maf. This word is followed by white-space-separated "variable=value pairs". There should be no white space surrounding the "=".
-Click here for more about MAF format.
+Multiple alignment format that is output by TBA and Multiz. The
+first line begins with ##maf. This word is followed by
+whitespace-separated "variable=value pairs". There should be no
+whitespace surrounding the "=".
+More information
- Can be converted to:
- BED
-Convert Formats→Maf to BED
- - Interval
-Convert Formats→Maf to Interval
+Convert Formats→MAF to BED
+ - interval
+Convert Formats→MAF to Interval
- FASTA
-Convert Formats→Maf to FASTA
+Convert Formats→MAF to FASTA
@@ -344,7 +404,7 @@ This is the binary version of the lped file format.
PSL
-format is for alignments, it is returned by
+format is used for alignments returned by
BLAT.
It does not include any sequence.
@@ -352,9 +412,10 @@ It does not include any sequence.
Scf
-A binary sequence file in 'scf' format with a '.scf' file extension. You must manually select this 'File Format' when uploading the file.
-Click here
-for more information.
+A binary sequence file in 'scf' format with a '.scf' file extension.
+You must manually select this file format when uploading the file.
+More information
Sff
@@ -373,58 +434,65 @@ Convert Formats→SFF converter
Table
-Text delimited into columns by something other than a tab.
+Text data separated into columns by something other than tabs.
-Tabular (tab delimited)
+Tabular (tab-delimited)
-Any data in tab delimited format (tabular)
+One or more columns of text data separated by tabs.
- Can be converted to:
- FASTA
Convert Formats→Tabular-to-FASTA
-Tabular file must have a title and sequence column.
+The tabular file must have a title and sequence column.
- interval
-If the tabular file has the chromosome, or is all on one chromosome, and
-a position you can create an interval file. If all one chromosome use
-Text Manipulation→Add column to add the chromosome. If the given position
-is a 1 based position use Text Manipulation→Compute and the position
-column minus 1 to get the start. Otherwise do plus 1 to get the end.
+If the tabular file has the chromosome, or is all on one chromosome,
+and has a position you can create an interval file (e.g. for SNPs).
+If it is all on one chromosome, use Text Manipulation→Add column
+to add a chromosome column. If the given position is 1-based, use
+Text Manipulation→Compute with the position column minus 1 to
+get the start, and use the original given column for the end.
+If the given position is 0-based, use it as the start, and compute
+that plus 1 to get the end.
Txtseq.zip
-A zipped archive consisting of flat text sequence files. All files in this archive must have the same file extension of '.txt'. You must manually select this 'File Format' when uploading the file.
-
+A zipped archive consisting of flat text sequence files. All files
+in this archive must have the same file extension of '.txt'. You
+must manually select this file format when uploading the file.
Wiggle custom track
-The wiggle format is line-oriented. Wiggle data is preceded by a track definition line, which gives the type of wiggle. There are 3 different types, each
-with their uses.
-More information here.
+The wiggle format is line-oriented. Wiggle data is preceded by a
+track definition line, which specifies the type of wiggle. There
+are three different types, for different uses.
+More information
- Can be converted to:
- interval
Convert Formats→Wiggle-to-Interval
-As a second step this could be converted to BED 3 or 4 by removing columns.
-Text Manipulation→Cut columns from a table
+As a second step this could be converted to BED-3 or BED-4 by removing
+columns, using Text Manipulation→Cut columns from a table.
Other text type
-Any text file
+Any text file.
- Can be converted to:
- tabular
-If this is space or some other delimiter separated fields it can be
-converted to tabular. Text Manipulations→Convert delimiters to TAB
+If this has fields separated by spaces, commas, or some other
+delimiter it can be converted to tabular using
+Text Manipulation→Convert delimiters to TAB
diff --git a/tools/evolution/add_scores.xml b/tools/evolution/add_scores.xml
index e4c6b5e9ac4..de272cf3607 100644
--- a/tools/evolution/add_scores.xml
+++ b/tools/evolution/add_scores.xml
@@ -40,10 +40,11 @@ This currently works only for build hg18.
**Dataset formats**
-The input dataset can be any interval_ format dataset.
-The output dataset is also interval format.
+The input can be any interval_ format dataset. The output is also in interval format.
+(`Dataset missing?`_)
.. _interval: ./static/formatHelp.html#interval
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -64,41 +65,42 @@ only the score for the first nucleotide is returned.
- input file, with SNPs::
- chr22 16440426 14440427 C/T
- chr22 15494851 14494852 A/G
- chr22 14494911 14494912 A/T
- chr22 14550435 14550436 A/G
- chr22 14611956 14611957 G/T
- chr22 14612076 14612077 A/G
- chr22 14668537 14668538 C
- chr22 14668703 14668704 A/T
- chr22 14668775 14668776 G
- chr22 14680074 14680075 A/T
+ chr22 16440426 14440427 C/T
+ chr22 15494851 14494852 A/G
+ chr22 14494911 14494912 A/T
+ chr22 14550435 14550436 A/G
+ chr22 14611956 14611957 G/T
+ chr22 14612076 14612077 A/G
+ chr22 14668537 14668538 C
+ chr22 14668703 14668704 A/T
+ chr22 14668775 14668776 G
+ chr22 14680074 14680075 A/T
etc.
- output file, showing conservation scores for primates::
- chr22 16440426 14440427 C/T 0.509
- chr22 15494851 14494852 A/G 0.427
- chr22 14494911 14494912 A/T NA
- chr22 14550435 14550436 A/G NA
- chr22 14611956 14611957 G/T -2.142
- chr22 14612076 14612077 A/G 0.369
- chr22 14668537 14668538 C 0.419
- chr22 14668703 14668704 A/T -1.462
- chr22 14668775 14668776 G 0.470
- chr22 14680074 14680075 A/T 0.303
+ chr22 16440426 14440427 C/T 0.509
+ chr22 15494851 14494852 A/G 0.427
+ chr22 14494911 14494912 A/T NA
+ chr22 14550435 14550436 A/G NA
+ chr22 14611956 14611957 G/T -2.142
+ chr22 14612076 14612077 A/G 0.369
+ chr22 14668537 14668538 C 0.419
+ chr22 14668703 14668704 A/T -1.462
+ chr22 14668775 14668776 G 0.470
+ chr22 14680074 14680075 A/T 0.303
etc.
-"NA" means that the phyloP score was not available.
+ "NA" means that the phyloP score was not available.
-----
**Reference**
-Siepel A, Pollard KS, and Haussler D. New methods for detecting
-lineage-specific selection. In Proceedings of the 10th International
-Conference on Research in Computational Molecular Biology (RECOMB
-2006), pp. 190-205.
+Siepel A, Pollard KS, Haussler D. (2006)
+New methods for detecting lineage-specific selection.
+In Proceedings of the 10th International Conference on Research in Computational
+Molecular Biology (RECOMB 2006), pp. 190-205.
+
diff --git a/tools/evolution/codingSnps.xml b/tools/evolution/codingSnps.xml
index 6c4fd8db8a5..3fde8063af5 100644
--- a/tools/evolution/codingSnps.xml
+++ b/tools/evolution/codingSnps.xml
@@ -53,12 +53,12 @@ Use the pencil icon to add the build to the files if necessary.
**Dataset formats**
The SNP dataset is in interval_ format, with a column of SNPs as described below.
-The gene dataset is in BED_ format with 12 columns.
-The output dataset is also interval.
+The gene dataset is in BED_ format with 12 columns. The output dataset is also interval.
+(`Dataset missing?`_)
.. _interval: ./static/formatHelp.html#interval
-
.. _BED: ./static/formatHelp.html#bed
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -87,51 +87,53 @@ or is synonymous then it is not included in the output file.
- first input file, with SNPs::
- chr22 15660821 15660822 A/G
- chr22 15825725 15825726 G/T
- chr22 15827035 15827036 G
- chr22 15827135 15827136 C/G
- chr22 15830928 15830929 A/G
- chr22 15830951 15830952 G
- chr22 15830955 15830956 C/T
- chr22 15848885 15848886 C/T
- chr22 15849048 15849049 A/C
- chr22 15919711 15919712 A/G
+ chr22 15660821 15660822 A/G
+ chr22 15825725 15825726 G/T
+ chr22 15827035 15827036 G
+ chr22 15827135 15827136 C/G
+ chr22 15830928 15830929 A/G
+ chr22 15830951 15830952 G
+ chr22 15830955 15830956 C/T
+ chr22 15848885 15848886 C/T
+ chr22 15849048 15849049 A/C
+ chr22 15919711 15919712 A/G
etc.
- alternatively indicating polymorphisms using ambiguous-nucleotide symbols:
- chr22 15660821 15660822 R
- chr22 15825725 15825726 K
- chr22 15827035 15827036 G
- chr22 15827135 15827136 S
- chr22 15830928 15830929 R
- chr22 15830951 15830952 G
- chr22 15830955 15830956 Y
- chr22 15848885 15848886 Y
- chr22 15849048 15849049 M
- chr22 15919711 15919712 R
+ or, indicating polymorphisms using ambiguous-nucleotide symbols::
+
+ chr22 15660821 15660822 R
+ chr22 15825725 15825726 K
+ chr22 15827035 15827036 G
+ chr22 15827135 15827136 S
+ chr22 15830928 15830929 R
+ chr22 15830951 15830952 G
+ chr22 15830955 15830956 Y
+ chr22 15848885 15848886 Y
+ chr22 15849048 15849049 M
+ chr22 15919711 15919712 R
etc.
- second input file, with UCSC annotations for human genes::
- chr22 15688363 15690225 uc010gqr.1 0 + 15688363 15688363 0 2 587,794, 0,1068,
- chr22 15822826 15869112 uc002zlw.1 0 - 15823622 15869004 0 10 940,105,97,91,265,86,251,208,304,282, 0,1788,2829,3241,4163,6361,8006,26023,29936,46004,
- chr22 15826991 15869112 uc010gqs.1 0 - 15829218 15869004 0 5 1380,86,157,304,282, 0,2196,21858,25771,41839,
- chr22 15897459 15919682 uc002zlx.1 0 + 15897459 15897459 0 4 775,128,103,1720, 0,8303,10754,20503,
- chr22 15945848 15971389 uc002zly.1 0 + 15945981 15970710 0 13 271,25,147,113,127,48,164,84,85,12,102,42,2193, 0,12103,12838,13816,15396,17037,17180,18535,19767,20632,20894,22768,23348,
+ chr22 15688363 15690225 uc010gqr.1 0 + 15688363 15688363 0 2 587,794, 0,1068,
+ chr22 15822826 15869112 uc002zlw.1 0 - 15823622 15869004 0 10 940,105,97,91,265,86,251,208,304,282, 0,1788,2829,3241,4163,6361,8006,26023,29936,46004,
+ chr22 15826991 15869112 uc010gqs.1 0 - 15829218 15869004 0 5 1380,86,157,304,282, 0,2196,21858,25771,41839,
+ chr22 15897459 15919682 uc002zlx.1 0 + 15897459 15897459 0 4 775,128,103,1720, 0,8303,10754,20503,
+ chr22 15945848 15971389 uc002zly.1 0 + 15945981 15970710 0 13 271,25,147,113,127,48,164,84,85,12,102,42,2193, 0,12103,12838,13816,15396,17037,17180,18535,19767,20632,20894,22768,23348,
etc.
- output file, showing non-synonymous substitutions in coding regions::
- chr22 15825725 15825726 G/T uc002zlw.1 Gln:Pro/Gln 469 T
- chr22 15827035 15827036 G uc002zlw.1 Glu:Asp 414 C
- chr22 15827135 15827136 C/G uc002zlw.1 Gly:Gly/Ala 381 C
- chr22 15830928 15830929 A/G uc002zlw.1 Ala:Ser/Pro 281 C
- chr22 15830951 15830952 G uc002zlw.1 Leu:Pro 273 A
- chr22 15830955 15830956 C/T uc002zlw.1 Ser:Gly/Ser 272 T
- chr22 15848885 15848886 C/T uc002zlw.1 Ser:Trp/Stop 217 G
- chr22 15848885 15848886 C/T uc010gqs.1 Ser:Trp/Stop 200 G
- chr22 15849048 15849049 A/C uc002zlw.1 Gly:Stop/Gly 163 C
+ chr22 15825725 15825726 G/T uc002zlw.1 Gln:Pro/Gln 469 T
+ chr22 15827035 15827036 G uc002zlw.1 Glu:Asp 414 C
+ chr22 15827135 15827136 C/G uc002zlw.1 Gly:Gly/Ala 381 C
+ chr22 15830928 15830929 A/G uc002zlw.1 Ala:Ser/Pro 281 C
+ chr22 15830951 15830952 G uc002zlw.1 Leu:Pro 273 A
+ chr22 15830955 15830956 C/T uc002zlw.1 Ser:Gly/Ser 272 T
+ chr22 15848885 15848886 C/T uc002zlw.1 Ser:Trp/Stop 217 G
+ chr22 15848885 15848886 C/T uc010gqs.1 Ser:Trp/Stop 200 G
+ chr22 15849048 15849049 A/C uc002zlw.1 Gly:Stop/Gly 163 C
etc.
+
diff --git a/tools/human_genome_variation/beam.xml b/tools/human_genome_variation/beam.xml
index 073fce636e4..29cffb25af0 100644
--- a/tools/human_genome_variation/beam.xml
+++ b/tools/human_genome_variation/beam.xml
@@ -56,12 +56,12 @@ single-SNP associations), please use the GPASS tool instead.**
**Dataset formats**
-The input dataset is in lped_ format. The output datasets are both
-tabular_.
+The input dataset must be in lped_ format. The output datasets are both tabular_.
+(`Dataset missing?`_)
.. _lped: ./static/formatHelp.html#lped
-
.. _tabular: ./static/formatHelp.html#tabular
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -77,9 +77,9 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD).
- input map file::
- 1 rs0 0 738547
- 1 rs1 0 5597094
- 1 rs2 0 9424115
+ 1 rs0 0 738547
+ 1 rs1 0 5597094
+ 1 rs2 0 9424115
etc.
- input ped file::
@@ -90,33 +90,33 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD).
- first output file, significance.txt::
- ID chr position results
- rs0 chr1 738547 10 20 score= 45.101397 , df= 8 , p= 0.000431 , N=1225
+ ID chr position results
+ rs0 chr1 738547 10 20 score= 45.101397 , df= 8 , p= 0.000431 , N=1225
- second output file, posterior.txt::
- id: chr position marginal + interaction = total posterior
- 0: 1 738547 0.0000 + 0.0000 = 0.0000
- 1: 1 5597094 0.0000 + 0.0000 = 0.0000
- 2: 1 9424115 0.0000 + 0.0000 = 0.0000
- 3: 1 13879818 0.0000 + 0.0000 = 0.0000
- 4: 1 13934751 0.0000 + 0.0000 = 0.0000
- 5: 1 16803491 0.0000 + 0.0000 = 0.0000
- 6: 1 17236854 0.0000 + 0.0000 = 0.0000
- 7: 1 18445387 0.0000 + 0.0000 = 0.0000
- 8: 1 21222571 0.0000 + 0.0000 = 0.0000
+ id: chr position marginal + interaction = total posterior
+ 0: 1 738547 0.0000 + 0.0000 = 0.0000
+ 1: 1 5597094 0.0000 + 0.0000 = 0.0000
+ 2: 1 9424115 0.0000 + 0.0000 = 0.0000
+ 3: 1 13879818 0.0000 + 0.0000 = 0.0000
+ 4: 1 13934751 0.0000 + 0.0000 = 0.0000
+ 5: 1 16803491 0.0000 + 0.0000 = 0.0000
+ 6: 1 17236854 0.0000 + 0.0000 = 0.0000
+ 7: 1 18445387 0.0000 + 0.0000 = 0.0000
+ 8: 1 21222571 0.0000 + 0.0000 = 0.0000
etc.
- id: chr position block_boundary | allele counts in cases and controls
- 0: 1 738547 1.000 | 156 93 251 | 169 83 248
- 1: 1 5597094 1.000 | 323 19 158 | 328 16 156
- 2: 1 9424115 1.000 | 366 6 128 | 369 11 120
- 3: 1 13879818 1.000 | 252 31 217 | 278 32 190
- 4: 1 13934751 1.000 | 246 64 190 | 224 58 218
- 5: 1 16803491 1.000 | 91 160 249 | 91 174 235
- 6: 1 17236854 1.000 | 252 43 205 | 249 44 207
- 7: 1 18445387 1.000 | 205 66 229 | 217 56 227
- 8: 1 21222571 1.000 | 353 9 138 | 352 8 140
+ id: chr position block_boundary | allele counts in cases and controls
+ 0: 1 738547 1.000 | 156 93 251 | 169 83 248
+ 1: 1 5597094 1.000 | 323 19 158 | 328 16 156
+ 2: 1 9424115 1.000 | 366 6 128 | 369 11 120
+ 3: 1 13879818 1.000 | 252 31 217 | 278 32 190
+ 4: 1 13934751 1.000 | 246 64 190 | 224 58 218
+ 5: 1 16803491 1.000 | 91 160 249 | 91 174 235
+ 6: 1 17236854 1.000 | 252 43 205 | 249 44 207
+ 7: 1 18445387 1.000 | 205 66 229 | 217 56 227
+ 8: 1 21222571 1.000 | 353 9 138 | 352 8 140
etc.
The "id" field is an internally used index.
@@ -125,8 +125,13 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD).
**References**
-Zhang Y and Liu JS (2007). Bayesian Inference of Epistatic Interactions in Case-Control Studies. Nature Genetics, 39:1167-1173
+Zhang Y, Liu JS. (2007)
+Bayesian inference of epistatic interactions in case-control studies.
+Nat Genet. 39(9):1167-73. Epub 2007 Aug 26.
+
+Zhang Y, Zhang J, Liu JS. (2010)
+Block-based bayesian epistasis association mapping with application to WTCCC type 1 diabetes data.
+Submitted.
-Zhang Y, Zhang J, Liu JS (2010) Block-based Bayesian Epistasis Association Mapping with Application to WTCCC Type 1 Diabetes Data. Submitted.
diff --git a/tools/human_genome_variation/ctd.xml b/tools/human_genome_variation/ctd.xml
index e5f745518c8..30c2909daa5 100644
--- a/tools/human_genome_variation/ctd.xml
+++ b/tools/human_genome_variation/ctd.xml
@@ -256,8 +256,10 @@
**Dataset formats**
The input and output datasets are tabular_.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -280,35 +282,39 @@ Website: http://ctd.mdibl.org/
**Examples**
- input data file:
- HBB
+ HBB
-- select column c1, Identifier type = Genes, and Data to extract = All disease relationships
+- select Column = c1, Identifier type = Genes, and Data to extract = All disease relationships
- output file::
#Input GeneSymbol GeneName GeneID DiseaseName DiseaseID GeneDiseaseRelation OmimIDs PubMedIDs
- hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Ethanol 17676605|18926900
- hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Valproic Acid 8875741
+ hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Ethanol 17676605|18926900
+ hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Valproic Acid 8875741
etc.
Another example:
- same input file:
- HBB
+ HBB
-- select column c1, Identifier type = Genes, Data to extract = Curated chemical-gene interactions, and Interaction type = ANY
+- select Column = c1, Identifier type = Genes, Data to extract = Curated chemical-gene interactions, and Interaction type = ANY
- output file::
#Input GeneSymbol GeneName GeneID ChemicalName ChemicalID CasRN Organism OrganismID Interaction InteractionTypes PubMedIDs
- hbb HBB hemoglobin, beta 3043 1-nitronaphthalene C016614 86-57-7 Macaca mulatta 9544 1-nitronaphthalene metabolite binds to HBB protein binding 16453347
- hbb HBB hemoglobin, beta 3043 2,6-diisocyanatotoluene C026942 91-08-7 Cavia porcellus 10141 2,6-diisocyanatotoluene binds to HBB protein binding 8728499
+ hbb HBB hemoglobin, beta 3043 1-nitronaphthalene C016614 86-57-7 Macaca mulatta 9544 1-nitronaphthalene metabolite binds to HBB protein binding 16453347
+ hbb HBB hemoglobin, beta 3043 2,6-diisocyanatotoluene C026942 91-08-7 Cavia porcellus 10141 2,6-diisocyanatotoluene binds to HBB protein binding 8728499
etc.
-----
**Reference**
-Davis AP, Murphy CG, Saraceni-Richards CA, Rosenstein MC, Wiegers TC, Mattingly CJ. Comparative Toxicogenomics Database: a knowledgebase and discovery tool for chemical.gene.disease networks. Nucleic Acids Res. 2009 Jan;37(Database issue):D786-92.
+Davis AP, Murphy CG, Saraceni-Richards CA, Rosenstein MC, Wiegers TC, Mattingly CJ. (2009)
+Comparative Toxicogenomics Database: a knowledgebase and discovery tool for
+chemical-gene-disease networks.
+Nucleic Acids Res. 37(Database issue):D786-92. Epub 2008 Sep 9.
+
diff --git a/tools/human_genome_variation/funDo.xml b/tools/human_genome_variation/funDo.xml
index 98c6d15b5f5..52fb4669315 100644
--- a/tools/human_genome_variation/funDo.xml
+++ b/tools/human_genome_variation/funDo.xml
@@ -65,32 +65,37 @@ Typing::
results in::
- 1. 2. 3. 4. 5. 6. 7.
- chr11 89507465 89565427 + NAALAD2 10003 Adenocarcinoma
- chr15 50189113 50192264 - BCL2L10 10017 Carcinoma
- chr7 150535855 150555250 - ABCF2 10061 Clear cell carcinoma
- chr7 150540508 150555250 - ABCF2 10061 Clear cell carcinoma
- chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
- chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
+ 1. 2. 3. 4. 5. 6. 7.
+ chr11 89507465 89565427 + NAALAD2 10003 Adenocarcinoma
+ chr15 50189113 50192264 - BCL2L10 10017 Carcinoma
+ chr7 150535855 150555250 - ABCF2 10061 Clear cell carcinoma
+ chr7 150540508 150555250 - ABCF2 10061 Clear cell carcinoma
+ chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
+ chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
etc.
where the column contents are as follows::
- 1. chromosome name.
- 2. start position of the gene.
- 3. end position of the gene.
- 4. strand.
- 4. gene name.
- 6. Entrez Gene ID.
- 7. disease term.
+ 1. chromosome name
+ 2. start position of the gene
+ 3. end position of the gene
+ 4. strand
+ 4. gene name
+ 6. Entrez Gene ID
+ 7. disease term
-----
**References**
-Pan Du, Gang Feng,Jared Flatow, Jie Song, Michelle Holko, Warren A. Kibbe1 and
-Simon M. Lin. From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose ontology for the test of gene-ontology associations. Bioinformatics (2009) 25 (12):i63-i68.
+Du P, Feng G, Flatow J, Song J, Holko M, Kibbe WA, Lin SM. (2009)
+From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose
+ontology for the test of gene-ontology associations.
+Bioinformatics. 25(12):i63-8.
+
+Osborne JD, Flatow J, Holko M, Lin SM, Kibbe WA, Zhu LJ, Danila MI, Feng G, Chisholm RL. (2009)
+Annotating the human genome with Disease Ontology.
+BMC Genomics. 10 Suppl 1:S6.
-Osborne JD, Flatow J, Holko M, Lin SM, Kibbe WA, Zhu LJ, Danila MI, Feng G, Chisholm RL. Annotating the human genome with Disease Ontology. BMC Genomics (2009) 10 S1:S6.
diff --git a/tools/human_genome_variation/gpass.xml b/tools/human_genome_variation/gpass.xml
index 62b7d8ff51d..e28aa203a1a 100644
--- a/tools/human_genome_variation/gpass.xml
+++ b/tools/human_genome_variation/gpass.xml
@@ -36,11 +36,12 @@
**Dataset formats**
-The input dataset must be lped_, and the output is tabular_.
+The input dataset must be in lped_ format, and the output is tabular_.
+(`Dataset missing?`_)
.. _lped: ./static/formatHelp.html#lped
-
.. _tabular: ./static/formatHelp.html#tab
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -80,9 +81,9 @@ Otherwise use permutation.
- input map file::
- 1 rs0 0 738547
- 1 rs1 0 5597094
- 1 rs2 0 9424115
+ 1 rs0 0 738547
+ 1 rs1 0 5597094
+ 1 rs2 0 9424115
etc.
- input ped file::
@@ -93,17 +94,19 @@ Otherwise use permutation.
- output dataset, showing significant SNPs and their p-values and FDR::
- #ID chr position Statistics adj-Pvalue FDR
- rs35 chr1 136606952 4.890849 0.991562 0.682138
- rs36 chr1 137748344 4.931934 0.991562 0.795827
- rs44 chr2 14423047 7.712832 0.665086 0.218776
+ #ID chr position Statistics adj-Pvalue FDR
+ rs35 chr1 136606952 4.890849 0.991562 0.682138
+ rs36 chr1 137748344 4.931934 0.991562 0.795827
+ rs44 chr2 14423047 7.712832 0.665086 0.218776
etc.
-----
**Reference**
-Zhang Y and Liu JS (2010). Fast and Accurate Significance Approximation for
-Genome-wide Association Studies, submitted.
+Zhang Y, Liu JS. (2010)
+Fast and accurate significance approximation for genome-wide association studies.
+Submitted.
+
diff --git a/tools/human_genome_variation/hilbertvis.xml b/tools/human_genome_variation/hilbertvis.xml
index 221d450fa21..15632915615 100644
--- a/tools/human_genome_variation/hilbertvis.xml
+++ b/tools/human_genome_variation/hilbertvis.xml
@@ -61,8 +61,10 @@
**Dataset formats**
The input format is interval_, and the output is an image in PDF format.
+(`Dataset missing?`_)
.. _interval: ./static/formatHelp.html#interval
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -107,6 +109,9 @@ Here are some examples from the HilbertVis homepage, using ChIP-Seq data.
**Reference**
-Anders S. (2009) Visualization of genomic data with the Hilbert curve. Bioinformatics, 25:1231-1235.
+Anders S. (2009)
+Visualization of genomic data with the Hilbert curve.
+Bioinformatics. 25(10):1231-5. Epub 2009 Mar 17.
+
diff --git a/tools/human_genome_variation/ldtools.xml b/tools/human_genome_variation/ldtools.xml
index 14f425cf0a0..07c65c8c90f 100644
--- a/tools/human_genome_variation/ldtools.xml
+++ b/tools/human_genome_variation/ldtools.xml
@@ -32,8 +32,10 @@
**Dataset formats**
The input and output datasets are tabular_.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -43,9 +45,8 @@ This tool can be used to analyze the patterns of linkage disequilibrium
(LD) between polymorphic sites in a locus. SNPs are grouped based on the
threshold level of LD as measured by r\ :sup:`2` (regardless of genomic
position), and a representative "tag SNP" is reported for each group.
-Note that the groups are generated by transitive closure: each SNP in the
-group is within r\ :sup:`2` of *some* other SNP in the group, but not
-necessarily all of them (and not necessarily the tag SNP).
+The other SNPs in the group are in LD with the tag SNP, but not necessarily
+with each other.
The underlying algorithm is the same as the one used in ldSelect (Carlson
et al. 2004). However, this tool is implemented to be much faster and more
@@ -61,49 +62,50 @@ two allele nucleotides.
- input file::
- rs2334386 NA20364 G T
- rs2334386 NA20363 G G
- rs2334386 NA20360 G G
- rs2334386 NA20359 G G
- rs2334386 NA20358 G G
- rs2334386 NA20356 G G
- rs2334386 NA20357 G G
- rs2334386 NA20350 G G
- rs2334386 NA20349 G G
- rs2334386 NA20348 G G
- rs2334386 NA20347 G G
- rs2334386 NA20346 G G
- rs2334386 NA20345 G G
- rs2334386 NA20344 G G
- rs2334386 NA20342 G G
- etc.
+ rs2334386 NA20364 G T
+ rs2334386 NA20363 G G
+ rs2334386 NA20360 G G
+ rs2334386 NA20359 G G
+ rs2334386 NA20358 G G
+ rs2334386 NA20356 G G
+ rs2334386 NA20357 G G
+ rs2334386 NA20350 G G
+ rs2334386 NA20349 G G
+ rs2334386 NA20348 G G
+ rs2334386 NA20347 G G
+ rs2334386 NA20346 G G
+ rs2334386 NA20345 G G
+ rs2334386 NA20344 G G
+ rs2334386 NA20342 G G
+ etc.
- output file::
- rs2238748 rs2793064,rs6518516,rs6518517,rs2283641,rs5993533,rs715590,rs2072123,rs2105421,rs2800954,rs1557847,rs807750,rs807753,rs5993488,rs8138035,rs2800980,rs2525079,rs5992353,rs712966,rs2525036,rs807743,rs1034727,rs807744,rs2074003
- rs2871023 rs1210715,rs1210711,rs5748189,rs1210709,rs3788298,rs7284649,rs9306217,rs9604954,rs1210703,rs5748179,rs5746727,rs5748190,rs5993603,rs2238766,rs885981,rs2238763,rs5748165,rs9605996,rs9606001,rs5992398
- rs7292006 rs13447232,rs5993665,rs2073733,rs1057457,rs756658,rs5992395,rs2073760,rs739369,rs9606017,rs739370,rs4493360,rs2073736
- rs2518840 rs1061325,rs2283646,rs362148,rs1340958,rs361956,rs361991,rs2073754,rs2040771,rs2073740,rs2282684
- rs2073775 rs10160,rs2800981,rs807751,rs5993492,rs2189490,rs5747997,rs2238743
- rs5747263 rs12159924,rs2300688,rs4239846,rs3747025,rs3747024,rs3747023,rs2300691
- rs433576 rs9605439,rs1109052,rs400509,rs401099,rs396012,rs410456,rs385105
- rs2106145 rs5748131,rs2013516,rs1210684,rs1210685,rs2238767,rs2277837
- rs2587082 rs2257083,rs2109659,rs2587081,rs5747306,rs2535704,rs2535694
- rs807667 rs2800974,rs756651,rs762523,rs2800973,rs1018764
- rs2518866 rs1206542,rs807467,rs807464,rs807462,rs712950
- rs1110661 rs1110660,rs7286607,rs1110659,rs5992917,rs1110662
- rs759076 rs5748760,rs5748755,rs5748752,rs4819925,rs933461
- rs5746487 rs5992895,rs2034113,rs2075455,rs1867353
- rs5748212 rs5746736,rs4141527,rs5748147,rs5748202
- etc.
+ rs2238748 rs2793064,rs6518516,rs6518517,rs2283641,rs5993533,rs715590,rs2072123,rs2105421,rs2800954,rs1557847,rs807750,rs807753,rs5993488,rs8138035,rs2800980,rs2525079,rs5992353,rs712966,rs2525036,rs807743,rs1034727,rs807744,rs2074003
+ rs2871023 rs1210715,rs1210711,rs5748189,rs1210709,rs3788298,rs7284649,rs9306217,rs9604954,rs1210703,rs5748179,rs5746727,rs5748190,rs5993603,rs2238766,rs885981,rs2238763,rs5748165,rs9605996,rs9606001,rs5992398
+ rs7292006 rs13447232,rs5993665,rs2073733,rs1057457,rs756658,rs5992395,rs2073760,rs739369,rs9606017,rs739370,rs4493360,rs2073736
+ rs2518840 rs1061325,rs2283646,rs362148,rs1340958,rs361956,rs361991,rs2073754,rs2040771,rs2073740,rs2282684
+ rs2073775 rs10160,rs2800981,rs807751,rs5993492,rs2189490,rs5747997,rs2238743
+ rs5747263 rs12159924,rs2300688,rs4239846,rs3747025,rs3747024,rs3747023,rs2300691
+ rs433576 rs9605439,rs1109052,rs400509,rs401099,rs396012,rs410456,rs385105
+ rs2106145 rs5748131,rs2013516,rs1210684,rs1210685,rs2238767,rs2277837
+ rs2587082 rs2257083,rs2109659,rs2587081,rs5747306,rs2535704,rs2535694
+ rs807667 rs2800974,rs756651,rs762523,rs2800973,rs1018764
+ rs2518866 rs1206542,rs807467,rs807464,rs807462,rs712950
+ rs1110661 rs1110660,rs7286607,rs1110659,rs5992917,rs1110662
+ rs759076 rs5748760,rs5748755,rs5748752,rs4819925,rs933461
+ rs5746487 rs5992895,rs2034113,rs2075455,rs1867353
+ rs5748212 rs5746736,rs4141527,rs5748147,rs5748202
+ etc.
-----
**Reference**
-Carlson CS, Eberle MA, Rieder MJ, Yi Q, Kruglyak L, Nickerson DA.
+Carlson CS, Eberle MA, Rieder MJ, Yi Q, Kruglyak L, Nickerson DA. (2004)
Selecting a maximally informative set of single-nucleotide polymorphisms for
-association analysis using linkage disequilibrium. Am J Hum Genet. 2004 Jan;
-74(1):106-20. Epub 2003 Dec 15.
+association analyses using linkage disequilibrium.
+Am J Hum Genet. 74(1):106-20. Epub 2003 Dec 15.
+
diff --git a/tools/human_genome_variation/linkToDavid.xml b/tools/human_genome_variation/linkToDavid.xml
index afe241daefe..451f4eb1439 100644
--- a/tools/human_genome_variation/linkToDavid.xml
+++ b/tools/human_genome_variation/linkToDavid.xml
@@ -74,10 +74,11 @@ The list is limited to 400 IDs.
The input dataset is tabular_ format. The output dataset is html_ format with
a link to the DAVID website as described below.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
-
.. _html: ./static/formatHelp.html#html
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -97,8 +98,13 @@ lists of genes.
**References**
-Huang DW, Sherman BT, Lempicki RA. Systematic and integrative analysis of large gene lists using DAVID Bioinformatics Resources. Nature Protoc. 2009;4(1):44-57.
+Huang DW, Sherman BT, Lempicki RA. (2009) Systematic and integrative analysis
+of large gene lists using DAVID bioinformatics resources.
+Nat Protoc. 4(1):44-57.
+
+Dennis G, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. (2003)
+DAVID: database for annotation, visualization, and integrated discovery.
+Genome Biol. 4(5):P3. Epub 2003 Apr 3.
-Dennis G Jr, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. DAVID: Database for Annotation, Visualization, and Integrated Discovery. Genome Biol. 2003;4(5):P3.
diff --git a/tools/human_genome_variation/linkToGProfile.xml b/tools/human_genome_variation/linkToGProfile.xml
index 5d7a2bb90be..50e72e0d8e1 100644
--- a/tools/human_genome_variation/linkToGProfile.xml
+++ b/tools/human_genome_variation/linkToGProfile.xml
@@ -36,15 +36,17 @@
+
**Dataset formats**
The input dataset is tabular_ with a column of identifiers.
The output dataset is html_ with a link to g:Profiler.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
-
.. _html: ./static/formatHelp.html#html
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -73,6 +75,9 @@ the results to run other g:Profiler tools using the same list of IDs.
**Reference**
-\J. Reimand, M. Kull, H. Peterson, J. Hansen, J. Vilo: g:Profiler -- a web-based toolset for functional profiling of gene lists from large-scale experiments (2007) NAR 35 W193-W200
+Reimand J, Kull M, Peterson H, Hansen J, Vilo J. (2007) g:Profiler -- a web-based
+toolset for functional profiling of gene lists from large-scale experiments.
+Nucleic Acids Res. 35(Web Server issue):W193-200. Epub 2007 May 3.
+
diff --git a/tools/human_genome_variation/lps.xml b/tools/human_genome_variation/lps.xml
index 9f048954c15..a0a98415590 100644
--- a/tools/human_genome_variation/lps.xml
+++ b/tools/human_genome_variation/lps.xml
@@ -172,15 +172,17 @@
+
**Dataset formats**
-The input and output datasets are tabular_. The columns are described
-below. There is a second output dataset (a log) that is in text_ format.
+The input and output datasets are tabular_. The columns are described below.
+There is a second output dataset (a log) that is in text_ format.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
-
.. _text: ./static/formatHelp.html#text
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -217,19 +219,19 @@ Website: http://pages.cs.wisc.edu/~swright/LPS/
- input file::
- +1 1 0 0 0 0 1 0 1 1 ...
- +1 1 1 1 0 0 1 0 1 1 ...
- +1 1 0 1 0 1 0 1 0 1 ...
+ +1 1 0 0 0 0 1 0 1 1 ...
+ +1 1 1 1 0 0 1 0 1 1 ...
+ +1 1 0 1 0 1 0 1 0 1 ...
etc.
- output results file::
- 0
- 0
- 0
- 0
- 0.025541
- etc.
+ 0
+ 0
+ 0
+ 0
+ 0.025541
+ etc.
- output log file::
@@ -243,16 +245,31 @@ Website: http://pages.cs.wisc.edu/~swright/LPS/
**References**
-Koh K, Kim S-J, and Boyd S. (2007) An Interior-Point Method for Large-Scale l1-Regularized Logistic Regression. Journal of Machine Learning Research, 8:1519-1555.
+Koh K, Kim S-J, Boyd S. (2007)
+An interior-point method for large-scale l1-regularized logistic regression.
+Journal of Machine Learning Research. 8:1519-1555.
-Shi W, Wahba G, Wright S, Lee K, Klein R, and Klein B. (2008) LASSO-Patternsearch Algorithm with Application to Ophthalmology and Genomic Data. Statistics And Its Interface, 1:137-153.
+Shi W, Wahba G, Wright S, Lee K, Klein R, Klein B. (2008)
+LASSO-Patternsearch algorithm with application to ophthalmology and genomic data.
+Stat Interface. 1(1):137-153.
-Wright S, Novak R, and Figueiredo M. (2009) Sparse reconstruction via separable approximation. IEEE Transactions on Signal Processing, 57:2479-2403.
+
-Wright S. (2010) Accelerated block-coordinate relaxation for regularized optimization. Technical Report. University of Wisconsin. August 10, 2010.
diff --git a/tools/human_genome_variation/pass.xml b/tools/human_genome_variation/pass.xml
index 86f3ed0ff8f..83bd9508e6a 100644
--- a/tools/human_genome_variation/pass.xml
+++ b/tools/human_genome_variation/pass.xml
@@ -6,10 +6,10 @@
-
+
-
+
@@ -21,9 +21,8 @@
sed
-
-----
**Hints**
-- ChIP-seq data:
+- ChIP-Seq data:
-If the data is from ChIP-seq, you need to convert each ChIP-seq value into z-scores before using this program. Also for ChIP-seq,
-it is recommended that you group read counts within a neighborhood together, e.g., sum of reads within a 30bp window, for
-every 30bp shifting window. As a result, the ChIP-seq data looks like a ChIP-chip data in format.
+ If the data is from ChIP-Seq, you need to convert the ChIP-Seq values
+ into z-scores before using this program. It is also recommended that
+ you group read counts within a neighborhood together, e.g. in tiled
+ windows of 30bp. In this way, the ChIP-Seq data will resemble
+ ChIP-chip data in format.
- Choosing window size options:
-The window size is related with probe tiling density, if probes are tiled at every 100bp, then smallest window = 2 largest window = 6 is good, because DNA fragment size is around 300~500bp.
+ The window size is related to the probe tiling density. For example,
+ if the probes are tiled at every 100bp, then setting the smallest
+ window = 2 and largest window = 6 is appropriate, because the DNA
+ fragment size is around 300-500bp.
-----
**Example**
-- input file ChIP-chip data file; Nimblegene GFF file format::
+- input file::
- chr7 Nimblegen ID 40307603 40307652 1.668944
- chr7 Nimblegen ID 40307703 40307752 0.8041307
- chr7 Nimblegen ID 40307808 40307865 -1.089931
- chr7 Nimblegen ID 40307920 40307969 1.055044
- chr7 Nimblegen ID 40308005 40308068 2.447853
- chr7 Nimblegen ID 40308125 40308174 0.1638694
- chr7 Nimblegen ID 40308223 40308275 -0.04796628
- chr7 Nimblegen ID 40308318 40308367 0.9335709
- chr7 Nimblegen ID 40308526 40308584 0.5143972
- chr7 Nimblegen ID 40308611 40308660 -1.089931
+ chr7 Nimblegen ID 40307603 40307652 1.668944 . . .
+ chr7 Nimblegen ID 40307703 40307752 0.8041307 . . .
+ chr7 Nimblegen ID 40307808 40307865 -1.089931 . . .
+ chr7 Nimblegen ID 40307920 40307969 1.055044 . . .
+ chr7 Nimblegen ID 40308005 40308068 2.447853 . . .
+ chr7 Nimblegen ID 40308125 40308174 0.1638694 . . .
+ chr7 Nimblegen ID 40308223 40308275 -0.04796628 . . .
+ chr7 Nimblegen ID 40308318 40308367 0.9335709 . . .
+ chr7 Nimblegen ID 40308526 40308584 0.5143972 . . .
+ chr7 Nimblegen ID 40308611 40308660 -1.089931 . . .
etc.
- Chromosome, start, end, and score are required fields.
+ In GFF, a value of dot '.' is used to mean "not applicable".
- output file::
- #ID Chr Start End WinSz PeakValue # of FPs FDR
- 1 chr7 40310931 40311266 4 1.663446 0.208435 0.208435
+ ID Chr Start End WinSz PeakValue # of FPs FDR
+ 1 chr7 40310931 40311266 4 1.663446 0.248817 0.248817
-----
**References**
-Zhang Y (2008) Poisson Approximation for Significance in Genome-wide ChIP-chip Tiling Arrays. Bioinformatics, 24(24):2825-2831
+Zhang Y. (2008)
+Poisson approximation for significance in genome-wide ChIP-chip tiling arrays.
+Bioinformatics. 24(24):2825-31. Epub 2008 Oct 25.
+
+Chen KB, Zhang Y. (2010)
+A varying threshold method for ChIP peak calling using multiple sources of information.
+Submitted.
-Chen KB and Zhang Y (2010) A Varying Threshold Method for ChIP Peak Calling Using Multiple Sources of Information. Submitted.
diff --git a/tools/human_genome_variation/sift.xml b/tools/human_genome_variation/sift.xml
index be9d92b21d5..443919e87fc 100644
--- a/tools/human_genome_variation/sift.xml
+++ b/tools/human_genome_variation/sift.xml
@@ -76,15 +76,17 @@
.. class:: warningmark
-This currently works for builds hg18 or hg19.
+This currently works only for builds hg18 or hg19.
-----
**Dataset formats**
The input and output datasets are tabular_.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -114,40 +116,47 @@ Website: http://sift.jcvi.org/
- input file::
- chr3 81780820 + T/C
- chr2 230341630 + G/A
- chr2 43881517 + A/T
- chr2 43857514 + T/C
- chr6 88375602 + G/A
- chr22 29307353 - T/A
- chr10 115912482 - G/T
- chr10 115900918 - C/T
- chr16 69875502 + G/T
+ chr3 81780820 + T/C
+ chr2 230341630 + G/A
+ chr2 43881517 + A/T
+ chr2 43857514 + T/C
+ chr6 88375602 + G/A
+ chr22 29307353 - T/A
+ chr10 115912482 - G/T
+ chr10 115900918 - C/T
+ chr16 69875502 + G/T
etc.
- output file::
- #Chrom Position Strand Allele Codons Transcript ID Protein ID Substitution Region dbSNP ID SNP Type Prediction Score Median Info Num seqs at position User Comment
- chr3 81780820 + T/C AGA-gGA ENST00000264326 ENSP00000264326 R190G EXON CDS rs2229519:C Nonsynonymous DAMAGING 0.04 3.06 149
- chr2 230341630 + G/T - ENST00000389045 ENSP00000373697 NA EXON CDS rs1803846:A Unknown Not scored NA NA NA
- chr2 43881517 + A/T ATA-tTA ENST00000260605 ENSP00000260605 I230L EXON CDS rs11556157:T Nonsynonymous TOLERATED 0.47 3.19 7
- chr2 43857514 + T/C TTT-TcT ENST00000260605 ENSP00000260605 F33S EXON CDS rs2288709:C Nonsynonymous TOLERATED 0.61 3.33 6
- chr6 88375602 + G/A GTT-aTT ENST00000257789 ENSP00000257789 V217I EXON CDS rs2307389:A Nonsynonymous TOLERATED 0.75 3.17 13
- chr22 29307353 + T/A ACC-tCC ENST00000335214 ENSP00000334612 T264S EXON CDS rs42942:A Nonsynonymous TOLERATED 0.4 3.14 23
- chr10 115912482 + C/A CGA-CtA ENST00000369285 ENSP00000358291 R179L EXON CDS rs12782946:T Nonsynonymous TOLERATED 0.06 4.32 2
- chr10 115900918 + G/A CAA-tAA ENST00000369287 ENSP00000358293 Q271* EXON CDS rs7095762:T Nonsynonymous N/A N/A N/A N/A
- chr16 69875502 + G/T ACA-AaA ENST00000338099 ENSP00000337512 T608K EXON CDS rs3096381:T Nonsynonymous TOLERATED 0.12 3.41 3
+ #Chrom Position Strand Allele Codons Transcript ID Protein ID Substitution Region dbSNP ID SNP Type Prediction Score Median Info Num seqs at position User Comment
+ chr3 81780820 + T/C AGA-gGA ENST00000264326 ENSP00000264326 R190G EXON CDS rs2229519:C Nonsynonymous DAMAGING 0.04 3.06 149
+ chr2 230341630 + G/T - ENST00000389045 ENSP00000373697 NA EXON CDS rs1803846:A Unknown Not scored NA NA NA
+ chr2 43881517 + A/T ATA-tTA ENST00000260605 ENSP00000260605 I230L EXON CDS rs11556157:T Nonsynonymous TOLERATED 0.47 3.19 7
+ chr2 43857514 + T/C TTT-TcT ENST00000260605 ENSP00000260605 F33S EXON CDS rs2288709:C Nonsynonymous TOLERATED 0.61 3.33 6
+ chr6 88375602 + G/A GTT-aTT ENST00000257789 ENSP00000257789 V217I EXON CDS rs2307389:A Nonsynonymous TOLERATED 0.75 3.17 13
+ chr22 29307353 + T/A ACC-tCC ENST00000335214 ENSP00000334612 T264S EXON CDS rs42942:A Nonsynonymous TOLERATED 0.4 3.14 23
+ chr10 115912482 + C/A CGA-CtA ENST00000369285 ENSP00000358291 R179L EXON CDS rs12782946:T Nonsynonymous TOLERATED 0.06 4.32 2
+ chr10 115900918 + G/A CAA-tAA ENST00000369287 ENSP00000358293 Q271* EXON CDS rs7095762:T Nonsynonymous N/A N/A N/A N/A
+ chr16 69875502 + G/T ACA-AaA ENST00000338099 ENSP00000337512 T608K EXON CDS rs3096381:T Nonsynonymous TOLERATED 0.12 3.41 3
etc.
-----
**References**
-Predicting Deleterious Amino Acid Substitutions, Genome Res. 2001 May; 11(5): 863.874.
-Accounting for Human Polymorphisms Predicted to Affect Protein Function, Genome Res. 2002 December; 436-446
+Ng PC, Henikoff S. (2001) Predicting deleterious amino acid substitutions.
+Genome Res. 11(5):863-74.
-SIFT: predicting amino acid changes that affect protein function, Nucleic Acids Research, 2003, Vol. 31, No. 13 3812-3814
-Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm, Nature Protocols 4, - 1073 - 1081 (2009)
+Ng PC, Henikoff S. (2002) Accounting for human polymorphisms predicted to affect protein function.
+Genome Res. 12(3):436-46.
+
+Ng PC, Henikoff S. (2003) SIFT: Predicting amino acid changes that affect protein function.
+Nucleic Acids Res. 31(13):3812-4.
+
+Kumar P, Henikoff S, Ng PC. (2009) Predicting the effects of coding non-synonymous variants
+on protein function using the SIFT algorithm.
+Nat Protoc. 4(7):1073-81. Epub 2009 Jun 25.
diff --git a/tools/human_genome_variation/snpFreq.xml b/tools/human_genome_variation/snpFreq.xml
index 3aaa6ef0149..c06defba07f 100644
--- a/tools/human_genome_variation/snpFreq.xml
+++ b/tools/human_genome_variation/snpFreq.xml
@@ -40,11 +40,12 @@
**Dataset formats**
-The input is tabular_, with six columns of allele
-counts. The output is also tabular, and includes all of the input data plus
-the additional columns described below.
+The input is tabular_, with six columns of allele counts. The output is also tabular,
+and includes all of the input data plus the additional columns described below.
+(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
+.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -70,15 +71,15 @@ group, the p-value, and the q-value.
- input file::
- chr1 210 211 38 4 15 56 0 1 x
- chr1 228 229 55 0 2 56 0 1 x
- chr1 230 231 46 0 11 55 0 2 x
- chr1 234 235 43 0 14 55 0 2 x
- chr1 236 237 55 0 2 13 10 34 x
- chr1 437 438 55 0 2 46 0 11 x
- chr1 439 440 56 0 1 55 0 2 x
- chr1 449 450 56 0 1 13 20 24 x
- chr1 518 519 56 0 1 38 4 15 x
+ chr1 210 211 38 4 15 56 0 1 x
+ chr1 228 229 55 0 2 56 0 1 x
+ chr1 230 231 46 0 11 55 0 2 x
+ chr1 234 235 43 0 14 55 0 2 x
+ chr1 236 237 55 0 2 13 10 34 x
+ chr1 437 438 55 0 2 46 0 11 x
+ chr1 439 440 56 0 1 55 0 2 x
+ chr1 449 450 56 0 1 13 20 24 x
+ chr1 518 519 56 0 1 38 4 15 x
Here the group 1 genotype counts are in columns 4 - 6, while those
for group 2 are in columns 7 - 9.
@@ -89,15 +90,15 @@ to see where the new columns are appended in the output.
- output file::
- chr1 210 211 38 4 15 56 0 1 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
- chr1 228 229 55 0 2 56 0 1 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
- chr1 230 231 46 0 11 55 0 2 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
- chr1 234 235 43 0 14 55 0 2 x 49 0 8 49 0 8 0.00210854461554067 0.000739840215979182
- chr1 236 237 55 0 2 13 10 34 x 34 5 18 34 5 18 6.14613878554783e-17 4.31307984950725e-17
- chr1 437 438 55 0 2 46 0 11 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
- chr1 439 440 56 0 1 55 0 2 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
- chr1 449 450 56 0 1 13 20 24 x 34.5 10 12.5 34.5 10 12.5 2.25757007974134e-18 2.37638955762246e-18
- chr1 518 519 56 0 1 38 4 15 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
+ chr1 210 211 38 4 15 56 0 1 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
+ chr1 228 229 55 0 2 56 0 1 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
+ chr1 230 231 46 0 11 55 0 2 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
+ chr1 234 235 43 0 14 55 0 2 x 49 0 8 49 0 8 0.00210854461554067 0.000739840215979182
+ chr1 236 237 55 0 2 13 10 34 x 34 5 18 34 5 18 6.14613878554783e-17 4.31307984950725e-17
+ chr1 437 438 55 0 2 46 0 11 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
+ chr1 439 440 56 0 1 55 0 2 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
+ chr1 449 450 56 0 1 13 20 24 x 34.5 10 12.5 34.5 10 12.5 2.25757007974134e-18 2.37638955762246e-18
+ chr1 518 519 56 0 1 38 4 15 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06