Updates to help text for Human Genome Variation tools and formatHelp.html

This commit is contained in:
Richard Burhans
2010-09-28 12:38:52 -04:00
parent e1876ef767
commit 0cd1db87a2
15 changed files with 568 additions and 405 deletions
+205 -137
View File
@@ -1,51 +1,68 @@
<html>
<head><title>Galaxy data formats</title>
<head><title>Galaxy Data Formats</title>
</head>
<body>
<h2>Galaxy data formats</h2>
<h2>Galaxy Data Formats</h2>
<p>
For problems with missing queries in the tool selection boxes the most common
reason is the tool only lists history items with data formats compatible with the
tool. Some formats are subsets of others and Galaxy should also list those
with compatible subformats as well. If the query is not showing up still and you
believe it is in the correct format you can click on the pencil icon and
manually change the format. This will not edit the file just change the
metadata for the file. Some cases you will need to actually change the
file format. For example, if the file is space delimited and a
tabular file is required; then the "Convert delimiters to TAB" tool under "Text
Manipulation" can be used to reformat the file.
<br>
<h3>Dataset missing?</h3>
<p>
If you have a dataset in your history that is not appearing in the
drop-down selector for a tool, the most common reason is that it has
the wrong format. Each Galaxy dataset has an associated file format
recorded in its metadata, and tools will only list datasets from your
history that have a format compatible with that particular tool. Of
course some of these datasets might not actually contain relevant
data, or even the correct columns needed by the tool, but filtering
by format at least makes the list to select from a bit shorter.
<p>
Some of the formats are defined hierarchically, going from very
general ones like <a href="#tab">tabular</a> (which includes any text
file with tab-separated columns), to more restrictive sub-formats
like <a href="#interval">interval</a> (where three of the columns
must be the chromosome, start position, and end position), and on
to even more specific ones such as <a href="#bed">BED</a> or
<a href="#gff">GFF</a> that have additional requirements. So for
example if a tool's required input format is tabular, then all of
your history items whose format is recorded as tabular will be
listed, along with those in all sub-formats that also qualify as
tabular (interval, BED, GFF, etc.).
<p>
There are two usual methods for changing a dataset's format in
Galaxy: if the file contents are already in the required format but
the metadata is wrong (perhaps because the Auto-detect feature of the
Upload File tool guessed it incorrectly), you can fix the metadata
manually by clicking on the pencil icon beside that dataset in your
history. Or, if the file contents really are in a different format,
Galaxy provides a number of format conversion tools (e.g. in the
Text Manipulation and Convert Formats categories). For instance,
if the tool you want to run requires tabular but your columns are
delimited by spaces or commas, you can use the "Convert delimiters
to TAB" tool under Text Manipulation to reformat your data. However
if your files are in a completely unsupported format, then you need
to convert them yourself before uploading.
<p>
Some of the most commonly used formats are very similar. Start with the
basic tabular file. It has few requirements other than 1 or more columns of
data separated by tabs. Next is intervals which are tabular but they have the
added requirement that 3 of the columns must be the chromosome, start point,
and end point. There is optionally a strand and header labelling the
columns. Next is BED or GFF, which are also tabular and intervals, but with more
restrictions. BED can vary between 3 and 12 columns, with each being
precisely defined. Here
the order of the columns also matters, and only the end columns can be skipped.
Some groups of the columns have to be all there or all left off. GFF is
similar in setup but with all 9 columns required and different definitions.
See more detailed descriptions below.
<hr>
<H2>Formats</H2>
<h3>Format Descriptions</h3>
<ul>
<li><a href="#ab1">Ab1</a>
<li><a href="#axt">AXT</a>
<li><a href="#bam">BAM</a>
<li><a href="#binseq">Binseq.zip</a>
<li><a href="#bed">BED</a>
<li><a href="#bedgraph">BedGraph</a>
<li><a href="#binseq">Binseq.zip</a>
<li><a href="#fasta">FASTA</a>
<li><a href="#fastqsolexa">FastqSolexa</a>
<li><a href="#fped">fped</a>
<li><a href="#gff">GFF</a>
<li><a href="#gff3">GFF3</a>
<li><a href="#gtf">GTF</a>
<li><a href="#html">Html</a>
<li><a href="#html">HTML</a>
<li><a href="#interval">Interval</a>
<li><a href="#lav">LAV</a>
<li><a href="#lped">Lped</a>
<li><a href="#lped">lped</a>
<li><a href="#maf">MAF</a>
<li><a href="#pbed">pbed</a>
<li><a href="#psl">PSL</a>
@@ -57,20 +74,26 @@ See more detailed descriptions below.
<li><a href="#wig">Wiggle custom track</a>
<li><a href="#text">Other text type</a>
</ul>
<br>
<p>
<hr>
<strong>Ab1</strong>
<a name="ab1"/>
<p>
A binary sequence file in 'ab1' format with a '.ab1' file extension. You must manually select this 'File Format' when uploading the file.
A binary sequence file in 'ab1' format with a '.ab1' file extension.
You must manually select this file format when uploading the file.
<hr/>
<strong>AXT</strong>
<a name="axt"/>
<p>
blastz pairwise alignment format. Each alignment block in an axt file contains three lines: a summary line and 2 sequence lines. Blocks are separated from one another by blank lines. The summary line contains chromosomal position and size information about the alignment. It consists of 9 required fields.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/axt.html">Click here</a> for more information about axt format.
Used for pairwise alignment output from BLASTZ, after post-processing.
Each alignment block contains three lines: a summary line and two
sequence lines. Blocks are separated from one another by blank lines.
The summary line contains chromosomal position and size information
about the alignment, and consists of nine required fields.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/axt.html"
>More information</a>
<dl><dt>Can be converted to:
<dd><ul>
<li>FASTA<br/>
@@ -80,12 +103,13 @@ Convert Formats&rarr;AXT to LAV
</ul></dl>
<hr/>
<strong>Bam</strong>
<strong>BAM</strong>
<a name="bam"/>
<p>
A binary file compressed in the BGZF format with a '.bam' file extension.
<a href="http://samtools.sourceforge.net/SAM1.pdf">SAM</a>
format is the human readable text version of these files.
A binary file compressed in the BGZF format with a '.bam' file
extension.
<a href="http://samtools.sourceforge.net/SAM1.pdf">SAM</a> format
is the human readable text version of these files.
<dl><dt>Can be converted to:
<dd><ul>
<li>pileup<br/>
@@ -96,23 +120,21 @@ NGS: SAM Tools&rarr;Pileup-to-Interval
</ul></dl>
<hr/>
<strong>Binseq.zip</strong>
<a name="binseq"/>
<p>
A zipped archive consisting of binary sequence files in either 'ab1' or 'scf' format. All files in this archive must have the same file extension which is one of '.ab1' or '.scf'. You must manually select this 'File Format' when uploading the file.
<hr/>
<strong>BED</strong>
<a name="bed"/>
<p>
<ul>
<li> also tabular
<li> also interval
<li> also qualifies as tabular
<li> also qualifies as interval
</ul>
This describes a genomic interval, but has strict field specifications for use
in browsers.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/hgTracksHelp.html#BED">Click here</a> for field specifications.<br/>
This tab-separated format describes a genomic interval, but has
strict field specifications for use in genome browsers. BED files
can have from 3 to 12 columns, but the order of the columns matters,
and only the end ones can be omitted. Some groups of columns must
be all present or all absent.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/hgTracksHelp.html#BED"
>Field specifications</a>
<p>
Example:
<pre>
chr22 1000 5000 cloneA 960 + 1000 5000 0 2 567,488, 0,3512
@@ -129,22 +151,35 @@ Convert Formats&rarr;BED-to-GFF
<a name="bedgraph"/>
<p>
<ul>
<li> also tabular
<li> also interval
<li> also BED
<li> also qualifies as tabular
<li> also qualifies as interval
<li> also qualifies as BED
</ul>
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/bedgraph.html">BedGraph</a>
is a BED file with the name column being a float value that is displayed
as a Wiggle in tracks. Unlike wiggles this score can be retrieved in its
exact value after being loaded as a track.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/bedgraph.html"
>BedGraph</a> is a BED file with the name column being a float value
that is displayed as a wiggle score in tracks. Unlike in Wiggle
format, the exact value of this score can be retrieved after being
loaded as a track.
<hr/>
<strong>Fasta</strong>
<strong>Binseq.zip</strong>
<a name="binseq"/>
<p>
A zipped archive consisting of binary sequence files in either
'ab1' or 'scf' format. All files in this archive must have the same
file extension which is one of '.ab1' or '.scf'. You must manually
select this file format when uploading the file.
<hr/>
<strong>FASTA</strong>
<a name="fasta"/>
<p>
A sequence in
<a href="http://www.ncbi.nlm.nih.gov/blast/fasta.shtml">FASTA format</a>
consists of a single-line description, followed by lines of sequence data. The first character of the description line is a greater-than ("&gt;") symbol in the first column. All lines should be shorter than 80 characters::
<a href="http://www.ncbi.nlm.nih.gov/blast/fasta.shtml">FASTA</a>
format consists of a single-line description, followed by lines of
sequence data. The first character of the description line is a
greater-than ("&gt;") symbol. All lines should be shorter than 80
characters.
<pre>
>sequence1
atgcgtttgcgtgc
@@ -164,7 +199,8 @@ Convert Formats&rarr;FASTA-to-Tabular
<a name="fastqsolexa"/>
<p>
<a href="http://maq.sourceforge.net/fastq.shtml">FastqSolexa</a>
is the Illumina (Solexa) variant of the Fastq format, which stores sequences and quality scores in a single file
is the Illumina (Solexa) variant of the Fastq format, which stores
sequences and quality scores in a single file.
<pre>
@seq1
GACAGCTTGGTTTTTAGTGAGTTGTTCCTTTCTTT
@@ -174,9 +210,9 @@ hhhhhhhhhhhhhhhhhhhhhhhhhhPW@hhhhhh
GCAATGACGGCAGCAATAAACTCAACAGGTGCTGG
+seq2
hhhhhhhhhhhhhhYhhahhhhWhAhFhSIJGChO
</pre>
Or
<pre>
@seq1
GAATTGATCAGGACATAGGACAACTGTAGGCACCAT
+seq1
@@ -196,19 +232,22 @@ Convert Formats&rarr;FASTQ to FASTA
<strong>fped</strong>
<a name="fped"/>
<p>
Also known as the FBAT format, for use in the <a href="http://biosun1.harvard.edu/~fbat/fbat.htm">FBAT</a> program.
It consists of a pedigree file and an phenotype file.
Also known as the FBAT format, for use with the
<a href="http://biosun1.harvard.edu/~fbat/fbat.htm">FBAT</a> program.
It consists of a pedigree file and a phenotype file.
<hr/>
<strong>Gff</strong>
<strong>GFF</strong>
<a name="gff"/>
<p>
<ul>
<li> also tabular
<li> also interval
<li> also qualifies as tabular
<li> also qualifies as interval
</ul>
GFF lines have
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format3">nine required fields</a> that must be tab-separated.
GFF is a tab-separated format somewhat similar to BED, but it has
different columns and is more flexible. There are
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format3"
>nine required fields</a>.
<dl><dt>Can be converted to:
<dd><ul>
<li>BED<br/>
@@ -220,66 +259,80 @@ Convert Formats&rarr;GFF-to-BED
<a name="gff3"/>
<p>
<ul>
<li> also tabular
<li> also interval
<li> also qualifies as tabular
<li> also qualifies as interval
</ul>
The
<a href="http://www.sequenceontology.org/gff3.shtml">GFF3</a>
format addresses the most common extensions to GFF, while preserving backward compatibility with previous formats.
The <a href="http://www.sequenceontology.org/gff3.shtml">GFF3</a>
format addresses the most common extensions to GFF, while preserving
backward compatibility with previous formats.
<hr/>
<strong>GTF</strong>
<a name="gtf"/>
<p>
<ul>
<li> also tabular
<li> also interval
<li> also qualifies as tabular
<li> also qualifies as interval
</ul>
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format4">GTF</a>
is a format for describing genes and other features associated with DNA, RNA and Protein sequences.
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format4"
>GTF</a> is a format for describing genes and other features
associated with DNA, RNA, and protein sequences.
<dl><dt>Can be converted to:
<dd><ul>
<li>BED graph<br/>
<li>BedGraph<br/>
Convert Formats&rarr;GTF-to-BEDGraph
</ul></dl>
<hr/>
<strong>Html</strong>
<strong>HTML</strong>
<a name="html"/>
<p>
This format is a html web page. Click the eye icon to view the dataset in
your browser.
This format is an HTML web page. Click the eye icon next to the
dataset to view it in your browser.
<hr/>
<strong>Interval (Genomic Intervals)</strong>
<strong>Interval</strong>
<a name="interval">
<p>
<ul>
<li> also tabular
<li> also qualifies as tabular
</ul>
Required fields
This Galaxy format represents genomic intervals. It is tab-separated,
but has the added requirement that three of the columns must be the
chromosome name, start position, and end position. An optional
strand column can also be specified, and an initial header row can
be used to label the columns, which do not have to be in any special
order. Arbitrary additional columns can also be present.
<p>
Required fields:
<ul>
<li>CHROM - The name of the chromosome (e.g. chr3, chrY, chr2_random) or contig (e.g. ctgY1).
<li>START - The starting position of the feature in the chromosome or contig. The first base in a chromosome is numbered 0.
<li>END - The ending position of the feature in the chromosome or contig. The chromEnd base is not included in the display of the feature. For example, the first 100 bases of a chromosome are defined as chromStart=0, chromEnd=100, and span the bases numbered 0-99.
<li>CHROM - The name of the chromosome (e.g. chr3, chrY, chr2_random)
or contig (e.g. ctgY1).
<li>START - The starting position of the feature in the chromosome or
contig. The first base in a chromosome is numbered 0.
<li>END - The ending position of the feature in the chromosome or
contig. This base is not included in the feature. For example,
the first 100 bases of a chromosome are described as START=0,
END=100, and span the bases numbered 0-99.
</ul>
Optional
Optional:
<ul>
<li>STRAND - Defines the strand - either '+' or '-'.
<li>Headers
<li>STRAND - Defines the strand, either '+' or '-'.
<li>Header row
</ul>
Example:
<pre>
#CHROM START END STRAND NAME COMMENT
chr1 10 100 + exon myExon
chrX 1000 10050 - gene myGene
#CHROM START END STRAND NAME COMMENT
chr1 10 100 + exon myExon
chrX 1000 10050 - gene myGene
</pre>
<dl><dt>Can be converted to:
<dd><ul>
<li>BED<br/>
The exact changes needed and tools to run can vary with what fields are in
the interval file and what size BED you are converting to. In general you
will likely use Text Manipulation&rarr;Compute, Cut or Merge Columns.
The exact changes needed and tools to run will vary with what fields
are in the interval file and what type of BED you are converting to.
In general you will likely use Text Manipulation&rarr;Compute, Cut,
or Merge Columns.
</ul></dl>
<hr/>
@@ -287,7 +340,8 @@ will likely use Text Manipulation&rarr;Compute, Cut or Merge Columns.
<a name="lav"/>
<p>
<a href="http://www.bx.psu.edu/miller_lab/dist/lav_format.html">LAV</a>
is the primary output format for BLASTZ. The first line of a .lav file begins with #:lav..
is the raw pairwise alignment format that is output by BLASTZ. The
first line begins with <code>#:lav</code>.
<dl><dt>Can be converted to:
<dd><ul>
<li>BED<br/>
@@ -295,18 +349,20 @@ Convert Formats&rarr;LAV to BED
</ul></dl>
<hr/>
<strong>Lped</strong>
<strong>lped</strong>
<a name="lped"/>
<p>
This is the linkage pedigree format (separate map and ped files).
These files together describe SNPs, the map file has the position and an
identifier for the SNP and the pedigree file has the alleles.
To upload this format into Galaxy do not use auto-detect for the file
format, instead select lped. You will then be given two sections for uploading
files, one for the pedigree file and one for the map file.
For more information see <a href="http://www.broadinstitute.org/science/programs/medical-and-population-genetics/haploview/input-file-formats-0">linkage pedigree</a>
or <a href="http://pngu.mgh.harvard.edu/~purcell/plink/data.shtml#map">map</a>
or <a href="http://pngu.mgh.harvard.edu/~purcell/plink/data.shtml#ped">ped</a>.
This is the linkage pedigree format, which consists of separate
<code>map</code> and <code>ped</code> files. Together these files
describe SNPs; the map file contains the position and an identifier
for the SNP, while the pedigree file has the alleles.
To upload this format into Galaxy, do not use auto-detect for the
file format; instead select <code>lped</code>. You will then be
given two sections for uploading files, one for the pedigree file
and one for the map file. For more information, see
<a href="http://www.broadinstitute.org/science/programs/medical-and-population-genetics/haploview/input-file-formats-0">linkage pedigree</a>,
<a href="http://pngu.mgh.harvard.edu/~purcell/plink/data.shtml#map">map</a>,
and/or <a href="http://pngu.mgh.harvard.edu/~purcell/plink/data.shtml#ped">ped</a>.
<dl><dt>Can be converted to:
<dd><ul>
<li>pbed<br/>Automatic
@@ -317,16 +373,20 @@ or <a href="http://pngu.mgh.harvard.edu/~purcell/plink/data.shtml#ped">ped</a>.
<strong>MAF</strong>
<a name="maf"/>
<p>
TBA and multiz multiple alignment format. The first line of a .maf file begins with ##maf. This word is followed by white-space-separated "variable=value pairs". There should be no white space surrounding the "=".
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format5">Click here</a> for more about MAF format.
Multiple alignment format that is output by TBA and Multiz. The
first line begins with <code>##maf</code>. This word is followed by
whitespace-separated "variable=value pairs". There should be no
whitespace surrounding the "=".
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format5"
>More information</a>
<dl><dt>Can be converted to:
<dd><ul>
<li>BED<br/>
Convert Formats&rarr;Maf to BED
<li>Interval<br/>
Convert Formats&rarr;Maf to Interval
Convert Formats&rarr;MAF to BED
<li>interval<br/>
Convert Formats&rarr;MAF to Interval
<li>FASTA<br/>
Convert Formats&rarr;Maf to FASTA
Convert Formats&rarr;MAF to FASTA
</ul></dl>
<hr/>
@@ -344,7 +404,7 @@ This is the binary version of the lped file format.
<a name="psl"/>
<p>
<a href="http://main.genome-browser.bx.psu.edu/FAQ/FAQformat#format2">PSL</a>
format is for alignments, it is returned by
format is used for alignments returned by
<a href="http://genome.ucsc.edu/cgi-bin/hgBlat?command=start">BLAT</a>.
It does not include any sequence.
<hr/>
@@ -352,9 +412,10 @@ It does not include any sequence.
<strong>Scf</strong>
<a name="scf"/>
<p>
A binary sequence file in 'scf' format with a '.scf' file extension. You must manually select this 'File Format' when uploading the file.
<a href="http://staden.sourceforge.net/manual/formats_unix_2.html">Click here</a>
for more information.
A binary sequence file in 'scf' format with a '.scf' file extension.
You must manually select this file format when uploading the file.
<a href="http://staden.sourceforge.net/manual/formats_unix_2.html"
>More information</a>
<hr/>
<strong>Sff</strong>
@@ -373,58 +434,65 @@ Convert Formats&rarr;SFF converter
<strong>Table</strong>
<a name="table"/>
<p>
Text delimited into columns by something other than a tab.
Text data separated into columns by something other than tabs.
<hr/>
<strong>Tabular (tab delimited)</strong>
<strong>Tabular (tab-delimited)</strong>
<a name="tab"/>
<p>
Any data in tab delimited format (tabular)
One or more columns of text data separated by tabs.
<dl><dt>Can be converted to:
<dd><ul>
<li>FASTA<br/>
Convert Formats&rarr;Tabular-to-FASTA<br/>
Tabular file must have a title and sequence column.
The tabular file must have a title and sequence column.
<li>interval<br/>
If the tabular file has the chromosome, or is all on one chromosome, and
a position you can create an interval file. If all one chromosome use
Text Manipulation&rarr;Add column to add the chromosome. If the given position
is a 1 based position use Text Manipulation&rarr;Compute and the position
column minus 1 to get the start. Otherwise do plus 1 to get the end.
If the tabular file has the chromosome, or is all on one chromosome,
and has a position you can create an interval file (e.g. for SNPs).
If it is all on one chromosome, use Text Manipulation&rarr;Add column
to add a chromosome column. If the given position is 1-based, use
Text Manipulation&rarr;Compute with the position column minus 1 to
get the start, and use the original given column for the end.
If the given position is 0-based, use it as the start, and compute
that plus 1 to get the end.
</ul></dl>
<hr/>
<strong>Txtseq.zip</strong>
<a name="txtseqzip"/>
<p>
A zipped archive consisting of flat text sequence files. All files in this archive must have the same file extension of '.txt'. You must manually select this 'File Format' when uploading the file.
A zipped archive consisting of flat text sequence files. All files
in this archive must have the same file extension of '.txt'. You
must manually select this file format when uploading the file.
<hr/>
<strong>Wiggle custom track</strong>
<a name="wig"/>
<p>
The wiggle format is line-oriented. Wiggle data is preceded by a track definition line, which gives the type of wiggle. There are 3 different types, each
with their uses.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/wiggle.html">More information here.</a>
The wiggle format is line-oriented. Wiggle data is preceded by a
track definition line, which specifies the type of wiggle. There
are three different types, for different uses.
<a href="http://main.genome-browser.bx.psu.edu/goldenPath/help/wiggle.html"
>More information</a>
<dl><dt>Can be converted to:
<dd><ul>
<li>interval<br/>
Convert Formats&rarr;Wiggle-to-Interval<br/>
As a second step this could be converted to BED 3 or 4 by removing columns.
Text Manipulation&rarr;Cut columns from a table
As a second step this could be converted to BED-3 or BED-4 by removing
columns, using Text Manipulation&rarr;Cut columns from a table.
</ul></dl>
<hr/>
<strong>Other text type</strong>
<a name="text"/>
<p>
Any text file
Any text file.
<dl><dt>Can be converted to:
<dd><ul>
<li>tabular<br/>
If this is space or some other delimiter separated fields it can be
converted to tabular. Text Manipulations&rarr;Convert delimiters to TAB
If this has fields separated by spaces, commas, or some other
delimiter it can be converted to tabular using
Text Manipulation&rarr;Convert delimiters to TAB
</ul></dl>
<!-- blank lines so internal links will jump farther to end -->
<br/><br/><br/><br/><br/><br/><br/><br/><br/><br/><br/><br/>
+29 -27
View File
@@ -40,10 +40,11 @@ This currently works only for build hg18.
**Dataset formats**
The input dataset can be any interval_ format dataset.
The output dataset is also interval format.
The input can be any interval_ format dataset. The output is also in interval format.
(`Dataset missing?`_)
.. _interval: ./static/formatHelp.html#interval
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -64,41 +65,42 @@ only the score for the first nucleotide is returned.
- input file, with SNPs::
chr22 16440426 14440427 C/T
chr22 15494851 14494852 A/G
chr22 14494911 14494912 A/T
chr22 14550435 14550436 A/G
chr22 14611956 14611957 G/T
chr22 14612076 14612077 A/G
chr22 14668537 14668538 C
chr22 14668703 14668704 A/T
chr22 14668775 14668776 G
chr22 14680074 14680075 A/T
chr22 16440426 14440427 C/T
chr22 15494851 14494852 A/G
chr22 14494911 14494912 A/T
chr22 14550435 14550436 A/G
chr22 14611956 14611957 G/T
chr22 14612076 14612077 A/G
chr22 14668537 14668538 C
chr22 14668703 14668704 A/T
chr22 14668775 14668776 G
chr22 14680074 14680075 A/T
etc.
- output file, showing conservation scores for primates::
chr22 16440426 14440427 C/T 0.509
chr22 15494851 14494852 A/G 0.427
chr22 14494911 14494912 A/T NA
chr22 14550435 14550436 A/G NA
chr22 14611956 14611957 G/T -2.142
chr22 14612076 14612077 A/G 0.369
chr22 14668537 14668538 C 0.419
chr22 14668703 14668704 A/T -1.462
chr22 14668775 14668776 G 0.470
chr22 14680074 14680075 A/T 0.303
chr22 16440426 14440427 C/T 0.509
chr22 15494851 14494852 A/G 0.427
chr22 14494911 14494912 A/T NA
chr22 14550435 14550436 A/G NA
chr22 14611956 14611957 G/T -2.142
chr22 14612076 14612077 A/G 0.369
chr22 14668537 14668538 C 0.419
chr22 14668703 14668704 A/T -1.462
chr22 14668775 14668776 G 0.470
chr22 14680074 14680075 A/T 0.303
etc.
"NA" means that the phyloP score was not available.
"NA" means that the phyloP score was not available.
-----
**Reference**
Siepel A, Pollard KS, and Haussler D. New methods for detecting
lineage-specific selection. In Proceedings of the 10th International
Conference on Research in Computational Molecular Biology (RECOMB
2006), pp. 190-205.
Siepel A, Pollard KS, Haussler D. (2006)
New methods for detecting lineage-specific selection.
In Proceedings of the 10th International Conference on Research in Computational
Molecular Biology (RECOMB 2006), pp. 190-205.
</help>
</tool>
+40 -38
View File
@@ -53,12 +53,12 @@ Use the pencil icon to add the build to the files if necessary.
**Dataset formats**
The SNP dataset is in interval_ format, with a column of SNPs as described below.
The gene dataset is in BED_ format with 12 columns.
The output dataset is also interval.
The gene dataset is in BED_ format with 12 columns. The output dataset is also interval.
(`Dataset missing?`_)
.. _interval: ./static/formatHelp.html#interval
.. _BED: ./static/formatHelp.html#bed
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -87,51 +87,53 @@ or is synonymous then it is not included in the output file.
- first input file, with SNPs::
chr22 15660821 15660822 A/G
chr22 15825725 15825726 G/T
chr22 15827035 15827036 G
chr22 15827135 15827136 C/G
chr22 15830928 15830929 A/G
chr22 15830951 15830952 G
chr22 15830955 15830956 C/T
chr22 15848885 15848886 C/T
chr22 15849048 15849049 A/C
chr22 15919711 15919712 A/G
chr22 15660821 15660822 A/G
chr22 15825725 15825726 G/T
chr22 15827035 15827036 G
chr22 15827135 15827136 C/G
chr22 15830928 15830929 A/G
chr22 15830951 15830952 G
chr22 15830955 15830956 C/T
chr22 15848885 15848886 C/T
chr22 15849048 15849049 A/C
chr22 15919711 15919712 A/G
etc.
alternatively indicating polymorphisms using ambiguous-nucleotide symbols:
chr22 15660821 15660822 R
chr22 15825725 15825726 K
chr22 15827035 15827036 G
chr22 15827135 15827136 S
chr22 15830928 15830929 R
chr22 15830951 15830952 G
chr22 15830955 15830956 Y
chr22 15848885 15848886 Y
chr22 15849048 15849049 M
chr22 15919711 15919712 R
or, indicating polymorphisms using ambiguous-nucleotide symbols::
chr22 15660821 15660822 R
chr22 15825725 15825726 K
chr22 15827035 15827036 G
chr22 15827135 15827136 S
chr22 15830928 15830929 R
chr22 15830951 15830952 G
chr22 15830955 15830956 Y
chr22 15848885 15848886 Y
chr22 15849048 15849049 M
chr22 15919711 15919712 R
etc.
- second input file, with UCSC annotations for human genes::
chr22 15688363 15690225 uc010gqr.1 0 + 15688363 15688363 0 2 587,794, 0,1068,
chr22 15822826 15869112 uc002zlw.1 0 - 15823622 15869004 0 10 940,105,97,91,265,86,251,208,304,282, 0,1788,2829,3241,4163,6361,8006,26023,29936,46004,
chr22 15826991 15869112 uc010gqs.1 0 - 15829218 15869004 0 5 1380,86,157,304,282, 0,2196,21858,25771,41839,
chr22 15897459 15919682 uc002zlx.1 0 + 15897459 15897459 0 4 775,128,103,1720, 0,8303,10754,20503,
chr22 15945848 15971389 uc002zly.1 0 + 15945981 15970710 0 13 271,25,147,113,127,48,164,84,85,12,102,42,2193, 0,12103,12838,13816,15396,17037,17180,18535,19767,20632,20894,22768,23348,
chr22 15688363 15690225 uc010gqr.1 0 + 15688363 15688363 0 2 587,794, 0,1068,
chr22 15822826 15869112 uc002zlw.1 0 - 15823622 15869004 0 10 940,105,97,91,265,86,251,208,304,282, 0,1788,2829,3241,4163,6361,8006,26023,29936,46004,
chr22 15826991 15869112 uc010gqs.1 0 - 15829218 15869004 0 5 1380,86,157,304,282, 0,2196,21858,25771,41839,
chr22 15897459 15919682 uc002zlx.1 0 + 15897459 15897459 0 4 775,128,103,1720, 0,8303,10754,20503,
chr22 15945848 15971389 uc002zly.1 0 + 15945981 15970710 0 13 271,25,147,113,127,48,164,84,85,12,102,42,2193, 0,12103,12838,13816,15396,17037,17180,18535,19767,20632,20894,22768,23348,
etc.
- output file, showing non-synonymous substitutions in coding regions::
chr22 15825725 15825726 G/T uc002zlw.1 Gln:Pro/Gln 469 T
chr22 15827035 15827036 G uc002zlw.1 Glu:Asp 414 C
chr22 15827135 15827136 C/G uc002zlw.1 Gly:Gly/Ala 381 C
chr22 15830928 15830929 A/G uc002zlw.1 Ala:Ser/Pro 281 C
chr22 15830951 15830952 G uc002zlw.1 Leu:Pro 273 A
chr22 15830955 15830956 C/T uc002zlw.1 Ser:Gly/Ser 272 T
chr22 15848885 15848886 C/T uc002zlw.1 Ser:Trp/Stop 217 G
chr22 15848885 15848886 C/T uc010gqs.1 Ser:Trp/Stop 200 G
chr22 15849048 15849049 A/C uc002zlw.1 Gly:Stop/Gly 163 C
chr22 15825725 15825726 G/T uc002zlw.1 Gln:Pro/Gln 469 T
chr22 15827035 15827036 G uc002zlw.1 Glu:Asp 414 C
chr22 15827135 15827136 C/G uc002zlw.1 Gly:Gly/Ala 381 C
chr22 15830928 15830929 A/G uc002zlw.1 Ala:Ser/Pro 281 C
chr22 15830951 15830952 G uc002zlw.1 Leu:Pro 273 A
chr22 15830955 15830956 C/T uc002zlw.1 Ser:Gly/Ser 272 T
chr22 15848885 15848886 C/T uc002zlw.1 Ser:Trp/Stop 217 G
chr22 15848885 15848886 C/T uc010gqs.1 Ser:Trp/Stop 200 G
chr22 15849048 15849049 A/C uc002zlw.1 Gly:Stop/Gly 163 C
etc.
</help>
</tool>
+35 -30
View File
@@ -56,12 +56,12 @@ single-SNP associations), please use the GPASS tool instead.**
**Dataset formats**
The input dataset is in lped_ format. The output datasets are both
tabular_.
The input dataset must be in lped_ format. The output datasets are both tabular_.
(`Dataset missing?`_)
.. _lped: ./static/formatHelp.html#lped
.. _tabular: ./static/formatHelp.html#tabular
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -77,9 +77,9 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD).
- input map file::
1 rs0 0 738547
1 rs1 0 5597094
1 rs2 0 9424115
1 rs0 0 738547
1 rs1 0 5597094
1 rs2 0 9424115
etc.
- input ped file::
@@ -90,33 +90,33 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD).
- first output file, significance.txt::
ID chr position results
rs0 chr1 738547 10 20 score= 45.101397 , df= 8 , p= 0.000431 , N=1225
ID chr position results
rs0 chr1 738547 10 20 score= 45.101397 , df= 8 , p= 0.000431 , N=1225
- second output file, posterior.txt::
id: chr position marginal + interaction = total posterior
0: 1 738547 0.0000 + 0.0000 = 0.0000
1: 1 5597094 0.0000 + 0.0000 = 0.0000
2: 1 9424115 0.0000 + 0.0000 = 0.0000
3: 1 13879818 0.0000 + 0.0000 = 0.0000
4: 1 13934751 0.0000 + 0.0000 = 0.0000
5: 1 16803491 0.0000 + 0.0000 = 0.0000
6: 1 17236854 0.0000 + 0.0000 = 0.0000
7: 1 18445387 0.0000 + 0.0000 = 0.0000
8: 1 21222571 0.0000 + 0.0000 = 0.0000
id: chr position marginal + interaction = total posterior
0: 1 738547 0.0000 + 0.0000 = 0.0000
1: 1 5597094 0.0000 + 0.0000 = 0.0000
2: 1 9424115 0.0000 + 0.0000 = 0.0000
3: 1 13879818 0.0000 + 0.0000 = 0.0000
4: 1 13934751 0.0000 + 0.0000 = 0.0000
5: 1 16803491 0.0000 + 0.0000 = 0.0000
6: 1 17236854 0.0000 + 0.0000 = 0.0000
7: 1 18445387 0.0000 + 0.0000 = 0.0000
8: 1 21222571 0.0000 + 0.0000 = 0.0000
etc.
id: chr position block_boundary | allele counts in cases and controls
0: 1 738547 1.000 | 156 93 251 | 169 83 248
1: 1 5597094 1.000 | 323 19 158 | 328 16 156
2: 1 9424115 1.000 | 366 6 128 | 369 11 120
3: 1 13879818 1.000 | 252 31 217 | 278 32 190
4: 1 13934751 1.000 | 246 64 190 | 224 58 218
5: 1 16803491 1.000 | 91 160 249 | 91 174 235
6: 1 17236854 1.000 | 252 43 205 | 249 44 207
7: 1 18445387 1.000 | 205 66 229 | 217 56 227
8: 1 21222571 1.000 | 353 9 138 | 352 8 140
id: chr position block_boundary | allele counts in cases and controls
0: 1 738547 1.000 | 156 93 251 | 169 83 248
1: 1 5597094 1.000 | 323 19 158 | 328 16 156
2: 1 9424115 1.000 | 366 6 128 | 369 11 120
3: 1 13879818 1.000 | 252 31 217 | 278 32 190
4: 1 13934751 1.000 | 246 64 190 | 224 58 218
5: 1 16803491 1.000 | 91 160 249 | 91 174 235
6: 1 17236854 1.000 | 252 43 205 | 249 44 207
7: 1 18445387 1.000 | 205 66 229 | 217 56 227
8: 1 21222571 1.000 | 353 9 138 | 352 8 140
etc.
The "id" field is an internally used index.
@@ -125,8 +125,13 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD).
**References**
Zhang Y and Liu JS (2007). Bayesian Inference of Epistatic Interactions in Case-Control Studies. Nature Genetics, 39:1167-1173
Zhang Y, Liu JS. (2007)
Bayesian inference of epistatic interactions in case-control studies.
Nat Genet. 39(9):1167-73. Epub 2007 Aug 26.
Zhang Y, Zhang J, Liu JS. (2010)
Block-based bayesian epistasis association mapping with application to WTCCC type 1 diabetes data.
Submitted.
Zhang Y, Zhang J, Liu JS (2010) Block-based Bayesian Epistasis Association Mapping with Application to WTCCC Type 1 Diabetes Data. Submitted.
</help>
</tool>
+15 -9
View File
@@ -256,8 +256,10 @@
**Dataset formats**
The input and output datasets are tabular_.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -280,35 +282,39 @@ Website: http://ctd.mdibl.org/
**Examples**
- input data file:
HBB
HBB
- select column c1, Identifier type = Genes, and Data to extract = All disease relationships
- select Column = c1, Identifier type = Genes, and Data to extract = All disease relationships
- output file::
#Input GeneSymbol GeneName GeneID DiseaseName DiseaseID GeneDiseaseRelation OmimIDs PubMedIDs
hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Ethanol 17676605|18926900
hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Valproic Acid 8875741
hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Ethanol 17676605|18926900
hbb HBB hemoglobin, beta 3043 Abnormalities, Drug-Induced MESH:D000014 inferred via Valproic Acid 8875741
etc.
Another example:
- same input file:
HBB
HBB
- select column c1, Identifier type = Genes, Data to extract = Curated chemical-gene interactions, and Interaction type = ANY
- select Column = c1, Identifier type = Genes, Data to extract = Curated chemical-gene interactions, and Interaction type = ANY
- output file::
#Input GeneSymbol GeneName GeneID ChemicalName ChemicalID CasRN Organism OrganismID Interaction InteractionTypes PubMedIDs
hbb HBB hemoglobin, beta 3043 1-nitronaphthalene C016614 86-57-7 Macaca mulatta 9544 1-nitronaphthalene metabolite binds to HBB protein binding 16453347
hbb HBB hemoglobin, beta 3043 2,6-diisocyanatotoluene C026942 91-08-7 Cavia porcellus 10141 2,6-diisocyanatotoluene binds to HBB protein binding 8728499
hbb HBB hemoglobin, beta 3043 1-nitronaphthalene C016614 86-57-7 Macaca mulatta 9544 1-nitronaphthalene metabolite binds to HBB protein binding 16453347
hbb HBB hemoglobin, beta 3043 2,6-diisocyanatotoluene C026942 91-08-7 Cavia porcellus 10141 2,6-diisocyanatotoluene binds to HBB protein binding 8728499
etc.
-----
**Reference**
Davis AP, Murphy CG, Saraceni-Richards CA, Rosenstein MC, Wiegers TC, Mattingly CJ. Comparative Toxicogenomics Database: a knowledgebase and discovery tool for chemical.gene.disease networks. Nucleic Acids Res. 2009 Jan;37(Database issue):D786-92.
Davis AP, Murphy CG, Saraceni-Richards CA, Rosenstein MC, Wiegers TC, Mattingly CJ. (2009)
Comparative Toxicogenomics Database: a knowledgebase and discovery tool for
chemical-gene-disease networks.
Nucleic Acids Res. 37(Database issue):D786-92. Epub 2008 Sep 9.
</help>
</tool>
+22 -17
View File
@@ -65,32 +65,37 @@ Typing::
results in::
1. 2. 3. 4. 5. 6. 7.
chr11 89507465 89565427 + NAALAD2 10003 Adenocarcinoma
chr15 50189113 50192264 - BCL2L10 10017 Carcinoma
chr7 150535855 150555250 - ABCF2 10061 Clear cell carcinoma
chr7 150540508 150555250 - ABCF2 10061 Clear cell carcinoma
chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
1. 2. 3. 4. 5. 6. 7.
chr11 89507465 89565427 + NAALAD2 10003 Adenocarcinoma
chr15 50189113 50192264 - BCL2L10 10017 Carcinoma
chr7 150535855 150555250 - ABCF2 10061 Clear cell carcinoma
chr7 150540508 150555250 - ABCF2 10061 Clear cell carcinoma
chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
chr10 134925911 134940397 - ADAM8 101 Adenocarcinoma
etc.
where the column contents are as follows::
1. chromosome name.
2. start position of the gene.
3. end position of the gene.
4. strand.
4. gene name.
6. Entrez Gene ID.
7. disease term.
1. chromosome name
2. start position of the gene
3. end position of the gene
4. strand
4. gene name
6. Entrez Gene ID
7. disease term
-----
**References**
Pan Du, Gang Feng,Jared Flatow, Jie Song, Michelle Holko, Warren A. Kibbe1 and
Simon M. Lin. From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose ontology for the test of gene-ontology associations. Bioinformatics (2009) 25 (12):i63-i68.
Du P, Feng G, Flatow J, Song J, Holko M, Kibbe WA, Lin SM. (2009)
From disease ontology to disease-ontology lite: statistical methods to adapt a general-purpose
ontology for the test of gene-ontology associations.
Bioinformatics. 25(12):i63-8.
Osborne JD, Flatow J, Holko M, Lin SM, Kibbe WA, Zhu LJ, Danila MI, Feng G, Chisholm RL. (2009)
Annotating the human genome with Disease Ontology.
BMC Genomics. 10 Suppl 1:S6.
Osborne JD, Flatow J, Holko M, Lin SM, Kibbe WA, Zhu LJ, Danila MI, Feng G, Chisholm RL. Annotating the human genome with Disease Ontology. BMC Genomics (2009) 10 S1:S6.
</help>
</tool>
+14 -11
View File
@@ -36,11 +36,12 @@
<help>
**Dataset formats**
The input dataset must be lped_, and the output is tabular_.
The input dataset must be in lped_ format, and the output is tabular_.
(`Dataset missing?`_)
.. _lped: ./static/formatHelp.html#lped
.. _tabular: ./static/formatHelp.html#tab
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -80,9 +81,9 @@ Otherwise use permutation.
- input map file::
1 rs0 0 738547
1 rs1 0 5597094
1 rs2 0 9424115
1 rs0 0 738547
1 rs1 0 5597094
1 rs2 0 9424115
etc.
- input ped file::
@@ -93,17 +94,19 @@ Otherwise use permutation.
- output dataset, showing significant SNPs and their p-values and FDR::
#ID chr position Statistics adj-Pvalue FDR
rs35 chr1 136606952 4.890849 0.991562 0.682138
rs36 chr1 137748344 4.931934 0.991562 0.795827
rs44 chr2 14423047 7.712832 0.665086 0.218776
#ID chr position Statistics adj-Pvalue FDR
rs35 chr1 136606952 4.890849 0.991562 0.682138
rs36 chr1 137748344 4.931934 0.991562 0.795827
rs44 chr2 14423047 7.712832 0.665086 0.218776
etc.
-----
**Reference**
Zhang Y and Liu JS (2010). Fast and Accurate Significance Approximation for
Genome-wide Association Studies, submitted.
Zhang Y, Liu JS. (2010)
Fast and accurate significance approximation for genome-wide association studies.
Submitted.
</help>
</tool>
+6 -1
View File
@@ -61,8 +61,10 @@
**Dataset formats**
The input format is interval_, and the output is an image in PDF format.
(`Dataset missing?`_)
.. _interval: ./static/formatHelp.html#interval
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -107,6 +109,9 @@ Here are some examples from the HilbertVis homepage, using ChIP-Seq data.
**Reference**
Anders S. (2009) Visualization of genomic data with the Hilbert curve. Bioinformatics, 25:1231-1235.
Anders S. (2009)
Visualization of genomic data with the Hilbert curve.
Bioinformatics. 25(10):1231-5. Epub 2009 Mar 17.
</help>
</tool>
+40 -38
View File
@@ -32,8 +32,10 @@
**Dataset formats**
The input and output datasets are tabular_.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -43,9 +45,8 @@ This tool can be used to analyze the patterns of linkage disequilibrium
(LD) between polymorphic sites in a locus. SNPs are grouped based on the
threshold level of LD as measured by r\ :sup:`2` (regardless of genomic
position), and a representative "tag SNP" is reported for each group.
Note that the groups are generated by transitive closure: each SNP in the
group is within r\ :sup:`2` of *some* other SNP in the group, but not
necessarily all of them (and not necessarily the tag SNP).
The other SNPs in the group are in LD with the tag SNP, but not necessarily
with each other.
The underlying algorithm is the same as the one used in ldSelect (Carlson
et al. 2004). However, this tool is implemented to be much faster and more
@@ -61,49 +62,50 @@ two allele nucleotides.
- input file::
rs2334386 NA20364 G T
rs2334386 NA20363 G G
rs2334386 NA20360 G G
rs2334386 NA20359 G G
rs2334386 NA20358 G G
rs2334386 NA20356 G G
rs2334386 NA20357 G G
rs2334386 NA20350 G G
rs2334386 NA20349 G G
rs2334386 NA20348 G G
rs2334386 NA20347 G G
rs2334386 NA20346 G G
rs2334386 NA20345 G G
rs2334386 NA20344 G G
rs2334386 NA20342 G G
etc.
rs2334386 NA20364 G T
rs2334386 NA20363 G G
rs2334386 NA20360 G G
rs2334386 NA20359 G G
rs2334386 NA20358 G G
rs2334386 NA20356 G G
rs2334386 NA20357 G G
rs2334386 NA20350 G G
rs2334386 NA20349 G G
rs2334386 NA20348 G G
rs2334386 NA20347 G G
rs2334386 NA20346 G G
rs2334386 NA20345 G G
rs2334386 NA20344 G G
rs2334386 NA20342 G G
etc.
- output file::
rs2238748 rs2793064,rs6518516,rs6518517,rs2283641,rs5993533,rs715590,rs2072123,rs2105421,rs2800954,rs1557847,rs807750,rs807753,rs5993488,rs8138035,rs2800980,rs2525079,rs5992353,rs712966,rs2525036,rs807743,rs1034727,rs807744,rs2074003
rs2871023 rs1210715,rs1210711,rs5748189,rs1210709,rs3788298,rs7284649,rs9306217,rs9604954,rs1210703,rs5748179,rs5746727,rs5748190,rs5993603,rs2238766,rs885981,rs2238763,rs5748165,rs9605996,rs9606001,rs5992398
rs7292006 rs13447232,rs5993665,rs2073733,rs1057457,rs756658,rs5992395,rs2073760,rs739369,rs9606017,rs739370,rs4493360,rs2073736
rs2518840 rs1061325,rs2283646,rs362148,rs1340958,rs361956,rs361991,rs2073754,rs2040771,rs2073740,rs2282684
rs2073775 rs10160,rs2800981,rs807751,rs5993492,rs2189490,rs5747997,rs2238743
rs5747263 rs12159924,rs2300688,rs4239846,rs3747025,rs3747024,rs3747023,rs2300691
rs433576 rs9605439,rs1109052,rs400509,rs401099,rs396012,rs410456,rs385105
rs2106145 rs5748131,rs2013516,rs1210684,rs1210685,rs2238767,rs2277837
rs2587082 rs2257083,rs2109659,rs2587081,rs5747306,rs2535704,rs2535694
rs807667 rs2800974,rs756651,rs762523,rs2800973,rs1018764
rs2518866 rs1206542,rs807467,rs807464,rs807462,rs712950
rs1110661 rs1110660,rs7286607,rs1110659,rs5992917,rs1110662
rs759076 rs5748760,rs5748755,rs5748752,rs4819925,rs933461
rs5746487 rs5992895,rs2034113,rs2075455,rs1867353
rs5748212 rs5746736,rs4141527,rs5748147,rs5748202
etc.
rs2238748 rs2793064,rs6518516,rs6518517,rs2283641,rs5993533,rs715590,rs2072123,rs2105421,rs2800954,rs1557847,rs807750,rs807753,rs5993488,rs8138035,rs2800980,rs2525079,rs5992353,rs712966,rs2525036,rs807743,rs1034727,rs807744,rs2074003
rs2871023 rs1210715,rs1210711,rs5748189,rs1210709,rs3788298,rs7284649,rs9306217,rs9604954,rs1210703,rs5748179,rs5746727,rs5748190,rs5993603,rs2238766,rs885981,rs2238763,rs5748165,rs9605996,rs9606001,rs5992398
rs7292006 rs13447232,rs5993665,rs2073733,rs1057457,rs756658,rs5992395,rs2073760,rs739369,rs9606017,rs739370,rs4493360,rs2073736
rs2518840 rs1061325,rs2283646,rs362148,rs1340958,rs361956,rs361991,rs2073754,rs2040771,rs2073740,rs2282684
rs2073775 rs10160,rs2800981,rs807751,rs5993492,rs2189490,rs5747997,rs2238743
rs5747263 rs12159924,rs2300688,rs4239846,rs3747025,rs3747024,rs3747023,rs2300691
rs433576 rs9605439,rs1109052,rs400509,rs401099,rs396012,rs410456,rs385105
rs2106145 rs5748131,rs2013516,rs1210684,rs1210685,rs2238767,rs2277837
rs2587082 rs2257083,rs2109659,rs2587081,rs5747306,rs2535704,rs2535694
rs807667 rs2800974,rs756651,rs762523,rs2800973,rs1018764
rs2518866 rs1206542,rs807467,rs807464,rs807462,rs712950
rs1110661 rs1110660,rs7286607,rs1110659,rs5992917,rs1110662
rs759076 rs5748760,rs5748755,rs5748752,rs4819925,rs933461
rs5746487 rs5992895,rs2034113,rs2075455,rs1867353
rs5748212 rs5746736,rs4141527,rs5748147,rs5748202
etc.
-----
**Reference**
Carlson CS, Eberle MA, Rieder MJ, Yi Q, Kruglyak L, Nickerson DA.
Carlson CS, Eberle MA, Rieder MJ, Yi Q, Kruglyak L, Nickerson DA. (2004)
Selecting a maximally informative set of single-nucleotide polymorphisms for
association analysis using linkage disequilibrium. Am J Hum Genet. 2004 Jan;
74(1):106-20. Epub 2003 Dec 15.
association analyses using linkage disequilibrium.
Am J Hum Genet. 74(1):106-20. Epub 2003 Dec 15.
</help>
</tool>
+9 -3
View File
@@ -74,10 +74,11 @@ The list is limited to 400 IDs.
The input dataset is tabular_ format. The output dataset is html_ format with
a link to the DAVID website as described below.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _html: ./static/formatHelp.html#html
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -97,8 +98,13 @@ lists of genes.
**References**
Huang DW, Sherman BT, Lempicki RA. Systematic and integrative analysis of large gene lists using DAVID Bioinformatics Resources. Nature Protoc. 2009;4(1):44-57.
Huang DW, Sherman BT, Lempicki RA. (2009) Systematic and integrative analysis
of large gene lists using DAVID bioinformatics resources.
Nat Protoc. 4(1):44-57.
Dennis G, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. (2003)
DAVID: database for annotation, visualization, and integrated discovery.
Genome Biol. 4(5):P3. Epub 2003 Apr 3.
Dennis G Jr, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. DAVID: Database for Annotation, Visualization, and Integrated Discovery. Genome Biol. 2003;4(5):P3.
</help>
</tool>
@@ -36,15 +36,17 @@
<output name="out_file1" file="linkToGProfile_1.out" />
</test>
</tests>
<help>
**Dataset formats**
The input dataset is tabular_ with a column of identifiers.
The output dataset is html_ with a link to g:Profiler.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _html: ./static/formatHelp.html#html
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -73,6 +75,9 @@ the results to run other g:Profiler tools using the same list of IDs.
**Reference**
\J. Reimand, M. Kull, H. Peterson, J. Hansen, J. Vilo: g:Profiler -- a web-based toolset for functional profiling of gene lists from large-scale experiments (2007) NAR 35 W193-W200
Reimand J, Kull M, Peterson H, Hansen J, Vilo J. (2007) g:Profiler -- a web-based
toolset for functional profiling of gene lists from large-scale experiments.
Nucleic Acids Res. 35(Web Server issue):W193-200. Epub 2007 May 3.
</help>
</tool>
+35 -18
View File
@@ -172,15 +172,17 @@
<output name="log_file" file="lps_arrhythmia_log.txt"/>
</test>
</tests>
<help>
**Dataset formats**
The input and output datasets are tabular_. The columns are described
below. There is a second output dataset (a log) that is in text_ format.
The input and output datasets are tabular_. The columns are described below.
There is a second output dataset (a log) that is in text_ format.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _text: ./static/formatHelp.html#text
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -217,19 +219,19 @@ Website: http://pages.cs.wisc.edu/~swright/LPS/
- input file::
+1 1 0 0 0 0 1 0 1 1 ...
+1 1 1 1 0 0 1 0 1 1 ...
+1 1 0 1 0 1 0 1 0 1 ...
+1 1 0 0 0 0 1 0 1 1 ...
+1 1 1 1 0 0 1 0 1 1 ...
+1 1 0 1 0 1 0 1 0 1 ...
etc.
- output results file::
0
0
0
0
0.025541
etc.
0
0
0
0
0.025541
etc.
- output log file::
@@ -243,16 +245,31 @@ Website: http://pages.cs.wisc.edu/~swright/LPS/
**References**
Koh K, Kim S-J, and Boyd S. (2007) An Interior-Point Method for Large-Scale l1-Regularized Logistic Regression. Journal of Machine Learning Research, 8:1519-1555.
Koh K, Kim S-J, Boyd S. (2007)
An interior-point method for large-scale l1-regularized logistic regression.
Journal of Machine Learning Research. 8:1519-1555.
Shi W, Wahba G, Wright S, Lee K, Klein R, and Klein B. (2008) LASSO-Patternsearch Algorithm with Application to Ophthalmology and Genomic Data. Statistics And Its Interface, 1:137-153.
Shi W, Wahba G, Wright S, Lee K, Klein R, Klein B. (2008)
LASSO-Patternsearch algorithm with application to ophthalmology and genomic data.
Stat Interface. 1(1):137-153.
Wright S, Novak R, and Figueiredo M. (2009) Sparse reconstruction via separable approximation. IEEE Transactions on Signal Processing, 57:2479-2403.
<!--
Wright S, Novak R, Figueiredo M. (2009)
Sparse reconstruction via separable approximation.
IEEE Transactions on Signal Processing. 57:2479-2403.
Shi J, Yin W, Osher S, and Sajda P. (2010) A fast hybrid algorithm for large scale l1-regularized logistic regression. Journal of Machine Learning Research, 11:713-741.
Shi J, Yin W, Osher S, Sajda P. (2010)
A fast hybrid algorithm for large scale l1-regularized logistic regression.
Journal of Machine Learning Research. 11:713-741.
Byrd R, Chin G, Neveitt W, and Nocedal J. (2010) On the use of stochastic Hessian information in unconstrained optimization. Technical Report. Northwestern University. June 16, 2010.
Byrd R, Chin G, Neveitt W, Nocedal J. (2010)
On the use of stochastic Hessian information in unconstrained optimization.
Technical Report. Northwestern University. June 16, 2010.
Wright S. (2010)
Accelerated block-coordinate relaxation for regularized optimization.
Technical Report. University of Wisconsin. August 10, 2010.
-->
Wright S. (2010) Accelerated block-coordinate relaxation for regularized optimization. Technical Report. University of Wisconsin. August 10, 2010.
</help>
</tool>
+56 -29
View File
@@ -6,10 +6,10 @@
</command>
<inputs>
<param format="gff" name="input" type="data" label="GFF dataset"/>
<param format="gff" name="input" type="data" label="Dataset"/>
<param name="min_window" label="Smallest window size (by # of probes)" type="integer" value="2" />
<param name="max_window" label="Largest window size (by # of probes)" type="integer" value="6" />
<param name="false_num" label="Expected number of false positive intervals to be called" type="float" value="0.05" />
<param name="false_num" label="Expected total number of false positive intervals to be called" type="float" value="5.0" help="N.B.: this is a &lt;em&gt;count&lt;/em&gt;, not a rate." />
</inputs>
<outputs>
@@ -21,9 +21,8 @@
<requirement type="binary">sed</requirement>
</requirements>
<!-- we need to be able to set the seed for the random number generator
<test>
<tests>
<test>
<param name="input" ftype="gff" value="pass_input.gff"/>
<param name="min_window" value="2"/>
@@ -37,63 +36,91 @@
<help>
**Dataset formats**
The input is GFF_ format, and the output is tabular_.
The input is in GFF_ format, and the output is tabular_.
(`Dataset missing?`_)
.. _GFF: ./static/formatHelp.html#gff
.. _tabular: ./static/formatHelp.html#tab
.. _tabular: ./static/formatHelp.html#tab
.. _Dataset missing?: ./static/formatHelp.html
-----
**What it does**
PASS (Poisson Approximation for Statistical Significance) detects significant transcription factor binding sites from ChIP data in the genome. The method is probably the only peak calling method that accurately controls the false positive rate and FDR in ChIP data, which is important given the huge discrepancy in results obtained from different peak calling algorithms. At the same time, the method achieves a similar or better power than existing methods. Another unique feature of the method is that it allows varying thresholds to be used at different genomic locations for peak calling. For example, if a position lies in a open chromatin region, is deplete of nucleosome positioning, or a co-binding protein has been detected within the neighborhood, then the position is more likely to be bound by the target protein of interest, and hence a lower threshold will be used to call significant peaks. As a result, weak but real binding sites can be detected.
PASS (Poisson Approximation for Statistical Significance) detects
significant transcription factor binding sites in the genome from
ChIP data. This is probably the only peak-calling method that
accurately controls the false-positive rate and FDR in ChIP data,
which is important given the huge discrepancy in results obtained
from different peak-calling algorithms. At the same time, this
method achieves a similar or better power than previous methods.
<!-- we don't have wrapper support for the "prior" file yet
Another unique feature of this method is that it allows varying
thresholds to be used for peak calling at different genomic
locations. For example, if a position lies in an open chromatin
region, is depleted of nucleosome positioning, or a co-binding
protein has been detected within the neighborhood, then the position
is more likely to be bound by the target protein of interest, and
hence a lower threshold will be used to call significant peaks.
As a result, weak but real binding sites can be detected.
-->
-----
**Hints**
- ChIP-seq data:
- ChIP-Seq data:
If the data is from ChIP-seq, you need to convert each ChIP-seq value into z-scores before using this program. Also for ChIP-seq,
it is recommended that you group read counts within a neighborhood together, e.g., sum of reads within a 30bp window, for
every 30bp shifting window. As a result, the ChIP-seq data looks like a ChIP-chip data in format.
If the data is from ChIP-Seq, you need to convert the ChIP-Seq values
into z-scores before using this program. It is also recommended that
you group read counts within a neighborhood together, e.g. in tiled
windows of 30bp. In this way, the ChIP-Seq data will resemble
ChIP-chip data in format.
- Choosing window size options:
The window size is related with probe tiling density, if probes are tiled at every 100bp, then smallest window = 2 largest window = 6 is good, because DNA fragment size is around 300~500bp.
The window size is related to the probe tiling density. For example,
if the probes are tiled at every 100bp, then setting the smallest
window = 2 and largest window = 6 is appropriate, because the DNA
fragment size is around 300-500bp.
-----
**Example**
- input file ChIP-chip data file; Nimblegene GFF file format::
- input file::
chr7 Nimblegen ID 40307603 40307652 1.668944
chr7 Nimblegen ID 40307703 40307752 0.8041307
chr7 Nimblegen ID 40307808 40307865 -1.089931
chr7 Nimblegen ID 40307920 40307969 1.055044
chr7 Nimblegen ID 40308005 40308068 2.447853
chr7 Nimblegen ID 40308125 40308174 0.1638694
chr7 Nimblegen ID 40308223 40308275 -0.04796628
chr7 Nimblegen ID 40308318 40308367 0.9335709
chr7 Nimblegen ID 40308526 40308584 0.5143972
chr7 Nimblegen ID 40308611 40308660 -1.089931
chr7 Nimblegen ID 40307603 40307652 1.668944 . . .
chr7 Nimblegen ID 40307703 40307752 0.8041307 . . .
chr7 Nimblegen ID 40307808 40307865 -1.089931 . . .
chr7 Nimblegen ID 40307920 40307969 1.055044 . . .
chr7 Nimblegen ID 40308005 40308068 2.447853 . . .
chr7 Nimblegen ID 40308125 40308174 0.1638694 . . .
chr7 Nimblegen ID 40308223 40308275 -0.04796628 . . .
chr7 Nimblegen ID 40308318 40308367 0.9335709 . . .
chr7 Nimblegen ID 40308526 40308584 0.5143972 . . .
chr7 Nimblegen ID 40308611 40308660 -1.089931 . . .
etc.
Chromosome, start, end, and score are required fields.
In GFF, a value of dot '.' is used to mean "not applicable".
- output file::
#ID Chr Start End WinSz PeakValue # of FPs FDR
1 chr7 40310931 40311266 4 1.663446 0.208435 0.208435
ID Chr Start End WinSz PeakValue # of FPs FDR
1 chr7 40310931 40311266 4 1.663446 0.248817 0.248817
-----
**References**
Zhang Y (2008) Poisson Approximation for Significance in Genome-wide ChIP-chip Tiling Arrays. Bioinformatics, 24(24):2825-2831
Zhang Y. (2008)
Poisson approximation for significance in genome-wide ChIP-chip tiling arrays.
Bioinformatics. 24(24):2825-31. Epub 2008 Oct 25.
Chen KB, Zhang Y. (2010)
A varying threshold method for ChIP peak calling using multiple sources of information.
Submitted.
Chen KB and Zhang Y (2010) A Varying Threshold Method for ChIP Peak Calling Using Multiple Sources of Information. Submitted.
</help>
</tool>
+33 -24
View File
@@ -76,15 +76,17 @@
<help>
.. class:: warningmark
This currently works for builds hg18 or hg19.
This currently works only for builds hg18 or hg19.
-----
**Dataset formats**
The input and output datasets are tabular_.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -114,40 +116,47 @@ Website: http://sift.jcvi.org/
- input file::
chr3 81780820 + T/C
chr2 230341630 + G/A
chr2 43881517 + A/T
chr2 43857514 + T/C
chr6 88375602 + G/A
chr22 29307353 - T/A
chr10 115912482 - G/T
chr10 115900918 - C/T
chr16 69875502 + G/T
chr3 81780820 + T/C
chr2 230341630 + G/A
chr2 43881517 + A/T
chr2 43857514 + T/C
chr6 88375602 + G/A
chr22 29307353 - T/A
chr10 115912482 - G/T
chr10 115900918 - C/T
chr16 69875502 + G/T
etc.
- output file::
#Chrom Position Strand Allele Codons Transcript ID Protein ID Substitution Region dbSNP ID SNP Type Prediction Score Median Info Num seqs at position User Comment
chr3 81780820 + T/C AGA-gGA ENST00000264326 ENSP00000264326 R190G EXON CDS rs2229519:C Nonsynonymous DAMAGING 0.04 3.06 149
chr2 230341630 + G/T - ENST00000389045 ENSP00000373697 NA EXON CDS rs1803846:A Unknown Not scored NA NA NA
chr2 43881517 + A/T ATA-tTA ENST00000260605 ENSP00000260605 I230L EXON CDS rs11556157:T Nonsynonymous TOLERATED 0.47 3.19 7
chr2 43857514 + T/C TTT-TcT ENST00000260605 ENSP00000260605 F33S EXON CDS rs2288709:C Nonsynonymous TOLERATED 0.61 3.33 6
chr6 88375602 + G/A GTT-aTT ENST00000257789 ENSP00000257789 V217I EXON CDS rs2307389:A Nonsynonymous TOLERATED 0.75 3.17 13
chr22 29307353 + T/A ACC-tCC ENST00000335214 ENSP00000334612 T264S EXON CDS rs42942:A Nonsynonymous TOLERATED 0.4 3.14 23
chr10 115912482 + C/A CGA-CtA ENST00000369285 ENSP00000358291 R179L EXON CDS rs12782946:T Nonsynonymous TOLERATED 0.06 4.32 2
chr10 115900918 + G/A CAA-tAA ENST00000369287 ENSP00000358293 Q271* EXON CDS rs7095762:T Nonsynonymous N/A N/A N/A N/A
chr16 69875502 + G/T ACA-AaA ENST00000338099 ENSP00000337512 T608K EXON CDS rs3096381:T Nonsynonymous TOLERATED 0.12 3.41 3
#Chrom Position Strand Allele Codons Transcript ID Protein ID Substitution Region dbSNP ID SNP Type Prediction Score Median Info Num seqs at position User Comment
chr3 81780820 + T/C AGA-gGA ENST00000264326 ENSP00000264326 R190G EXON CDS rs2229519:C Nonsynonymous DAMAGING 0.04 3.06 149
chr2 230341630 + G/T - ENST00000389045 ENSP00000373697 NA EXON CDS rs1803846:A Unknown Not scored NA NA NA
chr2 43881517 + A/T ATA-tTA ENST00000260605 ENSP00000260605 I230L EXON CDS rs11556157:T Nonsynonymous TOLERATED 0.47 3.19 7
chr2 43857514 + T/C TTT-TcT ENST00000260605 ENSP00000260605 F33S EXON CDS rs2288709:C Nonsynonymous TOLERATED 0.61 3.33 6
chr6 88375602 + G/A GTT-aTT ENST00000257789 ENSP00000257789 V217I EXON CDS rs2307389:A Nonsynonymous TOLERATED 0.75 3.17 13
chr22 29307353 + T/A ACC-tCC ENST00000335214 ENSP00000334612 T264S EXON CDS rs42942:A Nonsynonymous TOLERATED 0.4 3.14 23
chr10 115912482 + C/A CGA-CtA ENST00000369285 ENSP00000358291 R179L EXON CDS rs12782946:T Nonsynonymous TOLERATED 0.06 4.32 2
chr10 115900918 + G/A CAA-tAA ENST00000369287 ENSP00000358293 Q271* EXON CDS rs7095762:T Nonsynonymous N/A N/A N/A N/A
chr16 69875502 + G/T ACA-AaA ENST00000338099 ENSP00000337512 T608K EXON CDS rs3096381:T Nonsynonymous TOLERATED 0.12 3.41 3
etc.
-----
**References**
Predicting Deleterious Amino Acid Substitutions, Genome Res. 2001 May; 11(5): 863.874.
Accounting for Human Polymorphisms Predicted to Affect Protein Function, Genome Res. 2002 December; 436-446
Ng PC, Henikoff S. (2001) Predicting deleterious amino acid substitutions.
Genome Res. 11(5):863-74.
SIFT: predicting amino acid changes that affect protein function, Nucleic Acids Research, 2003, Vol. 31, No. 13 3812-3814
Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm, Nature Protocols 4, - 1073 - 1081 (2009)
Ng PC, Henikoff S. (2002) Accounting for human polymorphisms predicted to affect protein function.
Genome Res. 12(3):436-46.
Ng PC, Henikoff S. (2003) SIFT: Predicting amino acid changes that affect protein function.
Nucleic Acids Res. 31(13):3812-4.
Kumar P, Henikoff S, Ng PC. (2009) Predicting the effects of coding non-synonymous variants
on protein function using the SIFT algorithm.
Nat Protoc. 4(7):1073-81. Epub 2009 Jun 25.
</help>
</tool>
+22 -21
View File
@@ -40,11 +40,12 @@
**Dataset formats**
The input is tabular_, with six columns of allele
counts. The output is also tabular, and includes all of the input data plus
the additional columns described below.
The input is tabular_, with six columns of allele counts. The output is also tabular,
and includes all of the input data plus the additional columns described below.
(`Dataset missing?`_)
.. _tabular: ./static/formatHelp.html#tab
.. _Dataset missing?: ./static/formatHelp.html
-----
@@ -70,15 +71,15 @@ group, the p-value, and the q-value.
- input file::
chr1 210 211 38 4 15 56 0 1 x
chr1 228 229 55 0 2 56 0 1 x
chr1 230 231 46 0 11 55 0 2 x
chr1 234 235 43 0 14 55 0 2 x
chr1 236 237 55 0 2 13 10 34 x
chr1 437 438 55 0 2 46 0 11 x
chr1 439 440 56 0 1 55 0 2 x
chr1 449 450 56 0 1 13 20 24 x
chr1 518 519 56 0 1 38 4 15 x
chr1 210 211 38 4 15 56 0 1 x
chr1 228 229 55 0 2 56 0 1 x
chr1 230 231 46 0 11 55 0 2 x
chr1 234 235 43 0 14 55 0 2 x
chr1 236 237 55 0 2 13 10 34 x
chr1 437 438 55 0 2 46 0 11 x
chr1 439 440 56 0 1 55 0 2 x
chr1 449 450 56 0 1 13 20 24 x
chr1 518 519 56 0 1 38 4 15 x
Here the group 1 genotype counts are in columns 4 - 6, while those
for group 2 are in columns 7 - 9.
@@ -89,15 +90,15 @@ to see where the new columns are appended in the output.
- output file::
chr1 210 211 38 4 15 56 0 1 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
chr1 228 229 55 0 2 56 0 1 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
chr1 230 231 46 0 11 55 0 2 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
chr1 234 235 43 0 14 55 0 2 x 49 0 8 49 0 8 0.00210854461554067 0.000739840215979182
chr1 236 237 55 0 2 13 10 34 x 34 5 18 34 5 18 6.14613878554783e-17 4.31307984950725e-17
chr1 437 438 55 0 2 46 0 11 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
chr1 439 440 56 0 1 55 0 2 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
chr1 449 450 56 0 1 13 20 24 x 34.5 10 12.5 34.5 10 12.5 2.25757007974134e-18 2.37638955762246e-18
chr1 518 519 56 0 1 38 4 15 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
chr1 210 211 38 4 15 56 0 1 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
chr1 228 229 55 0 2 56 0 1 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
chr1 230 231 46 0 11 55 0 2 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
chr1 234 235 43 0 14 55 0 2 x 49 0 8 49 0 8 0.00210854461554067 0.000739840215979182
chr1 236 237 55 0 2 13 10 34 x 34 5 18 34 5 18 6.14613878554783e-17 4.31307984950725e-17
chr1 437 438 55 0 2 46 0 11 x 50.5 0 6.5 50.5 0 6.5 0.0155644201009862 0.00409590002657532
chr1 439 440 56 0 1 55 0 2 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474
chr1 449 450 56 0 1 13 20 24 x 34.5 10 12.5 34.5 10 12.5 2.25757007974134e-18 2.37638955762246e-18
chr1 518 519 56 0 1 38 4 15 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06
</help>
</tool>