diff --git a/tools/evolution/codingSnps.xml b/tools/evolution/codingSnps.xml index 4242bead233..6c4fd8db8a5 100644 --- a/tools/evolution/codingSnps.xml +++ b/tools/evolution/codingSnps.xml @@ -45,16 +45,21 @@ .. class:: infomark -The build must be defined for the input files and must be the same for both files. Use the pencil icon to add the build to the files if necessary. +The build must be defined for the input files and must be the same for both files. +Use the pencil icon to add the build to the files if necessary. ----- **Dataset formats** -The SNP dataset is in interval_ format with a column of SNPs as described below. +The SNP dataset is in interval_ format, with a column of SNPs as described below. The gene dataset is in BED_ format with 12 columns. The output dataset is also interval. +.. _interval: ./static/formatHelp.html#interval + +.. _BED: ./static/formatHelp.html#bed + ----- **What it does** @@ -76,10 +81,6 @@ The amino acids are listed with the reference amino acid first, then a colon, and then the amino acids for the alleles. If a SNP is not in a coding region or is synonymous then it is not included in the output file. -.. _BED: ./static/formatHelp.html#bed - -.. _interval: ./static/formatHelp.html#interval - ----- **Example** diff --git a/tools/human_genome_variation/beam.xml b/tools/human_genome_variation/beam.xml index 06c9113547f..073fce636e4 100644 --- a/tools/human_genome_variation/beam.xml +++ b/tools/human_genome_variation/beam.xml @@ -1,5 +1,5 @@ - detects both single and multi-locus SNP associations in case-control studies + significant single- and multi-locus SNP associations in case-control studies BEAM2_wrapper.sh map=${input.extra_files_path}/${input.metadata.base_name}.map ped=${input.extra_files_path}/${input.metadata.base_name}.ped $burnin $mcmc $pvalue significance=$significance posterior=$posterior @@ -44,7 +44,13 @@ .. class:: infomark -This tool can take a long time to run, depending on the number of SNPs, the sample size, and the number of MCMC steps specified. If you have hundreds of thousands of SNPs, it may take over a day. The main tasks that slow down this tool searching for interactions and dynamically partitioning the SNPs into blocks. Optimization is certainly possible, but hasn't been done yet. **If your only interest is to detect SNPs with primary effects (i.e., single-SNP associations), please use the GPASS tool instead.** +This tool can take a long time to run, depending on the number of SNPs, the +sample size, and the number of MCMC steps specified. If you have hundreds +of thousands of SNPs, it may take over a day. The main tasks that slow down +this tool are searching for interactions and dynamically partitioning the +SNPs into blocks. Optimization is certainly possible, but hasn't been done +yet. **If your only interest is to detect SNPs with primary effects (i.e., +single-SNP associations), please use the GPASS tool instead.** ----- @@ -113,7 +119,7 @@ This tool also partitions SNPs into blocks based on linkage disequilibrium (LD). 8: 1 21222571 1.000 | 353 9 138 | 352 8 140 etc. - The "id" field is an internally used index. + The "id" field is an internally used index. ----- diff --git a/tools/human_genome_variation/ctd.xml b/tools/human_genome_variation/ctd.xml index 21b9ab93d03..e5f745518c8 100644 --- a/tools/human_genome_variation/ctd.xml +++ b/tools/human_genome_variation/ctd.xml @@ -257,6 +257,8 @@ The input and output datasets are tabular_. +.. _tabular: ./static/formatHelp.html#tab + ----- **What it does** @@ -271,9 +273,7 @@ interaction type from the search-and-select box. The choices that start with '-' are a subset of a choice above them; you can chose either the general interaction type or a more specific one. -Home page: http://ctd.mdibl.org/ - -.. _tabular: ./static/formatHelp.html#tab +Website: http://ctd.mdibl.org/ ----- diff --git a/tools/human_genome_variation/funDo.xml b/tools/human_genome_variation/funDo.xml index 62fb87a3231..98c6d15b5f5 100644 --- a/tools/human_genome_variation/funDo.xml +++ b/tools/human_genome_variation/funDo.xml @@ -30,7 +30,7 @@ -**Data formats** +**Dataset formats** There is no input dataset. The output is in interval_ format. @@ -53,7 +53,7 @@ not commas. As a special case, entering the word "disease" returns all genes associated with any disease, even if that word does not actually appear in the term field. -home page: http://django.nubic.northwestern.edu/fundo/ +Website: http://django.nubic.northwestern.edu/fundo/ ----- diff --git a/tools/human_genome_variation/gpass.xml b/tools/human_genome_variation/gpass.xml index 1aab1965c82..62b7d8ff51d 100644 --- a/tools/human_genome_variation/gpass.xml +++ b/tools/human_genome_variation/gpass.xml @@ -1,5 +1,5 @@ - detects significant single-SNP associations in case-control studies + significant single-SNP associations in case-control studies gpass.pl ${input1.extra_files_path}/${input1.metadata.base_name}.map ${input1.extra_files_path}/${input1.metadata.base_name}.ped $output $fdr @@ -34,7 +34,7 @@ --> -**Data formats** +**Dataset formats** The input dataset must be lped_, and the output is tabular_. diff --git a/tools/human_genome_variation/hilbertvis.xml b/tools/human_genome_variation/hilbertvis.xml index 58c6cc79179..221d450fa21 100644 --- a/tools/human_genome_variation/hilbertvis.xml +++ b/tools/human_genome_variation/hilbertvis.xml @@ -60,24 +60,36 @@ **Dataset formats** -The input format is interval_, the output is an image in PDF format. +The input format is interval_, and the output is an image in PDF format. .. _interval: ./static/formatHelp.html#interval +----- + **What it does** -HilbertVis uses the Hilbert space-filling curve to visualize the structure of position-dependent data. It maps the traditional one-dimensional line visualization onto a two-dimensional square. For example, here is a diagram showing -the path of a level-2 Hilbert curve. +HilbertVis uses the Hilbert space-filling curve to visualize the structure of +position-dependent data. It maps the traditional one-dimensional line +visualization onto a two-dimensional square. For example, here is a diagram +showing the path of a level-2 Hilbert curve. .. image:: ../static/images/hilbertvisDiagram.png -The shade of each pixel represents the value for the corresponding bin of consecutive genomic positions, calculated according to the specified summarization mode. The pixels are arranged so that bins that are close to each other on the data vector are represented by pixels that are close to each other in the plot. In particular, adjacent bins are mapped to adjacent pixels. Hence, dark spots in a figure represent a peak; the area of the spot in the two-dimensional plot is proportional to the width of the peak in the one-dimensional data, and the darkness of the spot corresponds to the height of the peak. +The shade of each pixel represents the value for the corresponding bin of +consecutive genomic positions, calculated according to the specified +summarization mode. The pixels are arranged so that bins that are close +to each other on the data vector are represented by pixels that are close +to each other in the plot. In particular, adjacent bins are mapped to +adjacent pixels. Hence, dark spots in a figure represent a peak; the area +of the spot in the two-dimensional plot is proportional to the width of the +peak in the one-dimensional data, and the darkness of the spot corresponds to +the height of the peak. -The input file is interval and typically contains a column with scores or other -numbers. Examples of scores could be the coverage of aligned reads from -conservation scores, SNP density, ChIP-Seq data, etc. +The input file is in interval format, and typically contains a column with +scores or other numbers, such as conservation scores, SNP density, the +coverage of aligned reads from ChIP-Seq data, etc. -HilbertVis homepage: http://www.ebi.ac.uk/huber-srv/hilbert/ +Website: http://www.ebi.ac.uk/huber-srv/hilbert/ ----- diff --git a/tools/human_genome_variation/ldtools.xml b/tools/human_genome_variation/ldtools.xml index 145872d87d3..14f425cf0a0 100644 --- a/tools/human_genome_variation/ldtools.xml +++ b/tools/human_genome_variation/ldtools.xml @@ -1,12 +1,12 @@ - linkage disequilibrium + linkage disequilibrium and tag SNPs ldtools_wrapper.sh rsquare=$rsquare freq=$freq input=$input output=$output - + @@ -31,7 +31,7 @@ **Dataset formats** -The Genotype dataset and output are tabular_. +The input and output datasets are tabular_. .. _tabular: ./static/formatHelp.html#tab @@ -40,20 +40,27 @@ The Genotype dataset and output are tabular_. **What it does** This tool can be used to analyze the patterns of linkage disequilibrium -(LD) between polymorphic sites in a locus. SNPs are binned based on the -threshold level of LD as measured by r\ :sup:`2`. +(LD) between polymorphic sites in a locus. SNPs are grouped based on the +threshold level of LD as measured by r\ :sup:`2` (regardless of genomic +position), and a representative "tag SNP" is reported for each group. +Note that the groups are generated by transitive closure: each SNP in the +group is within r\ :sup:`2` of *some* other SNP in the group, but not +necessarily all of them (and not necessarily the tag SNP). -The algorithm employed by this tool is the same as used in ldSelect (Carlson -et al., 2004). However it is designed to be much faster and more efficient -than ldSelect. +The underlying algorithm is the same as the one used in ldSelect (Carlson +et al. 2004). However, this tool is implemented to be much faster and more +efficient than ldSelect. + +The input is a tabular file with genotype information for each individual +at each SNP site, in exactly four columns: site ID, sample ID, and the +two allele nucleotides. ----- **Example** -- The input genotype information should be a tab-delimited text file with the SNPs genotype information in the following columns:: +- input file:: - Site Sample Allele1 Allele2 rs2334386 NA20364 G T rs2334386 NA20363 G G rs2334386 NA20360 G G diff --git a/tools/human_genome_variation/pass.xml b/tools/human_genome_variation/pass.xml index ba66823dad4..86f3ed0ff8f 100644 --- a/tools/human_genome_variation/pass.xml +++ b/tools/human_genome_variation/pass.xml @@ -1,5 +1,5 @@ - detects significant transcription factor binding sites from ChIP data in the genome + significant transcription factor binding sites from ChIP data pass_wrapper.sh "$input" "$min_window" "$max_window" "$false_num" "$output" diff --git a/tools/human_genome_variation/snpFreq.xml b/tools/human_genome_variation/snpFreq.xml index c27105461b3..3aaa6ef0149 100644 --- a/tools/human_genome_variation/snpFreq.xml +++ b/tools/human_genome_variation/snpFreq.xml @@ -1,5 +1,5 @@ - identify significant SNPs in case-control data + significant SNPs in case-control data snpFreq2.pl $input $group1_1 $group1_2 $group1_3 $group2_1 $group2_2 $group2_3 0.05 $output @@ -7,12 +7,12 @@ - - - - - - + + + + + + @@ -37,29 +37,32 @@ + **Dataset formats** -The input is tabular_, with 6 columns with allele -counts. The output format is also tabular, with the input data plus the added -columns described below. +The input is tabular_, with six columns of allele +counts. The output is also tabular, and includes all of the input data plus +the additional columns described below. + +.. _tabular: ./static/formatHelp.html#tab ----- **What it does** -This identifies significant SNPs in case-control data, using R and the -Fisher's Exact test to see if there is a -significant difference in the allele counts. R's qvalue -package is used to correct for multiple testing. -The input file must be tabular_ and include counts -for each position and each allele (AA aa Aa) for the 2 groups. -The order of the alleles is not important as long as it is the same for both -groups. All other columns are kept for the output only. -The output adds 8 columns -to the input file, namely the minimum expected counts of the 3 alleles -for each group, the p-value, and the q-value. +This tool performs a basic analysis of bi-allelic SNPs in case-control +data, using the R statistical environment and Fisher's exact test to +identify SNPs with a significant difference in the allele frequencies +between the two groups. R's "qvalue" package is used to correct for +multiple testing. -.. _tabular: ./static/formatHelp.html#tab +The input file includes counts for each allele combination (AA aa Aa) +for each group at each SNP position. The assignment of codes (1 2 3) +to these genotypes is arbitrary, as long as it is consistent for both +groups. Any other input columns are ignored in the computation, but +are copied to the output. The output appends eight additional columns, +namely the minimum expected counts of the three genotypes for each +group, the p-value, and the q-value. ----- @@ -77,11 +80,12 @@ for each group, the p-value, and the q-value. chr1 449 450 56 0 1 13 20 24 x chr1 518 519 56 0 1 38 4 15 x -Group 1 allele counts are columns 4 - 6, and group 2 are 7 - 9. +Here the group 1 genotype counts are in columns 4 - 6, while those +for group 2 are in columns 7 - 9. -Note the x column has no meaning and was added to this example to show other -columns can be included and to make it easier to see where the new columns -are appended in the output. +Note that the "x" column has no meaning. It was added to this example +to show that extra columns can be included, and to make it easier +to see where the new columns are appended in the output. - output file:: @@ -94,5 +98,6 @@ are appended in the output. chr1 439 440 56 0 1 55 0 2 x 55.5 0 1.5 55.5 0 1.5 1 0.210526315789474 chr1 449 450 56 0 1 13 20 24 x 34.5 10 12.5 34.5 10 12.5 2.25757007974134e-18 2.37638955762246e-18 chr1 518 519 56 0 1 38 4 15 x 47 2 8 47 2 8 1.50219088598917e-05 6.32501425679652e-06 +