Commit Graph
31 Commits
Author SHA1 Message Date
John Chilton 51e269cf9b FASTQ Opt: Eliminate a few extra calls to is_ascii_encoded.
Kept 135848 of 250000 reads (54.34%).
         117127015 function calls (116877015 primitive calls) in 38.632 seconds

Down from 39.416 seconds on previous changeset. Main Difference

(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))

   385848   11.289    0.000   15.499    0.000 fastq.py:91(get_decimal_quality_scores)
  1271696    0.716    0.000    0.716    0.000 fastq.py:71(is_ascii_encoded)

   250000    0.185    0.000   10.434    0.000 fastq.py:107(get_decimal_quality_scores)
   385848   11.380    0.000   15.368    0.000 fastq.py:109(__get_decimal_quality_scores)
  1135848    0.649    0.000    0.649    0.000 fastq.py:71(is_ascii_encoded)
2013-11-12 22:02:58 -06:00
John Chilton 1151ed748d FASTQ Opt: Do not generate arrays just to check length.
... just compute what length would be.

Kept 135848 of 250000 reads (54.34%).
         117012863 function calls (116762863 primitive calls) in 38.412 seconds

Down from 39.416 seconds on previous changeset. Main Difference:

(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))

   500000    0.518    0.000    1.889    0.000 fastq.py:171(insufficient_quality_length)
   250000    0.322    0.000    1.103    0.000 fastq.py:173(assert_sequence_quality_lengths)

 - became -

   500000    0.306    0.000    0.932    0.000 fastq.py:187(insufficient_quality_length)
   250000    0.164    0.000    0.478    0.000 fastq.py:189(assert_sequence_quality_lengths)
2013-11-12 22:02:58 -06:00
John Chilton bc8939eec0 FASTQ Opt: Precompute difference ascii -> decimal difference.
Kept 135848 of 250000 reads (54.34%).
         117012863 function calls (116762863 primitive calls) in 39.416 seconds

Down from 46.897 seconds on previous changeset. Main Difference WRT:

(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))

   385848   18.215    0.000   22.866    0.000 fastq.py:90(get_decimal_quality_scores)

 -became-

   385848   11.415    0.000   15.546    0.000 fastq.py:91(get_decimal_quality_scores)

Looks like a math optimization but actual optimization is coming from reading local variable instead of dereferencing object variable twice per base per.
2013-11-12 22:02:58 -06:00
John Chilton 078d860877 FASTQ Opt: No need to map(str) over chr and join, can just join chr's.
Prevent an extra array creation and chr-> str map per base. Shaves 10% off remaining run time of filter on SSD.

Kept 135848 of 250000 reads (54.34%).
         117012863 function calls (116762863 primitive calls) in 46.897 seconds

Down from 48.272 seconds on previous changset. Main Difference:

(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))

   135848    0.502    0.000   22.979    0.000 fastq.py:99(convert_read_to_format)

 -became-

   135848    0.448    0.000   21.159    0.000 fastq.py:99(convert_read_to_format)

Not huge, but a consistent improvement.
2013-11-12 22:02:58 -06:00
John Chilton 99e59c3701 FASTQ Opt: Utilize optimized in place alternatives to restrict_scores_to_valid_range.
Separate out ascii vs. decimal encoding branches of convert_read_to_format and use these new transform_ alternatives operate "in place" (don't produce new lists/allocate memory). The ascii version transforms to ascii in place instead of requiring another call and creating another array.

Runtime Result:

Kept 135848 of 250000 reads (54.34%).
         117148711 function calls (116898711 primitive calls) in 48.272 seconds

Down from 85.375 seconds on previous changeset. Main Difference:

(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))

   135848    1.185    0.000   59.856    0.000 fastq.py:71(convert_read_to_format)

 -became-

   135848    0.502    0.000   22.979    0.000 fastq.py:99(convert_read_to_format)

About a third the time is spent convert reads to the correct format. This is a varaint of the core optimization I made when optimizing the FASTQ groomer for MSI - it has likewise a substantial impact on the performance of that tool.
2013-11-12 22:02:58 -06:00
John Chilton de38b2beaf FASTQ Opt: In convert_read_to_format, adjust when new_encoding logic is calculated.
Doesn't change the behavior or optimize anything, this is just done to simplify subsequent commits.

Baseline:

I took the first 100 megabytes of a 26 gigabyte example of a FASTQ file filtered with the FASTQ filter tool that I found in an "Important Galaxy User"'s history on main. I ran with the same command-line on my dev box with the start of this file and using Python's -m to profile function times and total amount of time to serve as a baseline as I optimized the FASTQ filter code. Here is the start of the output:

Kept 135848 of 250000 reads (54.34%).
         200015991 function calls (199765991 primitive calls) in 136.934 seconds

Extrapolating this out, that 26 gigabyte file would take roughly 10 hours to process on my laptop - this is slightly longer than what it took on main - indicating to me this is likely not disk bound since my SSD would probably outperform main?
2013-11-12 22:02:58 -06:00
Dannon Baker 0f120bd92c Strip trailing whitespace (and windows line endings) from all python files in lib 2013-08-29 23:39:52 -04:00
Florent Angly fcc56159e1 Avoid trailing whitespace 2011-11-30 11:38:52 +10:00
Florent Angly b382d19352 Paired-end code that properly ignores description part of FASTQ headers 2011-10-05 18:04:58 +10:00
Daniel Blankenberg 98c4351f28 Add basic VCF4.1 support. 2011-08-30 16:56:52 -04:00
Daniel Blankenberg 90e695c9aa Allow FASTQ Groomer tool to work on Color Space files that contain a fake/dummy quality score for the adapter base (e.g. files obtained from the SRA). The Groomer will remove the dummy/fake quality score from the read. 2011-08-22 16:07:23 -04:00
Florent Angly cdde306069 FASTQ interlacer and de-interlacer tools fully integrated in Galaxy and functional 2010-12-17 15:04:17 +10:00
Florent Angly 541554750e Little bug fix and more informative error message 2010-12-16 17:12:01 +10:00
Florent Angly 930e8df00d Added 2 Python scripts to deal with FASTQ mate pairs:
- the interlacer puts mate pairs present in 2 files into a single file
- the deinterlacer puts mate pairs present in a single file into 2 files
2010-12-16 18:26:59 +10:00
Kanwei Li 60ddbdcf78 VCF can take floats 2011-06-09 17:49:20 -04:00
Daniel Blankenberg 05242912c7 Minor reformatting for 5609:f530dbdde1f5 2011-06-01 11:40:38 -04:00
Daniel Blankenberg 0a47db9aa4 Allow VCF parser to accept unknown values for QUAL field. 2011-06-01 11:29:07 -04:00
Daniel Blankenberg 8bdda93ffc Add more verbose error reporting to FASTQ Groomer tool. 2011-03-24 11:45:50 -04:00
Daniel Blankenberg bde479e1f1 Do not print summary information in FASTQ groomer when grooming an empty file. 2011-02-18 12:07:27 -05:00
Jeremy Goecks 9c7065d898 Make VCF (variant call format) a Galaxy datatype and enable very basic VCF support in trackster. VCF datatype is sniffable and can be converted to summary tree and interval index. In trackster, VCF files are represented as single-base pair feature tracks. 2010-10-06 16:23:55 -04:00
Daniel Blankenberg 8b4aea49cb Bug fix for signature of lib.galaxy_utils.sequence.transform.?NA_reverse_complement method. 2010-10-01 10:45:17 -04:00
Kanwei Li a1f9e6a572 Support for VCFv4.0 and misc VCF fixes [Brad Chapman]
- Support for VCFv4.0, which should be identical to 3.3 support
- Correctly handle chromosome references when they start with 'chr' (instead of just numbers)
- Handle extra empty tabs on the header line which are present in GATK produced VCF and confuse the determination of how many sample states should be parsed.
2010-08-03 10:43:32 -04:00
Daniel Blankenberg ff18016e41 Add a VCF to MAF Custom Track converter tool. This tool converts a Variant Call Format (VCF) file into a Multiple Alignment Format (MAF) custom track file suitable for display at genome browsers.
This file should be used for display purposes only (e.g as a UCSC Custom Track). Performing an analysis using the output created by this tool as input is not recommended; the source VCF file should be used when performing an analysis.

Unknown nucleotides are represented as '*' as required to allow the display to draw properly; these include e.g. reference bases which appear before a deletion and are not available without querying the original reference sequence.
2010-06-14 15:07:46 -04:00
Daniel Blankenberg c4d5c8e0df Allow FASTQ Groomer/parser to work on tab-delimited decimal scores. 2010-05-24 15:33:55 -04:00
Daniel Blankenberg a333a57f9f Allow FASTQ parser to handle extra space padded decimal scores, e.g. '0 ' 2010-04-12 10:30:25 -04:00
Daniel Blankenberg 4ccb19f94b Add ability for reverse complement in e.g. FASTQ Manipulation tool to handle ambiguity codes. 2010-03-29 14:54:23 -04:00
Daniel Blankenberg 4ff8d400f3 Make fastqAggregator a new style class. 2010-03-05 10:13:36 -05:00
Daniel Blankenberg a19ae79b85 Change color space FASTA file type from fastqsolid to fastqcssanger.
Cripple accepted tool input formats for many of the FASTQ tools to only allow only fastqsanger and fastqcssanger to be used.
2010-03-02 10:47:23 -05:00
Daniel Blankenberg 77c4785eb7 Update Combine FASTA and QUAL tool to allow the quality score file to be optional.
When not provided, the output will be fastqsanger or fastqsolid (when a csfasta is provided) with each quality score being the maximal allowed value (93).
2010-02-24 16:50:06 -05:00
Daniel Blankenberg 1dcf92edde Move FASTA classes in galaxy_utils from fastq.py to fasta.py. 2010-02-24 11:57:35 -05:00
Daniel Blankenberg 8082c6f36f Add a new FASTQ tool suite. Four FASTQ variants are supported: sanger, illumina, solexa and solid.
Tools include:
	FASTQ Groomer convert between various FASTQ quality formats
	Combine FASTA and QUAL into FASTQ
	FASTQ joiner on paired end reads
	FASTQ splitter on joined paired end reads
	FASTQ to FASTA converter
	FASTQ Summary Statistics by column
	Filter FASTQ reads by quality score and length
	FASTQ Trimmer by column
	Manipulate FASTQ reads on various attributes
	Boxplot of quality statistics (Generic, with outliers)
2010-02-23 16:48:07 -05:00