... just compute what length would be.
Kept 135848 of 250000 reads (54.34%).
117012863 function calls (116762863 primitive calls) in 38.412 seconds
Down from 39.416 seconds on previous changeset. Main Difference:
(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))
500000 0.518 0.000 1.889 0.000 fastq.py:171(insufficient_quality_length)
250000 0.322 0.000 1.103 0.000 fastq.py:173(assert_sequence_quality_lengths)
- became -
500000 0.306 0.000 0.932 0.000 fastq.py:187(insufficient_quality_length)
250000 0.164 0.000 0.478 0.000 fastq.py:189(assert_sequence_quality_lengths)
Kept 135848 of 250000 reads (54.34%).
117012863 function calls (116762863 primitive calls) in 39.416 seconds
Down from 46.897 seconds on previous changeset. Main Difference WRT:
(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))
385848 18.215 0.000 22.866 0.000 fastq.py:90(get_decimal_quality_scores)
-became-
385848 11.415 0.000 15.546 0.000 fastq.py:91(get_decimal_quality_scores)
Looks like a math optimization but actual optimization is coming from reading local variable instead of dereferencing object variable twice per base per.
Prevent an extra array creation and chr-> str map per base. Shaves 10% off remaining run time of filter on SSD.
Kept 135848 of 250000 reads (54.34%).
117012863 function calls (116762863 primitive calls) in 46.897 seconds
Down from 48.272 seconds on previous changset. Main Difference:
(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))
135848 0.502 0.000 22.979 0.000 fastq.py:99(convert_read_to_format)
-became-
135848 0.448 0.000 21.159 0.000 fastq.py:99(convert_read_to_format)
Not huge, but a consistent improvement.
Separate out ascii vs. decimal encoding branches of convert_read_to_format and use these new transform_ alternatives operate "in place" (don't produce new lists/allocate memory). The ascii version transforms to ascii in place instead of requiring another call and creating another array.
Runtime Result:
Kept 135848 of 250000 reads (54.34%).
117148711 function calls (116898711 primitive calls) in 48.272 seconds
Down from 85.375 seconds on previous changeset. Main Difference:
(ncalls|tottime|percall|cumtime|percall|filename:lineno(function))
135848 1.185 0.000 59.856 0.000 fastq.py:71(convert_read_to_format)
-became-
135848 0.502 0.000 22.979 0.000 fastq.py:99(convert_read_to_format)
About a third the time is spent convert reads to the correct format. This is a varaint of the core optimization I made when optimizing the FASTQ groomer for MSI - it has likewise a substantial impact on the performance of that tool.
Doesn't change the behavior or optimize anything, this is just done to simplify subsequent commits.
Baseline:
I took the first 100 megabytes of a 26 gigabyte example of a FASTQ file filtered with the FASTQ filter tool that I found in an "Important Galaxy User"'s history on main. I ran with the same command-line on my dev box with the start of this file and using Python's -m to profile function times and total amount of time to serve as a baseline as I optimized the FASTQ filter code. Here is the start of the output:
Kept 135848 of 250000 reads (54.34%).
200015991 function calls (199765991 primitive calls) in 136.934 seconds
Extrapolating this out, that 26 gigabyte file would take roughly 10 hours to process on my laptop - this is slightly longer than what it took on main - indicating to me this is likely not disk bound since my SSD would probably outperform main?
- Support for VCFv4.0, which should be identical to 3.3 support
- Correctly handle chromosome references when they start with 'chr' (instead of just numbers)
- Handle extra empty tabs on the header line which are present in GATK produced VCF and confuse the determination of how many sample states should be parsed.
This file should be used for display purposes only (e.g as a UCSC Custom Track). Performing an analysis using the output created by this tool as input is not recommended; the source VCF file should be used when performing an analysis.
Unknown nucleotides are represented as '*' as required to allow the display to draw properly; these include e.g. reference bases which appear before a deletion and are not available without querying the original reference sequence.
When not provided, the output will be fastqsanger or fastqsolid (when a csfasta is provided) with each quality score being the maximal allowed value (93).
Tools include:
FASTQ Groomer convert between various FASTQ quality formats
Combine FASTA and QUAL into FASTQ
FASTQ joiner on paired end reads
FASTQ splitter on joined paired end reads
FASTQ to FASTA converter
FASTQ Summary Statistics by column
Filter FASTQ reads by quality score and length
FASTQ Trimmer by column
Manipulate FASTQ reads on various attributes
Boxplot of quality statistics (Generic, with outliers)