Constarium
← Search

Data · dataset · 2013

Turtle Software.

Listed in DataCite

We present a novel method that balances time, space and accuracy requirements to efficiently extract frequent k-mers even for high coverage libraries and large genomes such as human.

Description

Our method is designed to minimize cache-misses in a cache-efficient manner by using a Pattern-blocked Bloom filter to remove infrequent $k$-mers from consideration in combination with a novel sort-and-compact scheme, instead of a Hash, for the actual counting.

While this increases theoretical complexity, the savings in cache misses reduce the empirical running times. A variant can resort to a counting Bloom filter for even larger savings in memory at the expense of false negatives in addition to the false positives common to all Bloom filter based approaches. A comparison to the state-of-the-art shows reduced memory requirements and running times.

Links

Topics

Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
DataCite10.6084/m9.figshare.79157912 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
byte_sizesource · DataCiteconnector:datacite@1.0.0
concepts[field].fos:computer-and-information-sciencessource · DataCiteconnector:datacite@1.0.0
created_datesource · DataCiteconnector:datacite@1.0.0
descriptionsource · DataCiteconnector:datacite@1.0.0/data/attributes/descriptions
licensesource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
publication_datesource · DataCiteconnector:datacite@1.0.0/data/attributes/dates
titlesource · DataCiteconnector:datacite@1.0.0/data/attributes/titles/0/title
updated_datesource · DataCiteconnector:datacite@1.0.0