Dear all,

the next seminar will take place on Wednesday, please see the details below. We look forward to seeing you all!

Title: Masked superstrings as a unified framework for textual k-mer set representations
Speaker: Ondřej Sladký, Faculty of Mathematics and Physics, Charles University

Date and time: Wednesday 27/3/2023 - 17:20
Location: MFF UK, Malostranské nám. 25, lecture hall S4
https://bioinformatika.mff.cuni.cz/seminar/

Best wishes,
Petr Danecek

-------------------------------

Ondřej Sladký (Faculty of Mathematics and Physics, Charles University)
Masked superstrings as a unified framework for textual k-mer set representations

The popularity of k-mer-based methods has recently led to the development of compact k-mer-set representations, such as simplitigs/Spectrum-Preserving String Sets (SPSS), matchtigs, and eulertigs. These aim to represent k-mer sets via strings that contain individual k-mers as substrings more efficiently than the traditional unitigs. Here, we demonstrate that all such representations can be viewed as superstrings of input k-mers, and as such can be generalized into a unified framework that we call the masked superstring of k-mers. We study the complexity of masked superstring computation and prove NP-hardness for both k-mer superstrings and their masks. We then design local and global greedy heuristics for efficient computation of masked superstrings, implement them in a program called KmerCamel, and evaluate their performance using selected genomes and pan-genomes. Overall, masked superstrings unify the theory and practice of textual k-mer set representations and provide a useful framework for optimizing representations for specific bioinformatics applications.