Masked Minimizers: Unifying sequence sketching methods

Abstract

Minimizers and syncmers are sequence sketching methods that extract representative substrings from a long sequence. We show that both these sampling rules are different instantiations of a new unifying concept we call masked minimizers, which applies a sub-sampling binary mask on a minimizer sketch. This unification leads to the first formal procedure to meaningfully compare minimizers, syncmers and other comparable masked minimizers. We further demonstrate that existing sequence sketching metrics, such as density (which measures the sketch sparseness) and conservation (which measures the likelihood of the sketch being preserved under random mutations), should not be independently measured when evaluating masked minimizers. We propose a new metric that reflects the trade-off between these quantities called the generalized sketch score, or GSS. Finally, we introduce a sequence-specific and gradient-based learning objective that efficiently optimizes masked minimizer schemes with respect to the proposed GSS metric. We show that our method finds sketches with better overall density and conservation compared to existing expected and sequence-specific approaches, enabling more efficient and robust genomic analyses in the many settings where minimizers and syncmers are used.

Publication
bioRxiv