article
Clustering of genomic sequences is critical for elucidating biological relationships, yet unsupervised methods often compromise between accuracy, generalizability, and biological interpretability. We propose a Siamese Bidirectional Long ShortTerm Memory (BiLSTM) network with an attention mechanism, trained via contrastive loss to learn biologically meaningful representations of DNA sequences, entirely without supervision, augmentation, or hand-crafted <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$k$</tex>-mer features. We evaluate our model on a challenging benchmark of Betacoronavirus genomes. The resulting embeddings capture both sequence-level similarity and evolutionary patterns. When paired with standard clustering algorithms, these embeddings yield highquality partitions: silhouette score of 0.814, Adjusted Rand Index (ARI) of 0.713, and purity of 0.864. Our model also reduces intra-cluster embedding variance by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{2} \boldsymbol{-} \mathbf{4} \boldsymbol{\times}$</tex> compared to baselines, aligning closely with known taxonomic subgenera. Overall, this framework offers a scalable, interpretable, and unsupervised solution for genomic sequence clustering, with strong potential in comparative genomics and evolutionary biology.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/wincom65874.2025.11313377
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.