preprint · bioRxiv (Cold Spring Harbor Laboratory)
Abstract Below 30% pairwise sequence identity, alignment-based methods struggle to reliably distinguish true homologs from chance, and enzyme function prediction degrades accordingly: on proteins in this regime, even advanced methods (CLEAN) achieves only 55.1% accuracy at full EC specificity on the CARE benchmark. To this end Protein Language Models (PLMs) have gained favor as alternatives. However, PLMs often encode only a subset of the biology, whereas the understanding of enzyme function requires among other things a combination of sequence, structure and functional-context simultaneously. In this work, we describe FuncSeek, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities). Using SwissProt data, these 2816-D embeddings are labelled with Enzyme Commission numbers (EC) and are trained through a supervised contrastive head into a 256-D space. FuncSeek attains 64.6% nearest-neighbour EC4 accuracy on the CARE out-of-distribution benchmark set ( ood30 ; proteins below 30% identity to training set), outperforming CLEAN (55.1%) and Diamond BLASTp (51.4%), and obtains 93.7% nearest-neighbour EC4 accuracy on the promiscuous, multi-functional enzymes benchmark (CLEAN, 69.4%). We also show that the learned representations transfer without retraining to the TrEMBL database, achieving 97.3% nearest-neighbour EC4 accuracy on an 8,031 BRENDA-validated enzyme set, never seen during training. Because only projected embeddings are stored in the target index, and function is inferred from an annotated reference set, we propose this paradigm for rapidly searching extremely large metagenomic databases, bypassing costly sequence alignment and annotation pipelines. Author Summary Enzymes are the molecular machines that carry out the chemistry of life, and knowing which reaction an enzyme performs is essential for medicine, biotechnology, and understanding how organisms work. Yet new protein sequences are being discovered far faster than we can study them in the laboratory, and the most common computational shortcut, comparing a new sequence to well-studied ones, breaks down when the sequences are only distantly related. We asked whether recent artificial intelligence models that learn the “language” of proteins could close this gap. Rather than rely on a single model, we combined three that each capture a different aspect of protein biology: evolutionary patterns, three-dimensional shape, and functional context. We then trained a system, which we call FuncSeek, to arrange enzymes so that those performing the same reaction sit close together. FuncSeek predicted enzyme function more accurately than the leading existing tools, especially for distantly related and multi-functional enzymes, and this accuracy carried over to proteins it had never seen. Because it represents each protein as a compact numerical fingerprint, it can search enormous, unexplored collections of sequences quickly.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.64898/2026.08.07.738656
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.