An international team, including scientists at the Institute of Science Tokyo, has built a protein language model that combines two sources of information about proteins: their amino acid sequences and their three-dimensional structures, giving researchers a new way to chart relationships across the protein universe and study how proteins evolved over billions of years. The work was published in Proceedings of the National Academy of Sciences.
Thousands of protein families carry out nearly every function inside living cells, and a long-standing question in evolutionary biochemistry is how these proteins relate to each other and where they originated. Scientists have traditionally sorted proteins into hierarchical groups based on relatedness, much like genus and species classifications for organisms, but artificial intelligence is now opening additional ways to examine these connections.
Protein language models convert a protein into a numerical “embedding,” a kind of address where proteins with similar properties end up near each other, producing a protein world map. The complication is that sequence and structure don’t always align: unrelated sequences can fold into similar shapes, while similar sequences can produce very different structures. Most existing models handle the two separately, and even those using both do not necessarily place a protein’s sequence and structure at the same point on the map.
Search Antibodies Search Now Use our Antibody Search Tool to find the right antibody for your research. Filter
by Type, Application, Reactivity, Host, Clonality, Conjugate/Tag, and Isotype.
The team’s new model, Contrastive Learning Sequence-Structure (CLSS), generates closely matching embeddings for a protein’s sequence and its structure. Using contrastive learning, CLSS trains on paired sequences and structures, pulling matched pairs together and pushing unrelated ones apart, so a protein lands in roughly the same map location regardless of which form of data produced it. Compared with other leading protein language models, CLSS produced a more unified map, and its output closely tracked relationships already documented in the expert-curated ECOD and CATH classification systems, despite never being trained on those classifications.
“This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds,” said co-author Liam Longo. “What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognize using conventional approaches.”
Unlike most protein language models, which need a full sequence or structure to produce useful output, CLSS could often place short sequence fragments meaningfully alongside complete sequences and structures—useful for evolutionary study, since small protein pieces have been reused and rearranged repeatedly over evolutionary history, and similar fragments in otherwise unrelated proteins can hint at ancient evolutionary ties.
The maps also revealed broader patterns: proteins associated with organic cofactors clustered in specific regions, while metal-binding proteins spread more broadly across the space. The researchers see these unified representations opening possibilities for database searches, protein engineering, and reconstructing evolutionary trajectories—another way to trace how today’s protein diversity emerged over nearly four billion years of evolution.