Researchers from Florida Atlantic University have developed Deep Novel Mutation Search (DNMS), an AI-driven method to predict future SARS-CoV-2 spike protein mutations. Published in Communications Biology, this approach uses a protein language model fine-tuned on viral sequences to identify mutations most likely to emerge, offering a tool for proactive public health monitoring. 

DNMS analyzes mutations based on their adherence to structural "rules" learned by the ProtBERT language model, referred to as grammaticality. It also evaluates how similar the mutated sequence is to the original protein, measured by semantic change, and assesses shifts in amino-acid interaction patterns through attention change. Unlike traditional methods that compare mutations to a static reference sequence, DNMS employs a parent-child framework. It simulates single-point mutations on existing viral strains (parent sequences) from phylogenetic trees, then ranks candidates based on evolutionary plausibility.

"Our model ranks all possible mutations to find the ones that are most likely to occur in the future," said senior author Xingquan "Hill" Zhu. "Our study shows that mutations following the protein's grammars, with minimal changes compared to the original sequence and low attention differences, are considered the most likely future mutations."

Search Antibodies
Search Now Use our Antibody Search Tool to find the right antibody for your research. Filter
by Type, Application, Reactivity, Host, Clonality, Conjugate/Tag, and Isotype.

By simulating all possible single-point mutations for a given spike protein sequence, DNMS combines these metrics into a unified ranking system. Statistical tests showed it outperforms previous methods in predicting novel mutations, particularly those emerging from real-world evolutionary pathways.

The tool's predictive power stems from its focus on incremental adaptations. The study found that mutations with high grammaticality, small semantic change, and low attention change were associated with higher viral fitness. This suggests that mutations which fit well within the biological "rules" of the protein and cause minimal disruption to the protein's structure or function are more likely to be beneficial for the virus.

"We believe that using sequence data alone can help make these predictions, as proteins follow certain biological rules," said Zhu.

While focused on SARS-CoV-2, the method's framework could extend to other rapidly evolving pathogens, offering a faster, cost-effective alternative to lab-based mutation analysis.