New York University researchers have trained an AI model to learn chemical patterns associated with stability in drug-like molecules and accurately predict where their hydrogen atoms should be positioned.
Their research, published in the journal Chemical Science , addresses a longstanding challenge in molecular design and drug discovery: how to rapidly and reliably determine the stable form of molecules that share a molecular formula but readily convert into different forms.
Many drug-like molecules can exist in two or more closely related forms called tautomers. A hydrogen atom moves from one site to another, accompanied by a change in the bonding pattern.
"Although this may seem like a small change, different tautomers of the same molecule can alter how a molecule interacts with a protein target," explained Yingkai Zhang, professor of chemistry at NYU and the study's senior author. "Correct tautomer assignment is consequently important for molecular modeling and structure-based drug discovery."
Yet determining the correct tautomer remains a challenge—like finding a needle in a moving haystack. This is, in part, because of a scarcity of experimental data that characterizes tautomer structures. For instance, in the Protein Data Bank, a worldwide repository of the 3D structures of large biological molecules that is widely used in biological research, the location of hydrogen atoms that distinguish one tautomer from another are typically not available. The Protein Data Bank contains molecular structures that are largely determined using experimental methods such as X-ray crystallography, but the resolution of macromolecular X-ray structures is generally insufficient to locate the position of hydrogen atoms reliably. As a result, hydrogen positions—and therefore the tautomeric state of a molecule—often have to be inferred by scientists.
Other methods also fall short. Using quantum mechanics to assess tautomers is often too computationally demanding and expensive for screening full libraries, while machine learning is limited by the relatively small datasets of experimentally characterized tautomers in solution, which tend to contain only a few hundred molecules.
In contrast, many high-resolution small-molecule X-ray crystal structures in the Cambridge Structural Database—the largest repository of experimental crystal structures in the world—show the position of hydrogen atoms.
"We realized that experimentally resolved hydrogen positions in high-resolution small-molecule crystal structures provide a largely untapped source of experimental information about tautomer stability," said Zhang, who is also part of the NYU Simons Center for Computational Physical Chemistry.
Xiaolin Pan, a postdoc in Zhang's lab and the study's first author, systematically mined the Cambridge Structural Database to construct a dataset containing more than 1.1 million tautomeric states, orders of magnitude larger than existing experimental datasets. He then trained a graph neural network—a type of artificial intelligence that uses deep learning to find patterns between connected data points—to predict stable tautomers directly from 2D molecular forms, without requiring 3D structures or quantum-mechanical calculations.
When the researchers applied their model to 5,075 PDBbind ligands—biomolecular complexes found in the Protein Data Bank—with multiple possible tautomeric states, they identified 126 cases—approximately 2.5 percent—in which the assigned ligand tautomer was likely incorrect. In each of these cases, the model reassigned an alternative stable tautomer, which showed improved hydrogen bonding patterns.
"Reassigning these tautomers generally produced more chemically reasonable interactions. This does not mean that the experimentally determined protein structures themselves are incorrect; rather, our results suggest that the previously assigned chemical representation may warrant revision," noted Zhang.
For instance, in one of the molecules in the Protein Data Bank, the AI model predicted a tautomer with a differently positioned hydrogen atom than in the original assignment. As a result of this change, the ligand forms additional hydrogen bonding with nearby protein residues—illustrating how a seemingly small change can alter the molecule's interactions with its protein environment.
Using this model to more accurately identify a tautomer for drug discovery could not only help determine how it fits into a protein binding site, but could also improve subsequent computer simulations. When assessing drug candidates, molecular dynamics simulations follow a molecule and its protein target to see how they interact and move over time.
"These calculations require the molecule's hydrogen positions and chemical bonding to be correctly assigned, and using the wrong tautomer could alter the predicted interactions and dynamics," said Zhang.
The researchers released their method as an open-source tool, Tautomer-Predictor , which can rapidly analyze very large molecular libraries. In one test, it processed about 4.6 million compounds in 3.2 hours on a single GPU-enabled node.
Additional study authors include Chao Han and Fengyang Han of NYU's Department of Chemistry. This research is published in Chemical Science, the Royal Society of Chemistry's peer-reviewed flagship journal, and is free to read .
This work was supported by the National Institutes of Health (R35-GM127040). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.