Pharmaceutical Data Gives Protein-Folding AI a Major Performance Boost
A consortium of drug companies trained an OpenFold3 model on more than 20,000 proprietary protein structures, improving predictions of how proteins interact with potential medicines.
Protein-structure models based on artificial intelligence can be improved by incorporating data from pharmaceutical companies.
Credit: Miyako Nakamura/Getty
Artificial-intelligence models used for drug discovery, including protein-folding systems such as AlphaFold, face a fundamental data problem: public databases contain too few examples of how proteins interact with drugs and other molecules.
Thousands of additional protein structures are stored in pharmaceutical companies’ internal archives. A consortium of drug companies reports that using these proprietary data to train an AI protein-folding model significantly improves its performance.
The group developed a model based on OpenFold3, an open-source clone of AlphaFold 3. Trained on more than 20,000 unique protein structures, the system outperformed comparable models trained solely on public data, as well as systems trained on siloed datasets from individual companies.
The findings were described in a blog post and have not yet been peer-reviewed. The model is also not publicly available.
“When you add all this data, you get a huge performance boost,” says Mohamed Al-Quraishi, a computational biologist at Columbia University in New York City who participated in the effort.
The results strengthen the case for creating similar public datasets to support protein-folding AI, Al-Quraishi says. One project, called OpenBind, is backed by up to £8 million (US$10.8 million) in UK government funding. It published hundreds of new protein structures in the last month, with thousands more in development.
A largely unexplored source of protein data
The Protein Data Bank (PDB) is an open repository containing more than 200,000 experimentally determined protein structures. It provided the basis for AlphaFold 2’s training data, enabling the system to predict protein structures with remarkable accuracy. The breakthrough was recognized with the 2024 Nobel Prize in Chemistry.
Later systems, including AlphaFold 3, can predict how proteins interact with other molecules, including potential drugs. However, the PDB contains relatively few experimentally determined structures showing proteins interacting with drug-like molecules. There may be only about 10,000 such examples, says Paul Mortenson, vice president of computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, UK.
This shortage of data limits the usefulness of AI for drug discovery. Research suggests that the accuracy of AlphaFold 3 and other “cofolding” models, which predict the structures of interacting proteins or molecules, drops sharply when they are asked to predict interactions involving molecules that differ substantially from those in their training data.1
To make protein-folding tools more effective for drug discovery, researchers say they need access to more examples of protein–drug interactions. Pharmaceutical companies’ repositories could provide a valuable source of this information.
These structures are generated during drug-discovery programmes using techniques such as X-ray crystallography and cryo-electron microscopy. Many are linked to proprietary drug-development projects and have never been added to public databases.
The total size of these private collections is unknown, but some estimates suggest that they could contain more data than the PDB.
“The data that is missing from the PDB is exactly the data that exists in our internal data,” said John Karanikolas, head of computational drug discovery at AbbVie, a pharmaceutical company in Chicago, Illinois, to Nature last year.
Training AI on proprietary protein structures
To determine whether pharmaceutical-company data could improve protein-folding models, AbbVie, Astex and several other drug companies formed a collaboration last year called the AI Structural Biology (AISB) Network.
The group fine-tuned OpenFold3, which had previously been trained only on PDB data, using an additional 20,167 structures of proteins bound to potential drugs or other ligands. Five companies supplied the structures and contributed them to the training process in a way that kept their proprietary data private.
The AISB study found that the additional data improved the model’s predictions. When tested on 1,056 protein–ligand structures that were separate from the training data, the AISB model predicted more than half with a high level of accuracy.
By comparison, the public version of OpenFold3 achieved the same level of performance on only about one-third of the structures. Boltz-2, a competing open-source model, achieved a success rate of around 40%.
The team plans to submit a paper describing the research to a peer-reviewed journal.
The AISB model’s stronger performance compared with a cofolder trained on each company’s individual dataset highlights the potential benefits of pooling information, Karanikolas says.
Source: www.nature.com


