AI-designed proteins can carry hidden watermarks to improve biosecurity
Proteins designed by AI tools, illustrated by artists, can now include hidden digital tags that reveal their origins.
Credit: Christoph Burgstedt/SPL
Proteins designed by artificial intelligence could soon carry a hidden signature of their machine-generated origin, thanks to a watermarking system described today in Nature.1
Developed by researchers at Google DeepMind in London, the technology borrows techniques already used to identify AI-generated images, video, audio and text. It weaves subtle statistical information into both the amino acid sequence and the three-dimensional shape of computer-designed proteins without significantly compromising their function.
Known as SynthIDBio, the approach can distinguish AI-generated proteins from native proteins in biological repositories. This could help protect the integrity of databases used for scientific research and biosecurity screening.
Steph Guerra, a biosecurity scholar at RAND, a nonprofit research institute in Washington, D.C., sees value in an approach that aligns the interests of the life-sciences community. Watermarks can support innovation and scientific reproducibility, “and also have security benefits,” she says.
But the molecular stamp can be removed. Someone seeking to erase the “made by AI” tag could run a watermarked protein through another design tool to generate new sequences that preserve its structure and function while obscuring its synthetic origins.
For that reason, SynthIDBio is best viewed as one tool in a multilayered framework for defending against biological threats, says Tessa Alexanian, a former biosecurity researcher at the International Biosecurity and Biosafety Initiative for Science, a nonprofit organization in Geneva, Switzerland. “We are in a whole new world,” she says.
How AI protein watermarks could strengthen biosecurity
Potential users of watermarking systems include companies that create DNA sequences to order. Customers use this DNA to produce the desired protein.
These companies regularly check whether customer DNA sequences encode known toxins, proteins made by pathogens or related molecules. However, AI-designed proteins with sequences that have little or no similarity to those in reference databases could evade these checks, even when the resulting proteins perform nearly the same function after being produced and injected into cells.
One solution, outlined by Alexanian and her colleagues, involves screening not only for sequence matches but also for the biological functions that DNA sequences may encode.2–4 However, this requires researchers to determine the function of the resulting protein.
SynthIDBio offers a simpler alternative by inserting hidden clues into AI-designed proteins. Subtle changes in the pattern of amino acid building blocks are encoded at the sequence level, while subtle changes in the way atoms are arranged encode information in the protein’s three-dimensional structure.
Watermarks are automatically added to proteins designed by AI tools such as AlphaFold and RFdiffusion. Anyone with a secret detection key could potentially identify them, although the key would be shared only with trusted partners, such as DNA sequencers.
The watermarks reveal nothing beyond the fact that a protein was created using an AI tool, but they could still be useful. “The synthesis provider I talked to was like, ‘Any information you can give us to help us understand these orders is fine,’” Alexanian says.
Testing whether AI protein watermarks affect function
If watermarks interfere with the purpose of protein design, their usefulness is compromised. Computer scientist Pushmeet Kohli and his team at DeepMind therefore “stress-tested their approach to a number of difficult problems,” he says.
They found that watermarked proteins could bind to a variety of targets as efficiently as their unwatermarked counterparts, including targets involved in viral infection, angiogenesis and immune regulation.
Source: www.nature.com


