Google’s SynthID Bio Watermarks AI-Designed Proteins for Biosecurity
Google’s SynthID Bio system embeds a watermark across the full length of a protein sequence while preserving the protein’s ability to function. The approach could help researchers identify AI-designed proteins that do not closely resemble anything already known.
How SynthID Bio Watermarks Protein Sequences
One way to understand the system is to consider groups of chemically related amino acids, such as leucine, isoleucine, and valine, or serine and threonine. When multiple amino acids in one of these groups can work in a given position, the system attempts to use the amino acid that matches the watermark.
Another explanation, proposed by one of the people involved in developing the system, is that SynthID Bio searches the space of functional proteins for sequences containing enough watermark-associated amino acids.
As a result, the watermark is distributed randomly across the entire protein rather than appearing as a simple, obvious pattern. Detecting it is not a straightforward “yes” or “no” process. Researchers need the key, scan the complete protein sequence, and measure how often the amino acids suggested by SynthID Bio appear in the final sequence. Google has also developed software to perform this detection.
Do Watermarked Proteins Still Work?
The key question is whether adding a watermark affects a protein’s function. To test this, the team used SynthID Bio to design proteins that physically interact with important natural proteins previously targeted in AI-based design work.
The watermarked proteins worked as intended and bound to their target. This was not a rigorous test of whether the system could produce a catalyst, but it suggests that watermarks may not cause serious problems in more complex protein-design tasks.
How Protein Watermarking Could Improve Biosecurity
As long as a protein is long enough, the system can detect its watermark. That could be useful for biosecurity when researchers encounter AI-designed proteins that do not resemble known biological threats.
When someone orders a DNA sequence, the company producing it typically screens the sequence for genes that could encode parts of viruses, toxic proteins, or other threats. However, a protein that does not resemble anything already known—including a protein designed by AI—can be difficult to assess for potential risks.
A detectable watermark could provide an additional way to identify proteins created with AI systems, helping researchers distinguish novel designs from naturally occurring or previously documented sequences.
Source: arstechnica.com


