Google DeepMind debuts SynthID Bio for watermarking AI-designed proteins
Google DeepMind's SynthID Bio adds detectable signatures to AI-designed proteins, but its wet-lab testing was narrow and deliberate tampering remains unresolved.
Google DeepMind introduced SynthID Bio, a proof-of-concept family of methods that places detectable watermarks in AI-generated protein sequences and predicted biomolecular structures. The provenance signal could supplement DNA-synthesis screening and checks on biological databases, though DeepMind has not disclosed any production deployment by a synthesis provider or database.
The project sits alongside DeepMind’s work on open predicted viral protein-complex structures, but tackles a different question: whether a generated sequence or structure carries a detectable provenance signal.
For sequences, SynthID Bio adapts SynthID’s sampling technique to ProteinMPNN, a model that designs amino-acid sequences for a specified protein structure. A secret key influences which variants are selected during generation, while a watermark-score filter strengthens the signal. Detection requires the same key. The Nature methods paper frames this as function-preserving watermarking, not a way to identify every possible origin of a protein.
The disclosed laboratory evaluation was narrow. It covered redesigned binders against three targets: the SARS-CoV-2 receptor-binding domain, VEGF-A and PD-L1. Researchers resequenced 15 previously validated AlphaProteo backbones for each target rather than running a full de novo protein-design campaign. Excluding negative controls, the paper reports 222 non-watermarked sequences and 267 watermarked sequences for each of two watermark settings.
Across those three targets, the researchers reported no statistically significant difference between watermarked and non-watermarked designs in binder hit rates or binding-affinity distributions. The results come from the DeepMind-led study and have not been independently replicated in the opened evidence. Nor do they establish performance across broader protein classes.
A second method, SynthIDBio-structure, fine-tunes AlphaFold 3’s diffusion and confidence modules alongside a detector so that generated atomic coordinates carry a zero-bit watermark. That means the detector can test whether a mark is present, but cannot distinguish among users. The team reported near-perfect detectability on AlphaFold 3’s evaluation set, negligible impact on global accuracy, and robustness to rigid transformations, coordinate noise and cropping.
DeepMind released sequence-watermarking code and in vitro data and provided instructions for requesting the structure model’s weights. The repository lists the software under Apache 2.0 and the data under CC BY 4.0; AlphaFold 3 weights remain subject to separate terms.
Deliberate removal remains unresolved. The paper says resequencing attacks can largely strip the sequence watermark, including when an attacker starts with a known binder and generates a replacement sequence. It also says the structure watermark is not robust to structural relaxation, cannot attribute a design to a particular user and has not been tested for differentiation from other structure-watermark schemes. DeepMind calls the work a first step and identifies stronger resistance to deliberate tampering as a remaining challenge.
More news

Databricks launches Lakebase Search for vector and full-text retrieval in Postgres

OpenAI partners with America’s SBDC on small-business AI training

DoorDash opens waitlist for corporate AI ordering connector
