All known life on Earth uses the same four-letter genetic alphabet. Now, researchers at the University of California, San Diego, have shown that RNA polymerase—one of biology’s most important enzymes—can accurately read and transcribe an expanded genetic alphabet containing eight letters.
The discovery suggests that cells may be able to use existing molecular machinery to process synthetic genetic information. It also marks an important step toward a long-standing goal in synthetic biology: expanding DNA beyond the four natural bases. In the future, eight-letter DNA systems could help scientists develop biological machines with new capabilities and create compounds that do not occur in nature.
How RNA polymerase reads an eight-letter genetic alphabet
The researchers focused on RNA polymerase, the enzyme responsible for reading DNA and producing RNA during the first stage of gene expression. To understand how this enzyme processes synthetic genetic information, the team combined biochemical experiments with high-resolution cryo-electron microscopy.
The researchers captured a detailed structural view of RNA polymerase from Escherichia coli as it recognized and incorporated two synthetic base pairs. These artificial genetic letters are not found in natural DNA.
The structural images revealed that RNA polymerase relies on many of the same biochemical and molecular signals it uses to identify natural base pairs. This finding helps explain how the enzyme can accurately copy and transcribe DNA containing an expanded genetic alphabet.
In related research published in PNAS, the same team found that RNA polymerase can also recognize other synthetic base pairs, even when those pairs lack the hydrogen bonds that normally help stabilize natural DNA.
Synthetic DNA could enable new technologies
The potential applications extend far beyond understanding how DNA is processed. Previous studies have used expanded genetic alphabets to create synthetic DNA molecules capable of recognizing liver cancer cells.
By revealing how RNA polymerase reads and transcribes non-natural DNA letters, the new study provides a molecular foundation for technologies based on an expanded genetic code. Potential applications include advanced diagnostic tools, new therapeutics, and engineered biological systems with functions that do not exist in nature.
Two studies investigate the expanded genetic code
The study, led by Dr. Dong Wang, a professor at the University of California San Diego Skaggs School of Pharmacy, is titled “Structural basis of transcription of the eight-letter alphabet by Escherichia coli RNA polymerase.” It was published in the journal Science on September 2, 2026.
A related study published in PNAS on August 12, 2026, titled “Hydrophobic unnatural base pairs promote trigger loop closure and catalysis in cellular RNA polymerases independent of hydrogen bonding,” was also led by Wang.
Source: www.sciencedaily.com


