← Wonderland

The Central Dogma of Molecular Biology

In 1958, Francis Crick compressed the entire logic of life into a single arrow: DNA to RNA to protein. The exceptions found since have only sharpened the claim.

In 1958, a year before he shared the Nobel Prize for the double helix, Francis Crick delivered a lecture in London that laid out the rules of a game he could not yet play. He called it the central dogma. He later apologized for the word — “dogma,” he admitted, was exactly the wrong word for a scientific claim — but the name stuck because the claim behind it was so audacious. All of life, Crick was saying, runs on a single flow of information. From DNA, to RNA, to protein. One direction. No exceptions that mattered.

Sixty years later, biology has poked at every seam. Retroviruses copy RNA back into DNA. Prions replicate without any nucleic acid at all. Epigenetic marks pass from parent to child without touching the sequence. And still the dogma stands, because what Crick was really claiming was narrower than it sounded, and cleaner than it looked.

Crick's simple statement

When Crick sat down in 1958 he was not making a philosophical point. He was stating a rule about polymer sequences. Proteins are chains of amino acids. Nucleic acids are chains of bases. What Crick claimed was very specific: once information has passed into a protein — into the particular sequence of amino acids that defines it — it cannot get back out. Proteins cannot dictate the sequence of a new nucleic acid. Proteins cannot dictate the sequence of another protein. The arrow points one way.

This is subtler than it is often quoted. Crick did not claim that DNA → RNA → protein was the only possible flow; he claimed that protein → nucleic-acid was impossible. In his 1970 clarification he drew nine arrows: three general transfers (DNA → DNA, DNA → RNA, RNA → protein), three special transfers that can happen under the right circumstances (RNA → DNA, RNA → RNA, DNA → protein), and three that were forbidden (protein → DNA, protein → RNA, protein → protein). Only the last three carried the real content of the dogma. Everything else, he allowed, might happen somewhere in nature. Some of it does.

Transcription and translation

The machinery Crick was theorizing about was, at the time, still being pieced together. Today it is textbook. In the nucleus of every one of your cells, a stretch of double-stranded DNA is prized apart and an enzyme called RNA polymerase crawls along one strand, laying down a complementary single strand of RNA. This is transcription. The result — messenger RNA, mRNA — is a working copy, disposable, meant to travel.

DNA transcription — RNA polymerase A U G C U U G A C A A G C C U mRNA translation — ribosome, tRNAs Met Leu Asp Lys Ala protein information flows down. Never up.

The central pipeline. DNA is transcribed into mRNA; mRNA is translated into a chain of amino acids that folds into a protein.

The mRNA leaves the nucleus and finds a ribosome — a massive protein-RNA machine floating in the cytoplasm. There, translation happens. The ribosome reads the mRNA three bases at a time. Each triple of bases — a codon — corresponds to one of twenty amino acids. Transfer RNAs (tRNAs), tiny cloverleaf molecules carrying a codon on one end and its matching amino acid on the other, dock into the ribosome in the order the mRNA prescribes. The ribosome catalyzes a peptide bond between successive amino acids. A protein chain grows, one link every fifty milliseconds or so, until a stop codon ends the process. The chain then folds — usually within seconds — into a three-dimensional shape that determines what it does.

You are, right now, doing roughly a billion of these translations per second in every cell.

The genetic code

There are four bases in RNA (A, U, C, G), and 4³ = 64 possible codons. There are twenty standard amino acids. So the mapping is redundant: many codons code for the same amino acid, and three codons code for “stop.” The code was cracked between 1961 and 1965, largely by Marshall Nirenberg and Har Gobind Khorana, using synthetic RNA sequences fed to cell-free extracts.

A U G G C A U U C codon 1 codon 2 codon 3 Methionine Alanine Phenylalanine (start)

Reading frame. The ribosome walks three bases at a time; each codon selects one amino acid.

What is truly astonishing about the code is that it is nearly universal. The codon UGG codes for tryptophan in bacteria, in yeast, in wheat, in whales, in you. A gene from a jellyfish, spliced into a bacterium, produces jellyfish protein — that is how green fluorescent protein became a laboratory tool. The code was frozen — or almost frozen — before the last universal common ancestor, more than three billion years ago.

Every organism you have ever encountered runs the same instruction set.

Where the dogma bends

Almost as soon as Crick had drawn his diagram, someone found an exception. In 1970, David Baltimore and Howard Temin independently discovered reverse transcriptase — an enzyme, carried by certain viruses, that copies RNA back into DNA. This is how HIV integrates into the genome of a human cell. The arrow RNA → DNA, one of Crick's “special” transfers, turned out to be a routine trick in the viral world. It also turned out to be one of the most useful tools in molecular biology: reverse transcriptase lets you sequence RNA, engineer stable cell lines, and — a half-century later — verify what an mRNA vaccine actually delivers.

Then there are prions. A prion is a protein that, once misfolded, coaxes other copies of the same protein into misfolding in the same way. No nucleic acid is involved. Kuru, Creutzfeldt–Jakob disease, mad cow disease — all caused by protein propagating protein. This looks like a violation. But it isn't, quite: the prion's amino acid sequence is not being copied by another protein. It is being coded for by ordinary DNA, as always. Only the folded shape is copied protein-to-protein. Crick's dogma was about sequence, not shape.

Epigenetics — the pattern of chemical marks on DNA and histones that switches genes on or off, and that can, in some cases, pass across a generation — also looks like a subversion. It isn't; it is a layer above the sequence, not a rewriting of it. The dogma survives every serious attack because it was, from the start, a claim about sequence information, not about everything a cell does.

What Crick captured, in the phrase he later regretted, is that life is a code with a direction. Nature figured out, once, how to fold a self-replicating polymer into machinery capable of building its own descendants — and having found that trick, it never found another. From bacteria to elephants, the arrow points the same way. That is the central dogma. It is not a law. It is more like a discovery about what evolution stumbled into and then never left.


Further reading

  1. Crick, F. (1958). On Protein Synthesis. Symposia of the Society for Experimental Biology, XII.
  2. Crick, F. (1970). Central Dogma of Molecular Biology. Nature, 227, 561–563.
  3. Watson, J. D., et al. (2013). Molecular Biology of the Gene, 7th edition.
  4. Alberts, B., et al. (2014). Molecular Biology of the Cell, 6th edition, chapters 6–7.
  5. Baltimore, D. (1970). Viral RNA-dependent DNA polymerase. Nature, 226, 1209–1211.