Transcription (biology)
Process of copying DNA into RNA for gene expression.
Philip Cowie, Ruth Ross, and Alasdair MacKenzie · CC BY 3.0
Transcription is the biological process where a segment of DNA is copied into RNA, primarily to express genes. Some DNA segments are transcribed into messenger RNA (mRNA), which can then be translated into proteins, while others produce non-coding RNAs (ncRNAs). Both DNA and RNA are nucleic acids made of nucleotide sequences. During transcription, an enzyme called RNA polymerase reads a DNA sequence and builds a complementary RNA strand, known as the primary transcript.
In virology, transcription also refers to making mRNA from a viral RNA molecule. Many RNA viruses have a genome made of negative-sense RNA, which serves as a template for producing positive-sense viral mRNA—a crucial step for synthesizing viral proteins needed for replication. This process is driven by a viral RNA-dependent RNA polymerase.
A DNA transcription unit that codes for a protein includes both a coding sequence (translated into protein) and regulatory sequences that direct and control its synthesis. The regulatory region before the coding sequence is called the five prime untranslated region (5'UTR); the region after it is the three prime untranslated region (3'UTR). Unlike DNA replication, transcription produces an RNA strand that uses uracil (U) wherever thymine (T) would appear in a DNA copy.
Only one of the two DNA strands acts as a template for transcription. RNA polymerase reads the antisense strand from the 3' end to the 5' end (3' → 5'), and the complementary RNA is built in the opposite direction, 5' → 3', matching the sense strand except that uracil replaces thymine. This directionality exists because RNA polymerase can only add nucleotides to the 3' end of the growing RNA chain. Using only the 3' → 5' DNA strand avoids the need for Okazaki fragments seen in DNA replication, and also eliminates the requirement for an RNA primer to start synthesis. The non-template (sense) strand is called the coding strand, since its sequence matches the new RNA transcript (with uracil instead of thymine), and it is the strand conventionally used when presenting a DNA sequence. Transcription has some proofreading mechanisms, but they are fewer and less effective than those in DNA replication, resulting in lower copying fidelity.
Transcription proceeds through four major steps: initiation, promoter escape, elongation, and termination.
Setting up transcription in mammals involves many cis-regulatory elements, such as the core promoter and promoter-proximal elements near transcription start sites. Core promoters, along with general transcription factors, can direct initiation but usually have low basal activity. Other important cis-regulatory modules lie far from transcription start sites, including enhancers, silencers, insulators, and tethering elements. Among these, enhancers and their associated transcription factors play a leading role in initiating gene transcription. An enhancer located far from a gene’s promoter can dramatically boost transcription—some genes show up to 100-fold increases when an enhancer is activated.
Enhancers are major gene-regulatory regions that control cell-type-specific transcription, often by looping over long distances to physically contact the promoters of their target genes. Although there are hundreds of thousands of enhancer DNA regions, only specific enhancers are brought close to the promoters they regulate in a given tissue. In a study of brain cortical neurons, 24,937 such loops were found, connecting enhancers to their target promoters. Multiple enhancers, often tens or hundreds of thousands of nucleotides away from their target genes, loop to those promoters and can coordinate with each other to control transcription.
The schematic in the source shows an enhancer looping to come into close physical proximity with a target gene’s promoter. The loop is stabilized by a dimer of a connector protein (such as CTCF or YY1), with one member anchored to a binding motif on the enhancer and the other to a motif on the promoter (represented by red zigzags). Several cell-function-specific transcription factors (there are about 1,600 in a human cell) typically bind to specific motifs on an enhancer. When a small combination of these enhancer-bound factors is brought near the promoter by the DNA loop, they govern the transcription level of the target gene. Mediator, a complex of about 26 proteins, communicates regulatory signals from enhancer-bound transcription factors directly to RNA polymerase II (pol II) at the promoter. Active enhancers are generally transcribed from both DNA strands by RNA polymerases moving in opposite directions, producing two enhancer RNAs (eRNAs). An inactive enhancer may be bound by an inactive transcription factor; phosphorylation of that factor can activate it, and that activated transcription factor then helps drive transcription.
- field
- Molecular biology
- key_process
- Synthesis of RNA from DNA template
- enzyme_involved
- RNA polymerase
- direction_of_synthesis
- 5' → 3'
- template_strand
- Antisense (3' → 5')
- nucleotide_substitution
- Uracil replaces thymine
Lore & Background
Transcription is divided into initiation, promoter escape, elongation, and termination. In mammals, setting up for transcription is regulated by many cis-regulatory elements, including core promoter and promoter-proximal elements near transcription start sites. Enhancers, often located far from their target genes, loop through long distances to come into physical proximity with promoters, stabilized by connector proteins such as CTCF or YY1. Transcription factors bind to specific motifs on enhancers, and the Mediator complex communicates regulatory signals to RNA polymerase II.
Reader's Guide
Transcription is fundamental to gene expression, converting genetic information from DNA into RNA. It differs from DNA replication in several ways: only one DNA strand serves as template, uracil replaces thymine, and no RNA primer is needed. Transcription has fewer proofreading mechanisms than DNA replication, resulting in lower copying fidelity. In virology, transcription also refers to mRNA synthesis from viral RNA molecules, such as negative-sense RNA viruses using a viral RNA-dependent RNA polymerase. Regulation of transcription involves enhancers, silencers, insulators, and CpG island methylation, which can silence gene expression when methylated. About 60% of promoters contain CpG islands, and methylation of these regions reduces or silences transcription through methyl binding domain proteins.
Did You Know?
- Only one of the two DNA strands serves as a template for transcription; the antisense strand is read from 3' to 5'.
- Transcription has fewer and less effective proofreading mechanisms than DNA replication, leading to lower copying fidelity.
- Enhancers can loop over long distances to come into physical proximity with their target gene promoters, often stabilized by connector proteins like CTCF or YY1.
- About 60% of promoter sequences have a CpG island, and methylation of these islands can reduce or silence gene transcription.
The Chemical Architecture of Strand Orientation
The orientation of a single nucleic acid strand is defined by the numbering of carbon atoms within the pentose sugar ring of each nucleotide. At one terminus, the fifth carbon carries a phosphate group—this is the 5′ end, commonly spoken as "five-prime." At the opposite terminus, the third carbon bears a hydroxyl group, giving rise to the 3′ end, or "three-prime." This nomenclature is not arbitrary; it reflects the actual chemical structure of the ribose or deoxyribose ring. In a double helix, the two strands must run antiparallel to each other so that complementary bases can pair correctly, a geometric requirement that underpins both replication and transcription. The relative positioning of functional elements along a strand follows this same logic: regions closer to the 5′ terminus are termed upstream, while those nearer the 3′ terminus are called downstream. By universal convention, single-stranded DNA and RNA sequences are written from 5′ to 3′ unless the purpose is to display base-pairing geometry.
The Unidirectional Rule of Polymerization
A fundamental constraint governs all in-vivo nucleic acid synthesis: new strands are assembled exclusively in the 5′-to-3′ direction. The polymerases responsible for building RNA or DNA chains harness the energy released when nucleoside triphosphate bonds are cleaved to forge a phosphodiester bond between the incoming nucleotide's 5′-phosphate and the growing strand's 3′-hydroxyl group. This chemical mechanism makes reverse-direction synthesis impossible under normal cellular conditions. The 3′-hydroxyl is thus the critical reactive handle; without it, chain elongation simply cannot proceed. This principle has practical consequences in the laboratory: molecular biologists exploit it by introducing dideoxyribonucleotides—nucleotides that lack the 3′-hydroxyl—to deliberately terminate DNA replication, a strategy known as the Sanger chain-termination method for reading nucleotide sequences. Similarly, the 5′-phosphate can be enzymatically stripped with a phosphatase to block unwanted ligation events, such as the self-ligation of plasmid vectors during cloning experiments.
Maturing the Transcript: Capping and Polyadenylation
Once a nascent messenger RNA strand emerges from the transcription machinery, it undergoes two critical post-transcriptional modifications, one at each end. At the 5′ terminus, a methylated guanosine nucleotide is attached through an unusual 5′-to-5′ triphosphate linkage—a connection that is rare in nucleic acid chemistry. This cap shields the mRNA from exonucleases, thereby extending its functional lifespan during translation. At the 3′ terminus, a tail of roughly fifty to two hundred and fifty adenosine residues is appended in a process called polyadenylation. The length of this poly-A tail directly influences how long the mRNA persists in the cell and consequently how much protein it can encode. Flanking these modified regions, the 5′-untranslated region (from the cap site to the base just before the AUG initiation codon) and the 3′-untranslated region (from the stop codon to the poly-A tail) are transcribed but not translated; they harbor regulatory sequences such as the Kozak sequence, ribosome binding sites, and enhancer elements that modulate translation efficiency and mRNA stability.
Template, Sense, and the Reading Frame
Directionality and sense are related but distinct concepts in transcription. When a double-stranded DNA gene is transcribed, only one of the two strands serves as the direct template; RNA polymerase reads this template strand and assembles a complementary RNA strand. The opposite strand, though not copied, carries a sequence that mirrors the RNA product and is therefore called the sense strand. Transcription initiation sites are found on both strands of an organism's genome, each specifying where, in which direction, and under what conditions a gene will be transcribed. A concrete example illustrates the flow: the sense strand contains the sequence 5′-ATG-3′, while the template strand presents 3′-TAC-5′. The polymerase copies the template to produce 5′-AUG-3′ in the mRNA. The ribosome then scans this mRNA from its 5′ end, recognizes the AUG start codon, and begins incorporating amino acids at the N-terminus, extending the polypeptide toward the C-terminus. In bacteria, mitochondria, and plastids, the initiating amino acid is N-formylmethionine rather than plain methionine.
Gallery






Frequently Asked Questions
Who is Transcription (biology)?
Transcription is the molecular biology process that copies a segment of DNA into a complementary RNA strand, acting as the launch point for gene expression. It is carried out by the enzyme RNA polymerase reading the DNA template and assembling a primary RNA transcript.
What are Transcription (biology)'s powers/role?
Its signature ability is building an RNA strand in the 5′→3′ direction by reading the antisense (3′→5′) DNA template. Depending on the gene, it can produce messenger RNA destined for protein synthesis or a variety of non-coding RNAs with regulatory roles.
How does Transcription (biology)'s story end?
The arc wraps up when RNA polymerase hits a termination sequence on the DNA, releasing the finished primary transcript. That RNA molecule then heads off to splicing, translation, or other downstream cellular duties.
Why is Transcription (biology) important?
It is the critical step that converts the static genetic instructions stored in DNA into functional RNA molecules, without which no proteins could be made and no gene could be expressed. Every living cell depends on it to turn its genetic blueprint into working cellular machinery.
What's Transcription (biology)'s signature move?
Its most recognizable trait is swapping uracil in place of thymine when pairing with adenine on the template strand, a nucleotide substitution that clearly marks the product as RNA rather than DNA. Fans often cite this uracil-for-thymine swap as the tell that distinguishes a transcription product from the original DNA.
More in Cell & Molecular Biology 1-16
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
