Intron
Non-expressed nucleotide sequences within genes, removed during RNA processing.
An intron is a nucleotide sequence inside a gene that does not appear in the final RNA product. The name comes from "intragenic region," meaning a region within a gene. The term covers both the DNA sequence in the gene and the corresponding RNA sequence in the transcript. The parts that remain and are joined together after introns are removed are called exons.
Introns occur in most eukaryotes and many eukaryotic viruses, found in both protein-coding genes and noncoding RNA genes. There are four main types: tRNA introns, group I introns, group II introns, and spliceosomal introns. Introns are extremely rare in bacteria and archaea.
The discovery that genes are split by introns happened independently in several labs in 1977, including those of Phillip Sharp and Richard Roberts, who shared the 1993 Nobel Prize. Other contributors included Louise Chow and Thomas Broker; much of Sharp’s work was done by postdoc Susan Berget. The term "intron" was coined by Walter Gilbert in 1978, who proposed replacing the idea of a cistron with a transcription unit containing introns (lost from mature messenger RNA) alternating with exons (expressed regions). Though sometimes called intervening sequences, that term can also include inteins, untranslated regions, and nucleotides removed by RNA editing.
Intron frequency varies widely across organisms. They are extremely common in the nuclear genomes of all eukaryotes, where protein-coding genes almost always have multiple introns. They are rare in some eukaryotic microbes like baker’s yeast. Most vertebrate mitochondrial genomes lack introns, but some non-vertebrate mitochondrial genomes have introns. An extreme case is the human DMD gene, which contains an intron over 2.4 megabases long, taking about three days to transcribe. The shortest known metazoan intron is 30 base pairs in the human MST1L gene. Some heterotrich ciliates like *Stentor coeruleus* have introns as short as 15 or 16 base pairs, though this is not universally accepted as a canonical fact.
At least four distinct intron classes have been identified: spliceosomal introns in nuclear protein-coding genes; tRNA introns in nuclear and archaeal transfer RNA genes, removed by proteins; self-splicing group I introns; and self-splicing group II introns. A possible f
- distribution
- Common in all eukaryotes; extremely rare in bacteria and archaea; absent in most vertebrate mitochondrial genomes, but some non-vertebrate mitochondrial genomes have introns
- shortest_known_intron
- 30 base pairs (human MST1L gene); some heterotrich ciliates have introns as short as 15-16 bp, but this is not a universally accepted canonical fact
- longest_known_intron
- ~2.4 megabases (human DMD gene)
Lore & Background
Introns were first discovered in protein-coding genes of adenovirus, and subsequently identified in genes encoding transfer RNA and ribosomal RNA. The fact that genes were split or interrupted by introns was discovered independently in several labs in 1977, including those run by Phillip Allen Sharp and Richard J. Roberts, for which they shared the Nobel Prize in Physiology or Medicine in 1993. Other labs included those of Louise Chow and Thomas Broker, and much of the work in the Sharp lab was done by postdoc Susan Berget. The term "intron" was coined by Walter Gilbert in 1978, who proposed replacing the idea of a cistron with a transcription unit containing introns (lost from mature messenger RNA) alternating with exons (expressed regions). Note that "intracistron" is an archaic term that was proposed but not widely adopted. Though sometimes called intervening sequences, that term can also include inteins, untranslated regions, and nucleotides removed by RNA editing. Typical splicing error rates for spliceosomal introns are very low, often less than 0.1% per gene.
Reader's Guide
Introns are fundamental to understanding gene structure and expression in eukaryotes. Their discovery in 1977 overturned the classical view of genes as continuous sequences, revealing that genes are split into coding exons and non-coding introns. This finding has profound implications for molecular biology, including the mechanisms of RNA splicing, the evolution of genomes, and the generation of protein diversity through alternative splicing. Introns are classified into at least four types: spliceosomal introns (removed by spliceosomes), tRNA introns (removed by proteins), and self-splicing group I and group II introns (removed by RNA catalysis). The frequency and size of introns vary widely across organisms; for example, human protein-coding genes almost always contain multiple introns, while baker's yeast has few. Splicing accuracy is not perfect; error rates can be as high as 2-3% per gene, leading to aberrant transcripts that are often degraded. The persistence of suboptimal splice sites in genomes, particularly in humans, is attributed to the large number of possible mutations and small effective population sizes. Introns remain a key area of study in genetics, evolution, and disease research.
Did You Know?
- Introns were first discovered in protein-coding genes of adenovirus in 1977.
- The term 'intron' was coined by Walter Gilbert as a contraction of 'intragenic region'.
- The shortest known metazoan intron is 30 base pairs long, found in the human MST1L gene.
- Splicing error rates for spliceosomal introns can be as high as 2-3% per gene.
More in Genetics Fundamentals 1-24
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
