Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How to Learn RNA Processing and Alternative Splicing: From Pre-mRNA to Isoforms, Surveillance and Long-Read Transcriptomics

Reader safety: This is an educational molecular-biology guide. It explains RNA processing, measurement and current research without offering diagnosis or treatment advice.

Wait, What? A Gene Is Not Necessarily One Final Message

A beginner often learns a simple line:

DNA → RNA → protein

Useful. But incomplete.

In eukaryotic cells, the first RNA copy of many genes is a pre-mRNA. It contains exons that may remain in the mature message and introns that are usually removed. The cell must identify boundaries, cut the RNA, join selected pieces, add a 5′ cap, form the 3′ end, inspect the transcript and decide whether the message is ready to leave the nucleus.

Even more surprisingly, the same gene can be processed in more than one way. Different exon choices can create different RNA isoforms.

gene sequence ≠ one inevitable mature RNA

RNA processing is therefore not clerical cleanup after transcription. It is part of how cells construct usable information.

The One-Sentence Answer

Learn RNA processing by following one pre-mRNA from transcription through capping, splice-site recognition, lariat formation, exon joining, alternative splicing, 3′-end formation, surveillance and isoform measurement, while keeping the distinction between an observed RNA isoform and its actual biological function visible.

Stage 1: Start With Pre-mRNA, Not Mature mRNA

RNA polymerase II produces a primary transcript containing coding and non-coding regions. Many introns are removed before a transcript becomes a conventional mature mRNA.

The important learning move is to separate:

  • the genomic DNA sequence;
  • the newly transcribed RNA;
  • the processed RNA isoform;
  • the translated protein product.

Those are related representations of one information flow, not interchangeable objects.

Stage 2: The 5′ Cap Is Added Early

The growing RNA acquires a modified guanosine cap near its 5′ end. The cap helps with RNA stability, processing, nuclear export and later translation initiation.

Processing therefore begins while transcription is still happening.

Stage 3: Splicing Removes Introns and Joins Exons

The central machine is the spliceosome, a dynamic ribonucleoprotein assembly containing small nuclear RNAs and many proteins.

Its job sounds simple:

recognise the correct boundaries → remove intron → ligate exons

But the recognition problem is difficult because human introns can be very large while the sequence signals marking splice sites are comparatively short and context-dependent.

Stage 4: U1 and U2 Help Define the Intron

In the major spliceosome, U1 snRNP participates in recognising the 5′ splice site and U2 snRNP associates with the branch-point region. Additional components join and the complex is extensively rearranged before catalysis.

The mature active site is built largely from RNA interactions. Structural studies have shown the spliceosome to be a highly dynamic, protein-orchestrated RNA catalyst.

Stage 5: Splicing Uses Two Transesterification Reactions

The branch-point adenosine attacks the 5′ splice site, creating a branched lariat intermediate. The newly freed 5′ exon then attacks the 3′ splice site, joining the exons and releasing the intron lariat.

The sequence becomes much easier to remember when represented as a state change:

linear pre-mRNA → lariat intermediate + free exon → joined exons + excised lariat

Stage 6: Splicing Is a Kinetic Process, Not Just a Sequence Rule

Recognition depends on more than whether a short sequence resembles a consensus splice site. It can also depend on transcription rate, RNA structure, neighbouring regulatory elements and the concentrations or states of RNA-binding proteins.

This is why:

possible splice site ≠ selected splice site

Stage 7: Much Splicing Is Co-Transcriptional

Many introns are removed while transcription is still underway. But current work has also clarified that a substantial fraction of mammalian introns can remain after transcription terminates and be removed later while the transcript is still associated with chromatin.

A 2025 Nature Reviews Genetics review emphasised this mixed picture of co-transcriptional and post-transcriptional splicing.

Stage 8: Alternative Splicing Changes Which RNA Is Produced

Common patterns include:

  • exon skipping;
  • alternative 5′ splice sites;
  • alternative 3′ splice sites;
  • mutually exclusive exons;
  • intron retention.

One gene can therefore generate multiple transcripts.

But a critical scientific caution follows:

detected isoform ≠ abundant isoform ≠ translated isoform ≠ functional protein

Stage 9: Cis Elements and Trans Factors Regulate Choice

Regulatory sequences within the RNA can act as splicing enhancers or silencers. RNA-binding proteins such as SR proteins and heterogeneous nuclear ribonucleoproteins can alter splice-site recognition.

The same exon can therefore be treated differently in different cell types or physiological states.

Stage 10: Alternative Splicing Is Part of Cell Identity

Neurons, muscle cells, immune cells and developing tissues can express different combinations of splicing regulators. Their transcriptomes therefore differ not only in how much of each gene is expressed, but also in which isoforms are produced.

Gene-expression level and isoform identity are separate variables.

Stage 11: Intron Retention Is Not Automatically an Error

An intron retained in a transcript may lead to nuclear detention, degradation, altered translation or a regulated intermediate that can be processed later.

Some retained introns are mistakes. Others are regulated states.

So:

intron retained ≠ failed cell

Stage 12: The Minor Spliceosome Handles a Rare Intron Class

Most introns use the major spliceosome. A small fraction of U12-type introns use a distinct minor spliceosome.

Rarity does not mean biological irrelevance. Mutations affecting minor-spliceosome components can have major developmental consequences because the affected introns can occur in important genes.

Stage 13: RNA Processing Includes More Than Splicing

For many protein-coding transcripts, the final message also requires:

  • 5′ capping;
  • splicing;
  • 3′ cleavage;
  • poly(A) addition;
  • RNA surveillance;
  • export.

Alternative polyadenylation can change a transcript’s 3′ end, regulatory elements and sometimes coding sequence.

Stage 14: Nuclear Speckles Are Processing Environments

Splicing factors are enriched in nuclear speckles and related nuclear compartments. Modern cell biology increasingly treats these not simply as storage bins but as dynamic environments that can influence access, timing and coordination of RNA processing.

Recent work also connects multivalent interactions and condensate-like organisation to RNA-processing control.

Stage 15: The Exon Junction Complex Records a Processing Event

Splicing can leave protein complexes associated with exon junctions. These marks participate in downstream RNA metabolism, including surveillance.

The history of how an RNA was processed can therefore affect what happens later.

Stage 16: Nonsense-Mediated Decay Is RNA Quality Control

If translation encounters a premature termination context, the transcript may be targeted by nonsense-mediated mRNA decay.

This protects cells from some potentially harmful truncated proteins, but NMD also regulates many normal transcripts.

Surveillance is not merely a garbage-disposal step. It is part of gene-expression control.

Stage 17: Splice-Site Mutations Can Change the Message Without Changing an Encoded Amino Acid Directly

A DNA variant can disrupt:

  • a splice donor;
  • a splice acceptor;
  • a branch point;
  • an enhancer or silencer;
  • a regulatory protein-binding site.

The consequence may be exon skipping, intron retention or cryptic splice-site use.

This is why reading only the protein-coding triplets can miss important molecular effects.

Stage 18: RNA Structure Can Affect Accessibility

Pre-mRNA folds. Local structure can hide or expose sequence elements and alter access by splicing factors.

Sequence is therefore not interpreted in a flat one-dimensional line. Molecular geometry matters.

Stage 19: RNA Modifications Can Interact With Processing

RNA modifications such as m6A have been associated with multiple aspects of RNA metabolism, including processing and splicing in some contexts.

The correct model is not “one modification controls splicing”. It is that chemical state can alter interactions in a larger regulatory network.

Stage 20: Splicing Can Be Delayed Deliberately

Post-transcriptional splicing creates a temporal layer. A transcript may be made now but completed later.

That gives cells another control variable:

which RNA + how much RNA + which isoform + when processing finishes

Stage 21: Short-Read RNA-Seq Sees Fragments, Not Whole Transcripts

Conventional RNA sequencing produces many short reads. Reads spanning exon junctions are powerful evidence of splicing, but reconstructing full-length isoforms from fragments can be ambiguous when genes have many exons.

A set of locally correct junctions does not always identify one unique complete transcript.

Stage 22: Percent Spliced In Is a Useful Summary, Not the Whole Biology

A common metric, often called PSI, estimates how frequently a particular exon or splice event is included relative to alternatives.

It is useful for comparing conditions.

But PSI depends on:

  • read coverage;
  • mapping;
  • event definition;
  • cell mixture;
  • statistical model.

A precise-looking percentage can still carry substantial uncertainty.

Stage 23: Long-Read Sequencing Changes the Representation

Long-read technologies can observe much longer RNA molecules and therefore connect distant splice choices in the same transcript.

This is a major conceptual improvement for isoform biology:

local junction evidence → full-transcript evidence

But long-read sequencing has its own errors, depth constraints and analysis choices. A longer read is not automatically a true isoform.

Stage 24: Direct RNA Sequencing Can Preserve Additional State

Nanopore-based direct RNA approaches can sequence RNA molecules without first converting every molecule into amplified cDNA. This can preserve aspects of molecule-level information and reduce some amplification biases.

The trade-off is that raw signal interpretation and error correction become important parts of the evidence chain.

Stage 25: Single-Cell Splicing Is a Sparse-Data Problem

Single-cell transcriptomics can reveal cell-to-cell variation, but many transcripts are sampled incompletely in individual cells.

A missing junction may mean:

  • the isoform was absent;
  • the molecule was not captured;
  • coverage was insufficient.

Absence in sparse data is not the same as biological absence.

Stage 26: Protein Evidence Matters

Some RNA isoforms are degraded. Some are translated inefficiently. Some produce unstable proteins.

When a claim concerns functional protein diversity, proteomic or biochemical evidence can strengthen the inference beyond transcript detection alone.

Stage 27: Splicing Therapy Shows the Mechanism Can Be Redirected

Antisense oligonucleotides can be designed to alter access to RNA-processing elements and change exon usage in specific therapeutic contexts.

The educational point is not a treatment recipe. It is proof of principle:

splice-site choice is a manipulable molecular decision, not an immutable consequence of the DNA sequence.

Stage 28: Cryo-EM Revealed the Spliceosome as a Moving Machine

High-resolution structural studies have captured spliceosomes at multiple states. The catalytic core remains highly conserved while numerous proteins and RNA elements rearrange through assembly, activation, catalysis and disassembly.

One static structure therefore cannot represent the complete splice cycle.

Stage 29: Professional RNA Biology Separates Molecule, Measurement and Function

The advanced question becomes:

Which processing event produced this RNA isoform, how confidently was the complete molecule measured, what controls its abundance and timing, and what evidence shows that the isoform changes cellular function?

Evidence: How Do We Know Splicing Happened?

Evidence can come from:

  • junction-spanning RNA sequencing reads;
  • RT-PCR;
  • long-read transcript sequencing;
  • nascent RNA measurements;
  • spliceosome biochemistry;
  • cryo-electron microscopy;
  • genetic perturbation;
  • proteomics;
  • RNA localisation and stability measurements.

Strong claims converge across representations.

Misconceptions Worth Hunting

  • One gene always makes one mRNA.
  • Introns are simply useless DNA.
  • Alternative splicing always creates a functional protein.
  • Every retained intron is a processing error.
  • RNA-seq reconstructs full transcripts directly.
  • One splice-junction read proves a complete isoform.
  • More isoforms automatically mean more biological complexity.
  • A splice variant associated with disease is automatically causal.

Transfer Check

A gene produces two mature RNAs that differ by one exon. Did the DNA sequence necessarily change? No.

A long-read experiment detects a rare new full-length transcript. Is its function proven? No.

An intron remains in chromatin-associated RNA after transcription ends. Must splicing have failed permanently? No.

A DNA variant does not change a coding codon but disrupts a splice enhancer. Can protein output still change? Yes.

How We Know the Learning Has Held

A learner should be able to explain pre-mRNA, splice sites and branch points; describe the two catalytic steps; distinguish constitutive from alternative splicing; explain intron retention and the minor spliceosome; connect splicing to NMD; separate gene expression from isoform choice; compare short- and long-read evidence; explain PSI cautiously; and distinguish RNA detection from protein/function evidence.

Model Limits

Textbook exon boxes hide RNA structure and kinetics. Consensus splice sequences do not uniquely determine splice choice. Short-read reconstruction can be underdetermined. Long reads can contain sequencing and alignment errors. Single-cell data are sparse. RNA abundance does not establish translation. Protein detection does not establish biological necessity. Disease association does not establish causal mechanism.

Professional RNA science keeps sequence + processing state + molecule identity + timing + measurement method + function visible together.

Teaching Guide

Teach in this order:

DNA gene → pre-mRNA → 5′/3′ splice sites → branch point/lariat → exon joining → alternative splicing → surveillance → isoform measurement → long reads → function.

Begin with:

“If every cell has nearly the same genome, how can different cells produce different versions of the same gene’s RNA?”

Connect This to the eduKate Learning Estate

Research Foundations and Further Learning

The Quiet Ending

The beginner asks, “Which parts of RNA are removed?”

The developing biologist asks, “Why was this exon included?”

The advanced learner asks, “Which complete isoform was actually produced?”

And the professional asks: which processing pathway, measurement chain and functional evidence justify treating this RNA isoform as a real biological state rather than a plausible reconstruction?