Toolyard

How to Find Open Reading Frames (ORFs)

What an ORF is

An open reading frame is a run of codons that starts with a start codon and ends at the first stop codon (TAA, TAG or TGA) in the same frame. It is the part of a sequence that could be translated into protein. A real gene is an ORF, but not every ORF is a gene: short ones appear by chance everywhere.

Six reading frames

A sequence can be read in three frames on the forward strand, starting at base 1, 2 or 3, and three more on the reverse complement. A gene on the reverse strand is invisible in the forward frames, so a search should cover all six unless you know the orientation.

Choosing the settings

  • Start codons: ATG for eukaryotes and most searches. Bacteria also start more than one gene in ten with GTG and a few with TTG, so include them for bacterial DNA. "Any codon" finds stop-to-stop frames, useful for partial sequences with no start.
  • Minimum length: stop codons appear by chance about every 21 codons in random DNA, so a 100-codon cutoff removes most noise. Lower it for short peptides or viral genes; raise it for genome-scale scans.
  • Genetic code: mitochondrial and some bacterial sequences use different stop and start codons, so pick the right table.

Picking the real one

The longest ORF is usually the gene, but check it: a real coding sequence has a codon usage that fits the organism, a plausible protein with no long runs of one amino acid, and for eukaryotic mRNA a Kozak context (GCCACCATGG) around the start. In genomic DNA, introns break genes into pieces, so ORF finding works on cDNA and bacterial genomes, not on spliced genes.

Find and translate

  1. Paste the sequence into the ORF Finder, choose the frames, start codons and minimum length.
  2. Each ORF is listed with its frame, start, end, length and protein sequence.
  3. Translate a chosen region with Translate, or see all six frames at once in the Translation Map.

Tools in this guide

More guides