What an ORF is
An open reading frame is a run of codons that starts with a start codon and ends at the first stop codon (TAA, TAG or TGA) in the same frame. It is the part of a sequence that could be translated into protein. A real gene is an ORF, but not every ORF is a gene: short ones appear by chance everywhere.
Six reading frames
A sequence can be read in three frames on the forward strand, starting at base 1, 2 or 3, and three more on the reverse complement. A gene on the reverse strand is invisible in the forward frames, so a search should cover all six unless you know the orientation.
Choosing the settings
- Start codons: ATG for eukaryotes and most searches. Bacteria also start more than one gene in ten with GTG and a few with TTG, so include them for bacterial DNA. "Any codon" finds stop-to-stop frames, useful for partial sequences with no start.
- Minimum length: stop codons appear by chance about every 21 codons in random DNA, so a 100-codon cutoff removes most noise. Lower it for short peptides or viral genes; raise it for genome-scale scans.
- Genetic code: mitochondrial and some bacterial sequences use different stop and start codons, so pick the right table.
Picking the real one
The longest ORF is usually the gene, but check it: a real coding sequence has a codon usage that fits the organism, a plausible protein with no long runs of one amino acid, and for eukaryotic mRNA a Kozak context (GCCACCATGG) around the start. In genomic DNA, introns break genes into pieces, so ORF finding works on cDNA and bacterial genomes, not on spliced genes.
Find and translate
- Paste the sequence into the ORF Finder, choose the frames, start codons and minimum length.
- Each ORF is listed with its frame, start, end, length and protein sequence.
- Translate a chosen region with Translate, or see all six frames at once in the Translation Map.