Identity and similarity
Both are read off an alignment, column by column. Percent identity is the share of columns where the two sequences have the same residue. Percent similarity also counts columns where the residues differ but are chemically alike, such as leucine against isoleucine or aspartate against glutamate. Similarity is only meaningful for proteins; for DNA, identity is the number to use.
The denominator matters
Two sequences with 90 identical columns can be reported as 90% identical over a 100-column alignment, 75% over the 120-residue longer sequence, or 95% over the 95 columns that are not gaps. Tools differ:
- BLAST divides by the alignment length, including gaps.
- EMBOSS Needle does the same but reports gaps separately.
- Some tools divide by the shorter sequence, which inflates the number for a short fragment that matches part of a long protein.
When you quote a figure, say what it was divided by and whether the alignment was global or local.
Which residues are similar
Similarity groups usually follow the positive scores in the BLOSUM62 matrix. Common groups: I, L, V and M; F, Y and W; K, R and H; D and E; N and Q; S and T; A and G. Most tools let you see the rule they use; the Identity and Similarity tool here marks identical columns and similar ones separately.
What the numbers mean
Two proteins over 30% identical across their full length are almost always homologous and share a fold. Between 20 and 30% is the twilight zone, where an alignment can be real or chance, and the statistical E-value of a database search is the better guide. DNA diverges faster at silent positions, so two genes coding near-identical proteins can be only 70% identical at the DNA level.
Calculate it
- Align the sequences with Pairwise Align Protein or Pairwise Align DNA, or paste an alignment you already have.
- Paste the aligned sequences into Identity and Similarity. It counts identical and similar columns and reports both percentages with the alignment length.