The formats
- FASTA is a title line starting with >, followed by the sequence. Almost every program reads it.
- GenBank records, from NCBI, have a header, a FEATURES table describing genes and other parts, and the sequence after ORIGIN, with // at the end.
- EMBL records, from ENA, hold the same information with a two-letter code at the start of every line and the sequence after SQ.
Convert the whole sequence
- Open GenBank to FASTA, or EMBL to FASTA for EMBL files.
- Paste the record, or several one after another.
- Copy or download the FASTA. The accession and description become the title line.
Take out one gene
A genome record can hold thousands of features. The GenBank Feature Extractor writes the DNA of every CDS, gene, rRNA or other feature type you choose as its own FASTA entry. It follows join() for genes split into exons and complement() for genes on the other strand, so the sequence comes out in the right order and orientation.
Get the protein
For coding sequences, the protein is usually written in the record already, in the /translation qualifier. The GenBank Translation Extractor lists those proteins as FASTA without translating anything again.