A quick reference for DNA, RNA, codons and amino acid notation used in clinical genetics.
DNA to RNA to protein
Protein-coding DNA is transcribed into RNA. The RNA sequence is read in groups of three nucleotides, called codons. Each codon specifies an amino acid or a translation stop.
| Sequence | |
|---|---|
| DNA coding strand, 5′ → 3′ | ATG GAA TGG |
| mRNA, 5′ → 3′ | AUG GAA UGG |
| Amino acids | Met Glu Trp |
| One-letter code | M E W |
DNA → RNA → codons → amino acids → protein
For a DNA coding sequence, the corresponding mRNA sequence is the same except that RNA uses uracil (U) instead of thymine (T).
DNA and RNA bases
| DNA | RNA |
|---|---|
| A, adenine | A, adenine |
| C, cytosine | C, cytosine |
| G, guanine | G, guanine |
| T, thymine | U, uracil |
The example above refers to the DNA coding strand. The DNA template strand is complementary to the resulting RNA sequence.
Amino acid reference
Clinical genetics commonly uses both three-letter and one-letter amino acid abbreviations. Select an amino acid to inspect its three-dimensional chemical structure.
| Amino acid | 3-letter | 1-letter | DNA codons |
|---|---|---|---|
| Alanine | Ala | A | GCT, GCC, GCA, GCG |
| Arginine | Arg | R | CGT, CGC, CGA, CGG, AGA, AGG |
| Asparagine | Asn | N | AAT, AAC |
| Aspartic acid | Asp | D | GAT, GAC |
| Cysteine | Cys | C | TGT, TGC |
| Glutamine | Gln | Q | CAA, CAG |
| Glutamic acid | Glu | E | GAA, GAG |
| Glycine | Gly | G | GGT, GGC, GGA, GGG |
| Histidine | His | H | CAT, CAC |
| Isoleucine | Ile | I | ATT, ATC, ATA |
| Leucine | Leu | L | TTA, TTG, CTT, CTC, CTA, CTG |
| Lysine | Lys | K | AAA, AAG |
| Methionine | Met | M | ATG |
| Phenylalanine | Phe | F | TTT, TTC |
| Proline | Pro | P | CCT, CCC, CCA, CCG |
| Serine | Ser | S | TCT, TCC, TCA, TCG, AGT, AGC |
| Threonine | Thr | T | ACT, ACC, ACA, ACG |
| Tryptophan | Trp | W | TGG |
| Tyrosine | Tyr | Y | TAT, TAC |
| Valine | Val | V | GTT, GTC, GTA, GTG |
Genetic code
The standard genetic code is shown below. Select RNA or DNA to switch between mRNA codons and the corresponding coding-strand DNA codons.
Showing mRNA codons. RNA uses uracil (U).
| First base | Second base U | Second base C | Second base A | Second base G |
|---|---|---|---|---|
| U | UUU Phe (F) UUC Phe (F) UUA Leu (L) UUG Leu (L) |
UCU Ser (S) UCC Ser (S) UCA Ser (S) UCG Ser (S) |
UAU Tyr (Y) UAC Tyr (Y) UAA Stop UAG Stop |
UGU Cys (C) UGC Cys (C) UGA Stop UGG Trp (W) |
| C | CUU Leu (L) CUC Leu (L) CUA Leu (L) CUG Leu (L) |
CCU Pro (P) CCC Pro (P) CCA Pro (P) CCG Pro (P) |
CAU His (H) CAC His (H) CAA Gln (Q) CAG Gln (Q) |
CGU Arg (R) CGC Arg (R) CGA Arg (R) CGG Arg (R) |
| A | AUU Ile (I) AUC Ile (I) AUA Ile (I) AUG Met (M) |
ACU Thr (T) ACC Thr (T) ACA Thr (T) ACG Thr (T) |
AAU Asn (N) AAC Asn (N) AAA Lys (K) AAG Lys (K) |
AGU Ser (S) AGC Ser (S) AGA Arg (R) AGG Arg (R) |
| G | GUU Val (V) GUC Val (V) GUA Val (V) GUG Val (V) |
GCU Ala (A) GCC Ala (A) GCA Ala (A) GCG Ala (A) |
GAU Asp (D) GAC Asp (D) GAA Glu (E) GAG Glu (E) |
GGU Gly (G) GGC Gly (G) GGA Gly (G) GGG Gly (G) |
Start and stop codons
| Function | RNA codon | DNA coding sequence |
|---|---|---|
| Start, methionine | AUG | ATG |
| Stop | UAA | TAA |
| Stop | UAG | TAG |
| Stop | UGA | TGA |
AUG commonly initiates translation and also encodes methionine. Stop codons terminate translation and do not encode an amino acid.
Clinical genetics
A nucleotide change can alter the codon and therefore the resulting protein sequence.
| Reference | Alternate | |
|---|---|---|
| DNA codon | TGG |
TAG |
| RNA codon | UGG |
UAG |
| Protein | Trp | Stop |
TGG → TAG
Trp → Ter
A single nucleotide change can therefore leave the amino acid unchanged, change one amino acid to another, introduce a stop codon or remove an existing start or stop signal.
Variant consequences
A genomic variant can have different predicted consequences depending on the transcript and genomic feature it overlaps. Ensembl assigns Sequence Ontology (SO) consequence terms to each allele and transcript combination, so the same allele can have different consequences in different transcripts.
The table below is reproduced and reformatted from the Ensembl Variation calculated variant consequences reference (Ensembl release 116, June 2026), retaining Ensembl’s display colours, Sequence Ontology terms, descriptions, accessions, severity order and IMPACT labels.
IMPACT is not pathogenicity. Ensembl’s HIGH, MODERATE, LOW and MODIFIER labels describe predicted molecular consequence and are separate from clinical variant classification. Ensembl also notes that its severity ordering is necessarily subjective.
| Consequence | Description | SO accession | IMPACT | |
|---|---|---|---|---|
| Transcript ablation | A feature ablation whereby the deleted region includes a transcript feature | SO:0001893 | HIGH | |
| Splice acceptor variant | A splice variant that changes the 2 base region at the 3’ end of an intron | SO:0001574 | HIGH | |
| Splice donor variant | A splice variant that changes the 2 base region at the 5’ end of an intron | SO:0001575 | HIGH | |
| Stop gained | A sequence variant whereby at least one base of a codon is changed, resulting in a premature stop codon, leading to a shortened transcript | SO:0001587 | HIGH | |
| Frameshift variant | A sequence variant which causes a disruption of the translational reading frame, because the number of nucleotides inserted or deleted is not a multiple of three | SO:0001589 | HIGH | |
| Stop lost | A sequence variant where at least one base of the terminator codon (stop) is changed, resulting in an elongated transcript | SO:0001578 | HIGH | |
| Start lost | A codon variant that changes at least one base of the canonical start codon | SO:0002012 | HIGH | |
| Transcript amplification | A feature amplification of a region containing a transcript | SO:0001889 | HIGH | |
| Feature elongation | A sequence variant that causes the extension of a genomic feature, with regard to the reference sequence | SO:0001907 | HIGH | |
| Feature truncation | A sequence variant that causes the reduction of a genomic feature, with regard to the reference sequence | SO:0001906 | HIGH | |
| Inframe insertion | An inframe non synonymous variant that inserts bases into in the coding sequence | SO:0001821 | MODERATE | |
| Inframe deletion | An inframe non synonymous variant that deletes bases from the coding sequence | SO:0001822 | MODERATE | |
| Missense variant | A sequence variant, that changes one or more bases, resulting in a different amino acid sequence but where the length is preserved | SO:0001583 | MODERATE | |
| Protein altering variant | A sequence_variant which is predicted to change the protein encoded in the coding sequence | SO:0001818 | MODERATE | |
| Splice donor 5th base variant | A sequence variant that causes a change at the 5th base pair after the start of the intron in the orientation of the transcript | SO:0001787 | LOW | |
| Splice region variant | A sequence variant in which a change has occurred within the region of the splice site, either within 1-3 bases of the exon or 3-8 bases of the intron | SO:0001630 | LOW | |
| Splice donor region variant | A sequence variant that falls in the region between the 3rd and 6th base after splice junction (5’ end of intron) | SO:0002170 | LOW | |
| Splice polypyrimidine tract variant | A sequence variant that falls in the polypyrimidine tract at 3’ end of intron between 17 and 3 bases from the end (acceptor -3 to acceptor -17) | SO:0002169 | LOW | |
| Incomplete terminal codon variant | A sequence variant where at least one base of the final codon of an incompletely annotated transcript is changed | SO:0001626 | LOW | |
| Start retained variant | A sequence variant where at least one base in the start codon is changed, but the start remains | SO:0002019 | LOW | |
| Stop retained variant | A sequence variant where at least one base in the terminator codon is changed, but the terminator remains | SO:0001567 | LOW | |
| Synonymous variant | A sequence variant where there is no resulting change to the encoded amino acid | SO:0001819 | LOW | |
| Coding sequence variant | A sequence variant that changes the coding sequence | SO:0001580 | MODIFIER | |
| Mature miRNA variant | A transcript variant located with the sequence of the mature miRNA | SO:0001620 | MODIFIER | |
| 5 prime UTR variant | A UTR variant of the 5’ UTR | SO:0001623 | MODIFIER | |
| 3 prime UTR variant | A UTR variant of the 3’ UTR | SO:0001624 | MODIFIER | |
| Non coding transcript exon variant | A sequence variant that changes non-coding exon sequence in a non-coding transcript | SO:0001792 | MODIFIER | |
| Intron variant | A transcript variant occurring within an intron | SO:0001627 | MODIFIER | |
| NMD transcript variant | A variant in a transcript that is the target of NMD | SO:0001621 | MODIFIER | |
| Non coding transcript variant | A transcript variant of a non coding RNA gene | SO:0001619 | MODIFIER | |
| Coding transcript variant | A transcript variant of a protein coding gene | SO:0001968 | MODIFIER | |
| Upstream gene variant | A sequence variant located 5’ of a gene | SO:0001631 | MODIFIER | |
| Downstream gene variant | A sequence variant located 3’ of a gene | SO:0001632 | MODIFIER | |
| TFBS ablation | A feature ablation whereby the deleted region includes a transcription factor binding site | SO:0001895 | MODIFIER | |
| TFBS amplification | A feature amplification of a region containing a transcription factor binding site | SO:0001892 | MODIFIER | |
| TF binding site variant | A sequence variant located within a transcription factor binding site | SO:0001782 | MODIFIER | |
| Regulatory region ablation | A feature ablation whereby the deleted region includes a regulatory region | SO:0001894 | MODIFIER | |
| Regulatory region amplification | A feature amplification of a region containing a regulatory region | SO:0001891 | MODIFIER | |
| Regulatory region variant | A sequence variant located within a regulatory region | SO:0001566 | MODIFIER | |
| Intergenic variant | A sequence variant located in the intergenic region, between genes | SO:0001628 | MODIFIER | |
| Sequence variant | A sequence_variant is a non exact copy of a sequence_feature or genome exhibiting one or more sequence_alteration | SO:0001060 | MODIFIER |
Ensembl IMPACT labels
| IMPACT | Ensembl description |
|---|---|
| HIGH | The variant is assumed to have high (disruptive) impact in the protein, probably causing protein truncation, loss of function or triggering nonsense mediated decay. |
| MODERATE | A non-disruptive variant that might change protein effectiveness. |
| LOW | Assumed to be mostly harmless or unlikely to change protein behaviour. |
| MODIFIER | Usually non-coding variants or variants affecting non-coding genes, where predictions are difficult or there is no evidence of impact. |
HGVS notation
Protein changes are commonly represented using HGVS nomenclature.
For example:
p.Trp24Ter
means that tryptophan (Trp, W) at amino acid position 24 is replaced by a translation termination signal.
The corresponding coding DNA change is described separately using HGVS coding-sequence notation, for example:
c.71G>A
The exact relationship between genomic DNA, coding DNA and protein notation depends on the reference sequence, transcript and reading frame.
Interpretation
A codon or amino acid consequence describes what a sequence variant changes. It does not by itself establish whether the variant is pathogenic or whether it explains a patient’s phenotype.
Clinical variant interpretation additionally considers evidence such as population frequency, inheritance, phenotype, functional studies, segregation, gene-disease relationships and other relevant evidence.
Mitochondrial genetic code
Human mitochondrial DNA uses a genetic code that differs at several codons from the standard genetic code shown above. Mitochondrial variants should therefore be interpreted using the appropriate mitochondrial genetic code.
References
- National Center for Biotechnology Information. The genetic codes
- Human Genome Variation Society. HGVS nomenclature
- Ensembl. Calculated variant consequences. Ensembl release 116, June 2026.
- Sequence Ontology. Sequence Ontology
- National Library of Medicine. PubChem