Amino acids and codons

A quick reference for DNA, RNA, codons and amino acid notation used in clinical genetics.

DNA to RNA to protein

Protein-coding DNA is transcribed into RNA. The RNA sequence is read in groups of three nucleotides, called codons. Each codon specifies an amino acid or a translation stop.

  Sequence
DNA coding strand, 5′ → 3′ ATG GAA TGG
mRNA, 5′ → 3′ AUG GAA UGG
Amino acids Met Glu Trp
One-letter code M E W

DNA → RNA → codons → amino acids → protein

For a DNA coding sequence, the corresponding mRNA sequence is the same except that RNA uses uracil (U) instead of thymine (T).

DNA and RNA bases

DNA RNA
A, adenine A, adenine
C, cytosine C, cytosine
G, guanine G, guanine
T, thymine U, uracil

The example above refers to the DNA coding strand. The DNA template strand is complementary to the resulting RNA sequence.

Amino acid reference

Clinical genetics commonly uses both three-letter and one-letter amino acid abbreviations. Select an amino acid to inspect its three-dimensional chemical structure.

Amino acid 3-letter 1-letter DNA codons
Alanine Ala A GCT, GCC, GCA, GCG
Arginine Arg R CGT, CGC, CGA, CGG, AGA, AGG
Asparagine Asn N AAT, AAC
Aspartic acid Asp D GAT, GAC
Cysteine Cys C TGT, TGC
Glutamine Gln Q CAA, CAG
Glutamic acid Glu E GAA, GAG
Glycine Gly G GGT, GGC, GGA, GGG
Histidine His H CAT, CAC
Isoleucine Ile I ATT, ATC, ATA
Leucine Leu L TTA, TTG, CTT, CTC, CTA, CTG
Lysine Lys K AAA, AAG
Methionine Met M ATG
Phenylalanine Phe F TTT, TTC
Proline Pro P CCT, CCC, CCA, CCG
Serine Ser S TCT, TCC, TCA, TCG, AGT, AGC
Threonine Thr T ACT, ACC, ACA, ACG
Tryptophan Trp W TGG
Tyrosine Tyr Y TAT, TAC
Valine Val V GTT, GTC, GTA, GTG

Genetic code

The standard genetic code is shown below. Select RNA or DNA to switch between mRNA codons and the corresponding coding-strand DNA codons.

Showing mRNA codons. RNA uses uracil (U).

First base Second base U Second base C Second base A Second base G
U UUU Phe (F)
UUC Phe (F)
UUA Leu (L)
UUG Leu (L)
UCU Ser (S)
UCC Ser (S)
UCA Ser (S)
UCG Ser (S)
UAU Tyr (Y)
UAC Tyr (Y)
UAA Stop
UAG Stop
UGU Cys (C)
UGC Cys (C)
UGA Stop
UGG Trp (W)
C CUU Leu (L)
CUC Leu (L)
CUA Leu (L)
CUG Leu (L)
CCU Pro (P)
CCC Pro (P)
CCA Pro (P)
CCG Pro (P)
CAU His (H)
CAC His (H)
CAA Gln (Q)
CAG Gln (Q)
CGU Arg (R)
CGC Arg (R)
CGA Arg (R)
CGG Arg (R)
A AUU Ile (I)
AUC Ile (I)
AUA Ile (I)
AUG Met (M)
ACU Thr (T)
ACC Thr (T)
ACA Thr (T)
ACG Thr (T)
AAU Asn (N)
AAC Asn (N)
AAA Lys (K)
AAG Lys (K)
AGU Ser (S)
AGC Ser (S)
AGA Arg (R)
AGG Arg (R)
G GUU Val (V)
GUC Val (V)
GUA Val (V)
GUG Val (V)
GCU Ala (A)
GCC Ala (A)
GCA Ala (A)
GCG Ala (A)
GAU Asp (D)
GAC Asp (D)
GAA Glu (E)
GAG Glu (E)
GGU Gly (G)
GGC Gly (G)
GGA Gly (G)
GGG Gly (G)

Start and stop codons

Function RNA codon DNA coding sequence
Start, methionine AUG ATG
Stop UAA TAA
Stop UAG TAG
Stop UGA TGA

AUG commonly initiates translation and also encodes methionine. Stop codons terminate translation and do not encode an amino acid.

Clinical genetics

A nucleotide change can alter the codon and therefore the resulting protein sequence.

  Reference Alternate
DNA codon TGG TAG
RNA codon UGG UAG
Protein Trp Stop

TGG → TAG

Trp → Ter

A single nucleotide change can therefore leave the amino acid unchanged, change one amino acid to another, introduce a stop codon or remove an existing start or stop signal.

Variant consequences

A genomic variant can have different predicted consequences depending on the transcript and genomic feature it overlaps. Ensembl assigns Sequence Ontology (SO) consequence terms to each allele and transcript combination, so the same allele can have different consequences in different transcripts.

The table below is reproduced and reformatted from the Ensembl Variation calculated variant consequences reference (Ensembl release 116, June 2026), retaining Ensembl’s display colours, Sequence Ontology terms, descriptions, accessions, severity order and IMPACT labels.

IMPACT is not pathogenicity. Ensembl’s HIGH, MODERATE, LOW and MODIFIER labels describe predicted molecular consequence and are separate from clinical variant classification. Ensembl also notes that its severity ordering is necessarily subjective.

consequences.svg

  Consequence Description SO accession IMPACT
Transcript ablation A feature ablation whereby the deleted region includes a transcript feature SO:0001893 HIGH
Splice acceptor variant A splice variant that changes the 2 base region at the 3’ end of an intron SO:0001574 HIGH
Splice donor variant A splice variant that changes the 2 base region at the 5’ end of an intron SO:0001575 HIGH
Stop gained A sequence variant whereby at least one base of a codon is changed, resulting in a premature stop codon, leading to a shortened transcript SO:0001587 HIGH
Frameshift variant A sequence variant which causes a disruption of the translational reading frame, because the number of nucleotides inserted or deleted is not a multiple of three SO:0001589 HIGH
Stop lost A sequence variant where at least one base of the terminator codon (stop) is changed, resulting in an elongated transcript SO:0001578 HIGH
Start lost A codon variant that changes at least one base of the canonical start codon SO:0002012 HIGH
Transcript amplification A feature amplification of a region containing a transcript SO:0001889 HIGH
Feature elongation A sequence variant that causes the extension of a genomic feature, with regard to the reference sequence SO:0001907 HIGH
Feature truncation A sequence variant that causes the reduction of a genomic feature, with regard to the reference sequence SO:0001906 HIGH
Inframe insertion An inframe non synonymous variant that inserts bases into in the coding sequence SO:0001821 MODERATE
Inframe deletion An inframe non synonymous variant that deletes bases from the coding sequence SO:0001822 MODERATE
Missense variant A sequence variant, that changes one or more bases, resulting in a different amino acid sequence but where the length is preserved SO:0001583 MODERATE
Protein altering variant A sequence_variant which is predicted to change the protein encoded in the coding sequence SO:0001818 MODERATE
Splice donor 5th base variant A sequence variant that causes a change at the 5th base pair after the start of the intron in the orientation of the transcript SO:0001787 LOW
Splice region variant A sequence variant in which a change has occurred within the region of the splice site, either within 1-3 bases of the exon or 3-8 bases of the intron SO:0001630 LOW
Splice donor region variant A sequence variant that falls in the region between the 3rd and 6th base after splice junction (5’ end of intron) SO:0002170 LOW
Splice polypyrimidine tract variant A sequence variant that falls in the polypyrimidine tract at 3’ end of intron between 17 and 3 bases from the end (acceptor -3 to acceptor -17) SO:0002169 LOW
Incomplete terminal codon variant A sequence variant where at least one base of the final codon of an incompletely annotated transcript is changed SO:0001626 LOW
Start retained variant A sequence variant where at least one base in the start codon is changed, but the start remains SO:0002019 LOW
Stop retained variant A sequence variant where at least one base in the terminator codon is changed, but the terminator remains SO:0001567 LOW
Synonymous variant A sequence variant where there is no resulting change to the encoded amino acid SO:0001819 LOW
Coding sequence variant A sequence variant that changes the coding sequence SO:0001580 MODIFIER
Mature miRNA variant A transcript variant located with the sequence of the mature miRNA SO:0001620 MODIFIER
5 prime UTR variant A UTR variant of the 5’ UTR SO:0001623 MODIFIER
3 prime UTR variant A UTR variant of the 3’ UTR SO:0001624 MODIFIER
Non coding transcript exon variant A sequence variant that changes non-coding exon sequence in a non-coding transcript SO:0001792 MODIFIER
Intron variant A transcript variant occurring within an intron SO:0001627 MODIFIER
NMD transcript variant A variant in a transcript that is the target of NMD SO:0001621 MODIFIER
Non coding transcript variant A transcript variant of a non coding RNA gene SO:0001619 MODIFIER
Coding transcript variant A transcript variant of a protein coding gene SO:0001968 MODIFIER
Upstream gene variant A sequence variant located 5’ of a gene SO:0001631 MODIFIER
Downstream gene variant A sequence variant located 3’ of a gene SO:0001632 MODIFIER
TFBS ablation A feature ablation whereby the deleted region includes a transcription factor binding site SO:0001895 MODIFIER
TFBS amplification A feature amplification of a region containing a transcription factor binding site SO:0001892 MODIFIER
TF binding site variant A sequence variant located within a transcription factor binding site SO:0001782 MODIFIER
Regulatory region ablation A feature ablation whereby the deleted region includes a regulatory region SO:0001894 MODIFIER
Regulatory region amplification A feature amplification of a region containing a regulatory region SO:0001891 MODIFIER
Regulatory region variant A sequence variant located within a regulatory region SO:0001566 MODIFIER
Intergenic variant A sequence variant located in the intergenic region, between genes SO:0001628 MODIFIER
Sequence variant A sequence_variant is a non exact copy of a sequence_feature or genome exhibiting one or more sequence_alteration SO:0001060 MODIFIER

Ensembl IMPACT labels

IMPACT Ensembl description
HIGH The variant is assumed to have high (disruptive) impact in the protein, probably causing protein truncation, loss of function or triggering nonsense mediated decay.
MODERATE A non-disruptive variant that might change protein effectiveness.
LOW Assumed to be mostly harmless or unlikely to change protein behaviour.
MODIFIER Usually non-coding variants or variants affecting non-coding genes, where predictions are difficult or there is no evidence of impact.

HGVS notation

Protein changes are commonly represented using HGVS nomenclature.

For example:

p.Trp24Ter

means that tryptophan (Trp, W) at amino acid position 24 is replaced by a translation termination signal.

The corresponding coding DNA change is described separately using HGVS coding-sequence notation, for example:

c.71G>A

The exact relationship between genomic DNA, coding DNA and protein notation depends on the reference sequence, transcript and reading frame.

Interpretation

A codon or amino acid consequence describes what a sequence variant changes. It does not by itself establish whether the variant is pathogenic or whether it explains a patient’s phenotype.

Clinical variant interpretation additionally considers evidence such as population frequency, inheritance, phenotype, functional studies, segregation, gene-disease relationships and other relevant evidence.

Mitochondrial genetic code

Human mitochondrial DNA uses a genetic code that differs at several codons from the standard genetic code shown above. Mitochondrial variants should therefore be interpreted using the appropriate mitochondrial genetic code.

References