Peptide synthesis fails at the bench when sequence codes are ambiguous or when a chemist uses a free amino acid molecular weight instead of a residue mass. The 20 standard amino acids that constitute the protein alphabet are universally identified by a trivial name, a three-letter code, and a one-letter code, yet molecular weights range from 75.07 Da for glycine to 204.23 Da for tryptophan. Accurate design requires distinguishing between these values and understanding how side-chain ionization affects stability in solution.

The International Union of Pure and Applied Chemistry (IUPAC) and the International Union of Biochemistry and Molecular Biology (IUBMB) standardized these symbols to eliminate regional naming inconsistencies. As documented in IUPAC Table 1, each residue has a specific systematic name, such as "2-Aminopropanoic acid" for alanine, which provides unambiguous chemical identification. In laboratory practice, the one-letter code dominates sequence databases and synthesis orders, while the three-letter code remains standard in crystallography files and journal figures.

Misidentifying a residue, such as confusing arginine (R) with alanine (A), is one of the most consequential errors in peptide work because it alters the charge and hydrophobicity of the final construct. The following reference consolidates verified data from major biochemical suppliers to provide a single source for the 20 proteinogenic amino acids, excluding selenocysteine and pyrrolysine, which are genetically encoded in specific contexts but not part of the standard set used in most synthesis.

| Name | Code (1/3) | IUPAC Name | Formula | MW (Da) | pKa (Side Chain) | pI | | :--- | :--- | :--- | :--- | :--- | :--- | :--- | | Alanine | A / Ala | 2-Aminopropanoic acid | C3H7NO2 | 89.09 | – | 6.00 | | Arginine | R / Arg | 2-Amino-5-guanidinopentanoic acid | C6H14N4O2 | 174.20 | 12.48 | 10.76 | | Asparagine | N / Asn | 2-Amino-3-carbamoylpropanoic acid | C4H8N2O3 | 132.12 | – | 5.41 | | Aspartic Acid | D / Asp | 2-Aminobutanedioic acid | C4H7NO4 | 133.10 | 3.65 | 2.77 | | Cysteine | C / Cys | 2-Amino-3-mercaptopropanoic acid | C3H7NO2S | 121.16 | 8.18 | 5.07 | | Glutamine | Q / Gln | 2-Amino-4-carbamoylbutanoic acid | C5H10N2O3 | 146.15 | – | 5.65 | | Glutamic Acid | E / Glu | 2-Aminopentanedioic acid | C5H9NO4 | 147.13 | 4.25 | 3.22 | | Glycine | G / Gly | Aminoethanoic acid | C2H5NO2 | 75.07 | – | 5.97 | | Histidine | H / His | 2-Amino-3-(1H-imidazol-4-yl)propanoic acid | C6H9N3O2 | 155.16 | 6.00 | 7.59 | | Isoleucine | I / Ile | 2-Amino-3-methylbutanoic acid | C6H13NO2 | 131.18 | – | 6.02 | | Leucine | L / Leu | 2-Amino-4-methylpentanoic acid | C6H13NO2 | 131.18 | – | 5.98 | | Lysine | K / Lys | 2-Amino-6-aminohexanoic acid | C6H14N2O2 | 146.19 | 10.53 | 9.74 | | Methionine | M / Met | 2-Amino-4-(methylthio)butanoic acid | C5H11NO2S | 149.21 | – | 5.74 | | Phenylalanine | F / Phe | 2-Amino-3-phenylpropanoic acid | C9H11NO2 | 165.19 | – | 5.48 | | Proline | P / Pro | Pyrroline-2-carboxylic acid | C5H9NO2 | 115.13 | – | 6.30 | | Serine | S / Ser | 2-Amino-3-hydroxypropanoic acid | C3H7NO3 | 105.09 | – | 5.68 | | Threonine | T / Thr | 2-Amino-3-hydroxybutanoic acid | C4H9NO3 | 119.12 | – | 5.60 | | Tryptophan | W / Trp | 2-Amino-3-(1H-indol-3-yl)propanoic acid | C11H12N2O2 | 204.23 | – | 5.89 | | Tyrosine | Y / Tyr | 2-Amino-3-(4-hydroxyphenyl)propanoic acid | C9H11NO3 | 181.19 | 10.07 | 5.66 | | Valine | V / Val | 2-Amino-3-methylbutanoic acid | C5H11NO2 | 117.15 | – | 5.96 |

Data source: Alfa Chemistry and Genscript. Molecular weights are for free amino acids. For peptide mass calculations, subtract 18.02 Da per peptide bond formed.

Structural Impact of Side Chains

The hydrophobicity of a side chain determines how a peptide folds in aqueous solution. Amino acids such as leucine, isoleucine, and valine have branched, aliphatic side chains that drive to the interior of the protein core to minimize contact with water. Conversely, polar residues like serine and threonine contain hydroxyl groups that can form hydrogen bonds with water or other polar residues, often positioning them on the surface of the molecule. Understanding these classifications is essential for predicting the secondary and tertiary structure of a peptide, a process detailed in our guide on peptide structure. The pKa value indicates the pH at which a specific functional group loses a proton, while the isoelectric point (pI) is the pH at which the molecule carries no net electric charge, causing it to precipitate or move differently in electrophoresis. Among the 20 standard amino acids, only seven possess ionizable side chains: arginine, aspartic acid, cysteine, glutamic acid, histidine, lysine, and tyrosine. The others have non-ionizable side chains, meaning their net charge is determined solely by the terminal alpha-amino and alpha-carboxyl groups.

Histidine is unique among these because its imidazole ring has a pKa of approximately 6.00, which is close to physiological pH. This property allows histidine to act as a proton shuttle in enzyme active sites, making it a critical residue in catalytic mechanisms. In contrast, arginine has a side-chain pKa of 12.48, ensuring it remains positively charged at all physiological pH values. This permanent positive charge facilitates interactions with negatively charged phosphate groups in DNA or with other acidic residues in protein structures. These electrostatic interactions are fundamental to the stability of many biologically active peptides, where the precise arrangement of charged residues dictates the molecule's ability to bind to specific receptors or substrates.

Mass Calculation and Genetic Context

The table above lists molecular weights for free amino acids. However, when these residues are incorporated into a peptide chain, each linkage involves a condensation reaction that releases a water molecule. Therefore, the mass of a residue in a chain is 18.02 Da less than the free amino acid. Using the free amino acid weight to calculate the final mass of a synthetic peptide results in a significant error that grows with the length of the sequence. Researchers can use a peptide mass calculator to automate this adjustment, but manual verification remains a critical step in confirming the identity of the final product against high-resolution mass spectrometry data.

The 20 standard amino acids correspond to 61 codons in the genetic code, with three codons (UAA, UAG, and UGA) serving as stop signals. This code is nearly universal across life forms, though minor variations exist in mitochondria and certain protists. Each amino acid, except tryptophan and methionine, is encoded by multiple codons. This degeneracy buffers against mutations; a single nucleotide change may result in a synonymous codon that does not alter the protein sequence. For example, leucine is encoded by six different codons, the highest degeneracy alongside arginine and serine, according to Promega’s genetic code reference. In biological systems, amino acids are categorized as essential or non-essential based on the organism's ability to synthesize them. For adult humans, nine amino acids are essential, meaning they must be obtained from the diet. These include histidine, isoleucine, leucine, lysine, methionine, phenylalanine, threonine, tryptophan, and valine. The remaining 11 are non-essential, as the body can produce them through metabolic pathways. However, some non-essential amino acids, such as cysteine and glutamine, become conditionally essential during periods of high physiological stress, illness, or rapid growth, a distinction relevant for nutritional support in clinical settings but less so for synthetic peptide design. The standard nomenclature excludes selenocysteine and pyrrolysine, which are incorporated into proteins through specialized translation mechanisms that override the standard stop codons UGA and UAG, respectively. While these residues appear in specific enzymes, they are not part of the standard 20-residue set used in routine peptide synthesis and sequence analysis.

The precision of this reference table allows researchers to specify sequences with confidence. When ordering a custom peptide, the supplier relies on the one-letter or three-letter codes to select the correct resin or building block. A single letter error in the sequence input can result in a product that is chemically incorrect, rendering the entire synthesis batch useless. Verification of the sequence against the codes listed here is the final step before submitting a synthesis request, ensuring that the theoretical mass calculated during design matches the experimental mass obtained during quality control.