The Food and Drug Administration defines protein in 21 CFR 600.3(h)(6) as any alpha amino acid polymer with a specific, defined sequence greater than 40 amino acids in size. This definition became operative after the agency removed the parenthetical exception for polypeptides from the Public Health Service Act’s definition of biological product (FDA MAPP 600.3).
That statutory revision determines when a polypeptide chain crosses into protein territory for U.S. licensure.
A polypeptide chain is a continuous, unbranched linear polymer of amino acids joined by covalent peptide bonds. Each chain extends directionally from a free amino group at the N-terminus to a free carboxyl group at the C-terminus (GenScript). The linear sequence constitutes the primary structure that encodes all higher-order folding.
Figure 1: Linear polypeptide chain showing N-to-C directionality, peptide bonds, and the >40 amino acid FDA threshold separating polypeptides from proteins.
Structural Definition and Directionality
The monomeric unit of the chain is the amino acid. Each carries an amine group, a carboxyl group, and a unique side chain designated as the R group (GenScript). Twenty standard amino acids serve as the building blocks of proteins. Their specific arrangement within a chain determines the chemical properties of the resulting polymer.
Peptide bonds form through dehydration synthesis. This condensation reaction releases water as the carboxyl group of one amino acid reacts with the amino group of the next (JPT). The bond repeats along the backbone to produce a directional chain read strictly from N-terminus to C-terminus (GenScript).
Directionality is a non-negotiable attribute of identity. Regulatory specifications for peptide active pharmaceutical ingredients must designate the N-terminal and C-terminal residues to distinguish the intended sequence from isomers or degradation products.
Chain Length Classification Thresholds
Polypeptide chains occupy a middle position in a length-based hierarchy that carries regulatory weight.
| Classification | Residue Count (General) | Regulatory/Functional Note | | :--- | :--- | :--- | | Oligopeptide | <20 | Typically synthetic; distinct from biological products | | Polypeptide | ~50+ | Linear polymer; term denotes unfolded state (Britannica) | | Protein (FDA) | >40 | Statutory threshold in 21 CFR 600.3(h)(6) (FDA MAPP 600.3) | | Protein (Functional) | >100 | Folded tertiary structure required (Britannica) |
Biochemical references describe the boundary between polypeptide and protein as imprecise. Britannica places the conventional polypeptide threshold at approximately 50 residues while noting that most common proteins contain more than 100 amino acids (Britannica). Bachem clarifies the functional distinction: all proteins are polypeptides, but not all polypeptides are proteins, because proteins require folding into a specific three-dimensional structure (Bachem).
The 40-residue threshold in 21 CFR 600.3(h)(6) creates a concrete administrative consequence. A 48-residue sequence prepared as an active pharmaceutical ingredient falls within the FDA definition of protein and is regulated as a biological product. A 39-residue sequence does not meet this statutory threshold. The distinction alters the regulatory pathway and manufacturing oversight obligations.
Biosynthetic Formation via Translation
Cells synthesize polypeptide chains through translation on ribosomes. Messenger RNA transcribed from DNA carries the sequence instructions to the ribosome. Transfer RNA molecules deliver individual amino acids, matching each to its corresponding mRNA codon through complementary anticodon pairing (JPT).
Each elongation step is a condensation reaction catalyzed by the ribosomal peptidyl transferase center. The ribosome continues assembly until encountering a stop codon, releasing the completed chain. A freshly synthesized polypeptide chain is linear and non-functional. It must fold, undergo post-translational modification, and clear cellular quality control before functioning as a protein (Superpower).
Structural Hierarchy and Regulatory Identity
Polypeptide chain structure organizes across four levels. Primary structure is the linear amino acid sequence. Secondary structure refers to local folding patterns such as alpha-helices and beta-sheets stabilized by hydrogen bonds. Tertiary structure is the complete three-dimensional arrangement of a single chain. Quaternary structure arises when multiple chains assemble into a functional complex (GenScript).
Hemoglobin illustrates this hierarchy in a regulatory context. A hemoglobin molecule contains four polypeptide chains, each consisting of more than 140 amino acids, arranged as a tetramer with attached heme groups. The individual chains are polypeptides; the assembled, oxygen-carrying tetramer is the functional protein (Britannica).
Christian Anfinsen demonstrated in the early 1960s that denatured ribonuclease could spontaneously refold into its native structure when denaturing agents were removed. This established that amino acid sequence alone contains the complete structural information for folding (Superpower). For manufacturers, this principle means sequence verification serves as the primary proof of identity even before functional assays confirm tertiary structure.
Cross-Border Threshold Divergence
Structural definition carries direct consequences for compliance documentation. The FDA ties statutory classification to the numeric size threshold in 21 CFR 600.3(h)(6). Any alpha amino acid polymer greater than 40 amino acids is a protein regardless of folding state (FDA MAPP 600.3).
The European Medicines Agency applies a separate framework in its synthetic peptide guideline. The EMA states that polypeptide fragments undergoing limited further modifications under GMP, such as cyclisation or conjugation to non-peptide structural moieties, are generally not acceptable as starting materials (EMA). A chain that crosses into polypeptide territory therefore shifts the starting material designation in European submissions.
A sponsor filing in both jurisdictions must satisfy both definitions with the same characterization data. The U.S. framework demands proof of sequence length relative to the 40-residue statutory threshold. The European framework demands proof that the chain qualifies as an appropriate starting material under GMP. These parallel requirements establish the documentation standards for identity, purity, and potency that follow a polypeptide compound through development and review.
Evidence review due: 180 days.

