Peptides are chains of 2 to 100 amino acids, the same building blocks that form proteins. This guide explains how amino acids are built, how peptides are classified by length, how the 20 proteinogenic amino acids are named and written in three-letter and one-letter notation, and how enantiomer…
Peptides are short chains of amino acids , chemically the same building blocks that make up proteins, distinguished from them mainly by length. Most definitions place a peptide at 2 to 100 amino acids; some cap peptides at 50; and proteins are generally taken to exceed 100 amino acids. This article explains how amino acids are built, how peptides are classified and named, how sequences are written in standard notation, and why the handedness of amino acids matters for biological function.
Peptides and proteins are linear chains of amino acids joined by peptide bonds. The practical distinction is length, and the boundary is explicitly arbitrary. Common definitions set the peptide range at 2 to 100 amino acids, an alternative convention caps peptides at 50, and chains longer than 100 amino acids are called proteins. Researchers treat the cutoff as a matter of convenience rather than a molecular dividing line, and context decides which term fits a given chain.
Cells assemble proteins from the 20 proteinogenic amino acids , the standard set found in nature. These building blocks can be arranged in essentially any order, with any repetition frequency, like pearls on a string where the pearls come in 20 colors. Collagen, the most common protein in the body, shows how the arrangement itself carries meaning: its repeating sequence motif is -Xaa-Yaa-Gly-, with glycine at every third position and proline a common occupant of Xaa. That spacing lets collagen chains wind into the tight triple helix that gives skin, cartilage, and tendons their mechanical properties.
Because both identity and order count, sequence space expands quickly. From the 20 proteinogenic amino acids alone, 20 to the fifth power, or 3.2 million, distinct pentapeptides are possible. That figure excludes non-proteinogenic and modified amino acids, so it is a lower bound on the diversity a chemist can actually build; real peptide libraries reach far beyond it.
| Class | Length in amino acids |
|---|---|
| Dipeptide | 2 |
| Tripeptide | 3 |
| Tetrapeptide | 4 |
| Pentapeptide | 5 |
| Oligopeptide | 2 to 20 |
| Peptide | 2 to 100 |
| Protein | More than 100 |
The most common amino acids in nature are α-amino acids : a central α-carbon carries four different substituents, an amino group NH2 , a carboxylic acid group COOH , a variable side chain R , and a hydrogen atom. Because the α-carbon is attached to four different groups, the molecule is chiral, a property with direct consequences for how amino acids behave in biological systems.
The 20 proteinogenic amino acids differ only in their side chains, and the side chain dictates the chemical personality of each residue. The side-chain classes are compact enough to list.
| Side-chain class | Amino acids | Characteristic group |
|---|---|---|
| Simple | Glycine | Hydrogen |
| Acidic | Aspartic acid, glutamic acid | Carboxyl |
| Basic | Arginine, lysine, histidine | Amino or guanidino |
| Polar | Serine, threonine | Hydroxyl |
| Non-polar hydrocarbon | Alanine, phenylalanine, valine | Hydrocarbon chain or ring |
| Sulfur-containing | Cysteine, methionine | Thiol or thioether |
Side chains carry a limited set of functional groups, and recognizing them is the fastest way to predict how an amino acid will react during peptide chemistry. The main groups, with their structures and examples, are listed below.
| Functional group | Structure | Example |
|---|---|---|
| Amino | NH2 | All amino acids |
| Carboxyl | COOH | All amino acids |
| Hydroxyl | OH | Serine, threonine |
| Amide | CONH2 | Asparagine, glutamine |
| Thiol mercapto | SH | Cysteine |
| Guanidino | NH-C =NH -NH2 | Arginine |
The 20 proteinogenic amino acids are not the only ones that exist. Other natural α-amino acids occur free, as metabolic by-products, or as components of peptides and proteins: hydroxyproline Hyp appears in collagen, and ornithine Orn in urine. Norleucine Nle has only been produced by chemical synthesis. Such compounds are often called unnatural amino acids even when they occur in nature; non-proteinogenic is the more accurate label, and unusual is the term one major supplier uses. Beyond the α position, β-amino acids and γ-amino acids carry the amino group on a carbon other than the α-carbon, adding further versatility in peptide and protein design.
Aside from water and fat, proteins account for nearly the entire composition of the body. They earn that share by doing most of the mechanical and chemical work of the organism, with roles that follow directly from the type and number of amino acid building blocks.
| Protein | Occurrence | Function |
|---|---|---|
| Myosin, actin | Muscle | Flexible, contractile movement |
| Collagen | Connective tissue, tendons, skin | Stable shape, stretch resistance |
| Hemoglobin, albumins | Blood | Soluble transport |
| Trypsin | Digestive tract | Biological catalyst enzyme |
| TSH | Pituitary gland | Biological messenger hormone |
| Immunoglobulins | Immune system | Immune defense |
Hormones illustrate the functional spread of this list, because most hormones are peptides of varying length. The range is wide: TRH is a tripeptide of 3 amino acids, LHRH gonadotropin-releasing hormone, GnRH is a decapeptide of 10, calcitonin has 32, and PTH parathormone has 84, nearly protein length. Insulin consists of two peptide chains of 30 and 21 amino acids linked by disulfide bridges.
| Hormone | Chain length | Class |
|---|---|---|
| TRH | 3 | Tripeptide |
| LHRH GnRH | 10 | Decapeptide |
| Calcitonin | 32 | Peptide |
| PTH parathormone | 84 | Nearly a protein |
| Insulin | 30 and 21 | Two chains, disulfide-linked |
Peptide hormones work through a common mechanism. Specialized cells produce the peptide and release it into the bloodstream; the blood carries it to the target organ, where receptor proteins embedded in the cell membrane recognize and bind it. Binding generates a signal that triggers the biological effect. Because recognition is shape-based, the sequence of the hormone and the geometry of the receptor determine whether signaling happens at all.
Peptides in the body are generally produced by enzymatic splitting cleavage of proteins, and the same chemistry can be directed at designed sequences. An instructional protocol for designing matrix metalloproteinase 13-specific protease-sensitive linkers combines peptide-library cleavage screening with mass spectrometry, sequence optimization, and experimental validation of predicted cleavage sites PMID 38813796 . It shows that the proteolytic cleavage which generates natural peptides can also be engineered with precision. Not every active peptide is a fragment of a larger protein; some dipeptides act on their own. Leu-Trp and related dipeptides lower blood pressure, and N-acetyl-Asp-Glu NAAG acts as a neurotransmitter that carries signals between nerve cells.
The 20 proteinogenic amino acids are written in two ways. Three-letter codes are usually the first three letters of the name; one-letter codes are single letters, reserved exclusively for the proteinogenic set. In a peptide chain each unit becomes a residue and takes a residue name: alanine becomes alanyl, glycine becomes glycyl, and so on.
| Amino acid | Three-letter | One-letter | Residue form |
|---|---|---|---|
| Alanine | Ala | A | Alanyl |
| Arginine | Arg | R | Arginyl |
| Asparagine | Asn | N | Asparaginyl |
| Aspartic acid | Asp | D | Aspartyl |
| Cysteine | Cys | C | Cysteyl |
| Glutamine | Gln | Q | Glutaminyl |
| Glutamic acid | Glu | E | Glutamyl |
| Glycine | Gly | G | Glycyl |
| Histidine | His | H | Histidyl |
| Isoleucine | Ile | I | Isoleucyl |
| Leucine | Leu | L | Leucyl |
| Lysine | Lys | K | Lysyl |
| Methionine | Met | M | Methionyl |
| Phenylalanine | Phe | F | Phenylalanyl |
| Proline | Pro | P | Prolyl |
| Serine | Ser | S | Seryl |
| Threonine | Thr | T | Threonyl |
| Tryptophan | Trp | W | Tryptophyl |
| Tyrosine | Tyr | Y | Tyrosinyl |
| Valine | Val | V | Valyl |
A peptide sequence is written as the residue codes in order. The H- prefix marks a free amino terminus and the -OH suffix marks a free carboxyl terminus, so H-Ala-OH is unmodified alanine and H-Ala-Gly-OH is the dipeptide alanyl-glycine. An amidated C-terminus is written -NH2.
Conventions for handedness vary by supplier and journal. The system used by the commercial supplier Bachem is representative: because the L-form is ubiquitous, it is not marked; the less common D-form is written with a lowercase d, as in H-d-Ala-OH; and a racemate is written dl. In one-letter notation a D-amino acid is the lowercase letter, so f means D-Phe. Many non-proteinogenic amino acids have established three-letter codes, including Hyp for L-trans-hydroxyproline, Nle for L-norleucine, and Orn for L-ornithine. Not every abbreviation found in the literature is standard, and when a code is ambiguous the full name is safer, as with l-thiazolidine-4-carboxylic acid, sometimes abbreviated Thz.
The canonical 20 also function as a spectroscopic reference set. A teaching dataset of simple 1D and 2D NMR spectra of peptides containing all encoded amino acids is designed for introductory protein NMR instruction, with exercises in spin-system assignment, conformational analysis, hydrogen exchange, and proline cis-trans isomerism PMID 41812158 . Proline's cis-trans isomerism is a reminder that the codes encode sequence, not conformation; the same peptide can adopt several backbone shapes in solution.
The four substituents around the α-carbon sit at the corners of a tetrahedron, and that arrangement creates two mirror-image forms, stereoisomers called enantiomers . Enantiomers have nearly identical chemical and physical properties, but their biological effects can be very different, because biological targets recognize molecular shape. One enantiomer may bind a receptor and trigger a signaling pathway; its mirror image may fail to bind, or bind without effect, or produce a negative one.
Enantiomers rotate the plane of polarized light in opposite directions, a property called optical activity . It is widespread in nature: all proteinogenic amino acids except glycine are optically active, as are glucose and the building blocks of DNA. Measured rotations for alanine and tryptophan illustrate the pattern.
| Amino acid | Optical rotation |
|---|---|
| L-alanine | +14.3 degrees |
| D-alanine | -13.9 degrees |
| L-tryptophan | -31.8 degrees |
| D-tryptophan | +30.7 degrees |
Small differences between the absolute values of an L/D pair, as between alanine's +14.3 and -13.9, fall within the measurement method's range of accuracy.
The labels L and D, from Latin laevus left and dexter right , describe the spatial arrangement of the four substituents around the α-carbon, not the direction of optical rotation. An L-amino acid can have positive or negative rotation, and its sign is always opposite to that of the corresponding D-amino acid. All proteinogenic amino acids except glycine are L-enantiomers; D-amino acids are far less common in nature. A 1:1 mixture of the two enantiomers is called a racemate , and in it the rotations cancel.
Where D-amino acids do appear, dedicated enzymes remove them. The flavoprotein D-amino acid oxidase from the yeast Trigonopsis variabilis was shown to lose activity through distinct pathways: inhibition by micromolar copper or mercury ions, oxidative inactivation by iron and hydrogen peroxide targeting cysteine 298, and degradation by serine proteases, with FAD binding not involving cysteine residues PMID 8737570 . The work is mechanistic enzymology rather than clinical research, but it makes plain that the L/D distinction is chemically consequential enough that organisms keep specific machinery for the less common enantiomer.
Amino acid derivatives are made by modifying the amino group, the carboxyl group, or the side chain. When a modification can be reversed without changing the parent amino acid, it functions as a protecting group : it shields a reactive site while other operations, such as coupling the next amino acid, take place. Simple protected derivatives show the logic. Ac-Ala-OH blocks the amino group; H-Ala-NH2 blocks the carboxyl group; and Fmoc-Ala-OH blocks the amino group with a group that can be selectively removed under mild conditions, which is why Fmoc chemistry is a standard tool in peptide synthesis.
Unusual α-amino acids carrying nonstandard side-chain functionalities, and turn mimetics, are useful tools in peptide design, and they are often supplied as Nα-protected derivatives so they can be coupled directly into a growing chain.
Reading amino acid identity is also an engineering problem. A graphene field-effect transistor biosensor platform achieves electrochemical amino acid profiling at surface densities around 10^12 molecules per square centimeter, holds a Dirac point drift below 10 mV over 45 minutes, and supports direct on-graphene tripeptide synthesis PMID 41744702 . This is an early-stage analytical demonstration, not a validated clinical or manufacturing tool, but it shows that amino acid detection and even on-surface peptide assembly can be tracked electronically.
A few rules cover most day-to-day reading and writing of peptide sequences. H- at the left end and -OH at the right end mean free termini. A lowercase letter in one-letter notation means a D-amino acid. A lowercase d before a three-letter code marks the D-form, an unmarked code is the L-form, and dl means a racemate. When a code is ambiguous or nonstandard, write the full name. Optical rotation α is a characteristic value that suppliers report on analytical data sheets for amino acids, derivatives, and peptides.
The evidence assembled here is pedagogical and methodological, not clinical. No clinical trial registry entry bears on the peptide/protein length boundary; the boundary is a definitional convenience, not a molecular fact, and its arbitrariness matters mainly when comparing catalogs, literature claims, and regulatory documents. The NMR teaching dataset PMID 41812158 provides instructional NMR spectra of peptides, and the protease-linker protocol PMID 38813796 describes design of protease-sensitive linkers; neither addresses the peptide/protein length boundary. The graphene biosensor work PMID 41744702 is a proof of concept at the methods stage. The D-amino acid oxidase study PMID 8737570 is a study of enzyme kinetics and provides mechanistic detail.
What remains unresolved: whether a single peptide/protein threshold will ever be standardized; how the growing abbreviation list for non-proteinogenic amino acids should be unified; and whether D-amino acid sequences, whose enantiomer chemistry is now well understood, will find a place in therapeutic design. For a beginner, the practical answer to the question posed at the start is straightforward. Amino acids are the building blocks, peptides are short chains of them, length classes are conventional, and the notation system encodes sequence, termini, and handedness in a compact, learnable way.
PMID 38813796 - An Introductory Guide to Protease Sensitive Linker Design Using Matrix Metalloproteinase 13 as an Example. ACS Biomaterials Science & Engineering, 2024. https://pubmed.ncbi.nlm.nih.gov/38813796/
PMID 41812158 - A Data Set of Simple 1-D and 2-D NMR Spectra of Peptides, including All Encoded Amino Acids, for Introductory Instruction in Protein Biomolecular NMR Spectroscopy. Biochemistry, 2026. https://pubmed.ncbi.nlm.nih.gov/41812158/
PMID 41744702 - A Graphene Field-Effect Transistor-Based Biosensor Platform for the Electrochemical Profiling of Amino Acids. Biosensors, 2026. https://pubmed.ncbi.nlm.nih.gov/41744702/
PMID 8737570 - Studies on the inactivation of the flavoprotein D-amino acid oxidase from Trigonopsis variabilis. Applied Microbiology and Biotechnology, 1996. https://pubmed.ncbi.nlm.nih.gov/8737570/
Peptides referenced: Gonadorelin.
Related reading: Chemical Synthesis of Peptides: Protecting Groups and Side Reactions, Peptide Purification After Synthesis: From RP-HPLC to MCSGP, Peptide Storage and Reconstitution: A Practical Stability Guide, Peptide QC After Synthesis: Identity, Purity, and Net Peptide Content.