Enzyme classes for biocatalytic synthesis of short oligopeptides

Five enzyme classes can form peptide bonds outside the ribosome: non-ribosomal peptide synthetases, ATP-grasp ligases, α-amino acid ester acyltransferases, β-lactam acylases, and cyanophycinases. This article compares their activation chemistries, substrates, and product scopes, and assesses what…

Five enzyme classes for peptide bond formation

Peptide bonds do not require ribosomes. Five enzyme classes account for most biocatalytic routes to short oligopeptides: non-ribosomal peptide synthetases NRPS , ATP-grasp carboxylate-amine ligases , α-amino acid ester acyltransferases AETs , β-lactam acylases , and cyanophycinases Wang et al., Biomolecules 2019, 9 11 : 733 . The classes differ in how they activate the acyl donor, which substrates they accept, and which products they release. NRPS and ATP-grasp enzymes spend ATP to activate carboxylates. AETs and β-lactam acylases transfer an already activated acyl group without ATP. Cyanophycinases work in reverse, hydrolyzing a storage polymer into a defined dipeptide.

| Enzyme class | Activation chemistry | Representative substrates | Products | Practical use |

|---|---|---|---|---|

| Non-ribosomal peptide synthetases | ATP-dependent adenylation; some ATP-independent transacylation | Amino acids, aryl acids | Penicillin, bleomycin, cyclosporine | Therapeutic peptides |

| ATP-grasp ligases | ATP to acylphosphate intermediate | Carboxylates and amines | Glutathione, D-Ala-D-Ala, purine intermediates | Biosynthesis; defined ligations |

| α-amino acid ester acyltransferases | Ester-activated acyl donor, no ATP | Amino acid esters; unprotected amino acids | Oligopeptides, dipeptides | Synthesis from unprotected substrates |

| β-lactam acylases | Amide or ester acyl donor, no ATP | β-lactams, amino acid esters | 6-APA, semi-synthetic β-lactams | Industrial antibiotic manufacture |

| Cyanophycinases | Hydrolysis of poly Asp-Arg | Cyanophycin granule polypeptide | β-Asp-Arg dipeptide | Arginine-supplemented feed or food |

The distinction between ATP-dependent and ATP-independent activation is the first decision point for a researcher choosing a catalyst. ATP-dependent systems offer precise ligation but consume a co-substrate. ATP-independent transferases are cheaper to run but depend on a pre-activated acyl donor.

Non-ribosomal peptide synthetases: modular assembly lines

NRPS enzymes in bacteria and fungi produce some of the most important peptide therapeutics known, including penicillin, bleomycin, and cyclosporine. A typical NRPS is a modular assembly line in which each module adds one residue to a growing chain. A canonical module carries four functional elements: an adenylation A domain that selects and activates the incoming amino acid, a peptidyl carrier protein PCP that carries the growing chain as a thioester, a condensation C domain that forms the peptide bond, and a thioesterase Te domain that releases the finished product. The structure of a typical NRPS, PDB ID 2VSQ, shows the arrangement of the A, C, Te, and PCP domains within a module.

The activation step defines the energy budget. ATP-dependent biocatalysts, such as tRNA-dependent ligases, activate the substrate carboxylate as an aminoacyl-adenosine monophosphate aminoacyl-AMP intermediate. ATP-independent transacylases instead use aminoacyl phosphate. Both routes generate an electrophilic carbonyl that can be attacked by the amine of the next amino acid, but the leaving groups and cofactor requirements differ. Because NRPS modules can incorporate non-canonical residues, including D-amino acids and hydroxy acids, the products include structures that ribosomal synthesis cannot make, which is why the class is central to therapeutic peptide production.

The practical draw of NRPS for short peptides is specificity. Each bond in the product is installed by one module's A, PCP, and C domains, and the Te domain decides how the chain is released. The domain architecture visible in 2VSQ is therefore the map for any attempt to repurpose an NRPS for a new product, but it also means the system is modular rather than general: a new target usually requires module engineering, not a simple change of reaction conditions.

ATP-grasp enzymes: carboxylate-amine ligation

ATP-grasp enzymes form a second ATP-dependent route to peptide bonds, with different activation chemistry. These carboxylate-amine ligases activate the acid substrate as an acylphosphate intermediate : ATP phosphorylates the carboxylate, and the resulting mixed anhydride is attacked by the amine nucleophile. The family takes its name from the conserved three-domain architecture that grasps ATP, and it includes biotin carboxylase, D-alanine-D-alanine ligase Ddl , and glutathione synthetase. The same chemistry operates in de novo purine biosynthesis, where glycinamide ribonucleotide synthetase installs glycine; the structure of that enzyme, PDB ID 2IP4, is a representative ATP-grasp fold.

Three structural features distinguish the family Wang et al., Biomolecules 2019, 9 11 : 733 . The enzyme has three conserved domains and a nonclassical ATP-binding fold, and most family members require an Mg2+ ion coordinated by ATP in the active site. The metal ion positions the triphosphate for attack by the substrate carboxylate and stabilizes the developing charge in the transition state. The reaction logic resembles that of other ATP-dependent ligases, but the ATP-grasp fold itself is unrelated to the NRPS architecture.

For short oligopeptides, the ATP-grasp class matters most for two products. Ddl makes the dipeptide D-Ala-D-Ala, the crosslinking precursor in bacterial peptidoglycan, and glutathione synthetase completes the tripeptide glutathione. Both are research-relevant targets, and both show that ATP-grasp ligases can build specific short peptides from unprotected substrates. The trade-off is the same as in NRPS systems: ATP is consumed stoichiometrically, and the acylphosphate intermediate is intrinsically reactive, which places a premium on reaction conditions.

α-amino acid ester acyltransferases: synthesis from unprotected amino acids

AETs offer an ATP-free route to oligopeptides from unprotected amino acids, with high reported yields. Kenzo and colleagues developed an enzymatic method using Empedobacter brevis ATCC 14234 and identified the catalyst as an enzyme named carboxypeptidase Y. The significance of the approach is that the acyl donor is an amino acid ester, not a free carboxylic acid, so no activation cofactor is needed: the ester itself is the electrophile. That is the practical advantage that makes AET chemistry attractive for preparative peptide synthesis.

The molecular picture came later. Isao Abe and team reported the first cloning and expression of AETs from Empedobacter brevis and from a second strain, and measured the amino acid sequences as 35% and 36% identical to the α-amino acid ester hydrolase from Acetobacter pasteurianus. Sequence identity in that range indicates a shared ancestor and probably a shared fold, but also that the AETs are distinct enzymes with their own specificity. The enzymes display dual dipeptidyl peptidase and transferase activities, and they are specific for both acyl donors and nucleophiles. That double specificity is what allows a clean product: the enzyme rejects the wrong donor and the wrong acceptor, so the reaction does not drift into a mixture of oligomers.

The gap in the AET story is structural. The amino acid sequence, coding gene sequence, and three-dimensional crystal structure of the carboxypeptidase Y-like enzyme from ATCC 14234 have not been provided, and no study has yet reported the three-dimensional structure of an AET or its detailed reaction mechanism. A researcher who wants to engineer an AET for a new substrate is working without the two tools that make engineering feasible: a gene sequence and a structure. The class has demonstrated utility, but the mechanistic understanding lags behind the application.

β-lactam acylases: industrial workhorses

β-lactam acylases are the enzyme family with the deepest industrial track record. The family includes penicillin acylase, glutaryl acylase, and β-amino acid ester hydrolase. These enzymes were historically used to process β-lactam antibiotics, and they remain central to the current biosynthetic production of semi-synthetic β-lactams, a route described as environmentally friendly and cost-effective and increasingly applied in pharmaceutical production Wang et al., Biomolecules 2019, 9 11 : 733 .

Penicillin acylases are the most widely deployed members of the family. They are produced by many microorganisms and are categorized into two types based on substrate specificity. Their main industrial use is the production of 6-aminopenicillanic acid 6-APA , the active pharmaceutical intermediate from which semi-synthetic penicillins are made. The same enzymes are used for peptide synthesis, for resolution of racemic mixtures, and for production of chiral and achiral pharmaceutical intermediates. That range of uses explains why penicillin acylases are described as widely used industrially.

The chemistry is a transferase reaction run in a hydrolytic enzyme. Penicillin acylase normally hydrolyzes the amide bond of penicillin to release 6-APA, but under controlled conditions it can transfer the acyl group to a nucleophile other than water, forming a new amide or ester bond. The dual hydrolytic and synthetic behavior is shared with the AETs, and it is the reason both families can build peptide bonds rather than simply break them. The industrial record rests on this reversibility, managed through substrate choice and reaction conditions.

Cyanophycinases: recycling a storage polymer into a dipeptide

Cyanophycin granule polypeptide CGP is an intracellular storage polymer found in most cyanobacteria. It is not a random copolymer: it contains equimolar arginine and aspartic acid, and each arginine is linked through its α-amino group to the β-carboxyl group of an aspartic acid. The polymer is therefore a poly aspartic acid backbone with arginine residues attached to the side chains, and the linkage is a β-peptide bond rather than the canonical α-linkage.

Cyanophycinases are the enzymes that recycle this polymer. CphB degrades CGP intracellularly and CphE degrades it extracellularly, and both release the dipeptide β-Asp-Arg . Because the dipeptide is β-linked, it is not a substrate for the ordinary peptidases that cleave α-peptide bonds, which makes it an interesting product in its own right. The production strategy identified in the review Wang et al., Biomolecules 2019, 9 11 : 733 is simultaneous production of CGP and the CGPase, which synthesizes β-Asp-Arg efficiently and could provide arginine in feed or food.

CGP production has been established in recombinant strains of Escherichia coli, Nicotiana tabacum, Pseudomonas putida, and Pseudomonas alcaligenes DIP1. In one advance, co-expression of CGP and CGPase in Nicotiana tabacum was achieved, potentially allowing sufficient storage and efficient transport of arginine and β-Asp-Arg dipeptides. A crop that stores arginine in a stable polymer and then releases it as a small dipeptide could serve as a self-contained delivery vehicle for an essential amino acid. The review concludes that metabolic engineering of suitable hosts and chemo-enzymatic strategies are feasible for producing dipeptides such as β-Asp-Arg, though it stops short of reporting scaled production data.

What the primary evidence shows

The class-level picture above comes from the 2019 review. Primary studies on enzymatic peptide chemistry add both support and caution, and four findings are directly relevant to a researcher choosing among these biocatalysts.

First, enzyme behavior is modulated by substrate conformation. In a 2025 study of an engineered sortase, peptides with low helicity readily underwent intramolecular head-to-tail cyclization, whereas more rigid helical peptides tended to form cyclic dimers PMID 40289331 . Peptide rigidity redirected the enzymatic reaction from intramolecular cyclization to intermolecular dimerization. For anyone running a chemo-enzymatic cyclization, the lesson is to check substrate conformation before blaming the enzyme: a rigid peptide will dimerize even with a properly engineered ligase.

Second, enzymatic backbone modification is now a practical laboratory method. A methods chapter establishes in vitro protocols for enzymatic thioamidation of peptide backbones using recombinant enzymes, covering polypeptide expression, purification, reaction reconstitution, and mass spectrometry-based product analysis PMID 34325795 . The protocols show that post-synthetic backbone editing can be done with purified enzymes and standard biochemistry, rather than specialized equipment.

Third, the products justify the synthesis effort. A computational study used molecular docking to construct spatial models of DNA-peptide complexes for 19 short peptides and identified shared binding sites, such as KE/EDP binding to the motif 'agat' and KEDW/AED binding to 'acct' PMID 27909961 . The study supports the hypothesis that short peptides regulate gene expression by binding DNA. If short peptides act as signaling molecules, the enzymatic routes described above are a route to biologically meaningful products, not just a convenience.

Fourth, delivery remains a bottleneck. A review of gastrointestinal enzymatic barriers concluded that the small-intestine lumen and the brush border membrane, which contains at least 15 peptidases, are major obstacles to oral peptide delivery, and it identified enzyme-resistant peptide analogues as the most promising strategy PMID 7600588 . That finding connects directly to the β-Asp-Arg system: a β-linked dipeptide is by construction resistant to many α-peptidases, which is a plausible reason to pursue it for feed or food applications. The same logic argues for enzymatic synthesis of non-canonical linkages generally.

The limits of the primary evidence matter as much as its findings. None of the four studies is a head-to-head comparison of the five enzyme classes on the same target peptide, and none provides yield data for industrial-scale oligopeptide production. The class-level claims in the review are descriptive. The quantitative basis for choosing one class over another, on cost per gram or on final purity, is not established in the evidence reviewed here. A researcher should treat the class descriptions as a map of options, not as a ranking.

Practical guidance and limits

For a researcher or buyer, the practical question is which enzyme class fits the target molecule. The guidance below is drawn from the review and from the primary evidence.

The practical corollary of the delivery data is that non-canonical linkages are a feature, not a defect. The brush border contains a dense array of peptidases PMID 7600588 , and a β-linked dipeptide such as β-Asp-Arg resists the enzymes that demand an α-linkage. By the same logic, the helicity-dependent cyclization data PMID 40289331 are a warning: a product that is too rigid will not behave as intended in a one-pot enzymatic reaction. Conformation and linkage should be weighed at the design stage, alongside the choice of enzyme.

A buyer of custom peptides should weigh these enzyme classes against the specific target and should request the gene sequence, structure, and reaction-mechanism evidence that is missing for the AET class before committing to that route for a new product. For the established classes, the industrial record of penicillin acylase in 6-APA production is the strongest evidence of scalability. For the newer systems, the scalability evidence is thinner, and claims of high yield should be read against the specific substrates and conditions reported, not generalized.

Open questions and unresolved details

The largest gaps are structural. The enzyme catalyst named carboxypeptidase Y from Empedobacter brevis ATCC 14234 has been tied to efficient oligopeptide production from unprotected amino acids, yet its amino acid sequence, coding gene sequence, and three-dimensional crystal structure have not been provided. Without the gene, the enzyme cannot be recombinantly produced, evolved, or scaled by a third party. Similarly, no studies have yet been conducted on the three-dimensional structure of α-amino acid ester acyltransferase or on its reaction mechanism. The 35% and 36% sequence identities to the Acetobacter pasteurianus hydrolase place the AETs in a known family, but identity at that level is insufficient to predict substrate specificity or to guide engineering.

A second gap is comparative. The review describes each enzyme class on its own terms, and the primary studies discussed here address specific enzymes rather than the five classes head to head. In the evidence reviewed here, no dataset directly compares AET yield, penicillin acylase yield, and cyanophycinase yield on the same dipeptide target under matched conditions. Until such data exist, practical guidance rests on qualitative class descriptions and on the industrial track record of the β-lactam acylases.

A third gap is mechanism at the atomic level. For NRPS, the domain architecture is known from 2VSQ, and for ATP-grasp enzymes the fold is known from 2IP4, but those structures describe the scaffold rather than the full reaction trajectory. For the AETs, even the scaffold is missing. The result is that a class with an attractive synthesis property, high yield from unprotected amino acids, is also the class with the least molecular information, and the enzyme that makes the β-Asp-Arg dipeptide is understood from its substrate and product rather than from its structure.

What remains clear despite the gaps is the range of available chemistry. Peptide bonds can be made by ATP-driven modular assembly lines, by ATP-grasp ligases, by ester-activated transferases, and by acylases working in synthetic mode, and they can be recovered from storage polymers by specific hydrolases. Each route has a distinct substrate profile, product scope, and industrial record. The choice among them is a matter of matching the target molecule to the chemistry, and of recognizing where the evidence ends and engineering must begin.

References

Peptides referenced: Glutathione.

Related reading: Choosing Coupling Reagents for Solid-Phase Peptide Synthesis, AI-Guided Peptide Library Design: Capabilities and Outcomes, How Chameleon Cyclic Peptides Cross Membranes for Oral Drugs, Five Enzymatic Routes to Oligopeptides and Short Peptide Synthesis.