Flow matching model NCFlow places unseen non-canonical amino acids

Non-canonical amino acids broaden peptide chemistry beyond the 20 canonical residues, but most structure-based design tools cannot model them. NCFlow, a flow matching generative model described in Bioinformatics, places arbitrary ncAAs into protein backbones from atom types and bond connectivity…

NCFlow places non-canonical amino acids that other tools cannot model

NCFlow , a flow matching generative model described in the journal Bioinformatics , places arbitrary non-canonical amino acids into protein backbones using only their atom types and bond connectivity. The method generalizes to non-canonical amino acids never seen during training, and in validation it outperformed AlphaFold3 -based methods at predicting the structures of those unseen residues. Existing structure-based design tools either cannot model non-canonical amino acids at all or are restricted to a small, fixed vocabulary of residue types encountered during training; NCFlow carries neither limitation.

The gap the model addresses is basic. The 20 canonical amino acids define the chemical space that most protein and peptide design software can manipulate, and that limited vocabulary constrains what researchers can build. Expanding to the hundreds of non-canonical amino acids available opens properties that canonical residues do not offer: proteolytic stability, membrane permeability and improved immunogenicity profiles, all of which matter for therapeutic peptides such as macrocycles .

In an application to peptide design across four protein-peptide complex test cases, incorporating non-canonical amino acids improved predicted binding affinity by up to -7.0 kcal/mol relative to all-canonical designs. NCFlow is freely available at https://github.com/mjslee0921/ncflow, with supplementary data published alongside the Bioinformatics online record.

A vocabulary-free representation of an amino acid

The design decision that distinguishes NCFlow is representational. Most structure prediction and protein design tools encode residues as discrete tokens drawn from a fixed vocabulary learned during training. A model built that way cannot say anything useful about a residue type that never appeared in its training set, because there is no token to represent it. NCFlow instead describes each amino acid by its constituent atoms and the bonds between them, so the model can, in principle, accept any molecule that can be drawn as an amino acid derivative attached to a backbone.

NCFlow was pretrained on millions of small molecule structures plus a large set of protein-ligand complexes, then finetuned on native non-canonical amino acids taken from proteins. The pretraining on general small molecule chemistry supplies the model with a broad sense of which geometries are chemically plausible. The finetuning step then adapts that knowledge to the specific context of a residue joined to a protein backbone and packed against neighboring side chains.

The generalization follows from this construction. Because the conditioning signal is the chemical graph itself rather than an inventory of learned tokens, the model is not capped by the residue types present in its training data. The binding constraint shifts from vocabulary coverage to the quality of the model's underlying chemistry knowledge, which is exactly where pretraining on small molecules is supposed to pay off.

A purely computational study with two endpoints

The study is entirely in silico. There is no patient population, no animal cohort, and no synthesis or binding assay anywhere in the reported work. The pretraining data comprised millions of small molecule structures; the peptide design pipeline was applied to four protein-peptide complex test cases. The study duration is not stated, which is expected for a computational methods paper.

The workflow has three stages. First, pretraining on general small molecule chemistry and protein-ligand complexes. Second, finetuning on native non-canonical amino acids from the Protein Data Bank , followed by validation against AlphaFold3-based methods. Third, application as an in silico deep mutational scanning -style peptide design pipeline, in which peptide positions are systematically varied and the resulting variants are ranked by scoring functions that combine deep learning-based and molecular dynamics-based alchemical binding free energy calculations .

Two endpoints were reported:

The design determines what the work can and cannot show. It can show that a generative model conditioned on atom types and bond connectivity produces more plausible coordinates for novel side chains than vocabulary-based methods, and it can rank peptide variants by predicted binding energy. It cannot show that any designed variant actually binds with the predicted affinity, because binding was never measured. The description of the work does not name the four complexes, state the study duration, or identify authorship.

Why unseen side chains defeat vocabulary-based predictors

Flow matching, the generative framework behind NCFlow, learns a continuous transformation, expressed as a velocity field, that transports samples from a simple prior distribution into the target data distribution. For a molecule, the model is conditioned on the chemical graph, the list of atoms and the bonds between them, and must produce three-dimensional coordinates consistent with that graph. Applied to a protein backbone carrying a non-canonical amino acid, the model must generate side chain coordinates that are simultaneously consistent with the backbone conformation, the local chemical environment and the residue's own bond topology.

Vocabulary-based tools fail at this task for a straightforward reason. AlphaFold-family models tokenize residues, and a residue type absent from training has no embedding to feed the network. The training data is the problem as much as the architecture: the Protein Data Bank is dominated by the 20 canonical amino acids, so the supervised signal for anything else is sparse. The non-canonical residues that do appear are a biased sample, drawn mostly from modified residues that serve crystallographic purposes, such as selenomethionine, or common post-translational modifications, rather than the chemically diverse building blocks a peptide chemist would want to deploy.

That bias is why the finetuning step alone would be insufficient. Native non-canonical amino acid data in the Protein Data Bank is scarce and heavily biased, so finetuning cannot teach a model the full chemistry of the hundreds of available non-canonical amino acids. The compensating signal is the small molecule pretraining: crystal structures of small molecules cover enormous chemical diversity, including functional groups, ring systems and stereochemistry that recur in non-canonical amino acids. That coverage, in the design of NCFlow, is what makes an unseen residue addressable.

What NCFlow changes in peptide design practice

For a research group designing a peptide inhibitor, the concrete change is ordering. A deep mutational scanning-style run with non-canonical amino acid substitutions produces a ranked shortlist of variants before anything is synthesized. The experimental workflow then shifts from synthesizing and assaying a large analog panel to concentrating synthesis on the top-ranked candidates. That reordering has real cost consequences, because non-canonical amino acid building blocks are more expensive than the 20 canonical residues.

The property rationale for making those substitutions is well established in peptide medicinal chemistry. Non-canonical residues can protect peptides from protease cleavage, reduce the polar surface and desolvation penalty that limit membrane passage, and lower the risk of unwanted immune recognition. For macrocyclic peptides, the combination can yield molecules that bind protein surfaces with small molecule-like pharmacokinetics, a major reason the class attracts therapeutic interest. NCFlow expands the searchable space for those molecules.

The supply chain consequence follows from the economics. If computational prioritization works, demand shifts from broad libraries of canonical analogs toward fewer, higher-value peptides containing exotic monomers. Suppliers of Fmoc-protected non-canonical amino acids become the gating resource for taking such designs forward, and the practical value of the tool will depend as much on monomer availability as on model accuracy. For clinicians, the relevance is indirect but real: stability, permeability and immunogenicity are the properties that decide dosing frequency, route of administration and safety profile. None of those properties, however, is established by a structure prediction. Candidates ranked by NCFlow still require synthesis, biophysical binding confirmation, selectivity and permeability assays, and pharmacokinetic and toxicology studies before any regulatory filing is plausible. An in silico score does not substitute for those data.

What the affinity gain does not prove

The headline figure, up to -7.0 kcal/mol, is an upper bound and a prediction. The phrasing matters: the maximum predicted improvement across four test cases is not an average, and the scoring pipeline that produced it is a combination of deep learning scorers and molecular dynamics-based alchemical calculations, each with its own error. Alchemical free energy calculations are among the most rigorous computational estimates of binding available, but their typical error is on the order of 1 to 2 kcal/mol per system, and they are estimates, not measurements.

The chemical scale of the claim deserves scrutiny. A more negative binding free energy means stronger binding, and roughly 1.4 kcal/mol corresponds to an order of magnitude in dissociation constant at physiological temperature. Taken at face value, -7.0 kcal/mol implies a shift of about five orders of magnitude in predicted affinity. Gains that large, even partially realized, would be remarkable, which is exactly why the number needs experimental confirmation. The history of computational affinity prediction contains many large predicted gains that did not survive contact with a binding assay, and an upper bound across four unnamed complexes is a thin basis for confidence.

The data limitations compound the problem. The non-canonical amino acid content of the Protein Data Bank is scarce and heavily biased, so the finetuning distribution is narrow by construction. Four protein-peptide complexes is a small sample from which to generalize, and the complexes are not named. The pretraining description is also imprecise: the number of small molecule structures is given only as millions, and the protein-ligand complex set is described only as large. These gaps do not invalidate the method, but they define how much weight the predicted affinity gains can reasonably carry.

Open questions that only experiments can settle

The most direct test is synthesis and measurement. Have any NCFlow-designed variants been made and assayed? Surface plasmon resonance or isothermal titration calorimetry would confirm whether the predicted ranking holds and whether anything close to -7.0 kcal/mol is realizable. The reported work offers no experimental confirmation, and no claim of synthesis appears in the description.

The second question is distributional. Which four protein-peptide complexes were used, and how was the predicted improvement spread among them? An upper bound across four cases is consistent with one large gain and three modest ones, or with four moderate gains. Reporting per-case binding predictions would sharpen the claim and make the method easier to evaluate.

The third set of questions concerns whether predicted binding gains translate to the properties that motivated the non-canonical substitutions in the first place. Binding free energy does not predict membrane permeability, proteolytic…

Related reading: SinCAA: a two-task pretraining model for non-canonical amino acid peptides, Semaglutide hsCRP cut of 37.8% in SELECT suggests anti-inflammatory role, GLP-1 Drugs May Cut Protein Intake, Raising Sarcopenia Risk, Ac 2-26 preserves retinal neurons via FPR2-p38-NOX4 suppression.