# LiteFold > LiteFold is an applied AI4Science lab and AI drug discovery platform. Its autonomous research agent, Rosalind, works across therapeutic R&D (literature research, protein and molecule design, docking, molecular dynamics, ADMET and toxicity prediction, CMC and process development) on private, customer-controlled infrastructure. LiteFold also publishes open research: the LiteMol-1 molecular foundation model, the AminoWeb protein dataset and the BenchPLM protein language model benchmark. Key facts: - Company: LiteFold, an applied AI4Science lab based in Bengaluru, India, working with enterprise life-sciences R&D teams. - Website: https://www.lite.bio · Platform: https://prod.litefold.ai · Docs: https://docs.litefold.ai · Contact: contact@litefold.ai - Flagship product: Rosalind, an autonomous AI co-scientist for drug discovery. It reads private context (ELNs, assay databases, sequences, structures, papers, patents), writes and runs code in isolated sandboxes, and calls LiteFold scientific compute without proprietary data leaving the customer's environment. - Scientific compute: biomolecule and binder design (peptides, antibodies, nanobodies, small molecules, PROTACs), protein structure prediction, protein–ligand docking and virtual screening, pocket and hotspot detection, molecular dynamics, binding affinity, ADMET, toxicity, inverse folding and model fine-tuning on customer data. - Benchmarks: Rosalind scores 90.24% on BixBench (Future House's benchmark of 200+ real bioinformatics tasks), ahead of K-Dense (90.0%), Biomni Lab (88.7%), Edison (78.0%) and Claude Code with Opus 4.6 (65.3%). LiteFold's MD engine runs the STMV benchmark (1.07M atoms) at 136 ns/day on one RTX 5090, compared with 62 ns/day on an H100. MitoTox predicts mitochondrial toxicity with 82.6% accuracy. - Security: SOC 2 compliant, workloads run in private enclaves, self-hosted or private-cloud deployment is available, and customers own all IP, models and designs. Customer data is never used to train shared models. Trust center: https://trust.litefold.ai - Enterprise offerings: a Platform License (1–3 years, self-hosted), a Computation and AI Service (LiteFold designs 10–20 binder candidates with a full workup report) and an End-to-End Service (adds wet-lab SPR validation). Every engagement includes forward-deployed engineers and scientists. - Backed and supported by: NVIDIA Inception and Mercatus Center (George Mason University). - Team: Anindya Sannigrahi (CEO & Co-Founder), Cory Kornowicz (CSO & Co-Founder), Siddhant Prateek Mahanayak (Founding Infrastructure Engineer), Nabajit Borah and Aditi Sinha (Founding Computational Biologists), Pannalal Aich (Director of Business Development). --- # LiteFold use cases - Rare Disease Variant Interpretation (https://www.lite.bio/use-cases/rare-disease-variant-interpretation): Interpret rare disease variants faster with literature, population, and structural evidence in one place. - Hypothesis Generation (https://www.lite.bio/use-cases/hypothesis-generation): Generate and rank research hypotheses backed by literature, variant, and structural evidence. - Analytical Comparability Planning (https://www.lite.bio/use-cases/analytical-comparability-planning): Plan comparability studies and flag critical quality attributes before you touch the bench. - Biosimilarity Assessment (https://www.lite.bio/use-cases/biosimilarity-assessment): Build a defensible biosimilarity case grounded in regulatory guidance and analytical evidence. - Cell Culture Process Optimization (https://www.lite.bio/use-cases/cell-culture-process-optimization): Diagnose process deviations and design the next experiment to close the gap faster. - Process Characterization Study Design (https://www.lite.bio/use-cases/process-characterization-study-design): Design statistically sound characterization studies and identify your critical process parameters. - Formulation Development Advisor (https://www.lite.bio/use-cases/formulation-development-advisor): Select excipients and forecast stability to converge on a robust formulation faster. - Patent Analysis (https://www.lite.bio/use-cases/patent-analysis): Map the patent landscape, assess freedom to operate, and surface white-space opportunities. - Lead Optimization (https://www.lite.bio/use-cases/lead-optimization): Rank candidates by binding quality, resistance coverage, and ADMET before you synthesize. - Preclinical to IND (https://www.lite.bio/use-cases/preclinical-to-ind): Organize pharmacology, toxicology, and CMC evidence into an IND-ready development package. - Experimental Protocol Generation & Optimization (https://www.lite.bio/use-cases/experimental-protocol-generation): Generate a complete purification SOP with CPPs, CQAs, and a scale-up strategy built in. - Clinical Trial Development (https://www.lite.bio/use-cases/clinical-trial-development): Design randomized trial protocols with biomarker-driven endpoints and eligibility criteria. - Bioprocess Scale up Strategy (https://www.lite.bio/use-cases/bioprocess-scale-up-strategy): Optimize media, feeding strategy, and CPPs to scale production without losing quality. - Docking Campaign and Design Interpretation (https://www.lite.bio/use-cases/docking-campaign-design-interpretation): Rank experimental vs. AI-predicted binding models to identify the most reliable complex. - Molecular Dynamics Simulation & Stability Analysis (https://www.lite.bio/use-cases/molecular-dynamics-simulation): Validate protein–ligand binding stability with RMSD, RMSF, and contact persistence over a full MD trajectory. - Analytical Comparability Package for a Biosimilar Monoclonal Antibody (https://www.lite.bio/use-cases/biosimilar-analytical-comparability-package): A structural, impurity, and stability comparability dossier with a three-tier statistical similarity framework, aligned to FDA, EMA, and WHO biosimilar guidance. - Comprehensive Biosimilarity Assessment Strategy (https://www.lite.bio/use-cases/biosimilarity-assessment-strategy): An end-to-end biosimilarity assessment spanning reference product characterization, analytical and functional comparability, immunogenicity risk, and indication extrapolation. - Formulation Development Strategy for a High-Concentration Biosimilar (https://www.lite.bio/use-cases/formulation-development-strategy): A QTPP, formulation strategy, stability program, and risk-based development plan for a high-concentration antibody intended for prefilled syringe or autoinjector delivery. - Glycosylation Comparability Assessment of a Biosimilar Candidate (https://www.lite.bio/use-cases/glycosylation-comparability-assessment): A tiered statistical equivalence framework comparing glycoform distributions, functional impact, and immunogenicity risk between a biosimilar candidate and its reference product. - Technical Investigation into Protein A Chromatography Yield Decline (https://www.lite.bio/use-cases/protein-a-recovery-investigation): A fishbone and ranked FMEA analysis of a progressive Protein A step-yield decline across commercial batches, with confirmatory experiments and CAPA recommendations. - Process Characterization Study for a Protein A Capture Step (https://www.lite.bio/use-cases/protein-a-process-characterization): A characterization study linking critical process parameters to critical quality attributes for a commercial-scale Protein A affinity chromatography step. - CMC Scale-Up Strategy for Upstream Bioreactor Manufacturing (https://www.lite.bio/use-cases/upstream-scaleup-cmc-strategy): The engineering and QbD-based justification for a large bioreactor volume scale-up, including FMEA risk assessment and a comparability testing plan. - Upstream Process Optimization Strategy for Fed-Batch Cell Culture (https://www.lite.bio/use-cases/upstream-process-optimization): A prioritized, risk-ranked set of interventions spanning feeding strategy, temperature shift, and media supplementation to improve titer and quality attributes. - Comprehensive CMC Development Strategy for a Monoclonal Antibody Biosimilar (https://www.lite.bio/use-cases/biosimilar-cmc-development-strategy): An integrated CMC strategy spanning cell line development, upstream and downstream process design, formulation, analytical comparability, and scale-up across dual filing pathways. - Structural Analysis of an Antibody-Receptor Interaction via Docking (https://www.lite.bio/use-cases/antibody-receptor-docking-analysis): A comparison of an experimental antibody-antigen co-crystal structure against an ab initio cofolding prediction, validating interface residues and binding accuracy. --- # LiteFold benchmarks ## LiteMol-1 LiteMol-1 is LiteFold's first molecular foundation model for agent-driven design. The benchmark panel summarizes peptide co-fold confidence across ipSAE, ipTM, and pLDDT, plus small-molecule docking strength against public baselines. - LiteMol-1 + MCTS: 8.02 score - ProtoBind-Diff: 8.17 score - LiteMol-1: 7.77 score - PocketXMol: 7.18 score ## MitoTox MitoTox is LiteFold's state-of-the-art model for mitochondrial toxicity prediction, part of our pre-clinical stack (Signal). It helps teams screen liabilities early and run toxicity-aware lead design and optimization. - MitoTox (LiteFold): 82.6 % - Atom-Pair + GB Baseline: 80.9 % - Mammoth (LiteFold): 39.5 % ## Molecular Dynamics STMV benchmark (1.07M atoms, explicit PME), higher is better. Our optimized engine brings datacenter-class molecular dynamics throughput to a single RTX 5090, measured in nanoseconds simulated per day. Representative benchmark comparing our engine against stock OpenMM and datacenter GPUs. - RTX 5090 (LiteFold): 136 ns/day - B200: 118 ns/day - RTX 5090 (Stock): 101 ns/day - H200: 78 ns/day - H100: 62 ns/day - RTX 4080: 46 ns/day - A100: 32 ns/day ## BixBench BixBench is the Future House benchmark of 200+ real bioinformatics tasks built from published research notebooks. Agents have to navigate complex datasets, execute Python and R, generate testable hypotheses, and defend their answers. New re-evaluations are ongoing right now, and the leaderboard will be refreshed as results finish. - Rosalind (LiteFold): 90.24 % - K-Dense: 90.0 % - Biomni Lab: 88.7 % - Edison: 78.0 % - Claude Code (Opus 4.6): 65.3 % --- # LiteMol-1: A Multi Molecule Foundation Diffusion Language Model for Agents URL: https://www.lite.bio/research/litemol1 Author: LiteFold Research Date: 2026-08-17 ***Let the Agents dream, explore and optimize*** ***LiteMol-1** generates across molecular classes — small molecules, peptides, macrocycles, and more. Drag a tile to rotate.* Today, we are releasing the first foundation model from LiteFold: LiteMol-1. LiteMol-1 is a masked diffusion language model, pre-trained from scratch on molecular sequence tokens. A single set of weights can generate across multiple molecular modalities: small molecules, linear peptides, cyclic peptides, peptides containing non-canonical amino acids, depsipeptides, macrocycles, and PROTACs. The model can generate molecules unconditionally or conditioned on a protein target. Its generations can also be steered toward desired properties and constrained during sampling. For multi-objective optimization, we use **Monte Carlo Tree Search**,[^mcts] where partially denoised sequences are explored as a search tree and candidates are evaluated across multiple objectives. Rather than collapsing these objectives into a single weighted score, the search maintains a **Pareto non-dominated set**: a group of designs that represent different trade-offs. A design stays in this set when no other candidate clearly improves one objective without giving up something elsewhere, which is closer to how medicinal design decisions are usually made.[^peptune] The part we find most interesting is not that LiteMol-1 can generate molecules, it is that the model is unusually well suited to agents. Sequence provides a compact and editable substrate. An agent can inspect a molecule, preserve parts of it, modify others, generate alternatives, evaluate them with external models or physics-based tools, and feed those observations back into another round of generation. In this setting, LiteMol-1 becomes something closer to an infinite molecular canvas. The model proposes possibilities. The agent explores them, reasons about them, runs experiments, and decides what should change next. When LiteMol-1 is placed inside an automated research loop, this creates a simple but powerful cycle: *The agent loop LiteMol-1 is built for — propose, check, change, and go again.* We strongly believe this auto-research loop is where some of the more interesting consequences of sequence-first molecular design begin. The results below are our first experiments in that direction. > **What this release shows:** LiteMol-1 performs strongly on the benchmarks reported below and, to our knowledge, is the only model in this comparison that covers this range of molecular classes with a single set of weights. The benchmarks below shows that sequence-first design can be competitive while staying fast, editable, and agent-friendly. > > This research post then goes beyond benchmark numbers to study the core capabilities of the model: **Represent, Edit, and Steer**. We look at what happens when LiteMol-1 is placed inside an Auto Research loop, and how quickly that loop can produce useful candidates across multiple targets and molecular modalities. > > This post does not include wet-lab validation yet. We are currently evaluating the model across designing peptides, PROTACs, and small molecules, and these reports will be released in the coming months. ## Why We Build a Sequence Model Molecular design has increasingly been framed as a structure problem. RFdiffusion,[^rfdiffusion] BoltzGen,[^boltzgen] PX-Design,[^pxdesign] and the generation of co-folding models that followed AlphaFold 3[^af3] all operate primarily in coordinate space. As a design choice, this is intuitive. If molecular function ultimately emerges from three-dimensional interactions, then learning directly over structure gives a model an explicit representation of *where* interactions happen and *how* a molecule might bind. But there are some practical limitations to structure-based molecular design. 1. **Good binders often require enormous search:** Systems such as BoltzGen and O-Design[^odesign] recommend generating and filtering 10-50K designs before arriving at a small number of strong candidates.[^design-scale] 2. **We eventually come back to sequence anyway:** A typical structure-design pipeline generates a backbone, inverse-folds it into a sequence, and then folds that sequence again to verify that it recovers the intended structure. This raises a natural question: for many design problems, why not operate directly in sequence space first? 3. **Structure models are expensive to run at campaign scale:** Many modern architectures rely on Pairformer-like stacks and triangular operations with unfavorable scaling, making large generation-and-filtering campaigns computationally expensive and slow. 4. **Structures are also harder interfaces for agents:** An agent can inspect a PDB/CIF file, or interact with a visual 3D environment, but both approaches rapidly consume context. Long coordinate files and repeated multimodal observations make sustained auto-research loops increasingly inefficient. Sequence models, meanwhile, have mostly been used as representation learners: pretrain a model, extract embeddings, and apply them to downstream tasks. We think they are also useful as generative models for molecular design. A large fraction of experimental biology is naturally recorded alongside sequence: expression, binding, mutational scans, activity, toxicity, stability, and other molecular properties. Sequence data is also substantially more abundant and easier to scale than experimentally resolved structure. Work such as ESM-2 has shown that models trained purely on sequence can still recover meaningful structural and evolutionary information.[^esm2] That is why we built Lite-Mol. Sequence models like Lite-Mol operate slightly earlier in the design loop. They provide a cheap and flexible substrate for hypothesis generation, editing, and iteration. A model trained on enough molecular sequence data can learn not just syntax or local chemistry, but parts of the accumulated footprint of medicinal chemistry, peptide engineering, and synthesis: which motifs are common, which combinations are plausible, and which regions of chemical space have repeatedly survived experimental selection. The other advantage is speed and editability. With masked diffusion, parts of a molecule can remain fixed while only the region we want to change is corrupted and reconstructed. Because the working representation is a sequence, these operations are also unusually natural for an agent to specify and repeat. For instance consider this prompt: ***Agent-native molecular design.** The model proposes editable sequence-level changes; the agent evaluates, filters, and sends the next design instruction.* That creates a different interface to molecular generation. The goal is not to replace structural reasoning, but to make the design loop cheap enough, editable enough, and programmable enough for an agent to continuously explore molecular hypotheses and bring in structural models when deeper verification is actually needed. ## Primer to Discrete Diffusion Language Model At it’s core, LiteMol-1 builds on top of Discrete Diffusion Language Model. During inference, we start from a fully masked SMILES sequence and the model denoises into a complete valid molecule SMILE. ![**Figure:** Absorbing-state discrete diffusion on SMILES. Tokens begin as `[MASK]` and are gradually revealed until a complete molecule appears — the reverse process LiteMol-1 learns to run.](images/diff.gif) In general, diffusion models are built around two processes: a **forward process** that gradually adds noise to data, and a **reverse process** that learns to remove that noise. For continuous data such as images, coordinates, or latent representations, the forward process usually adds a small amount of Gaussian noise at each step: $$ q(x_t \mid x_{t-1}) = \mathcal{N}\left(x_t;\sqrt{1-\beta_t}\,x_{t-1},\,\beta_t I\right) $$ Here, $\beta_t$ controls how much noise is added at step $t$. After many steps, the original structure in $x_0$ is gradually destroyed and $x_t$ approaches noise. We then train a neural network, typically written as $\epsilon_\theta(x_t,t)$, to look at the noisy sample $x_t$ and estimate the noise that was added. Once the model learns this denoising process, we can start from random noise and repeatedly apply the learned reverse process: $$ p_\theta(x_{t-1} \mid x_t) $$ Eventually, the noise is transformed back into a structured sample. This formulation works naturally for **continuous data**, where adding Gaussian noise is well defined, such as pixels, 3D coordinates, or continuous latent vectors. Smiles strings are not continuous. It’s a sequence of discrete tokens from fixed vocabulary. Adding Gaussian noise to a one-hot token embedding and rounding back doesn't work. That’s why Discrete Diffusion (D3PM; Austin et al., 2021)[^d3pm] fixes this by defining the forward process directly on the categorical distribution over tokens, via a transition matrix $Q_t$: $$ q(z_t \mid z_{t-1}) = \operatorname{Cat}(z_t; z_{t-1}Q_t) $$ $Q_t$ defines **how tokens are corrupted at each step**. It could replace tokens randomly, swap them with similar tokens, or, in our case, simply replace tokens with a special `[MASK]` token. Once a token becomes `[MASK]`, the forward process keeps it masked. So instead of gradually adding continuous noise, we gradually hide more and more of the sequence. This is also called **absorbing-state diffusion**. This makes the diffusion process much simpler mathematically and turns the learning problem into something very similar to masked language modeling: given a partially masked sequence, predict the original tokens. Another reason, we call it discrete is because we sample over from a finite set of vocabulary tokens. ## What LiteMol-1 is LiteMol-1 is a Masked Diffusion Language Model (MDLM).[^mdlm] The forward process is fairly simple to state. We start with a clean token $x$ and gradually replace it with `[MASK]` as diffusion progresses. At each time $t$, the schedule $\alpha_t$ controls how much the original sequence is still preserved. $$ q(z_t \mid x) = \operatorname{Cat}\left(z_t;\alpha_t x + (1-\alpha_t)m\right) $$ Here $m$ is the special `[MASK]` token. So for every token: - With probability $\alpha_t$, we keep the original token. - With probability $1 - \alpha_t$, we replace it with `[MASK]` token. At the beginning $\alpha_t$ is close to 1, so almost nothing is masked. As $t$ increases, $\alpha_t$ decreases and more of the sequences disappears behind `[MASK]`. The model then learns the reverse problems, where given a partially masked sequence: $z_t$, we predict the original tokens. You can play with the slider below to watch a real SMILES string absorb into `[MASK]` under a linear schedule $\alpha_t = 1 - t$. Once a token masks, it stays masked which is the absorbing-state rule LiteMol-1 uses. ***Forward process on SMILES.** Each token draws a uniform threshold and becomes `[MASK]` once diffusion time passes it. Play to run the schedule; Resample draws a new masking order. Training teaches the reverse: recover the molecule from whatever is still visible.* ### Structure aware masking schedule An observation from PepTune (Tang et al., 2024)[^peptune] which worked with a masked diffusion language model to unconditionally generate peptide sequence shows us that not every tokens in the SMILES string should carry the same weight. A peptide bond or a disulfide linkage defines the molecule’s identity. Getting that wrong will give us completely invalid molecule. In PepTune’s (Tang et al., 2024) paper, the peptide-bond tokens were marked by an indicator vector **b**. Non-bond tokens keep the standard linear schedule, where as the bond tokens (marked with **b**) get a slower growing polynomial masking scheduling rate. Mathematically: $$ \alpha_t(x_0)=\begin{cases}1-t^w, & x_0=b \quad \text{(peptide-bond token)}\\1-t, & x_0\neq b \quad \text{(everything else)}\end{cases} $$ With $w$ found empirically around 3. Since $t^w < t$ for $t \in (0,1)$ and $w > 1$, a bond token has a lower probability of already being masked at any given time: $t$. It survives longer in the forward process, and correspondingly gets unmasked earlier in the learned reverse process. At LiteMol-1, we generalized this idea beyond peptide-specific bonds to a broader set of molecular classes and structural motifs. Importantly, this is a **scheduling strategy, not additional supervision**. The model still predicts the same tokens; what changes is *when* different classes of tokens are exposed to corruption during training. This biases generation toward establishing the molecular scaffold first and filling in more editable regions later. The sampler does not need explicit knowledge of these roles at inference time. A masked token has no identity until it is predicted. The effect therefore comes from the model training itself and act as a structural bias that encourages scaffold-first generation, rather than a hand-crafted per-token reverse schedule during sampling. ### Training LiteMol-1 LiteMol-1 is trained as a single masked diffusion model across all of the molecular classes we support. We use a **SMILES-native tokenizer (SMILE Pair Encoding)**[^spe] designed around our training distribution, with lossless fallback so that uncommon fragments, stereochemistry, ring closures, and other chemistry are preserved rather than converting into unknown tokens. The goal here is simple: the representation should be compact enough for a language model, without losing information from the original molecule. We created a structure-aware training paradigm where we do not treat every token in a molecule as equally important. Chemically meaningful regions and structural motifs influence how corruption and reconstruction should be learned. Alongside the main diffusion objective, the model receives additional supervised signals about molecular class, structural roles, and physicochemical properties. Together, these biases help one shared backbone learn across very different objects (from small molecules and peptides to macrocycles and PROTACs), without turning LiteMol-1 into a collection of separate specialist models. For target-conditioned generation, we provide the model with a compressed representation of the target protein and train **conditional and unconditional generation together within the same network**. During training, the protein condition is deliberately removed from a fraction of paired examples, allowing the model to learn both $p_\theta(x)$ and $p_\theta(x \mid y)$ without maintaining separate checkpoints. At inference time, this enables **classifier-free guidance (CFG)**:[^cfg] we compare the model’s conditional and unconditional predictions and amplify the difference between them to control how strongly generation follows the protein target. In practice, stronger guidance pushes samples further toward what the target condition implies, while weaker guidance preserves more of the model’s unconditional molecular prior. We leave the exact training mixture, loss weights, masking configuration, and conditioning architecture outside the scope of this release. What matters here is the overall design: LiteMol-1 learns **generation, structural priors, molecular properties, and protein conditioning jointly within a single model**, while retaining a controllable separation between its unconditional and target-conditioned behavior. ### Multi-Objective Constraint Sampling with MCTS Molecular design is not a single-objective problem. Currently every go-to molecule design model is optimized for better binding. However, that does not guarantee a good therapeutic index overall. A useful molecule will need to bind well while remaining synthesizable, staying within a desired molecular-weight or LogP range, avoiding off-targets, and preserving parts of an existing scaffold. For this, we use **Monte Carlo Tree Search (MCTS)**[^mcts] as an inference-time search layer over the reverse diffusion process. A partially masked molecule naturally becomes a state in the search tree, while each denoising step creates possible branches. The model itself provides the prior over plausible continuations, and MCTS decides which of those branches are worth exploring further. At each step, the search balances what has already worked with what still looks promising using the usual exploration–exploitation rule:[^puct] $$ U(s,a)=Q(s,a)+c_{\mathrm{puct}}\,p_\theta(a \mid s)\,\frac{\sqrt{N(s)}}{1+N(s,a)} $$ Here, $Q(s,a)$ captures the quality observed from previous rollouts, while $p_\theta(a\mid s)$ keeps the search close to what LiteMol-1 considers a plausible molecular continuation. The loop is straightforward: **select, expand, roll out, score, and backpropagate**. See the animation below for a more intuitive understanding. ***Monte Carlo Tree Search over reverse diffusion.** Each node is a partially masked sequence and every edge is a denoising move. The search repeatedly *selects* a promising state by UCB, *expands* it one reverse-diffusion step, *rolls out* to a scored terminal, and *backpropagates* the reward — steering sampling toward sequences that jointly satisfy a weighted set of objectives.* For multiple objectives, we do not necessarily collapse everything into one weighted reward. Instead, candidates can be compared through **Pareto dominance**,[^peptune] allowing the search to maintain a set of molecules representing different trade-offs between competing properties. This is useful because many molecular properties are better treated as constraints or desired ranges rather than quantities that should simply be maximized. The expensive checks stay outside the inner search loop. MCTS can operate on fast property models and proxy objectives, while docking, co-folding, structure prediction, or molecular dynamics are used later on the smaller set of promising candidates. In this setup, **LiteMol-1 defines the molecular prior, while search decides where inside that space we should spend compute.** ## Evaluations We evaluated LiteMol-1 across several settings. We start with peptide design against frontier structure-based models: BoltzGen,[^boltzgen] RFdiffusion3,[^rfdiffusion3] and O-Design.[^odesign] We then compare LiteMol-1 with peptide sequence models on chemical validity, non-canonical residue coverage, and protein-conditioned docking. Finally, we report small-molecule evaluations for the generative prior, protein-conditioned generation, docking, and inference-time search before moving into case studies. Before going further, two decoding terms are worth clarifying because they appear in both peptide and small-molecule results. **Best-of-*N* (BoN)** means sampling *N* complete candidates, scoring them after generation, and keeping the best ones. **MCTS** uses the scorer during generation: it expands promising partially denoised strings more often, then rolls them out into complete candidates. Put simply, BoN ranks finished samples; MCTS steers where the sampling budget goes before the sample is finished. ### Comparing LiteMol-1 with Frontier Peptide Design Models For each target, we ran RFdiffusion3,[^rfdiffusion3] BoltzGen,[^boltzgen] and O-Design[^odesign] using their recommended structure-design workflows. In practice, we generated approx. 300-500 backbones per target, inverse folded them and re-folded them according to individual model’s protocol. That sweep used two H100s and took more than 24 hours to return the full set of designs. For LiteMol-1, we generated 600 sequences per target (200 sequences across each of 3 seeds) and then folded those sequences under the same evaluation setup. All structure folds in this panel are done with AlphaFold 3.[^af3] The figure below reports ipSAE, ipTM, and pLDDT across targets and models. Filled markers are individual designs; open markers are the per-model mean. ipSAE is shown by default, and the toggle switches to ipTM or pLDDT. *AF3 co-fold ipSAE* For LiteMol-1, sequence generation ran on a single L40S. We treat refolding as a downstream evaluation job, separate from generation. Even with that separation, the overall workflow is five to six times faster. We are not treating speed or throughput as the benchmark here, but the difference is large enough that it should be stated clearly. ### Comparison with other Peptide Sequence Models The comparison above is against structure-based binder designers. Peptide specialists are a different family: they generate sequences directly rather than backbones. Here we compare LiteMol-1 with PepTune,[^peptune] PepInvent,[^pepinvent] and PepMLM[^pepmlm] on chemical validity, non-canonical residue coverage, sequence diversity, and protein-conditioned docking. We generated 3,000 sequences per model. We also ran MCTS on LiteMol-1 to test whether search can improve the raw prior. PepInvent leads on chemical validity, ncAA rate, and Morgan internal diversity, which is expected for a model specialized for non-canonical peptide generation. PepTune covers ncAAs at much lower validity (0.34). LiteMol-1 sits between them: chemically usable (0.77), with an ncAA rate of 0.59. MCTS on the same checkpoint lifts validity to 0.98 and ncAA coverage to 0.85, with a small drop in diversity. Sequence uniqueness is saturated for every arm. *Unconditional generation quality* We also tested whether invalid LiteMol-1 outputs can be repaired after generation. A lightweight in-painting pass can target chemically invalid regions, and validity-aware MCTS scoring can steer the model away from invalid strings during search. Across five PPI targets, these steps lift chemical validity by 12-20 points while keeping the copy-like rate at 0. *Repair lifts validity* In the third peptide evaluation, we docked the designed peptides into five PPI targets using our matched docking protocol.[^unidock] LiteMol-1 is compared with PepMLM,[^pepmlm] PepTune,[^peptune] and the crystal peptide for each complex. We also compare the raw LiteMol-1 prior with repair, best-of-*N*, and MCTS, all using the same underlying checkpoint. These are median docking scores, not a co-fold leaderboard. Repair helps on Keap1, BCL-xL, MCL-1, and PD-L1; MCTS is strongest among the LiteMol-1 arms on MDM2, BCL-xL, and MCL-1. PepTune still edges Keap1 and MCL-1. *Docking strength by target* ### Small Molecule Evaluation The peptide results above test whether LiteMol-1 can work in a modality where sequence validity, non-canonical residues, and co-folding matter. We also evaluate the small-molecule side of the same model: first as an unconditional molecular prior, then as a protein-conditioned generator, and finally under best-of-*N* and MCTS. These are computational metrics, not wet-lab outcomes, and each comparison stays inside the task family it is meant to measure. **Unconditional generation.** We sample small molecules without a protein target and score validity, uniqueness among valid molecules, and internal diversity. GenMol,[^genmol] MolGPT,[^molgpt] and DrugGen[^druggen] are strong small-molecule baselines. LiteMol-1 gives up a few points of raw validity, but keeps uniqueness saturated while staying highly diverse. With MCTS on the same checkpoint, validity reaches 1.00. *Small-molecule generation quality* Drug-likeness is a second axis. Mean QED versus internal diversity shows the trade-off more directly: GenMol sits highest on QED but in a narrower region, DrugGen spreads out but loses drug-likeness, and LiteMol-1 stays high on both axes. *QED vs diversity* **Protein-conditioned generation.** Next we evaluate held-out protein targets under a shared docking protocol. Pocket2Mol,[^pocket2mol] PocketXMol,[^pocketxmol] TamGen,[^tamgen] DrugGen,[^druggen] and ProtoBind-Diff[^protobinddiff] are included as context. This is not one universal task for every model: some baselines are structure-conditioned, some are sequence-conditioned, and LiteMol-1 is evaluated as an editable SMILES generator. Docking is reported as `-median docking`, so taller bars are better. LiteMol-1 and its search variants clear most pocket baselines, while ProtoBind-Diff is the closest competitor on raw median docking. *Docking strength* ## Case Studies We divide each experiment into two parts. The first is a **high-level agentic campaign**, where the agent studies the target, goes through the relevant literature, understands what kind of molecule should be designed, and decides which scoring functions and constraints are useful for that particular problem. The second is the **Lite-Mol campaign** itself. Here, LiteMol-1 continuously samples candidates, steers generation using the selected scoring functions, and edits or in-paints molecules along the way. This can mean repairing an invalid SMILES, preserving part of an existing molecule while replacing another region, or iteratively modifying functional groups to improve a candidate. Once a campaign finishes, we take the top-*N* candidates—typically around 20 for peptides and up to 100 for small molecules—and move them into more expensive evaluation. Depending on the modality, this includes co-folding for peptides[^protenix] and molecular dynamics simulations for both peptides and small molecules. We also track the optimization trajectory throughout the campaign and look for a **cumulative net-positive improvement** over successive iterations. Because these are multi-objective campaigns, the winner is defined by the full optimization stack, not by one cherry-picked metric. For a fixed target, scorer set, constraints, and search budget, the goal is to objectively find the best molecule under that design specification: the candidate that best satisfies the properties we asked the campaign to optimize. The following case studies cover several targets and molecular modalities: **MDM2** with linear and non-canonical peptides, **Nav1.8** and **OX1R** with small molecules, **FXIa / Milvexian** with macrocycles and cyclic peptides, and **BRD4 / MZ1** with PROTACs. ### MDM2 **Target Description:** MDM2 is a RING-type E3 ubiquitin ligase and a major negative regulator of the tumour suppressor **p53**. Its N-terminal domain binds the p53 transactivation domain, inhibiting p53-dependent transcription, while its C-terminal RING domain mediates p53 ubiquitination and promotes its degradation. In the 2.6 Å crystal structure of the canonical MDM2–p53 complex (PDB `1YCR`), p53 residues 18–26 adopt an amphipathic α-helix within a deep hydrophobic cleft on MDM2. Three key p53 residues, **F19**, **W23**, and **L26**, project into complementary hydrophobic pockets.[^mdm2-1ycr] This well-defined interaction makes MDM2–p53 a useful benchmark for peptide design. **Therapeutic Rationale:** Amplification or overexpression of MDM2 is a common mechanism by which tumours suppress otherwise functional p53 signalling.[^mdm2-amp] Disrupting the MDM2–p53 interaction can therefore restore p53 activity without directly modifying p53 itself. Peptides, particularly constrained or stapled helices, are a natural modality for this interface because they can reproduce the geometry of the native p53 transactivation helix while occupying the same hydrophobic cleft. We keep two reference co-folds on the same MDM2 construct, under the same Protenix protocol used later for every designed peptide. **Native p53 (`1YCR`).** This is the main reference for the experiment.[^mdm2-1ycr] The crystal contains the MDM2 binding domain together with the 15-residue p53 peptide `SQETFSDLWKLLPEN`. It is the native interaction we want the model to recover: a helical pose, with hydrophobic side chains in the same pockets occupied by p53 residues **F19**, **W23**, and **L26**. Under this protocol, native p53 scores **0.915 ipTM**. Every generated peptide below is co-folded the same way and compared against this number. **Stapled peptide (`5AFG`).** This is a 1.90 Å structure of MDM2 bound to a chemically stapled helix, built from the parent sequence `LTFAEYWAQLAS`.[^mdm2-5afg] A staple is a covalent linker that locks the peptide into a helix, so the bound shape does not have to be held together by sequence alone. That is why it scores **0.959** here, well above native p53, and why it is a harder bar for the constrained designs later. A disulfide or a head-to-tail cycle on one of our peptides is not the same chemistry. `5AFG` is never used as a generation seed. ***Two matched co-folds, colored by pLDDT.** Native p53 is the interaction we are trying to recover. The stapled `5AFG` peptide is a chemically locked helix used only as a harder comparison. Switch tabs; drag to rotate.* To test this, we ran two generation arms: **linear peptides** and **peptides containing non-canonical amino acids**, each with Pareto-guided MCTS[^peptune] and a matched-budget best-of-*N* baseline. Generation was conditioned on MDM2, with peptide length restricted to 8–16 residues. The search optimized several fast objectives simultaneously, including an MDM2-binding predictor, pharmacophore recovery, solubility, hemolysis, nonfouling, and toxicity. Structure Prediction model such as: Protenix[^protenix] was deliberately kept outside the search loop and used only for final structural evaluation. During search we only used those fast scores. Each round proposes new peptides, scores them, and keeps the ones that look promising. Protenix co-folding is applied later, on a smaller set. *Search progress on linear peptides* | Method | Peptides Scored | Passed Filters | Best Search Score | | --- | ---: | ---: | ---: | | Search, linear | 2048 | 99 | 0.956 | | Best-of-*N*, linear | 2048 | 137 | 0.956 | | Search, ncAA | 2048 | 73 | 0.808 | | Best-of-*N*, ncAA | 2048 | 34 | 0.838 | \*Passed Filters is the number of unique peptides that passed every validation check and property window used during search. After generation and filtering, we selected **50 candidates** for co-folding. Every candidate was evaluated against the same MDM2 construct used for the p53 reference, with the same MSA, templates, Protenix configuration, and sampling protocol. *Co-folded peptides compared with native p53* An interesting result was that the peptide with the highest search-time reward was **not** the best structure after co-folding. The fast scoring models were effective at pushing the search toward plausible MDM2-binding peptides, but structural evaluation still changed the final ranking. *ipTM, PTM, and ranking score against native p53* ***Matched co-folds on the same MDM2 construct, colored by pLDDT.** Left: native p53 (ipTM 0.915). Right: the best LiteMol-1 peptide (ipTM 0.935). Drag either tile.* | | Sequence | Search score | ipTM | PTM | Ranking score | Similarity to the best known binder | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Best search peptide | `SQWSEDWLRLASQESP` | 0.956 | 0.888 | 0.888 | 0.888 | 0.31 | | Best co-folded peptide | `GADLEAAFKDEYFEL` | — | **0.935** | 0.898 | **0.928** | **0.07** | Several de novo peptides reached the same ipTM regime as the native p53 peptide, with the strongest design exceeding the `1YCR` reference while remaining highly dissimilar in sequence to the best known binder. | | p53 | Rank 0 | Rank 1 | Rank 2 | Rank 3 | Rank 4 | Rank 5 | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | Sequence | `SQETFSDLWKLLPEN` | `GADLEAAFKDEYFEL` | `FFTXLLXTS` | `SKEWLNTHQWGLELGL` | `SSTSNWKAFHHNLLLE` | `SYLSEEFNSQYTDLLT` | `PAIILLLYA` | | ipTM | 0.915 | **0.935** | 0.926 | 0.919 | 0.911 | 0.908 | 0.905 | | PTM | 0.893 | 0.898 | 0.917 | 0.900 | 0.895 | 0.891 | 0.899 | | Ranking score | 0.910 | **0.928** | 0.924 | 0.915 | 0.908 | 0.904 | 0.904 | | vs native p53 (ipTM) | 0 | **+0.020** | +0.011 | +0.004 | −0.004 | −0.007 | −0.010 | | Similarity to the best known binder | 1.00 | **0.07** | 0.08 | 0.19 | 0.06 | 0.12 | 0.08 | | Pocket occupancy | 0.933 | **0.933** | 0.933 | 0.800 | 0.800 | 0.867 | 0.733 | | Hydrophobic contacts in pocket | 5 | **5** | ligand | 5 | 4 | 6 | 5 | | Length | 15 | 15 | 9 | 16 | 16 | 16 | 9 | | Method | crystal | Search, linear | Best-of-*N*, ncAA | Best-of-*N*, linear | Search, linear | Best-of-*N*, linear | Search, ncAA | ***The next five peptides**, same Protenix protocol, colored by pLDDT. `FFTXLLXTS` is an ncAA co-fold; the others are ordinary peptide chains. Drag to rotate.* The strongest linear peptide recovered the same hydrophobic cleft used by p53 without reproducing the p53 sequence. Its hydrophobic side chains occupy the F19/W23/L26 region of the interface, suggesting that the search recovered the **interaction pattern** rather than simply rediscovering a known peptide. We then locked those peptides and added constraints. A first pass used head-to-tail cyclisation and disulfides. Most got worse. One disulfide, `PCIILCLYA`, rose from 0.905 to **0.932**. That still left the actual `5AFG` chemistry unused. So we put the same **P07** staple from `5AFG` onto our own sequences: two solvent-facing alanines, seven residues apart, linked through the bis-triazolyl staple. Pocket residues and the two staple sites were kept. The remaining solvent-facing positions were regenerated. Nineteen variants were co-folded on the same MDM2 construct: **14 stapled**, **5** head-to-tail. All 19 stayed in the pocket. **8** beat their linear parent. **0** beat `5AFG` (0.959). The best staple is the Rank 0 peptide with alanines at positions 3 and 10: `GADLEAAFKDEYFEL` → `GAALEAAFKAEYFEL`. ipTM moves from **0.935** to **0.940**. A painted variant of the same staple, `GSALEASFKAEYFEL`, is **0.939**. The largest lift is on `SYLSEEFNSQYTDLLT`, which goes from 0.908 to **0.932** once stapled. *Stapled LiteMol-1 peptides against native p53 and 5AFG* ***Same MDM2 construct, colored by pLDDT.** Linear Rank 0, the P07-stapled Rank 0 peptide, and the 5AFG stapled helix. Switch tabs; drag to rotate.* | | Sequence | Kind | ipTM | Ranking score | vs parent | vs `5AFG` | | --- | --- | --- | ---: | ---: | ---: | ---: | | Rank 0 | `GAALEAAFKAEYFEL` | Stapled parent | **0.940** | **0.936** | +0.005 | −0.019 | | Rank 1 | `GSALEASFKAEYFEL` | Painted + staple | 0.939 | 0.935 | +0.003 | −0.020 | | Rank 2 | `SYLSAEFNSQYADLLT` | Stapled parent | 0.932 | 0.930 | +0.024 | −0.027 | | Rank 3 | `SYLAAEFHAQYAKLLT` | Painted + staple | 0.931 | 0.929 | +0.023 | −0.028 | | Rank 4 | `SYLSAEFNAHYADLLT` | Painted + staple | 0.930 | 0.929 | +0.022 | −0.029 | ***The five stapled peptides above**, same protocol, colored by pLDDT. Switch tabs; drag to rotate.* Blank protein-conditioned search found the peptides. The P07 staple then locked the helix on those sequences. Rank 0 ends at **0.940**, above native p53 and above its linear parent, and still short of the `5AFG` crystal peptide. *All ipTM, ranking-score, and pocket-occupancy values here are protocol-relative computational proxies rather than experimental binding measurements. MDMX selectivity, cellular p53 reactivation, and pose stability under molecular dynamics were not evaluated in this campaign.* ### Nav1.8 Nav1.8 is a voltage-gated sodium channel encoded by **SCN10A** and expressed mainly in peripheral sensory neurons.[^nav18-scn10a] It opens at more depolarized voltages than the other neuronal sodium channels, and it helps carry the current that keeps pain-sensing neurons firing. The channel has four voltage-sensing domains. Small-molecule analgesics that bind **voltage-sensing domain 2 (VSD2)** can hold the channel in a non-conducting state without blocking the central pore, which is how the approved drug **suzetrigine** (VX-548 / Journavx) works. **Therapeutic Rationale:** Suzetrigine was approved in January 2025 after showing that selective Nav1.8 inhibition can reduce acute pain.[^suzetrigine] The VSD2 site we dock into was mapped by a **peptide toxin**. The structure we use is PDB `9DBN`: a 2.76 Å cryo-EM complex of human Nav1.8 bound to the tarantula peptide **Protoxin-I**, which sits on VSD2 and shifts activation.[^nav18-9dbn] We crop to the extracellular VSD2 cleft (residues 680–770) and dock into that box. The Nav1.8-unique motif in this site is **V746 / K748 / K749 / G750 / S751** (KKGS). Contacts with this motif provide a check for selectivity. A-803467 is included as a pore-binder control from a separate Nav1.8 structure (PDB `7WE4`), not from `9DBN` and not as a VSD2 ligand.[^nav18-a803467] ***Nav1.8 VSD2 crop** from PDB `9DBN`. This is the receptor used for every docking score below. Drag to rotate.* Below are some important details about the PDB we chose and the filters we used. | Item | Detail | | --- | --- | | PDB | `9DBN` · Nav1.8 + Protoxin-I (VSD2 crop) | | Clinical comparator | Suzetrigine (VX-548 / Journavx) | | Docking control | Same molecule, **−5.75 kcal/mol** on this protocol | | Selectivity motif | V746 · K748 · K749 · G750 · S751 | | Filters used | Morgan Tanimoto ≥ 0.5 to 917 known Nav1.8 compounds is excluded | Suzetrigine serves as the clinical comparator: an approved analgesic that selectively inhibits Nav1.8 through VSD2. The **−5.75 kcal/mol** docking score was obtained under this protocol within the peptide-defined docking box, not as a clinical efficacy measure. Because the approved comparator is not the top-ranked docked molecule, docking scores are interpreted as a prioritization metric. The agent first ran **13 short pilots** (16 search rounds each) to decide what the campaign should optimize. Extra protein guidance, shorter or longer SMILES, hotter sampling, and dropping medchem filters did not help. Taking docking **out** of the search loop produced a lucky single best (−7.73) with a weaker top-5. The setting that held up was **docking plus novelty**, with medchem alerts kept on, docking inside the loop, and length 28. We then ran two matched 60-round searches, 8 candidates per round (**480** molecules each): one protein-conditioned on the VSD2 sequence, and one unconditional. *Search progress on Nav1.8* | Arm | Molecules scored | Passed filters | Unique docked | Best in-loop | Mean top-5 | | --- | ---: | ---: | ---: | ---: | ---: | | Search, conditioned | 480 | 221 | 172 | −6.85 | **−6.75** | | Search, unconditional | 480 | 425 | 380 | **−7.09** | −6.67 | \*Passed filters is the number of unique molecules that passed every validation check and property window used during search. After search, the top **30** conditioned molecules were re-docked more carefully (20 poses). Nine of the thirty touched the KKGS/V746 motif. Unconditional confirmation reached **−7.27** but it had nitro groups, which inflate docking scores and are a mutagenicity liability, so we reject that molecule on chemistry. The next chart docks those results against known molecules in the **same VSD2 box**, with the same software, so the numbers can be compared: - **Generated molecules** — after more careful docking, the strongest generated molecule reached a score of **−6.85**. The molecule we pick scored **−6.68**: a bit weaker on dock, kept because of the predicted-safety picture later. - **LTGO-33** — a published Nav1.8 research ligand that binds VSD2. Not an approved drug; a known-site control. - **A-803467** — another known Nav1.8 inhibitor, but it binds the pore (the ion hole), not VSD2. It comes from a separate Nav1.8 pore-binder structure (PDB `7WE4`), and a strong score in the VSD2 box is not proof that the molecule is in the right place. - **Suzetrigine** — the approved VSD2 pain drug. It has the weakest docking score here. *Docking strength across five molecules* Same lesson as MDM2, in a different metric: the molecule that won the computer score is not automatically the one you take forward. *Re-docked molecules vs the selectivity motif* We do not recommend the generated molecule with the best docking score. The one we keep is a chloroindole with a benzoic acid (−6.68 kcal/mol). It makes two contacts with the Nav1.8-unique motif. Predicted liver injury (DILI) is flagged; suzetrigine has the same flag. Predicted hERG risk and predicted brain entry are both lower than suzetrigine, which matters for a peripheral pain target. *Predicted liabilities against suzetrigine* ***Same VSD2 crop and box.** Left: suzetrigine (−5.75 kcal/mol). Right: the recommended molecule (−6.68 kcal/mol). Drag either tile.* | | Suzetrigine | Best dock | **Recommended** | | --- | ---: | ---: | ---: | | Docking strength (kcal/mol) | −5.75 | −6.85 | **−6.68** | | Motif contacts | 6 | 0 | **2** | | ADMET flags | 1 (DILI) | 2 | **1 (DILI)** | | hERG (↓) | 0.43 | 0.74 | **0.38** | | BBB (↓) | 0.49 | 0.38 | **0.18** | | QED | 0.65 | 0.51 | 0.50 | | LogP | 3.55 | 4.13 | 5.58 | | Recommended | reference | no | **yes** | The recommended molecule is new relative to known Nav1.8 ligands (Tanimoto stays below 0.5). It is not a better oral package: suzetrigine still wins on QED, LogP, predicted absorption, and PAMPA. What we have is a starting point in the same predicted-liability neighborhood as the approved drug, not a replacement for it. ***Suzetrigine and five re-docked generated molecules**, same VSD2 crop and box. The recommended molecule is the chloroindole. Switch tabs; drag to rotate.* The search was conditioned on the VSD2 sequence and did not use suzetrigine as a template. Docking had to be part of the search for the scores to improve. We chose the recommended molecule on predicted safety, not on the best docking score. It is dissimilar to known Nav1.8 ligands, contacts the selectivity motif, and shares suzetrigine’s predicted DILI flag, with lower predicted hERG risk and brain entry, but poorer oral physicochemical properties. We then ran **100 ns** of all-atom OpenMM on the recommended molecule in the VSD2 crop, as five 20 ns continuations. ***Recommended molecule in the VSD2 crop.** All-atom OpenMM trajectory, 100 ns.* The ligand stays in the site. Ligand RMSD stays in a 2–3 Å band for the full trajectory. Around 40–50 ns the backbone rearranges: protein RMSD peaks near 4.5 Å and the radius of gyration briefly expands, then both settle. After that rearrangement the ligand moves closer to the pocket, from ~4.5 Å to ~2.5 Å ligand–protein distance, and holds there for the remaining ~40 ns. Ligand–protein hydrogen bonds drop from three to one; what remains is a stable core of ten residues that stay in contact for the entire 100 ns. This is pose retention on a cropped site, not a binding free energy, and not experimental activity. ![**100 ns OpenMM on the recommended molecule.** Ligand RMSD stays in a 2–3 Å band. After a 40–50 ns rearrangement the ligand settles closer to the pocket and remains there.](data/nav18_litemol_case_study/md_100ns_dashboard.png) ### FXIa Factor XIa (FXIa) is the active form of factor XI. It sits in the intrinsic clotting pathway and amplifies thrombin once clotting has already begun.[^fxia] The catalytic domain is a trypsin-like protease. **Asp189** sits at the bottom of the specificity pocket. **His57 / Asp102 / Ser195** is the catalytic triad. **Trp215** sits on the wall of the specificity pocket. **Therapeutic Rationale:** Blocking FXIa can reduce pathological clots while leaving enough of the tissue-factor pathway for ordinary haemostasis. **Milvexian** is a reversible oral FXIa inhibitor: a **non-peptide 12-membered macrocycle**.[^milvexian] We use it as the clinical comparator, and as a test of whether LiteMol-1 can make a new ring for this pocket. The reference structure is PDB `7MBO`, in which milvexian binds the FXIa catalytic domain. Milvexian extends one end of the molecule into the primary specificity pocket, with Asp189 at its base. A credible pose should reproduce this placement. ***FXIa catalytic domain** (PDB `7MBO`). This is the receptor used for every score below. Drag to rotate.* This is a different design problem from MDM2 and Nav1.8. The clinical molecule is a synthetic macrocycle. We asked LiteMol-1 for two classes on the same target: **non-peptide macrocycles** and **cyclic peptides**. | Item | Detail | | --- | --- | | PDB | `7MBO` · FXIa catalytic domain + milvexian | | Clinical comparator | Milvexian | | Docking control | Same molecule, crystal redock **−9.83 kcal/mol** | | Pocket residues | His57 · Asp189 · Lys192 · Ser195 · Trp215 | | Filters used | Tanimoto ≥ 0.85 to milvexian is excluded | Milvexian is the clinical comparator: an oral FXIa inhibitor. The **−9.83 kcal/mol** is this protocol's crystal redock of that same drug, not a clinical number. We first ran short pilots to decide how to generate. LiteMol-1 writes a SMILES string of a chosen length. That length decides the chemistry: the shortest setting made 12–15 membered synthetic rings, and longer settings drifted toward natural-product-like macrolides. For cyclic peptides, a shorter setting produced nothing that passed filters. Length 64 did. Each bar represents a pilot batch: **8 molecules** for the three shorter settings and **12** for the longest. **Chemically valid** means the structure parsed successfully. **Passed filters** additionally requires a 12–24-membered ring, non-peptidic character, novelty from milvexian, and basic medicinal-chemistry criteria. Bar height is the passing fraction, not a binding score. *Generation length controls ring size* We then generated both classes without using milvexian as a template: 192 macrocycles and 192 cyclic peptides. | Class | Molecules scored | Passed filters | Unique | | --- | ---: | ---: | ---: | | Macrocycles | 192 | 107 | 25 | | Cyclic peptides | 192 | 137 | 21 | \*Passed filters is the number of unique molecules that passed every validation check and property window used during generation. After generation we co-folded a smaller set on the same FXIa construct, with the same protocol used for milvexian. The molecule we keep is a **15-membered brominated macrocycle**. Interface confidence is **0.983**, next to milvexian at **0.989**. Bromine sits on Asp189. Similarity to milvexian is **0.13**. The chart below is that comparison. We also show a cyclic peptide we generated, another generated macrocycle, and a macrolide decoy, so you can see where the pick sits. - **Milvexian**: the clinical molecule. - **Recommended**: 15-membered brominated ring. Bromine on Asp189. - **Best co-fold macro**: strongest generated pose besides the pick. - **Cyclic peptide**: best peptide co-fold, `TRDYTTCCG`. - **Macrolide decoy**: not a designed molecule. High confidence, no Asp189 contact. *Matched co-folds on FXIa* We then docked the macros in the same box, with the same software. The recommended molecule scores **−6.37 kcal/mol**. The pose holds across repeats (RMSD **0.15–0.22 Å**). Bromine stays 3.80–3.91 Å from Asp189 every time. *Predicted permeability versus docking* The next chart docks those results in the **same box**: - **Recommended**: 15-membered brominated ring. Pose holds. Bromine on Asp189. - **Nitro 18-membered**: strongest generated dock. The pose moves (RMSD 8.57 Å). - **Best co-fold macro**: looks good in the structure model. In docking it sits 7.5–9.6 Å from Asp189, so we do not take it. - **Macrolide decoy**: not a designed molecule. - **Milvexian**: clinical molecule, crystal redock **−9.83 kcal/mol**. *Docking strength across five molecules* Same lesson as Nav1.8: we do not take the strongest generated dock. We take the ring that is new, pose-stable, and on Asp189. ***Matched co-folds on the same FXIa construct, colored by pLDDT.** Left: milvexian (ipTM 0.989). Right: the recommended 15-membered brominated macrocycle (ipTM 0.983). Drag either tile.* | | Milvexian | Recommended | Best co-fold macro | Cyclic peptide | | --- | ---: | ---: | ---: | ---: | | Chemistry | 12-membered clinical | **15-membered brominated** | generated macrocycle | TRDYTTCCG | | Co-fold ipTM | 0.989 | **0.983** | 0.928 | 0.910 | | Docking (kcal/mol) | −9.83 | **−6.37** | −4.44 | n/a | | Pose RMSD (Å) | 0.38 | **0.15–0.22** | 0.59–0.97 | n/a | | Asp189 | reference | **bromine** | misses in dock | Arg/Lys in co-fold | | Similarity to milvexian | 1.00 | **0.13** | 0.36 | n/a | | Recommended | reference | **yes** | no | no | ***Four co-folds, same protocol, colored by pLDDT.** The recommended molecule is the 15-membered brominated ring. Switch tabs; drag to rotate.* LiteMol-1 made a new synthetic ring for this pocket without milvexian as a template. The 15-membered brominated macrocycle co-folds next to the clinical molecule, docks with a pose that holds, and puts bromine on Asp189. That is the starting point. It is not a better milvexian. ### BRD4 BRD4 is a BET-family bromodomain protein. Its tandem bromodomains read acetylated lysines on histones and transcription factors, which is why it is a well-studied oncology target.[^jq1] **Therapeutic Rationale:** A PROTAC is a single heterobifunctional molecule comprising a target-binding ligand, an E3-ligase ligand, and a connecting linker. By bringing the target and E3 ligase into a ternary complex, it promotes target ubiquitination and subsequent proteasomal degradation. **MZ1** is the crystallographic example: the BRD4 warhead **JQ1** joined to the VHL ligand **VH032**.[^mz1] The ternary complex is PDB `5T35`.[^mz1-5t35] ***BRD4 BD2–VHL ternary** (PDB `5T35`). MZ1 bridges the bromodomain and VHL. Drag to rotate.* LiteMol-1 is not good at generating full PROTACs today, and we should say that plainly. Asked to write the whole chimera, most samples are invalid (25% valid in a matched probe). The one molecule that looked like a PROTAC was a VHL recruiter glued to a kinase-like fragment, not a BRD4 warhead. Conditioning on BRD4 does not fix this: it pulls the model back toward MZ1. Holding both ends of MZ1 fixed and rewriting the middle is linker variation (Tanimoto **0.84**). Holding only VH032 fixed recovers JQ1 (Tanimoto **0.91**). That is analogue search. For an agent, it is also the useful behaviour. A PROTAC is three pieces. The agent can keep the E3 ligand that already works, keep a warhead that already docks, and ask LiteMol-1 to change only the span that should change. The model is weak as an end-to-end PROTAC generator and strong as an editor of one piece at a time. *Full PROTAC generation recovers MZ1. Piecewise editing is the usable mode.* If the goal is a new BRD4 warhead rather than a new linker, that is what we did. LiteMol-1 generated a small molecule conditioned on BRD4. The agent installed a coupling handle and attached a known linker and VH032. Novelty was scored on the warhead, not on the assembled chimera: any molecule that still contains VH032 looks similar to MZ1. The warhead is a benzofuran–imidazothiazole–pyrrole, not JQ1. It docks into the same BRD4 BD2 box at **−9.74 kcal/mol**, against JQ1 at **−8.34**. It is larger (40 versus 28 heavy atoms), so ligand efficiency is worse (**0.244** versus **0.298**). Similarity to JQ1 is **0.17**. Two stronger-looking docks were dropped because they were analogues of mivebresib. We treat MZ1-like *chimeras* as analogue design. We do not treat a known BRD4 chemotype as a new warhead. *Warhead docking against JQ1* ***Matched ternary co-folds, colored by pLDDT.** Left: MZ1 (ipTM 0.655). Right: the benzofuran warhead assembled onto VH032 (ipTM 0.644). Drag either tile.* | | JQ1 | Benzofuran warhead | Assembled chimera | MZ1 | | --- | ---: | ---: | ---: | ---: | | Role | reference warhead | designed warhead | agent-assembled PROTAC | crystal PROTAC | | Vina (kcal/mol) | −8.34 | **−9.74** | — | — | | Ligand efficiency | **0.298** | 0.244 | — | — | | Similarity to JQ1 | 1.00 | **0.17** | — | — | | Ternary ipTM | — | — | 0.644 | 0.655 | The assembled chimera sits **0.011** ipTM below MZ1. That number is not evidence of a degrader. A phenyl ring attached to VH032 lands in the same ternary confidence band. PDB `5T35` is already in the co-folding training neighbourhood, so the complex shape is easy to recover. The claim stops at a dockable, non-JQ1 warhead with a handle an agent can couple. LiteMol-1 does not yet write a PROTAC as one object. Used with an agent, it does not have to. Keep the E3 ligand, rewrite the linker, or generate the warhead as a small molecule and assemble. That is the honest advantage. *Docking scores and ternary ipTM are protocol-relative computational proxies rather than experimental binding or degradation measurements. BRD4 binding, VHL recruitment, ternary cooperativity, and cellular degradation were not evaluated in this campaign.* ## Representation We took the same **4,046 molecules** and embedded them using LiteMol-1, Uni-Mol (a 3D conformation-based molecular encoder)[^unimol] and Morgan circular fingerprints using radius 2 and 2,048 bits.[^ecfp] Each representation was projected separately into two dimensions with UMAP. Most points are unlabelled. Selected chemical and therapeutic classes are coloured as visual guides; these projections are not clustering tests or quantitative rankings of the representations. ![**Same molecules, three spaces.** LiteMol-1 splits known chemotypes into islands (steroids, β-lactams, kinase inhibitors, benzodiazepines, and so on). Uni-Mol is a smoother cloud. Morgan fragments into many small fingerprint islands.](data/litemol_embedding_observations/figures/01_umap_chemotype.png) In these projections, LiteMol-1 and Morgan both place related scaffolds in compact regions, while Uni-Mol appears more continuous. This is a qualitative description of these particular projections, not evidence that one representation is intrinsically better. The next figure uses the same UMAP coordinates, now coloured by QED, a quantitative estimate of drug-likeness. Dark green indicates higher QED and pale yellow lower QED. LiteMol-1 shows a visible gradient across the projection, Uni-Mol shows a weaker pattern, and Morgan shows more localized patches. UMAP axes are arbitrary, so the direction of a gradient has no intrinsic meaning. ![**QED on the same maps.** Pale yellow is less drug-like; dark green is more. LiteMol-1 shows a clear top-to-bottom gradient. Uni-Mol is weaker. Morgan has no global trend.](data/litemol_embedding_observations/figures/02_umap_qed.png) UMAPs are qualitative visualizations. To quantify how accessible basic chemical information is, we froze each representation and fitted ridge regressions under five-fold cross-validation. Higher $R^2$ means that a property is more linearly recoverable from the representation. LiteMol-1 produced the highest $R^2$ values for all five tested properties. Molecular weight, logP, and ring count were strongly recoverable (0.93, 0.91, and 0.95). Topological polar surface area was similar between LiteMol-1 (0.88) and Uni-Mol (0.85). QED showed the largest observed difference: 0.65 for LiteMol-1, 0.25 for Uni-Mol, and 0.07 for Morgan. These values measure linear accessibility under this dataset and split; they do not measure the total information contained in each representation. | Property | LiteMol-1 | Uni-Mol | Morgan | | --- | ---: | ---: | ---: | | Molecular weight | **0.93** | 0.79 | 0.49 | | logP | **0.91** | 0.69 | 0.52 | | Polar surface area | **0.88** | 0.85 | 0.59 | | Ring count | **0.95** | 0.65 | 0.64 | | Drug-likeness (QED) | **0.65** | 0.25 | 0.07 | ![**Linear readout of chemistry.** 5-fold ridge $R^2$. The generator is frozen. Polar surface area is the closest race. Ring count and QED are the largest gaps.](data/litemol_embedding_observations/figures/06_property_probe.png) Known ligand series also occupy compact regions in LiteMol-1 space. BindingDB actives for CCR9, GRIK1, and HCRTR1 separate from property-matched DUD-E decoys. Because DUD-E decoys are deliberately topologically dissimilar from known ligands, this demonstrates organization by ligand-series chemistry rather than target-binding prediction. HCRTR1 is the largest series in the panel, so its cloud also reflects the larger sample size. We next tested whether a small classifier trained on the same frozen representation could predict peptide labels, using the official PeptiVerse splits. The reported values are AUROC. Toxicity is high for both LiteMol-1 and Morgan (0.92 versus 0.90). The larger difference is in hemolysis (0.80 versus 0.68), while non-fouling performance is similar. ![**Peptide labels, frozen readout.** Official PeptiVerse splits. Hemolysis is where LiteMol-1 pulls away from Morgan.](data/litemol_embedding_observations/figures/08_peptide_auroc.png) On Tox21, using one scaffold split, logistic regression on frozen LiteMol-1 embeddings reaches a macro-AUROC of **0.76**, compared with **0.66** for Morgan. LiteMol-1 is higher on all 12 assays, with SR-p53 the strongest individual task at approximately 0.85. Because this result uses one scaffold split, it characterizes performance under that split. The dashed literature results use different models and evaluation protocols and are shown only for context, not as directly comparable baselines. ![**Tox21, one scaffold split.** Frozen logistic. LiteMol-1 is above Morgan on every assay. SR-p53 is the strongest single task (~0.85).](data/litemol_embedding_observations/figures/09_tox21_auroc.png) We do not fine-tune the encoders in these experiments; only shallow downstream models are trained on the frozen representations. Together, the results show that basic chemical properties and several biological labels are accessible from LiteMol-1 embeddings under the datasets and splits tested. ## Conclusion LiteMol-1 is our first proof that sequence space is a practical interface for molecular design. It is not a claim that sequence models will replace structure models, or that they are better on every target. The result is narrower and more useful: with the right training objective, conditioning, and search loop, an alternate representation can reach competitive outcomes at a fraction of the computational cost, while remaining easy to edit and iterate. It also points to a different way of using generative models. These systems are becoming agent-native. Humans are not good at reading thousands of molecular strings, comparing small sequence edits, or running many careful rounds of filtering. Agents are. LiteMol-1 works best when it is treated as a fast prior inside a loop: propose, repair, score, filter, search, and then verify with heavier structure or simulation tools only when the candidate set is small enough. The next bottleneck is therefore not only generation. It is verification. Docking, ADMET, toxicity, solubility, hemolysis, and other property models are still proxies, and their calibration now matters more than ever. Better scoring functions will make sequence design more controllable: not just producing molecules that look plausible, but steering them toward properties that survive stronger computational checks and, eventually, experimental and pre-clinical validation. In the next release, we will go further in that direction: more wet-lab validation for selected designs, and the next step toward a generalized framework for steerable biomolecule design across modalities. ## Citation ```bibtex @misc{litefold2026litemol1, title = {LiteMol-1: A Multi Molecule Foundation Diffusion Language Model for Agents}, author = {{LiteFold Research}}, year = {2026}, howpublished = {\url{https://litefold.ai/research/litemol1}}, note = {Correspondence: anindya@litefold.ai, cory@litefold.ai} } ``` ## References 1. Coulom R. Efficient selectivity and backup operators in Monte-Carlo tree search. In: Computers and Games (CG 2006). LNCS 4630. Springer; 2007:72-83. 2. Tang S, Zhang Y, Chatterjee P. PepTune: de novo generation of therapeutic peptides with multi-objective-guided discrete diffusion. ICML 2025. PMLR;267. [arXiv](https://arxiv.org/abs/2412.17780) 3. Watson JL, Juergens D, Bennett NR, et al. De novo design of protein structure and function with RFdiffusion. Nature. 2023;620(7976):1089-1100. [DOI](https://doi.org/10.1038/s41586-023-06415-8) 4. Stark H, Faltings F, Choi M, et al. BoltzGen: toward universal binder design. bioRxiv. 2025. [DOI](https://doi.org/10.1101/2025.11.20.689494) 5. ByteDance Seed. PXDesign. 2025. [protenix.github.io/pxdesign](https://protenix.github.io/pxdesign/) 6. Abramson J, Adler J, Dunger J, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630(8016):493-500. [DOI](https://doi.org/10.1038/s41586-024-07487-w) 7. ODesign Team. ODesign: a world model for biomolecular interaction design. 2025. [odesign1.github.io](https://odesign1.github.io/) 8. Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123-1130. [DOI](https://doi.org/10.1126/science.ade2574) 9. Austin J, Johnson DD, Ho J, Tarlow D, van den Berg R. Structured denoising diffusion models in discrete state-spaces. NeurIPS 2021. [arXiv](https://arxiv.org/abs/2107.03006) 10. Sahoo SS, Agrawal A, Arriola M, et al. Simple and effective masked diffusion language models. NeurIPS 2024. [arXiv](https://arxiv.org/abs/2406.07524) 11. Li X, Fourches D. SMILES pair encoding: a data-driven substructure tokenization algorithm for deep learning. J Chem Inf Model. 2021;61(4):1560-1569. [DOI](https://doi.org/10.1021/acs.jcim.0c01127) 12. Ho J, Salimans T. Classifier-free diffusion guidance. NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications. [arXiv](https://arxiv.org/abs/2207.12598) 13. Silver D, Schrittwieser J, Simonyan K, et al. Mastering the game of Go without human knowledge. Nature. 2017;550(7676):354-359. [DOI](https://doi.org/10.1038/nature24270) 14. ByteDance AML Team. Protenix: advancing structure prediction through a comprehensive AlphaFold 3 reproduction. bioRxiv. 2025. [DOI](https://doi.org/10.1101/2025.01.08.631967) 15. Oliner JD, Kinzler KW, Meltzer PS, George DL, Vogelstein B. Amplification of a gene encoding a p53-associated protein in human sarcomas. Nature. 1992;358:80-83. [DOI](https://doi.org/10.1038/358080a0) 16. Kussie PH, Gorina S, Marechal V, et al. Structure of the MDM2 oncoprotein bound to the p53 tumor suppressor transactivation domain. Science. 1996;274(5289):948-953. [DOI](https://doi.org/10.1126/science.274.5289.948) 17. Lau YH, Wu Y, Rossmann M, et al. Double strain-promoted macrocyclization for the rapid selection of cell-active stapled peptides. Angew Chem Int Ed. 2015;54:15410-15413. [DOI](https://doi.org/10.1002/anie.201508416) 18. Akopian AN, Sivilotti L, Wood JN. A tetrodotoxin-resistant voltage-gated sodium channel expressed by sensory neurons. Nature. 1996;379:257-262. [DOI](https://doi.org/10.1038/379257a0) 19. Jones J, Correll DJ, Lechner SM, et al. Selective inhibition of NaV1.8 with VX-548 for acute pain. N Engl J Med. 2023;389(5):393-405. [DOI](https://doi.org/10.1056/NEJMoa2209870) 20. Neumann B, McCarthy S, Gonen S. Structural basis of inhibition of human NaV1.8 by the tarantula venom peptide Protoxin-I. Nat Commun. 2025;16:1459. [DOI](https://doi.org/10.1038/s41467-024-55764-z) 21. Huang X, Jin X, Huang G, et al. Structural basis for high-voltage activation and subtype-specific inhibition of human Nav1.8. Proc Natl Acad Sci USA. 2022;119(30):e2208211119. [DOI](https://doi.org/10.1073/pnas.2208211119) 22. Weitz JI, Chan NC. Advances in antithrombotic therapy. Arterioscler Thromb Vasc Biol. 2019;39(1):7-12. [DOI](https://doi.org/10.1161/ATVBAHA.118.310960) 23. Dilger AK, Pabbisetty KB, Corte JR, et al. Discovery of milvexian, a high-affinity, orally bioavailable inhibitor of factor XIa in clinical studies for antithrombotic therapy. J Med Chem. 2022;65(3):1770-1785. [DOI](https://doi.org/10.1021/acs.jmedchem.1c00613) 24. Filippakopoulos P, Qi J, Picaud S, et al. Selective inhibition of BET bromodomains. Nature. 2010;468:1067-1073. [DOI](https://doi.org/10.1038/nature09504) 25. Zengerle M, Chan KH, Ciulli A. Selective small molecule induced degradation of the BET bromodomain protein BRD4. ACS Chem Biol. 2015;10(8):1770-1777. [DOI](https://doi.org/10.1021/acschembio.5b00216) 26. Gadd MS, Testa A, Lucas X, et al. Structural basis of PROTAC cooperative recognition for selective protein degradation. Nat Chem Biol. 2017;13:514-521. [DOI](https://doi.org/10.1038/nchembio.2329) 27. Zhou G, Gao Z, Ding Q, et al. Uni-Mol: a universal 3D molecular representation learning framework. ICLR 2023. [OpenReview](https://openreview.net/forum?id=6K2RM6wVqKu) 28. Rogers D, Hahn M. Extended-connectivity fingerprints. J Chem Inf Model. 2010;50(5):742-754. [DOI](https://doi.org/10.1021/ci100050t) 29. Butcher J, Krishna R, Mitra R, et al. De novo design of all-atom biomolecular interactions with RFdiffusion3. bioRxiv. 2025. [DOI](https://doi.org/10.1101/2025.09.18.676967) 30. Chen T, Dumas M, Watson R, et al. PepMLM: target sequence-conditioned generation of therapeutic peptide binders via span masked language modeling. Nat Biotechnol. 2025. [DOI](https://doi.org/10.1038/s41587-025-02761-2) · [arXiv](https://arxiv.org/abs/2310.03842) 31. Geylan G, Janet JP, Tibo A, et al. PepINVENT: generative peptide design beyond natural amino acids. Chem Sci. 2025. [DOI](https://doi.org/10.1039/d4sc07642g) · [arXiv](https://arxiv.org/abs/2409.14040) 32. Yu Y, Cai C, Wang J, et al. Uni-Dock: GPU-accelerated docking enables ultralarge virtual screening. J Chem Theory Comput. 2023;19(11):3336-3345. [DOI](https://doi.org/10.1021/acs.jctc.2c01145) 33. Lee S, Kreis K, Veccham SP, et al. GenMol: a drug discovery generalist with discrete diffusion. ICML 2025; PMLR 267:33205-33226. [PMLR](https://proceedings.mlr.press/v267/lee25o.html) 34. Bagal V, Aggarwal R, Vinod PK, Priyakumar UD. MolGPT: molecular generation using a transformer-decoder model. J Chem Inf Model. 2022;62(9):2064-2076. [DOI](https://doi.org/10.1021/acs.jcim.1c00600) 35. Sheikholeslami M, Mazrouei N, Gheisari Y, et al. DrugGen enhances drug discovery with large language models and reinforcement learning. Sci Rep. 2025;15:13445. [DOI](https://doi.org/10.1038/s41598-025-98629-1) 36. Peng X, Luo S, Guan J, et al. Pocket2Mol: efficient molecular sampling based on 3D protein pockets. ICML 2022; PMLR 162. [arXiv](https://arxiv.org/abs/2205.07249) 37. Peng X, Guo R, Guo F, et al. Unified modeling of 3D molecular generation via atomic interactions with PocketXMol. Cell. 2026. [DOI](https://doi.org/10.1016/j.cell.2026.01.003) 38. Wu K, Xia Y, Deng P, et al. TamGen: drug design with target-aware molecule generation through a chemical language model. Nat Commun. 2024;15:9360. [DOI](https://doi.org/10.1038/s41467-024-53632-4) 39. Mistryukova L, Manuilov V, Avchaciov K, Fedichev P. ProtoBind-Diff: a structure-free diffusion language model for protein sequence-conditioned ligand design. bioRxiv. 2025. [DOI](https://doi.org/10.1101/2025.06.16.659955) [^mcts]: Coulom R. "Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search." *Computers and Games* 2006. The search operators PepTune later adapted as Monte Carlo Tree Guidance over masked diffusion. [DOI](https://doi.org/10.1007/978-3-540-75538-8_7) [^peptune]: Tang S, Zhang Y, Chatterjee P. "PepTune: De Novo Generation of Therapeutic Peptides with Multi-Objective-Guided Discrete Diffusion." *ICML* 2025. Peptide-only MDLM with bond-dependent masking, inference-time MCTG, and a Pareto archive over competing objectives. [arXiv](https://arxiv.org/abs/2412.17780) [^pepmlm]: Chen T, Dumas M, Watson R, et al. "PepMLM: Target Sequence-Conditioned Generation of Therapeutic Peptide Binders via Span Masked Language Modeling." *Nat Biotechnol* 2025. ESM-2 fine-tuned to reconstruct a peptide binder from the target sequence alone. [DOI](https://doi.org/10.1038/s41587-025-02761-2) [^pepinvent]: Geylan G, Janet JP, Tibo A, et al. "PepINVENT: generative peptide design beyond natural amino acids." *Chem Sci* 2025. REINVENT-style peptide generator that can emit non-natural amino acids as SMILES. [DOI](https://doi.org/10.1039/d4sc07642g) [^unidock]: Yu Y, Cai C, Wang J, et al. "Uni-Dock: GPU-Accelerated Docking Enables Ultralarge Virtual Screening." *J Chem Theory Comput* 2023;19(11):3336-3345. GPU-accelerated molecular docking used inside our matched docking protocol. [DOI](https://doi.org/10.1021/acs.jctc.2c01145) [^genmol]: Lee S, Kreis K, Veccham SP, et al. "GenMol: A Drug Discovery Generalist with Discrete Diffusion." *ICML* 2025. Discrete diffusion generalist for small-molecule generation and optimization. [PMLR](https://proceedings.mlr.press/v267/lee25o.html) [^molgpt]: Bagal V, Aggarwal R, Vinod PK, Priyakumar UD. "MolGPT: Molecular Generation Using a Transformer-Decoder Model." *J Chem Inf Model* 2022;62(9):2064-2076. GPT-style SMILES generator for de novo molecules. [DOI](https://doi.org/10.1021/acs.jcim.1c00600) [^druggen]: Sheikholeslami M, Mazrouei N, Gheisari Y, et al. "DrugGen enhances drug discovery with large language models and reinforcement learning." *Sci Rep* 2025;15:13445. Protein-sequence-conditioned SMILES generator based on DrugGPT with RL feedback. [DOI](https://doi.org/10.1038/s41598-025-98629-1) [^pocket2mol]: Peng X, Luo S, Guan J, et al. "Pocket2Mol: Efficient Molecular Sampling Based on 3D Protein Pockets." *ICML* 2022. E(3)-equivariant structure-based generator conditioned on protein pockets. [arXiv](https://arxiv.org/abs/2205.07249) [^pocketxmol]: Peng X, Guo R, Guo F, et al. "Unified modeling of 3D molecular generation via atomic interactions with PocketXMol." *Cell* 2026. Pocket-interacting 3D molecular generation foundation model. [DOI](https://doi.org/10.1016/j.cell.2026.01.003) [^tamgen]: Wu K, Xia Y, Deng P, et al. "TamGen: drug design with target-aware molecule generation through a chemical language model." *Nat Commun* 2024;15:9360. Target-aware chemical language model for molecule generation and refinement. [DOI](https://doi.org/10.1038/s41467-024-53632-4) [^protobinddiff]: Mistryukova L, Manuilov V, Avchaciov K, Fedichev P. "ProtoBind-Diff: A Structure-Free Diffusion Language Model for Protein Sequence-Conditioned Ligand Design." *bioRxiv* 2025. Protein-sequence-conditioned masked diffusion model for ligand design. [DOI](https://doi.org/10.1101/2025.06.16.659955) [^rfdiffusion]: Watson JL, Juergens D, Bennett NR, et al. "De novo design of protein structure and function with RFdiffusion." *Nature* 2023;620(7976):1089–1100. Coordinate-space binder design that made structure the default generative interface. [DOI](https://doi.org/10.1038/s41586-023-06415-8) [^rfdiffusion3]: Butcher J, Krishna R, Mitra R, et al. "De novo Design of All-atom Biomolecular Interactions with RFdiffusion3." *bioRxiv* 2025. All-atom diffusion for proteins in the context of ligands, nucleic acids, and other non-protein atoms. [DOI](https://doi.org/10.1101/2025.09.18.676967) [^boltzgen]: Stark H, Faltings F, Choi M, et al. "BoltzGen: Toward Universal Binder Design." *bioRxiv* 2025. All-atom generative design unified with structure prediction. [DOI](https://doi.org/10.1101/2025.11.20.689494) [^pxdesign]: ByteDance Seed. "PXDesign." 2025. Structure-based protein-binder design built around Protenix. [protenix.github.io/pxdesign](https://protenix.github.io/pxdesign/) [^af3]: Abramson J, Adler J, Dunger J, et al. "Accurate structure prediction of biomolecular interactions with AlphaFold 3." *Nature* 2024;630(8016):493–500. The cofolding model that made biomolecular complexes a default 3D prior. [DOI](https://doi.org/10.1038/s41586-024-07487-w) [^odesign]: ODesign Team. "ODesign: A World Model for Biomolecular Interaction Design." 2025. All-atom generative world model for biomolecular interaction design. [odesign1.github.io](https://odesign1.github.io/) [^design-scale]: The BoltzGen pipeline recommends generating 10,000–60,000 designs (`--num_designs`) before filtering to a small `--budget`. [GitHub](https://github.com/HannesStark/boltzgen). ODesign reports large-scale inference over 10k+ designs (typically 100 seeds × 20 samples, scaled further in campaigns). [pipeline](https://github.com/OTeam-AI4S/ODesign-pipeline) [^esm2]: Lin Z, Akin H, Rao R, et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model." *Science* 2023;379(6637):1123–1130. Sequence-only language models that recover structural and evolutionary information. [DOI](https://doi.org/10.1126/science.ade2574) [^d3pm]: Austin J, Johnson DD, Ho J, Tarlow D, van den Berg R. "Structured Denoising Diffusion Models in Discrete State-Spaces." *NeurIPS* 2021. Discrete categorical diffusion with a token transition matrix $Q_t$; absorbing-state masking is the special case used here. [arXiv](https://arxiv.org/abs/2107.03006) [^mdlm]: Sahoo SS, Agrawal A, Arriola M, et al. "Simple and Effective Masked Diffusion Language Models." *NeurIPS* 2024. Absorbing-state discrete diffusion with a `[MASK]` token; the public generative primitive LiteMol-1 uses. [arXiv](https://arxiv.org/abs/2406.07524) [^spe]: Li X, Fourches D. "SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep Learning." *J Chem Inf Model* 2021;61(4):1560–1569. Subword SMILES tokens learned from frequent character n-grams, with character-level fallback so uncommon fragments are not mapped to `[UNK]`. [DOI](https://doi.org/10.1021/acs.jcim.0c01127) [^cfg]: Ho J, Salimans T. "Classifier-Free Diffusion Guidance." *NeurIPS Workshop* 2021. Mix conditional and unconditional scores at inference without a separate classifier. [arXiv](https://arxiv.org/abs/2207.12598) [^puct]: Silver D, Schrittwieser J, Simonyan K, et al. "Mastering the game of Go without human knowledge." *Nature* 2017;550:354–359. The PUCT rule: $Q$ plus an exploration bonus weighted by the prior $p_\theta$. [DOI](https://doi.org/10.1038/nature24270) [^protenix]: ByteDance AML Team. "Protenix: Advancing Structure Prediction Through a Comprehensive AlphaFold 3 Reproduction." *bioRxiv* 2025. Open AF3-class cofolder used as the matched peptide and macrocycle structure protocol. [DOI](https://doi.org/10.1101/2025.01.08.631967) [^mdm2-amp]: Oliner JD, Kinzler KW, Meltzer PS, George DL, Vogelstein B. "Amplification of a gene encoding a p53-associated protein in human sarcomas." *Nature* 1992;358:80–83. MDM2 amplification as a route to p53 inactivation. [DOI](https://doi.org/10.1038/358080a0) [^mdm2-1ycr]: Kussie PH, Gorina S, Marechal V, et al. "Structure of the MDM2 oncoprotein bound to the p53 tumor suppressor transactivation domain." *Science* 1996;274(5289):948–953. The F19 / W23 / L26 helix in the MDM2 cleft (PDB `1YCR`). [DOI](https://doi.org/10.1126/science.274.5289.948) [^mdm2-5afg]: Lau YH, Wu Y, Rossmann M, et al. "Double Strain-Promoted Macrocyclization for the Rapid Selection of Cell-Active Stapled Peptides." *Angew Chem Int Ed* 2015;54:15410–15413. Stapled MDM2 peptide crystal used as the harder co-fold ceiling (PDB `5AFG`). [DOI](https://doi.org/10.1002/anie.201508416) [^nav18-scn10a]: Akopian AN, Sivilotti L, Wood JN. "A tetrodotoxin-resistant voltage-gated sodium channel expressed by sensory neurons." *Nature* 1996;379:257–262. The original description of SNS / PN3, later named Nav1.8 (SCN10A). [DOI](https://doi.org/10.1038/379257a0) [^suzetrigine]: Jones J, Correll DJ, Lechner SM, et al. "Selective Inhibition of NaV1.8 with VX-548 for Acute Pain." *N Engl J Med* 2023;389(5):393–405. Clinical evidence that selective Nav1.8 inhibition can treat acute pain; suzetrigine (Journavx) was approved in January 2025. [DOI](https://doi.org/10.1056/NEJMoa2209870) [^nav18-9dbn]: Neumann B, McCarthy S, Gonen S. "Structural basis of inhibition of human NaV1.8 by the tarantula venom peptide Protoxin-I." *Nat Commun* 2025;16:1459. Cryo-EM of Nav1.8 with Protoxin-I on VSD2 (PDB `9DBN`). [DOI](https://doi.org/10.1038/s41467-024-55764-z) [^nav18-a803467]: Huang X, Jin X, Huang G, et al. "Structural basis for high-voltage activation and subtype-specific inhibition of human Nav1.8." *Proc Natl Acad Sci USA* 2022;119(30):e2208211119. A-803467 occupies the central pore, not VSD2 (PDB `7WE4`). [DOI](https://doi.org/10.1073/pnas.2208211119) [^fxia]: Weitz JI, Chan NC. "Advances in Antithrombotic Therapy." *Arterioscler Thromb Vasc Biol* 2019;39(1):7–12. Factor XI as a target for anticoagulants that may spare haemostasis. [DOI](https://doi.org/10.1161/ATVBAHA.118.310960) [^milvexian]: Dilger AK, Pabbisetty KB, Corte JR, et al. "Discovery of Milvexian, a High-Affinity, Orally Bioavailable Inhibitor of Factor XIa in Clinical Studies for Antithrombotic Therapy." *J Med Chem* 2022;65(3):1770–1785. The 12-membered macrocycle and the FXIa complex (PDB `7MBO`). [DOI](https://doi.org/10.1021/acs.jmedchem.1c00613) [^jq1]: Filippakopoulos P, Qi J, Picaud S, et al. "Selective inhibition of BET bromodomains." *Nature* 2010;468:1067–1073. JQ1, the acetyl-lysine mimetic warhead for BRD4. [DOI](https://doi.org/10.1038/nature09504) [^mz1]: Zengerle M, Chan KH, Ciulli A. "Selective Small Molecule Induced Degradation of the BET Bromodomain Protein BRD4." *ACS Chem Biol* 2015;10(8):1770–1777. MZ1: JQ1 linked to VH032, selective BRD4 degradation. [DOI](https://doi.org/10.1021/acschembio.5b00216) [^mz1-5t35]: Gadd MS, Testa A, Lucas X, et al. "Structural basis of PROTAC cooperative recognition for selective protein degradation." *Nat Chem Biol* 2017;13:514–521. The BRD4 BD2–MZ1–VHL ternary complex (PDB `5T35`). [DOI](https://doi.org/10.1038/nchembio.2329) [^unimol]: Zhou G, Gao Z, Ding Q, et al. "Uni-Mol: A Universal 3D Molecular Representation Learning Framework." *ICLR* 2023. SE(3)-equivariant 3D conformation encoder used here as the structure-based embedding baseline. [OpenReview](https://openreview.net/forum?id=6K2RM6wVqKu) [^ecfp]: Rogers D, Hahn M. "Extended-Connectivity Fingerprints." *J Chem Inf Model* 2010;50(5):742–754. Circular (Morgan) fingerprints; radius 2, 2048 bits in this comparison. [DOI](https://doi.org/10.1021/ci100050t) --- # BenchPLM: Proteins Aren't Sentences: Why Bigger Protein Models Don't Win URL: https://www.lite.bio/research/benchplm Author: LiteFold Research Date: 2026-06-26 A protein language model (PLM) turns an amino-acid sequence into vectors that can be reused for downstream protein tasks: stability, localization, binding, mutation fitness, evolutionary structure, and more. In the frozen-embedding setting, the pretrained PLM is not fine-tuned. We extract a sequence embedding from it, train a small supervised probe on top, and ask how much useful biological signal was already present in the representation. ![](images/image.png) Several protein language models have been developed in recent years. Among the earliest, Meta (formerly Facebook) FAIR introduced the ESM-2 family of models.[^esm2] Since then, a growing number of PLMs, including ESM2, ProGen,[^progen2] and DPLM,[^dplm] have demonstrated strong performance across various protein tasks. Each of these models learns its representations in a fundamentally different way, and that difference propagates into downstream research and the decisions built on it. To make the comparison concrete, we benchmark the following models: | **Model** | **Parameters** | **Objective family** | **Readout direction** | **Frozen embedding interface** | | --- | --- | --- | --- | --- | | ESM-C 300M | 300M | masked language model | bidirectional | mean-pooled encoder hidden state | | DPLM 150M | 150M | discrete diffusion | bidirectional | teacher-forced hidden-state pooling | | ESM-2 150M | 150M | masked language model | bidirectional | mean-pooled encoder hidden state | | ProGen3 219M | 219M | autoregressive language model | left-to-right | teacher-forced hidden-state pooling | | ProGen2-small | 151M | autoregressive language model | left-to-right | teacher-forced hidden-state pooling | ## BenchPLM Protein sequences read a lot like text: residues behave like words and local motifs like phrases. That resemblance is what led researchers to borrow architectures from natural language processing. That observation led researchers to borrow ideas from natural language processing and develop different architectures for learning representations of proteins. How a model reads a protein determines what its embeddings can represent. Masked models, causal models, diffusion models all see sequence context differently, so they inevitably learn different representations. However those variations naturally leads to the following questions: - Which protein language models learn the best representations, and why? - Which protein language model should be used for downstream tasks? - What kind of representations have these models learned, and how much evolutionary knowledge do they capture? To answer these, we introduce PLMBench, a collection of 9 tasks designed to systematically evaluate protein language model representations. | **Task** | **Split size (train / validation / test)** | **Primary metric** | **What it probes** | | --- | --- | --- | --- | | Thermostability Regression | 5310 / 706 / 706 | Spearman rho | Sequence-level stability signal. | | DeepLoc multiclass localization | 10414 / 1368 / 1368 | Accuracy | Subcellular localization across multiple classes. | | DeepLoc binary localization | 6707 / 698 / 807 | Accuracy | A simpler binary localization ablation. | | Metal ion binding | 5797 / 719 / 719 | Accuracy | Compact function and binding-site signal. | | FLIP2 Alpha Amylase | 2574 / 644 / 488 | Spearman rho | Mutation-fitness ranking under a one-to-many split. | | FLIP2 Hydrophobic Core | 9974 / 2493 / 12468 | Spearman rho | Mutation-fitness ranking under a larger low-to-high shift. | Let’s understand the tasks in more details. In case if you are more interested to know the results, you can skip this section. 1. **Thermostability Regression:** Given a protein sequence, predict its thermostability score. We use Spearman's ρ, which measures how well the model ranks more-stable proteins above less-stable ones. 2. **DeepLoc Tasks:**[^deeploc] Two classification tasks. In multi-class classification, we predict the subcellular localization of a protein (e.g., nucleus, cytoplasm, mitochondrion, membrane, extracellular). In binary classification, we predict whether a protein is membrane-bound or soluble. 3. **Metal Ion Binding:** Predict whether a given protein sequence binds metal ions. We do not predict the specific metal or binding site. 4. **FLIP2 Alpha Amylase:**[^flip] Given mutant alpha-amylase sequences, the model ranks variants from worse to better function. We use Spearman's ρ, where higher values indicate better ranking accuracy. 5. **FLIP2 Hydrophobic Core:** Similar to above, but using mutant sequences from the hydrophobic core dataset. The task is to predict mutation fitness or structural packing quality. As a regression task, we again use Spearman's ρ, where higher values indicate stronger predictive ranking. Every model was evaluated on the same six tasks with the same train/validation/splits, the same embedding policy and the same shallow probe family. The protocol was: 1. Freeze the pre-trained PLM backbone 2. Extract sequence embeddings with sliding window mean pooling 3. Train a probe on the training split. 4. Select probe hyperparameters on the validation split. 5. Report the held-out test metric. One probe family is held fixed across all models, so every score reflects the representation and that probe jointly. We keep the probe shallow to stay close to the raw embedding, but absolute numbers would shift under a different probe; the cross-model ranking is the durable signal, not the decimal. ### Interpretability We also ran three diagnostic analyses. The first two are mechanistic probes; the third is a representation-geometry check. They are useful for interpretation, but they are not included in the benchmark score. | **Diagnostic** | **Models covered** | **Question** | | --- | --- | --- | | Evolutionary layer localization | ESM-2 scales, DPLM 150M, ProGen2-small | At which layer is evolutionary distance most readable? | | Context occlusion | ESM-2 150M, DPLM 150M, ProGen2-small | Does a residue-level score depend on left or right context? | | Embedding PCA and same-family retrieval | ESM-2 150M, DPLM 150M, ProGen2-small; larger cached check with ESM-2 650M, ESM-2 3B, ProGen2-medium | Do frozen embeddings form clean ortholog-family neighborhoods? | ## ESM-C and DPLM Leads Overall The overall result is close at the top. ESM-C 300M[^esmc] and DPLM 150M both have a mean task rank of `1.83` across the six tasks. ESM-C wins more individual tasks (3 versus 2), while DPLM has the marginally higher average normalized primary score (0.938 versus 0.924). With six tasks and a single probe seed, a gap this small is inside the noise floor; the honest reading is that ESM-C and DPLM are tied at the top, not that one edges the other. ESM-2 150M is the next strongest model. Both ProGen checkpoints fall below the bidirectional models in this frozen-probe setting. ESM-C is also the largest backbone here, so this aggregate alone cannot separate architecture from scale. The same-size comparison does: ESM-2 and DPLM at 150M beat both ProGen checkpoints of equal or larger size, which isolates the effect we return to in the conclusion. *Frozen linear-probe benchmark, six tasks* The per-task view explains why the aggregate is close: | **Task** | **Winner** | **Winning score** | **Runner up** | | --- | --- | --- | --- | | Thermostability | ESM-C 300M | 0.664 Spearman | ESM-2 150M at 0.645 | | DeepLoc multiclass | ESM-C 300M | 0.814 accuracy | DPLM 150M at 0.812 | | DeepLoc binary | ESM-2 150M | 0.914 accuracy | DPLM 150M at 0.912 | | Metal ion binding | DPLM 150M | 0.711 accuracy | ESM-2 150M at 0.702 | | FLIP2 Alpha Amylase | ESM-C 300M | 0.676 Spearman | DPLM 150M at 0.613 | | FLIP2 Hydrophobic Core | DPLM 150M | 0.371 Spearman | ESM-C 300M at 0.355 | ProGen3[^progen3] remains usable on some stability and classification tasks, but it does not win any task in this suite. ProGen2-small is weaker across the benchmark, with the clearest gap on FLIP2 Alpha Amylase: `0.105` Spearman, compared with `0.676` for ESM-C and `0.613` for DPLM. *Per-task test scores from the same frozen protocol* The pattern is clear: on downstream tasks, models that condition each residue on the full sequence outperform models restricted to a left-to-right prefix. This is not a verdict on the ProGen family. ProGen models were trained to generate viable protein sequences. However the pattern is biologically plausible. Many protein properties depend on residues that are distant in sequence but coupled through structure, family constraints, or functional motifs. A frozen embedding that can integrate both upstream and downstream context gives a shallow probe more of the relevant signal. ## Evolutionary Understanding When Protein Language Models were trained, researchers saw that the models were learning evolutionary relationship. This is one reason ESMFold dropped the MSA:[^esmfold] the language model was meant to stand in for it. The payoff is generalization to sequences with few or no homologs. In this section, we tried to quantify how much evolutionary understanding does different protein language models carries. ### TimeTree TimeTree[^timetree] is a database of species divergence times. Given two species, it tells you how many millions of years ago they last shared a common ancestor. So for three proteins drawn from species X, Y, and Z, TimeTree can say whether X branched off closer to Y or to Z. This probe asks a single question: **at what layer depth does a protein language model start encoding evolutionary history rather than surface sequence?** For each model we sample a set of layers, pull the hidden states, mean-pool over residues to get one vector per protein, and check whether cosine distances between those vectors reproduce TimeTree's branching order. The test set is built from *identity-discordant triplets*. Each triplet has three proteins: - **A**, the anchor - **B**, the history-correct neighbor (TimeTree says A and B diverged more recently) - **C**, the identity-misleading decoy (raw sequence identity says A looks more like C) ![](images/image-3.png) The two signals are deliberately put in conflict. By sequence identity alone, A resembles C more than B. By evolutionary history, A is closer to B. A model that only tracks surface similarity will be pulled toward C and fail. A model that has internalized deeper phylogenetic structure will still place B nearer. Triplet accuracy is just the fraction of triplets where the embedding agrees with history: distance(A, B) < distance(A, C) Chance is 50%. Scores above that mean the layer carries evolutionary signal beyond raw identity; scores at or below mean it does not. *Evolutionary signal becomes most readable in later layers* The same result is easier to inspect as a heatmap. ESM-2 and DPLM become more readable with depth, while ProGen2-small does not show the same late-layer gain. *Layer readout heatmap* Not only accuracy, we also ask another question: Across many protein pairs, do embedding distances increase as TimeTree divergence times increase? Now consider this table: | **Model** | **Objective** | **Direction** | **Best layer** | **Best triplet acc** | **Best Spearman rho** | | --- | --- | --- | --- | --- | --- | | ESM-2 8M | Masked LM | Bidirectional | 6 / 6 | 0.602 | 0.374 | | ESM-2 35M | Masked LM | Bidirectional | 12 / 12 | 0.630 | 0.435 | | ESM-2 150M | Masked LM | Bidirectional | 30 / 30 | 0.609 | 0.412 | | DPLM 150M | Discrete diffusion | Bidirectional | 30 / 30 | 0.627 | 0.468 | | ProGen2-small | Causal LM | Left-to-right | 0 / 12 | 0.460 | 0.158 | Scale is not the axis here. ESM-2 35M matches 150M on this triplet probe, and DPLM 150M holds the strongest TimeTree-distance correlation despite no size advantage. What moves the signal is how the model reads context, not how many parameters it has. For ESM-2 and DPLM, the signal gets stronger in deeper layers. That suggests the model gradually builds a more evolution-aware representation as sequence context is processed. For ProGen2, the best TimeTree signal was at the input/early layer, and deeper causal layers did not improve it, which supports the blog’s argument that left-to-right models are weaker at forming bidirectional evolutionary geometry. *Layer readout tradeoff* ### Context Occlusion To test whether bidirectional models genuinely use context from both directions, we run a simple occlusion experiment inspired by input-perturbation methods common in neural network interpretability. The setup is straightforward. Pick a target residue and mask it. Record the model's log-probability for the correct amino acid at that position. Then mask one additional context residue at a known offset from the target and record the log-probability again. The difference between the two measurements tells us how much that context position contributed to the prediction. A large change means the context residue mattered; a small or zero change means the model was not relying on it. For ESM-2 and DPLM, both masked language models, we mask the target and measure perturbations symmetrically on both sides. For ProGen2, a causal language model, we score the target from its prefix only, since by construction it never sees residues to the right. *Context occlusion heatmap* The binned heatmap reports mean absolute log-probability change by relative position. Two patterns stand out. 1. First, all three models show strongest sensitivity in the immediate neighborhood of the target (offsets -20 to +20), with effects decaying at greater distances. This is expected: local sequence context carries the strongest signal for residue identity. 2. Second, and more telling, is the asymmetry. ESM-2 and DPLM show measurable sensitivity on both sides of the target, with the right-side effect at +20 (0.050 for ESM-2, 0.044 for DPLM) comparable to or exceeding the left-side effect at -20. ProGen2-small shows a sharp spike at -20 (0.104) and exactly zero for every right-side bin. The right side is blank because those residues simply do not exist in the causal prefix used to score the target. *Context directionality* The bar plot compresses the full heatmap into a single left-versus-right summary. For each model, we sum the absolute effects from all left-side bins and all right-side bins, then report the right-context fraction. - ESM-2 150M: left 0.021, right 0.026, right fraction 0.552 - DPLM 150M: left 0.018, right 0.022, right fraction 0.556 - ProGen2-small: left 0.031, right 0.000, right fraction 0.000 ESM-2 and DPLM split their context dependence nearly evenly across both sides, with a slight lean toward the right. ProGen2's entire context mass sits on the left, exactly as a causal architecture dictates. Proteins are not sentences. In natural language, left-to-right context is a reasonable inductive bias because meaning largely builds sequentially. In proteins, the constraints on a residue's identity are fundamentally non-sequential. A residue may be constrained by a disulfide partner dozens of positions downstream, by residues it packs against in the folded structure, by compensatory mutations elsewhere in the sequence, or by a distant active-site motif it co-evolves with. None of these relationships respect sequence order. The occlusion results confirm that ESM-2 and DPLM capture this: their representations at any given position are shaped by the full sequence context. ProGen2 behaves correctly for its architecture, but its representations at each position are blind to everything that follows. This asymmetry likely explains part of the performance gap we observe in downstream tasks that require whole-sequence understanding. ### Do Ortholog Families Stay Together? As a final representation-level check, we reused the same TimeTree panel containing `1,024` proteins from `128` ortholog families, with exactly eight sequences per family. For ESM-2 150M, DPLM 150M, and ProGen2-small, we used the cached final mean-pooled embeddings from the phylogenetic-geometry runs. ![](images/image-9.png) The visualization below fits PCA separately for each model after L2-normalizing the embeddings, matching the cosine geometry used in the TimeTree benchmark. Colored points mark eight evenly spaced ortholog families; gray points are the remaining families. The plotted colored points use a tiny deterministic display jitter so duplicated or near-identical PCA coordinates remain visible. The panels are qualitative and independently scaled, so the quantitative readout is the bar plot underneath: same-family retrieval across all 128 families. Consider this table: | **Source** | **Same-family NN@1** | **Same-family P@7** | | --- | --- | --- | | Position identity | 0.661 | 0.431 | | 3-mer Jaccard | 0.850 | 0.787 | | ESM-2 150M | 0.855 | 0.798 | | DPLM 150M | 0.859 | 0.820 | | ProGen2-small | 0.641 | 0.397 | DPLM and ESM-2 produce the cleanest same-family neighborhoods, with DPLM slightly ahead by precision@7. ProGen2-small still retrieves a same-family nearest neighbor well above random chance, but its local neighborhoods decay quickly: across the seven nearest non-self neighbors, fewer than half are from the same ortholog family. We also repeated the same diagnostic for the larger cached checkpoints available on this panel: ESM-2 650M, ESM-2 3B, and ProGen2-medium as well. ![](images/image-10.png) The larger ESM-2 checkpoints improve family-neighborhood purity on this diagnostic, with ESM-2 3B reaching `0.875` P@7. ProGen2-medium does not show the same scaling behavior here: it is below the 3-mer baseline and below ProGen2-small by P@7. That result is local to this ortholog-family retrieval probe, but it reinforces the broader pattern that the causal ProGen embeddings are less useful as frozen neighborhood representations in these experiments. ## The Curse of Sparse Data Most of the real world protein engineering datasets are small. A useful embedding should not only perform well with thousand of labels; it should also be sample-efficient when labels are scarce. Measured variants could be in the terms of 20, 100 or 700. To benchmark this, the setup is simple. 1. We take a frozen PLM Embedding. 2. Train a small regression probe on K-labeled examples. 3. Predict fitness on held-out variants. 4. Repeat this on different label budgets. 5. Plot it as the curve and compare the differences. | **Task** | **Training budgets** | **Test size used** | **What it stresses** | | --- | --- | --- | --- | | FLIP2 GB1 one-vs-rest | 8 / 16 / 24 / 28 | 4096 | Extremely low-data fitness transfer. | | FLIP2 Alpha Amylase close-to-far | 64 / 128 / 256 / 512 / 1024 / 1782 | 1924 | Extrapolation from closer mutations to farther mutations. | | FLIP2 Rhodopsin by-wild-type | 32 / 64 / 128 / 256 / 512 / 700 | 184 | Membrane-protein fitness transfer. | So this means, if a model gets a good Spearman score with only 16 or 64 labels, that means its embedding is sample-efficient: the probe does not need much supervision because the representation already carries useful signal. *Frozen sample efficiency on sparse mutation-fitness tasks — Spearman rho vs number of labeled training examples. Only frozen-curve rows are plotted. Values digitized from the original figure; error bars omitted.* *GB1 one-vs-rest* *Alpha amylase* *Rhodopsin* The sparse-transfer result largely agrees with the full frozen benchmark. ESM-C has the best mean normalized curve score (`0.771`) and wins two of the three sparse tasks. ESM-2 is second by normalized curve score (`0.691`). DPLM is third overall (`0.657`) but is the strongest model on Rhodopsin at full budget. | **Task** | **Best frozen model** | **Full-budget score** | **Interpretation** | | --- | --- | --- | --- | | GB1 one-vs-rest | ESM-C 300M | 0.535 Spearman | ESM-C separates early, especially after 16 labels. | | Alpha Amylase close-to-far | ESM-C 300M | 0.299 Spearman | ESM-C is strongest under this mutation-distance shift. | | Rhodopsin by-wild-type | DPLM 150M | 0.688 Spearman | DPLM is the clearest winner on this membrane-protein fitness task. | The ProGen models do not close the frozen-transfer gap under sparse labels. ProGen3 improves on Rhodopsin as the budget grows, reaching `0.472` at 700 examples, but remains behind DPLM and ESM-2. ProGen2-small remains weak across the sparse suite. The limited-data conclusion is therefore aligned with our initial study: **ESM-C is the strongest broad frozen representation, while DPLM deserves special attention on Rhodopsin-like fitness transfer.** ## Conclusion Across every probe in this study, the most consistent predictor of frozen-representation quality was not parameter count or pretraining scale. It was the conditioning set the model reads. Bidirectional encoders, ESM-2 and ESM-C as masked models and DPLM as a diffusion model, beat the causal ProGen checkpoints on every task that rewards whole-sequence understanding: downstream probes, evolutionary geometry, and same-family retrieval alike. The reason is visible if we write down what each model computes. A causal model factorizes the sequence left to right, $$ p(x) = \prod_{i=1}^{L} p(x_i \mid x_{ "Rosalind seems to reason better than other AI co-scientists and always > performed deeper analysis, retrieving more findings and generating more > hypotheses. It performed similarly to K-dense in literature review and > analysis, but wins in agentic orchestration of bioinformatic tools." That same depth of reasoning surfaced a real limitation. Rosalind pursued a full analog-design workthrough for a peptide before Molecule's team established that the target mechanism wasn't scientifically resolved enough to justify it: > "One can argue this was also my error." Rather than treat that as a dead end, Molecule closed the gap on their side — adding an internal "druggability" check to their own workflow to filter out targets where the mechanism isn't resolved enough to reason about further, avoiding wasted compute on both sides. On day-to-day usefulness, the split between human and agentic use was clear: > "For human use it's the ease of visualizing designs and working together > with Rosalind like a scientific co-pilot, able to jump from idea to > execution. For an agent, retrieval of information, sequences, production of > high quality scientific papers." ## What's next Molecule's next ask points at a gap in the field rather than in LiteFold specifically — an agonist prediction tool, which needs training data from functional assays to build: > "I have already started defining the problem and working on the solution, > and once we get data on Q3 this could be accelerated." With the KISS1R and OX2R validation results expected by the end of Q3, Molecule and LiteFold are positioned to turn that functional assay data into exactly the training set an agonist prediction tool would need. --- # Introducing Hybrid Scientific Intelligence Runtime URL: https://www.lite.bio/blogs/hybrid-scientific-intelligence-runtime Author: Siddhant Prateek Mahanayak Date: 2026-09-23 Every enterprise conversation we have had eventually ends at the same place: **the data cannot leave the organization.** As intelligence systems get more capable, general-purpose intelligence is quickly turning into a commodity. Pharma and biotech, however, have always adopted new technology more cautiously, and for good reason: proprietary data, sensitive scientific workflows and the risk of IP leakage. Once intelligence is cheap and widely available, the question is no longer just how to *use AI*. It becomes how to build AI systems that are sovereign: systems the organization fully owns and controls, that run inside its own infrastructure, and that keep improving from its own data and workflows. That is why today at LiteFold we are introducing the **Hybrid Scientific Intelligence Runtime**, or **HSIR**. Modern AI-native systems and workflows are generally built from five core layers: 1. **Independent GPU and CPU workflows** 2. **An agentic harness** connected to those workflows 3. **Secure sandboxes** for isolated code execution 4. **A core LLM inference engine** 5. **A governance, policy and authorization layer** ***HSIR in one picture.** The five layers sit inside the customer's perimeter, with the control plane as the single entry point. Cloud pools are optional, join the same scheduler, and exchange only what policy allows.* HSIR packages all five layers into a single compute layer that runs on infrastructure the customer owns. Underneath, it maintains a **warm worker pool and an image cache**, so workloads start quickly while the entire execution stack stays inside the customer's environment. To see how the runtime holds up in practice, we ran two public sandbox benchmarks against HSIR on a single machine shaped like an on-premise deployment. We compared it with the public results for **Modal, Google Cloud Run, Cloudflare, Daytona, E2B, Vercel and Runloop**, and with **plain Docker on the exact same machine**. On the **DAX full-build benchmark**, HSIR finished the workload in **55.8 seconds**, ahead of Daytona, Vercel, Modal and Runloop, the providers in that set that completed all seven phases. The goal of this runtime is simple to state: **keep the control and security properties of private infrastructure without giving up the developer experience and elasticity of a modern AI cloud.** # Why we built this runtime Most infrastructure for scientific computing and AI is still delivered as a hosted service. If you want to run a structure prediction model, analyze SAR data, optimize lead candidates, or let an agent work across experimental results, the data usually has to leave your environment first. For a drug program, that is a serious constraint. Sequences, SAR tables, assay results, internal reports, lab notebooks, and the relationships between them are often among the most valuable IP a company has. For many teams, the requirement from security and legal is not "encrypt the data in transit". It is **"the data does not leave our environment."** This matters even more as AI companies move deeper into biology and drug discovery. The companies providing foundation models and AI infrastructure are increasingly running their own research programs and building their own biological models and datasets. [Isomorphic Labs](https://www.prnewswire.com/news-releases/isomorphic-labs-secures-2-1-billion-funding-to-scale-its-ai-drug-design-engine-302769674.html)[^isomorphic], the Alphabet drug discovery company, runs wholly-owned programs alongside its partnerships, OpenAI now ships [GPT-Rosalind](https://openai.com/index/introducing-gpt-rosalind/)[^gptrosalind], a reasoning model built for life sciences, and Anthropic launched [Claude Science](https://www.anthropic.com/news/claude-science-ai-workbench)[^claudescience] and has [reportedly started its own drug programs](https://endpoints.news/anthropic-debuts-claude-science-an-ai-product-for-bioscience/)[^endpoints]. At the same time, model policies, access rules, product terms and infrastructure keep changing. GPT-Rosalind, for example, is available only through a [gated trusted-access program](https://openai.com/index/introducing-gpt-rosalind/)[^gptrosalind], and consumer AI products have [changed their data-training defaults](https://www.anthropic.com/news/updates-to-our-consumer-terms)[^antdefaults] after launch. Zero-data-retention (ZDR) agreements help, but they do not remove the dependency. ZDR is an arrangement you apply for, it covers specific endpoints and features, and it still allows retention for flagged content or when the law requires it.[^zdr] In 2025, a US court [ordered OpenAI to preserve consumer ChatGPT and API data](https://openai.com/index/response-to-nyt-data-demands/)[^nytdata] as part of the New York Times lawsuit. ZDR customers were excluded, but it showed how much of a customer's data posture ultimately depends on the provider's architecture, contracts and legal situation. For many scientific organizations the cleanest answer is much simpler: **run the intelligence where the data already lives.** Large organizations already have substantial compute: private cloud environments, Kubernetes clusters, internal storage, identity systems and, more and more, their own GPUs. The missing piece is usually not more hardware. It is a unified runtime that connects all of it and runs modern scientific AI workflows inside the organization's own environment: models, agents, scientific tools, code execution and internal data access, all through one system, with governance and authorization around everything. That is why we built HSIR from the ground up. HSIR can run on a single GPU server, an existing cluster, or inside the customer's own cloud account. The same runtime handles model inference, scientific workloads, agent sandboxes, batch jobs and training, without the underlying data ever moving to LiteFold. For scientists, it should still feel like a modern cloud platform. For the infrastructure team, everything stays under the organization's control. Concretely, these are some of the workloads the runtime is built to handle: | Workload | Example | What HSIR does | | --- | --- | --- | | **Model serving** | Running a structure prediction or protein language model | Keeps models loaded and serves concurrent requests without reloading weights for every call | | **Sandboxes** | An agent writing and executing analysis code against assay data | Creates an isolated environment with the required tools, executes the job, then destroys it | | **Batch compute** | Embedding 100,000 sequences or scoring 50,000 compounds | Distributes the workload across parallel workers and collects the results | | **Training and fine-tuning** | Fine-tuning a model on internal binding data | Runs long-lived GPU jobs with persistent datasets, checkpoints, retries and timeouts | | **Continuous learning** | Updating a model as new experimental data arrives | Runs scheduled training jobs and deploys updated checkpoints back into the runtime | Each workload runs in its own isolated [OCI container](https://github.com/opencontainers/runtime-spec)[^oci] with explicit CPU, memory and GPU limits. Jobs cannot see each other's files or memory. Temporary files disappear when a job finishes, and anything that needs to persist, such as datasets or checkpoints, lives in explicitly mounted volumes. The runtime also scales capacity with demand. Frequently used workers stay warm for fast startup, while idle workers shut down instead of sitting on expensive GPU capacity. # Hybrid intelligence: open and closed models over gated contexts *A task starts against public context on a frontier model, and crosses into the private environment when proprietary data is required. The workflow carries forward instead of restarting.* On-premise infrastructure matters, but we are not saying everything should run on-premise. Most organizations already operate across private infrastructure, public cloud, external databases, SaaS systems and internal networks. Scientific AI systems will have to work across the same boundaries. For agentic workflows in particular, running only on local infrastructure often does not make sense: 1. **Some tasks need the best available frontier models.** Deep research, long-context synthesis, planning and reasoning-heavy work can still benefit from closed models that are significantly stronger than what can be deployed locally. 2. **Agents often need information from outside the organization.** Literature, patents, clinical trial registries, public databases, regulatory documents, vendor systems and other external APIs. 3. **Some workloads are highly bursty.** A team may occasionally need hundreds of parallel workers, or far more compute than the local cluster has. Those workloads can burst into approved cloud infrastructure without changing how the task itself runs. 4. **Different parts of the same task have different sensitivity levels.** Searching public literature and reasoning over a proprietary assay table should not have to happen inside the same trust boundary. This is where the hybrid architecture in HSIR becomes useful. A task can start with a frontier model working on public or non-sensitive information. That model does the expensive part of the work: searching, reasoning, generating hypotheses, building a plan, and reducing a large amount of information into a small working context. That task state, including retrieved information, intermediate outputs, tool results and any other allowed context, is then carried forward. When the workflow reaches a point where proprietary information is needed, execution moves into the private environment. There, an internally hosted model works with institutional knowledge such as assay results, electronic lab notebook (ELN) records, molecular data, internal reports or manufacturing data, without any of it being sent back to an external model. The key point is that the workflow does not restart. The public half does the heavy lifting, and the private model receives only the reduced context it needs to finish the sensitive part. In practice, the execution boundary looks something like this: *Public context* The boundary is enforced by policy, not by the agent. Organizations decide which data sources may leave the private environment, which models can access which datasets, which tools can make outbound requests, and where each part of a workflow is allowed to run. For example, an organization could allow an agent to: - search [PubMed](https://pubmed.ncbi.nlm.nih.gov/)[^pubmed], patents and [ClinicalTrials.gov](https://clinicaltrials.gov/)[^clinicaltrials] using an external model, - summarize those findings into a structured research context, - move that context into the private runtime, - combine it with an internal target dossier and assay history, - run proprietary models and scientific workflows locally, - and generate the final analysis without exposing any private data externally. Every crossing between these boundaries is logged, audited and traceable. The goal is not to choose between open and closed models, or between cloud and on-premise. It is to use each where it makes sense, while the organization stays in control of **which intelligence can see which context, and where that computation is allowed to happen.** *Deploy HSIR inside your own environment* # Benchmarks for the runtime With these benchmarks we wanted to answer two simple questions: 1. **Once a sandbox is running, how much performance do we lose compared with plain Docker?** 2. **How quickly can the runtime create many sandboxes at once?** We used two public sandbox benchmarks maintained by ComputeSDK, **DAX** and **Burst TTI**,[^bench] and ran them on a single machine shaped like a typical private deployment: a 30-vCPU Intel Xeon server with 222 GB of memory and one warm HSIR worker. We also ran the same workload directly in Docker on the same machine as a baseline. ## Full workload DAX creates a fresh sandbox, installs a toolchain, clones a real repository ([opencode](https://opencode.ai/)[^opencode]), installs its dependencies and runs a full typecheck. Across 10 runs, HSIR completed the workload in a median of **55.8 seconds**. *DAX build benchmark, median total build time* Broken down by phase, our container setup is in line with the fastest providers, and the phases that do real work are decided by the host CPU rather than by the runtime: *DAX build benchmark, time per phase* Among the providers we compared against that completed the full workload, HSIR was the fastest. More importantly, when we compared the actual compute with plain Docker on the same machine, the numbers were almost identical: *Sandbox overhead on the compute-bound phases* In other words, once the sandbox is running, **we are effectively operating at bare-container speed.** The one visible gap is install, which writes tens of thousands of small files and pays roughly a second of filesystem overhead. Clone and typecheck are within noise of plain Docker. ## 100 sandboxes at once The second benchmark stresses a very different part of the runtime. Instead of one long workload, it asks the system to create **100 isolated sandboxes at the same instant** and measures time-to-interactive (TTI): how long from `create()` until each sandbox runs its first command successfully. This is much closer to what happens when an agent fans out across many candidates, analyses or experiments in parallel. On our single 30-vCPU machine, all **100 out of 100 sandboxes came up successfully**. Median TTI was **2.66 seconds**, with the first sandbox ready in 1.49 seconds and the entire burst ready in under four seconds. There is one important constraint here. A 30-vCPU machine cannot reserve a full vCPU for each of 100 sandboxes without overcommitting, and HSIR intentionally does not overcommit. So the 100-way test used 0.2 vCPU per sandbox. For reference, when we ran 24 sandboxes at the full 1 vCPU configuration, median startup dropped to **1.69 seconds**. Even so, the 100-way result is still slower than the stronger systems on the public leaderboard: *Burst TTI, 100 sandboxes created at the same instant* This is the part of the runtime that still needs work. Right now every sandbox waits roughly **one extra second** because of a reconnect delay in our readiness path, and the container itself starts much faster than that number suggests. Removing that delay alone should move the burst median much closer to the middle of the leaderboard. The takeaway is fairly simple. **Execution is already fast. Startup under heavy concurrency is not yet where we want it.** So the next round of work focuses on sandbox startup, keeping warm workers alive for longer, and cutting the remaining filesystem overhead for workloads that create large numbers of small files. For us, that is a more useful outcome than a leaderboard number: it tells us exactly where the runtime is already competitive and where the next engineering effort needs to go. # Conclusions and next steps In this post we introduced the current version of our **Hybrid Scientific Intelligence Runtime**: where it fits, the workloads it supports, how hybrid intelligence works across public and private contexts, and how the runtime performs today. There is a bigger question we are working on internally: **how much intelligence actually needs to live inside the organization?** We are building evaluation sets around real pharmaceutical and life-sciences workflows to understand what level of model capability, and what model size, is enough to reliably handle **80 to 90% of day-to-day scientific tasks**. The goal is not to run the largest possible model privately. It is to find the smallest capable intelligence layer that works well once it is connected to the right tools, institutional context and scientific workflows. That is likely what the next engineering log will be about. At LiteFold, our broader goal is to help life-sciences organizations become **AI-native and agentic** without giving up control of their data, infrastructure or scientific IP. HSIR is the infrastructure layer we are building toward that. If you are thinking about deploying scientific agents, private models or hybrid AI infrastructure inside your organization, [reach out to us](/contact). *Thinking about running this inside your organization?* [^zdr]: See [OpenAI's data controls](https://developers.openai.com/api/docs/guides/your-data) and [Anthropic's API data retention docs](https://platform.claude.com/docs/en/manage-claude/api-and-data-retention). Both providers require approval for ZDR. Without it, OpenAI keeps abuse-monitoring logs for up to 30 days. With it, Anthropic still excludes features such as batch processing, the Files API and code execution, and may retain content flagged by its trust and safety systems for up to two years. [^bench]: Public sandbox benchmarks from ComputeSDK: [DAX](https://www.computesdk.com/benchmarks/sandboxes/dax/) and [Burst TTI](https://www.computesdk.com/benchmarks/sandboxes/burst-tti). Harness commit `92fbbc9`; leaderboard run of 2026-09-18 used throughout. The Burst TTI composite blends median (60%), p95 (25%) and p99 (15%) against a 10-second ceiling and scales by success rate. Upstream numbers are one iteration per provider per day, so the charts are best read as "which third of the board" rather than a precise rank. [^isomorphic]: [Isomorphic Labs secures $2.1 billion funding to scale its AI drug design engine](https://www.prnewswire.com/news-releases/isomorphic-labs-secures-2-1-billion-funding-to-scale-its-ai-drug-design-engine-302769674.html), PR Newswire. [^gptrosalind]: [Introducing GPT-Rosalind](https://openai.com/index/introducing-gpt-rosalind/), OpenAI. Access is gated behind a trusted-access program rather than the general API. [^claudescience]: [Claude Science](https://www.anthropic.com/news/claude-science-ai-workbench), Anthropic. [^endpoints]: [Anthropic debuts Claude Science, an AI product for bioscience](https://endpoints.news/anthropic-debuts-claude-science-an-ai-product-for-bioscience/), Endpoints News. [^antdefaults]: [Updates to our Consumer Terms](https://www.anthropic.com/news/updates-to-our-consumer-terms), Anthropic. [^nytdata]: [Response to NYT data demands](https://openai.com/index/response-to-nyt-data-demands/), OpenAI. ZDR customers were excluded from the preservation order. [^oci]: [OCI Runtime Specification](https://github.com/opencontainers/runtime-spec), Open Container Initiative. [^pubmed]: [PubMed](https://pubmed.ncbi.nlm.nih.gov/), National Library of Medicine. [^clinicaltrials]: [ClinicalTrials.gov](https://clinicaltrials.gov/), U.S. National Library of Medicine. [^opencode]: [opencode](https://opencode.ai/), the open-source coding agent the DAX benchmark clones and builds. --- # Proteins Aren't Sentences: Why Bigger Protein Models Don't Win URL: https://www.lite.bio/blogs/benchplm Author: Anindyadeep, Nabajit Borah Date: 2026-06-23 A protein language model (PLM) turns an amino-acid sequence into vectors that can be reused for downstream protein tasks: stability, localization, binding, mutation fitness, evolutionary structure, and more. In the frozen-embedding setting, the pretrained PLM is not fine-tuned. We extract a sequence embedding from it, train a small supervised probe on top, and ask how much useful biological signal was already present in the representation. ![Image](images/8d79bca5ef01ac6dc990d586324a1e4424a57235-700x253.png) Several protein language models have been developed in recent years. Among the earliest, Meta (formerly Facebook) FAIR introduced the ESM-2 family of models. Since then, a growing number of PLMs, including: ESM2, ProGen, and DPLM, have demonstrated strong performance across various protein tasks. Each of these models learns its representations in a fundamentally different way, and that difference propagates into downstream research and the decisions built on it. To make the comparison concrete, we benchmark the following models: | Model | Parameters | Objective family | Readout direction | Frozen embedding interface | | --- | --- | --- | --- | --- | | ESM-C 300M | 300M | masked language model | bidirectional | mean-pooled encoder hidden state | | DPLM 150M | 150M | discrete diffusion | bidirectional | teacher-forced hidden-state pooling | | ESM-2 150M | 150M | masked language model | bidirectional | mean-pooled encoder hidden state | | ProGen3 219M | 219M | autoregressive language model | left-to-right | teacher-forced hidden-state pooling | | ProGen2-small | 151M | autoregressive language model | left-to-right | teacher-forced hidden-state pooling | # BenchPLM Protein sequences read a lot like text: residues behave like words and local motifs like phrases. That resemblance is what led researchers to borrow architectures from natural language processing. That observation led researchers to borrow ideas from natural language processing and develop different architectures for learning representations of proteins. How a model reads a protein determines what its embeddings can represent. Masked models, causal models, diffusion models all see sequence context differently, so they inevitably learn different representations. However those variations naturally leads to the following questions: - Which protein language models learn the best representations, and why? - Which protein language model should be used for downstream tasks? - What kind of representations have these models learned, and how much evolutionary knowledge do they capture? To answer these, we introduce PLMBench, a collection of 9 tasks designed to systematically evaluate protein language model representations. | Task | Split size (train / validation / test) | Primary metric | What it probes | | --- | --- | --- | --- | | Thermostability Regression | 5310 / 706 / 706 | Spearman rho | Sequence-level stability signal. | | DeepLoc multiclass localization | 10414 / 1368 / 1368 | Accuracy | Subcellular localization across multiple classes. | | DeepLoc binary localization | 6707 / 698 / 807 | Accuracy | A simpler binary localization ablation. | | Metal ion binding | 5797 / 719 / 719 | Accuracy | Compact function and binding-site signal. | | FLIP2 Alpha Amylase | 2574 / 644 / 488 | Spearman rho | Mutation-fitness ranking under a one-to-many split. | | FLIP2 Hydrophobic Core | 9974 / 2493 / 12468 | Spearman rho | Mutation-fitness ranking under a larger low-to-high shift. | Let’s understand the tasks in more details. In case if you are more interested to know the results, you can skip this section. 1. **Thermostability Regression:** Given a protein sequence, predict its thermostability score. We use Spearman's ρ, which measures how well the model ranks more-stable proteins above less-stable ones. 1. **DeepLoc Tasks:** Two classification tasks. In multi-class classification, we predict the subcellular localization of a protein (e.g., nucleus, cytoplasm, mitochondrion, membrane, extracellular). In binary classification, we predict whether a protein is membrane-bound or soluble. 1. **Metal Ion Binding:** Predict whether a given protein sequence binds metal ions. We do not predict the specific metal or binding site. 1. **FLIP2 Alpha Amylase:** Given mutant alpha-amylase sequences, the model ranks variants from worse to better function. We use Spearman's ρ, where higher values indicate better ranking accuracy. 1. **FLIP2 Hydrophobic Core:** Similar to above, but using mutant sequences from the hydrophobic core dataset. The task is to predict mutation fitness or structural packing quality. As a regression task, we again use Spearman's ρ, where higher values indicate stronger predictive ranking. Every model was evaluated on the same six tasks with the same train/validation/splits, the same embedding policy and the same shallow probe family. The protocol was: 1. Freeze the pre-trained PLM backbone 1. Extract sequence embeddings with sliding window mean pooling 1. Train a probe on the training split. 1. Select probe hyperparameters on the validation split. 1. Report the held-out test metric. One probe family is held fixed across all models, so every score reflects the representation and that probe jointly. We keep the probe shallow to stay close to the raw embedding, but absolute numbers would shift under a different probe; the cross-model ranking is the durable signal, not the decimal. ## Interpretability We also ran three diagnostic analyses. The first two are mechanistic probes; the third is a representation-geometry check. They are useful for interpretation, but they are not included in the benchmark score. | Diagnostic | Models covered | Question | | --- | --- | --- | | Evolutionary layer localization | ESM-2 scales, DPLM 150M, ProGen2-small | At which layer is evolutionary distance most readable? | | Context occlusion | ESM-2 150M, DPLM 150M, ProGen2-small | Does a residue-level score depend on left or right context? | | Embedding PCA and same-family retrieval | ESM-2 150M, DPLM 150M, ProGen2-small; larger cached check with ESM-2 650M, ESM-2 3B, ProGen2-medium | Do frozen embeddings form clean ortholog-family neighborhoods? | # ESM-C and DPLM Leads Overall The overall result is close at the top. ESM-C 300M and DPLM 150M both have a mean task rank of `1.83` across the six tasks. ESM-C wins more individual tasks (3 versus 2), while DPLM has the marginally higher average normalized primary score (0.938 versus 0.924). With six tasks and a single probe seed, a gap this small is inside the noise floor; the honest reading is that ESM-C and DPLM are tied at the top, not that one edges the other. ESM-2 150M is the next strongest model. Both ProGen checkpoints fall below the bidirectional models in this frozen-probe setting. ESM-C is also the largest backbone here, so this aggregate alone cannot separate architecture from scale. The same-size comparison does: ESM-2 and DPLM at 150M beat both ProGen checkpoints of equal or larger size, which isolates the effect we return to in the conclusion. ![Image](images/eb604344c08d3670f21fef6e98f84bbb89c55599-1280x620.png) The per-task view explains why the aggregate is close: | Task | Winner | Winning score | Runner up | | --- | --- | --- | --- | | Thermostability | ESM-C 300M | 0.664 Spearman | ESM-2 150M at 0.645 | | DeepLoc multiclass | ESM-C 300M | 0.814 accuracy | DPLM 150M at 0.812 | | DeepLoc binary | ESM-2 150M | 0.914 accuracy | DPLM 150M at 0.912 | | Metal ion binding | DPLM 150M | 0.711 accuracy | ESM-2 150M at 0.702 | | FLIP2 Alpha Amylase | ESM-C 300M | 0.676 Spearman | DPLM 150M at 0.613 | | FLIP2 Hydrophobic Core | DPLM 150M | 0.371 Spearman | ESM-C 300M at 0.355 | ProGen3 remains usable on some stability and classification tasks, but it does not win any task in this suite. ProGen2-small is weaker across the benchmark, with the clearest gap on FLIP2 Alpha Amylase: 0.105 Spearman, compared with 0.676 for ESM-C and 0.613 for DPLM. ![Image](images/065163a4ab3a412d93c7e313342ca432b54d87be-1280x650.png) The pattern is clear: on downstream tasks, models that condition each residue on the full sequence outperform models restricted to a left-to-right prefix. This is not a verdict on the ProGen family. ProGen models were trained to generate viable protein sequences. However the pattern is biologically plausible. Many protein properties depend on residues that are distant in sequence but coupled through structure, family constraints, or functional motifs. A frozen embedding that can integrate both upstream and downstream context gives a shallow probe more of the relevant signal. # Evolutionary Understanding When Protein Language Models were trained, researchers saw that the models were learning evolutionary relationship. This is one reason ESMFold dropped the MSA: the language model was meant to stand in for it. The payoff is generalization to sequences with few or no homologs. In this section, we tried to quantify how much evolutionary understanding does different protein language models carries. ## TimeTree TimeTree is a database of species divergence times. Given two species, it tells you how many millions of years ago they last shared a common ancestor. So for three proteins drawn from species X, Y, and Z, TimeTree can say whether X branched off closer to Y or to Z. This probe asks a single question: **at what layer depth does a protein language model start encoding evolutionary history rather than surface sequence?** For each model we sample a set of layers, pull the hidden states, mean-pool over residues to get one vector per protein, and check whether cosine distances between those vectors reproduce TimeTree's branching order. The test set is built from *identity-discordant triplets*. Each triplet has three proteins: - **A**, the anchor - **B**, the history-correct neighbor (TimeTree says A and B diverged more recently) - **C**, the identity-misleading decoy (raw sequence identity says A looks more like C) ![Image](images/a6d266eac13a56134226af26ac900ae2951ac294-1334x784.png) The two signals are deliberately put in conflict. By sequence identity alone, A resembles C more than B. By evolutionary history, A is closer to B. A model that only tracks surface similarity will be pulled toward C and fail. A model that has internalized deeper phylogenetic structure will still place B nearer. Triplet accuracy is just the fraction of triplets where the embedding agrees with history: distance(A, B) < distance(A, C) Chance is 50%. Scores above that mean the layer carries evolutionary signal beyond raw identity; scores at or below mean it does not. ![Image](images/516d4aed38fc8b3663499d82a853ea93e1d63384-1280x650.png) The same result is easier to inspect as a heatmap. ESM-2 and DPLM become more readable with depth, while ProGen2-small does not show the same late-layer gain. ![Image](images/3042fb8a60abe8d7dba0585bf4b153bd0c851559-1280x610.png) Not only accuracy, we also ask another question: Across many protein pairs, do embedding distances increase as TimeTree divergence times increase? Now consider this table: | Model | Objective | Direction | Best layer | Best triplet acc | Best Spearman rho | | --- | --- | --- | --- | --- | --- | | ESM-2 8M | Masked LM | Bidirectional | 6 / 6 | 0.602 | 0.374 | | ESM-2 35M | Masked LM | Bidirectional | 12 / 12 | 0.630 | 0.435 | | ESM-2 150M | Masked LM | Bidirectional | 30 / 30 | 0.609 | 0.412 | | DPLM 150M | Discrete diffusion | Bidirectional | 30 / 30 | 0.627 | 0.468 | | ProGen2-small | Causal LM | Left-to-right | 0 / 12 | 0.460 | 0.158 | Scale is not the axis here. ESM-2 35M matches 150M on this triplet probe, and DPLM 150M holds the strongest TimeTree-distance correlation despite no size advantage. What moves the signal is how the model reads context, not how many parameters it has. For ESM-2 and DPLM, the signal gets stronger in deeper layers. That suggests the model gradually builds a more evolution-aware representation as sequence context is processed. For ProGen2, the best TimeTree signal was at the input/early layer, and deeper causal layers did not improve it, which supports the blog’s argument that left-to-right models are weaker at forming bidirectional evolutionary geometry. ![Image](images/9621ad9f93a30f496fcbd998614683f6758d06c6-1280x650.png) ## Context Occlusion To test whether bidirectional models genuinely use context from both directions, we run a simple occlusion experiment inspired by input-perturbation methods common in neural network interpretability. The setup is straightforward. Pick a target residue and mask it. Record the model's log-probability for the correct amino acid at that position. Then mask one additional context residue at a known offset from the target and record the log-probability again. The difference between the two measurements tells us how much that context position contributed to the prediction. A large change means the context residue mattered; a small or zero change means the model was not relying on it. For ESM-2 and DPLM, both masked language models, we mask the target and measure perturbations symmetrically on both sides. For ProGen2, a causal language model, we score the target from its prefix only, since by construction it never sees residues to the right. ![Image](images/2ab3970ebf654ad52cf4683f0ff3a698a62e222f-1280x500.png) The binned heatmap reports mean absolute log-probability change by relative position. Two patterns stand out. 1. First, all three models show strongest sensitivity in the immediate neighborhood of the target (offsets -20 to +20), with effects decaying at greater distances. This is expected: local sequence context carries the strongest signal for residue identity. 1. Second, and more telling, is the asymmetry. ESM-2 and DPLM show measurable sensitivity on both sides of the target, with the right-side effect at +20 (0.050 for ESM-2, 0.044 for DPLM) comparable to or exceeding the left-side effect at -20. ProGen2-small shows a sharp spike at -20 (0.104) and exactly zero for every right-side bin. The right side is blank because those residues simply do not exist in the causal prefix used to score the target. ![Image](images/738490c88ab7b9b8dd1b08625cc4a6d94285ef67-1120x440.png) The bar plot compresses the full heatmap into a single left-versus-right summary. For each model, we sum the absolute effects from all left-side bins and all right-side bins, then report the right-context fraction. - ESM-2 150M: left 0.021, right 0.026, right fraction 0.552 - DPLM 150M: left 0.018, right 0.022, right fraction 0.556 - ProGen2-small: left 0.031, right 0.000, right fraction 0.000 ESM-2 and DPLM split their context dependence nearly evenly across both sides, with a slight lean toward the right. ProGen2's entire context mass sits on the left, exactly as a causal architecture dictates. Proteins are not sentences. In natural language, left-to-right context is a reasonable inductive bias because meaning largely builds sequentially. In proteins, the constraints on a residue's identity are fundamentally non-sequential. A residue may be constrained by a disulfide partner dozens of positions downstream, by residues it packs against in the folded structure, by compensatory mutations elsewhere in the sequence, or by a distant active-site motif it co-evolves with. None of these relationships respect sequence order. The occlusion results confirm that ESM-2 and DPLM capture this: their representations at any given position are shaped by the full sequence context. ProGen2 behaves correctly for its architecture, but its representations at each position are blind to everything that follows. This asymmetry likely explains part of the performance gap we observe in downstream tasks that require whole-sequence understanding. ## Do Ortholog Families Stay Together? As a final representation-level check, we reused the same TimeTree panel containing 1,024 proteins from 128 ortholog families, with exactly eight sequences per family. For ESM-2 150M, DPLM 150M, and ProGen2-small, we used the cached final mean-pooled embeddings from the phylogenetic-geometry runs. ![Image](images/ac2a24faa4bc639790bfa59e827d226cdc725a9e-1440x1120.png) The visualization below fits PCA separately for each model after L2-normalizing the embeddings, matching the cosine geometry used in the TimeTree benchmark. Colored points mark eight evenly spaced ortholog families; gray points are the remaining families. The plotted colored points use a tiny deterministic display jitter so duplicated or near-identical PCA coordinates remain visible. The panels are qualitative and independently scaled, so the quantitative readout is the bar plot underneath: same-family retrieval across all 128 families. Consider this table: | Source | Same-family NN@1 | Same-family P@7 | | --- | --- | --- | | Position identity | 0.661 | 0.431 | | 3-mer Jaccard | 0.850 | 0.787 | | ESM-2 150M | 0.855 | 0.798 | | DPLM 150M | 0.859 | 0.820 | | ProGen2-small | 0.641 | 0.397 | DPLM and ESM-2 produce the cleanest same-family neighborhoods, with DPLM slightly ahead by precision@7. ProGen2-small still retrieves a same-family nearest neighbor well above random chance, but its local neighborhoods decay quickly: across the seven nearest non-self neighbors, fewer than half are from the same ortholog family. We also repeated the same diagnostic for the larger cached checkpoints available on this panel: ESM-2 650M, ESM-2 3B, and ProGen2-medium as well. ![Image](images/dd11dafa5847be0caee7d448fade0531cd1e0b7b-1440x1120.png) The larger ESM-2 checkpoints improve family-neighborhood purity on this diagnostic, with ESM-2 3B reaching 0.875 P@7. ProGen2-medium does not show the same scaling behavior here: it is below the 3-mer baseline and below ProGen2-small by P@7. That result is local to this ortholog-family retrieval probe, but it reinforces the broader pattern that the causal ProGen embeddings are less useful as frozen neighborhood representations in these experiments. # The Curse of Sparse Data Most of the real world protein engineering datasets are small. A useful embedding should not only perform well with thousand of labels; it should also be sample-efficient when labels are scarce. Measured variants could be in the terms of 20,100 or 700. To benchmark this, the setup is simple. 1. We take a frozen PLM Embedding. 1. Train a small regression probe on K-labeled examples. 1. Predict fitness on held-out variants. 1. Repeat this on different label budgets. 1. Plot it as the curve and compare the differences. | Task | Training budgets | Test size used | What it stresses | | --- | --- | --- | --- | | FLIP2 GB1 one-vs-rest | 8 / 16 / 24 / 28 | 4096 | Extremely low-data fitness transfer. | | FLIP2 Alpha Amylase close-to-far | 64 / 128 / 256 / 512 / 1024 / 1782 | 1924 | Extrapolation from closer mutations to farther mutations. | | FLIP2 Rhodopsin by-wild-type | 32 / 64 / 128 / 256 / 512 / 700 | 184 | Membrane-protein fitness transfer. | So this means, if a model gets a good Spearman score with only 16 or 64 labels, that means its embedding is sample-efficient: the probe does not need much supervision because the representation already carries useful signal. ![Image](images/38aa1080ddd91e86bc82a8a08b25786a6fc7d63a-1280x720.png) The sparse-transfer result largely agrees with the full frozen benchmark. ESM-C has the best mean normalized curve score (0.771) and wins two of the three sparse tasks. ESM-2 is second by normalized curve score (0.691). DPLM is third overall (0.657) but is the strongest model on Rhodopsin at full budget. | Task | Best frozen model | Full-budget score | Interpretation | | --- | --- | --- | --- | | GB1 one-vs-rest | ESM-C 300M | 0.535 Spearman | ESM-C separates early, especially after 16 labels. | | Alpha Amylase close-to-far | ESM-C 300M | 0.299 Spearman | ESM-C is strongest under this mutation-distance shift. | | Rhodopsin by-wild-type | DPLM 150M | 0.688 Spearman | DPLM is the clearest winner on this membrane-protein fitness task. | The ProGen models do not close the frozen-transfer gap under sparse labels. ProGen3 improves on Rhodopsin as the budget grows, reaching 0.472 at 700 examples, but remains behind DPLM and ESM-2. ProGen2-small remains weak across the sparse suite. The limited-data conclusion is therefore aligned with our initial study: **ESM-C is the strongest broad frozen representation, while DPLM deserves special attention on Rhodopsin-like fitness transfer.** # Conclusion Across every probe in this study, the most consistent predictor of frozen-representation quality was not parameter count or pretraining scale. It was the conditioning set the model reads. Bidirectional encoders, ESM-2 and ESM-C as masked models and DPLM as a diffusion model, beat the causal ProGen checkpoints on every task that rewards whole-sequence understanding: downstream probes, evolutionary geometry, and same-family retrieval alike. The reason is visible if we write down what each model computes. A causal model factorizes the sequence left to right: $$ p(x) = \prod_{i=1}^{L} p\!\left(x_i \mid x_{[huggingface.co/LiteFold](https://huggingface.co/LiteFold/datasets) # Licensing AminoWeb does not relicense upstream data. Each dataset retains its source license: UniProt and AlphaFoldDB under CC BY 4.0, PDB coordinates under CC0 with mixed terms for derived metadata, ProteinGym constituents under their respective per-paper licenses, and benchmark releases under their original terms. The processing scripts and AminoWeb-specific metadata are released under Apache 2.0. Users adopting AminoWeb for commercial training should review the per-dataset license fields exposed in each dataset card. # **References** - Tsuboyama, K. et al. *Mega-scale experimental analysis of protein folding stability in biology and design.* Nature 620, 434-444 (2023). [https://doi.org/10.1038/s41586-023-06328-6](https://doi.org/10.1038/s41586-023-06328-6) - Notin, P. et al. *ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design.* NeurIPS Datasets and Benchmarks (2023). [https://www.proteingym.org/](https://www.proteingym.org/) - ByteDance Protenix authors. *Protenix: Advancing Structure Prediction Through a Comprehensive AlphaFold3 Reproduction.* bioRxiv (2025). [https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1](https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1) # **Citation** `@misc{litefold2026aminoweb, title={AminoWeb: A Curated Protein-Data Atlas for Biology and Protein Machine Learning}, author={LiteFold Team}, year={2026}, publisher={Hugging Face}, url={https://huggingface.co/LiteFold} }` --- # Rational Design of a Covalent EGFR T790M Inhibitor Using LiteFold URL: https://www.lite.bio/blogs/rational-design-of-a-covalent-egfr-t790m-inhibitor-using-litefold Author: Aditi Sinha, Cory Kornowicz Date: 2026-04-27 ### **The Clinical Imperative: The Evolutionary Arms Race in Non-Small Cell Lung Cancer** The rise of precision oncology has been driven by the understanding that specific genetic mutations can directly control cancer growth and survival. Targeted cancer therapy has greatly improved outcomes in non-small cell lung cancer (NSCLC), especially through inhibition of the Epidermal Growth Factor Receptor (EGFR). Activating mutations in EGFR, such as exon 19 deletions and L858R, lead to continuous signaling through growth pathways like RAS–RAF–MEK–ERK and PI3K–AKT, resulting in uncontrolled cell proliferation. However, in many lung adenocarcinomas, activating mutations in the EGFR kinase domain disrupt this regulation. The most common mutations, including exon 19 deletions and the L858R substitution, lock the receptor in a constitutively active state. As a result, cells receive continuous growth signals independent of external stimuli, leading to uncontrolled proliferation and tumor progression. First-generation inhibitors such as Gefitinib and Erlotinib initially show strong clinical response by reversibly binding to the ATP-binding pocket. However, resistance develops in most patients within a year. The primary cause of this resistance is the T790M mutation, where threonine is replaced by methionine at position 790. This mutation increases ATP affinity and introduces a bulkier hydrophobic side chain, making it difficult for reversible inhibitors to bind effectively. Third-generation inhibitors like Osimertinib address this issue by forming a covalent bond with Cys797. However, design challenges and resistance mechanisms still remain. In this study, we aim to address these challenges through a structure-guided, AI-assisted design approach. Using the LiteFold platform, we perform de novo generation and optimization of small molecules tailored specifically to the EGFR T790M binding pocket. By integrating molecular docking, covalent design principles, and molecular dynamics simulations, we evaluate not only the binding affinity of the designed molecules but also their dynamic stability within the receptor environment. This approach allows us to move beyond static predictions and establish whether the designed inhibitor can maintain a stable pre-covalent state, a critical requirement for effective covalent inhibition. ![RAS-RAF-MEK-ERK and PI3K-AKT pathways Involving EGFR transmembrane Receptor Tyrosine Kinase](images/23da2fd024cab29266df57b16e1d9d05d2f3d371-1000x500.png) First-generation reversible inhibitors Gefitinib and Erlotinib produce strong initial clinical responses, but acquired resistance develops in the majority of patients within approximately twelve months. The dominant resistance mechanism is the T790M gatekeeper mutation, in which threonine at position 790 is replaced by the bulkier methionine. This substitution simultaneously increases ATP-binding affinity and sterically hinders accommodation of reversible inhibitors, effectively restoring kinase activity despite drug treatment. Third-generation inhibitors, most notably Osimertinib, circumvent this resistance by forming an irreversible covalent bond with Cysteine 797 (Cys797) in the ATP-binding pocket, and by doing so are selective for the T790M mutant over wild-type EGFR. Yet even this class is not immune to further resistance most notably through the C797S mutation, which ablates the nucleophilic cysteine entirely. This study describes a structure-guided, AI-assisted workflow for de novo design and dynamic validation of a new covalent inhibitor candidate specifically tailored to the EGFR T790M binding pocket, using the LiteFold computational platform. ### The Rationale for Covalent Inhibition The pharmaceutical industry was historically cautious about intentionally reactive (electrophilic) molecules, owing to concerns over idiosyncratic toxicity, haptenization, and off-target covalent modification of unintended cellular proteins or nucleic acids. For decades, the prevailing view held that the risks of covalent targeting outweighed any benefits, a position complicated by the fact that several landmark drugs, including aspirin (cyclooxygenase inhibition) and the penicillin class (bacterial transpeptidase inhibition), owe their efficacy to serendipitous covalent mechanisms. Advances in structural bioinformatics, chemoproteomics, and structure-guided design have substantially rehabilitated the field. The modern approach to targeted covalent inhibition (TCI) employs two integrated molecular features: a selective non-covalent recognition scaffold that precisely positions the molecule within the target site, and a mildly electrophilic warhead that reacts irreversibly only once the correct geometry is achieved. In the context of EGFR T790M, the target nucleophile is Cys797, a non-catalytic residue at the solvent-exposed edge of the ATP-binding cleft. This residue is relatively rare across the broader kinome, making it an attractive orthogonal handle for achieving both potency and selectivity. Second-generation pan-HER inhibitors (Afatinib, Dacomitinib) and the mutant-selective third-generation inhibitor Osimertinib all exploit an acrylamide warhead to alkylate Cys797, extending progression-free survival in T790M-positive patients. **Inhibitor Generation Overview** | Class | Binding | Selectivity | Limitation | | --- | --- | --- | --- | | 1st gen (Erlotinib) | Reversible | WT + mutants | T790M resistance | | 2nd gen (Afatinib) | Covalent (pan-HER) | WT + mutants | WT toxicity | | 3rd gen (Osimertinib) | Covalent | T790M mutant | C797S resistance | **Computational Design Workflow** To address EGFR T790M resistance, a three-phase computational workflow was conducted using the LiteFold platform: 1. De novo molecular design — pocket-guided generation of candidate molecules 2. Molecular docking — evaluation of binding affinity and interaction geometry 3. Molecular dynamics (MD) simulation — assessment of dynamic stability in an explicit solvent environment ![Computational Discovery Pipeline: TIHK based Inhibitor for EGFR T790M](images/c8cffb8b10f881b863e3e00feba403d0788f2c82-1000x501.png) **Phase 1: De Novo Design** LiteFold's de novo design module generates small molecules directly from the geometry and chemical character of a specified binding pocket, rather than by decorating known scaffolds. The EGFR T790M pocket (PDB ID: 3IKA) was used as the design template, with explicit attention to four key residues: • Met790 — the gatekeeper mutation site, which introduces increased hydrophobicity and reduced pocket volume • Met793 — the hinge region, a critical hydrogen-bond donor/acceptor anchor for kinase inhibitors • Lys745 and Glu762 — catalytic residues that define the electrostatic environment of the ATP pocket • Cys797 — the intended covalent target, positioned at the solvent-exposed lip of the binding cleft The platform generated a library of candidate molecules that were immediately subjected to preliminary docking against the EGFR T790M crystal structure. The top five ranked compounds are shown below. | Rank | Molecule | Score | Assessment | Selected | | --- | --- | --- | --- | --- | | 1 | mol_88 | −9.71 | Outstanding | — | | 2 | mol_65 | −8.52 | Excellent | — | | 3 | mol_43 | −8.43 | Excellent | — | | 4 | mol_85 | −8.20 | Very strong | — | | 5 | mol_99 | −8.18 | Very strong | ✓ Selected | *Docking scores below −8.0 kcal/mol are generally associated with nanomolar-range binding affinities, though docking scores alone are not reliable quantitative predictors of experimental Ki values and should be interpreted with appropriate caution.* Despite not achieving the top docking score, mol_99 was selected as the primary scaffold for further development for three reasons: its interaction geometry placed the reactive vector of the nascent warhead in close proximity to the Cys797 sulfur; the scaffold demonstrated structural compatibility with acrylamide incorporation without predicted steric penalties; and its docking score spread across poses was narrow, suggesting a well-defined and reproducible binding mode rather than pose ambiguity. ![mol99 OV4 lead molecule.](images/d6eb1840bfd79867dc1cb8c69b00090de56a7576-671x614.png) **Phase 2: Molecular Docking and Binding Evaluation** The mol_99 scaffold, modified to carry an acrylamide warhead at the appropriate exit vector (mol99_OV4_lead_covalent.sdf), was docked against the active conformation of EGFR T790M (PDB ID: 3IKA) to generate a detailed binding profile. Across ten independently generated docking poses, the binding scores ranged from −6.84 to −6.35 kcal/mol a spread of approximately 0.5 kcal/mol. In computational drug design, a tight energetic cluster of this magnitude is a meaningful signal: it indicates that the molecule consistently samples a single binding mode within the ATP pocket rather than adopting multiple, potentially artifactual orientations. This pose consistency implies that the shape and electronic properties of the molecule are well-matched to the receptor's topography a hallmark of a robust lead candidate suitable for progression to dynamic simulation. ![visualization of docking of mol99 with the EGFR kinase domain.](images/a7f67ae63fef8dec3e125d2cefca499463d5d249-887x660.png) ******** **The Two-Step Kinetics of Covalent Inhibition** A critical conceptual point for interpreting both the docking and MD results is the two-step kinetic mechanism that governs all targeted covalent inhibitors. **Step 1 — Non-Covalent Association** The inhibitor must first enter the binding pocket and form a stable, reversible non-covalent complex (E·I). This step is governed entirely by classical intermolecular forces hydrogen bonding, van der Waals contacts, and hydrophobic packing and is described by the equilibrium dissociation constant Ki. Without a thermodynamically stable pre-covalent complex, the warhead is never positioned close enough to the target nucleophile for chemistry to occur. **Step 2 — Covalent Bond Formation** Only once the pre-covalent complex is stabilized can the electrophilic acrylamide warhead undergo Michael addition with the Cys797 thiolate, forming the irreversible covalent adduct (E–I). The rate of this irreversible step is governed by kinact, the maximal rate of enzyme inactivation. The overall potency of a targeted covalent inhibitor is therefore determined by the kinact/Ki ratio not by the warhead reactivity alone. This two-step paradigm explains why static docking alone is insufficient to validate a covalent inhibitor candidate. A molecule may appear optimally positioned in a frozen crystal-structure pose, yet rapidly drift or dissociate when exposed to the thermodynamic realities of an aqueous environment: protein conformational flexibility, entropic penalties, and solvent competition. Establishing that the designed inhibitor maintains a stable pre-covalent orientation over biologically relevant timescales is therefore a prerequisite before proceeding to synthesis or advanced covalent modelling. **Phase 3: Molecular Dynamics Validation** **Simulation Setup** Molecular dynamics (MD) simulation treats each atom as a classical particle obeying Newtonian mechanics. By computing interatomic forces and integrating Newton's equations of motion across femtosecond time steps, MD generates a continuous, high-resolution trajectory of atomic behaviour effectively serving as a computational microscope capable of resolving protein ligand dynamics that static docking cannot capture. The solvation model is a particularly important methodological choice. Protein structure and function are inextricably coupled to the solvent environment: water molecules mediate hydrogen bonding networks, drive the hydrophobic effect that contributes to ligand binding, and determine the desolvation penalty that the Cys797 thiolate must pay before covalent attack. For this study, the system was solvated using the explicit TIP3P (Transferable Intermolecular Potential with 3 Points) water model a three-site rigid representation with partial charges on oxygen and hydrogen, combined with a Lennard–Jones potential centered on the oxygen. Although four- and five-site water models offer improved reproduction of certain bulk-water properties, TIP3P remains the standard for atomistic protein–ligand simulations in conjunction with AMBER or CHARMM force fields, where it reliably reproduces the electrostatic environment of kinase binding pockets. **Ensemble Strategy** To avoid the well-recognised risk of a single MD trajectory becoming trapped in a minor conformational sub-state, an ensemble approach was employed: 1. Three independent 10 ns replicate simulations, each initialised with a distinct set of randomised atomic velocities, were executed to sample divergent thermodynamic pathways and verify reproducibility of the binding pose. 2. One extended 100 ns production simulation was conducted to probe the long-term kinetic stability of the complex, including any slow-onset conformational transitions or delayed dissociation events that fall outside the window of shorter simulations. All simulations were run following standard preparation: the protein–ligand complex was placed in a periodic boundary box, solvated with TIP3P water molecules, charge-neutralised with counter-ions, energy-minimised to remove steric clashes, and thermodynamically equilibrated to physiological conditions (300 K, 1 atm). ![Image](images/7c7929f6542324508a2bf1eac288d52fc548db23-907x382.png) **RMSD Analysis and Thermodynamic Stability** Root Mean Square Deviation (RMSD) the average atomic displacement relative to the initial reference structure was used to quantify both global protein stability and local ligand mobility throughout the 100 ns trajectory. The protein backbone RMSD stabilised and plateaued at approximately 3.0–4.0 Å (0.3–0.4 nm) within the first tens of nanoseconds. This magnitude of fluctuation is expected and physically meaningful for a multidomain kinase: proteins are not rigid crystals, and the conformational breathing required for function is captured here as it would be in a real biological setting. The plateau rather than continued drift confirms that the simulated EGFR T790M structure maintained overall structural integrity throughout the run. More critically, the ligand heavy-atom RMSD remained tightly constrained throughout the simulation. A sustained ligand RMSD below 2.0–3.0 Å is the widely accepted criterion for a stable, pharmacologically viable binding pose in MD-based drug design. A molecule with poor complementarity to its binding site will rapidly destabilise, typically producing RMSD excursions exceeding 5.0 Å as it tumbles into the bulk solvent. No such behaviour was observed for the OV4 lead compound. Global thermodynamic indicators were equally reassuring: potential energy stabilised around −936,000 kJ/mol and kinetic energy held steady at approximately 204,800 kJ/mol across the entire trajectory, with no evidence of sudden steric clashes, energetic discontinuities, or partial unfolding events. ![Image](images/e452cf288f3ca063bb068b58bde0268910d822da-898x360.png) **Interpreting the 35 ns Structural Transition** The most informative event in the extended trajectory was a discrete upward shift in ligand RMSD occurring at approximately 35 ns, at which point the ligand deviation rose to parallel the protein backbone at approximately 3.5 Å. A superficial reading of this event might suggest ligand instability. The extended 100 ns window provided the context necessary to interpret it correctly. EGFR and other kinases are intrinsically plastic macromolecules: over extended nanosecond timescales it is routine for local side-chain rotamer flips, activation loop breathing motions, or solvent network reorganisations to occur in response to a bound ligand a phenomenon known as induced fit. Critically, following the 35 ns positional translation, the ligand RMSD did not continue to rise. Instead, it immediately established a new horizontal plateau at approximately 3.5 Å that persisted for the remaining 65 ns of the trajectory. This behaviour is diagnostic of induced-fit accommodation rather than incipient dissociation. Had the shift reflected unfavourable binding energetics, the RMSD would have continued an unabated upward trajectory towards solvent escape. Instead, the system transitioned into a slightly reorganised thermodynamic equilibrium, confirming that the pre-covalent complex is robust against the receptor's natural conformational dynamics. **Interaction Fingerprints: Confirming Pre-Covalent Geometry** Two additional metrics from the 100 ns trajectory directly support the molecule's readiness for covalent chemistry: **1. Ligand–Binding Site Distance** The distance between the molecule and the core binding site fluctuated within a constrained range of 0.4–0.5 nm (4–5 Å) throughout the full trajectory. Successful Michael addition requires the acrylamide β-carbon to approach within approximately 3.6–4.0 Å of the Cys797 sulfur. The MD data confirm that the ligand does not drift away from the reactive zone at any point during the simulation. ![Image](images/edeeea9535a43ed6e292a945182e60cde748865e-899x370.png) **2. Sustained Hydrogen Bond Network** The ligand maintained an uninterrupted network of three hydrogen bonds with the surrounding pocket residues across the entire 100 ns run. While the broader protein structure naturally flexed sustaining approximately 280–300 internal hydrogen bonds with normal fluctuation the three ligand-specific anchor contacts never broke. This persistent electrostatic tethering provides the geometric constraint necessary for the acrylamide warhead to remain properly oriented relative to Cys797. ![Image](images/6c8dd7c2b860790c2054cb27175370154cf19b19-894x365.png) **Summary and Comparative Context** The table below places the designed OV4 lead compound in context relative to established inhibitors and the initial de novo scaffold from which it was derived | Molecule | Binding | Score | Strength | Limitation | | --- | --- | --- | --- | --- | | Gefitinib | Reversible | ~−6 to −7 | Hinge binding (Met793) | Clash with Met790 | | Osimertinib | Covalent (Cys797) | ~−8 to −9 | Mutant selective | C797S resistance | | mol_99 | Non-covalent | −8.18 | Stable pocket fit | No covalent action | | OV4 (lead) | Covalent-ready | −9+ | Fits Met790 + aligns Cys797 | Needs validation | **Conclusions** The results of this computational study support three main conclusions: The LiteFold de novo design module is capable of generating acrylamide-warhead inhibitor candidates with docking profiles comparable to established third-generation EGFR inhibitors, without starting from a known scaffold. The mol_99/OV4 lead compound maintains a stable pre-covalent binding state under dynamic conditions across a 100 ns MD trajectory, as evidenced by constrained ligand RMSD, persistent hydrogen bonding, and sustained proximity to Cys797. The 35 ns structural transition is consistent with induced-fit pocket accommodation rather than dissociation. The integration of AI-guided de novo generation with MD-based dynamic validation represents a meaningful advance over static docking workflows for covalent inhibitor design, where pre-covalent geometric fidelity is a mandatory prerequisite for effective warhead chemistry. These computational findings establish a strong rationale for progression to experimental validation including synthesis of the OV4 candidate, biochemical Ki and kinact/Ki determination, and cellular potency profiling in T790M-positive NSCLC models. The extent to which computationally predicted binding geometries translate to measured inhibitory activity will be the definitive test of this workflow. --- # Improving Binding Precision of Therapeutic Antibodies with Rosalind by LiteFold URL: https://www.lite.bio/blogs/improving-binding-precision-of-therapeutic-antibodies-with-rosalind-by-litefold Author: Aditi Sinha, Cory Kornowicz Date: 2026-04-22 **Disclaimer:** This is a purely in-silico case study intended to demonstrate the computational capabilities of the LiteFold platform and its in-house AI co-scientist, Rosalind. None of the designs reported here have been experimentally validated. All claims about "improved" metrics refer to model-internal confidence scores, not to empirical binding affinity. Rosalind, as referenced in this document, is LiteFold's proprietary AI co-scientist and is unrelated to OpenAI's GPT-Rosalind release. ## From Discovery to Design For decades, monoclonal antibody discovery depended heavily on biological chance. Researchers immunized mice, generated hybridomas, screened thousands of clones, and hoped that one antibody would show the right combination of affinity, specificity, and stability. This approach produced some of the most successful therapies in history, including pembrolizumab (brand name Keytruda). The traditional pipeline is slow, expensive, and bounded by what natural immune selection happens to generate. The emerging paradigm of generative biologics shifts the work from *discovering* antibodies to *designing* them with computational tools. At the center of this shift is LiteFold, a biomolecular AI platform that brings advanced protein design tools into the hands of researchers. In this case study, LiteFold's AI co-scientist Rosalind was given a focused test: could we computationally redesign the CDR loops of one of the most successful therapeutic antibodies ever developed, while preserving its overall framework and binding mode? ## **Understanding Pembrolizumab and the PD-1 Axis** Pembrolizumab is a humanized IgG4κ monoclonal antibody used in oncology as an immune checkpoint inhibitor. Rather than directly killing tumor cells, it modulates the immune system. Its target is Programmed Cell Death Protein 1 (PD-1), a receptor expressed on activated T cells. Its target is **Programmed cell death protein 1 (PD-1)**, a receptor expressed on activated T cells and other immune cells. PD-1 plays a regulatory role in maintaining immune balance. Under normal physiological conditions, it prevents excessive immune activation and protects healthy tissues from immune-mediated damage. PD-1 interacts with two ligands: - **Programmed death-ligand 1 (PD-L1)** - **Programmed death-ligand 2 (PD-L2)** These ligands can be expressed by tumor cells or cells within the tumor microenvironment. When PD-1 binds to PD-L1 or PD-L2, it sends an inhibitory signal into the T cell. This signal reduces T-cell activation and limits its ability to attack Many cancers exploit this pathway by overexpressing PD-L1, effectively pressing the immune system’s “brake pedal” and avoiding immune destruction. ## Mechanism of Action: Releasing the Immune Brake Pembrolizumab binds to the extracellular domain of PD-1 specifically the surface that normally engages PD-L1 and PD-L2. By occupying this interface, it prevents PD-1 from interacting with its ligands As a result: - The inhibitory signal is blocked. - T cells remain active. - Anti-tumor immune responses are restored. Rather than killing tumor cells directly, pembrolizumab reactivates cytotoxic T cells, allowing them to perform their natural anti-tumor function. This strategy has transformed oncology and has led to durable responses in multiple cancer types. In simple terms, cancer hides by suppressing immune activity. Pembrolizumab removes that suppression Using LiteFold and Rosalind, we generated novel antigen-binding loop variants while preserving the overall antibody framework. The objective was to demonstrate the platform's ability to maintain structural integrity while exploring alternative interaction surfaces. The generated variants achieved high structural confidence scores and predicted PD-1 engagement at the interface, suggesting that LiteFold can serve as a rapid candidate generation engine for antibody engineering pipelines. Rather than relying solely on immune selection, researchers can now generate and evaluate new antibody variants computationally, with experimental validation as the next step. ## **The Challenge: Redesigning a Blockbuster** Pembrolizumab targets the Programmed Cell Death Protein 1 (PD-1) receptor, a critical immune checkpoint. By blocking the interaction between PD-1 and its ligand PD-L1, the antibody releases the "brakes" on the immune system, allowing T-cells to attack tumors. The specificity of an antibody lies in its Complementarity-Determining Regions (CDRs) six flexible loops at the tip of the Y-shaped molecule. Of these, the heavy chain CDR3 (CDR-H3) is the most critical and the most diverse. Redesigning these loops is biologically perilous; a single amino acid change can abolish binding or destabilize the protein structure. **The Experiment:** Using the LiteFold platform, we tasked the Rosalind AI Co-Scientist with a specific objective: 1. **Input:** The crystal structure of the PD-1/Pembrolizumab complex. 1. **Task:** Retain the antibody framework but use generative diffusion models to redesign the heavy chain CDR regions. 1. **Goal:** Generate a novel sequence that is chemically distinct from the wild-type but possesses high predicted structural compatibility with the PD-1 antigen. This experiment used the Boltz-2 co-folding model for structure prediction and the BoltzGen diffusion model for generative design, both integrated into LiteFold. Boltz-2 is a state-of-the-art co-folding model for protein-protein and protein-ligand complex structure prediction. ![Image](images/60f106ff2f6da26571acbaeccd94a64fe87258e6-637x806.png) ## **From Prompt to Prediction** One of the defining features of LiteFold is the seamless integration of complex computation into an intuitive interface. This experiment did not require setting up local GPU clusters or managing command-line dependencies. The process began with Rosalind. Acting as a "digital medicinal chemist," Rosalind analyzed the target PDB file to identify the binding interface. Unlike traditional methods that rely on rigid docking (fitting a key into a lock), Rosalind used a diffusion-based generative approach. It denoised a cloud of atoms into new loop geometries conditioned on the energetic and geometric requirements of the PD-1 pocket. We started by uploading PDB entry 5GGS, the 2.0 Å crystal structure of the pembrolizumab Fab in complex with PD-1. Rosalind parsed the chains automatically: chains A and C correspond to the pembrolizumab heavy chain, chains B and D to the light chain, and chains Y and Z to the PD-1 extracellular IgV domain. For this experiment we worked with one copy of the complex (chains A, B, Z). ![Image](images/91d44628fd8fc876d1154da1a2b043b0ea36ff36-577x367.png) After this, Rosalind's runs an initial assessment of the interface, calculating the buried surface area (BSA) and identifying "hotspot" residues. correctly identifies residues 99-108 (RDYRFDMGFD) as the CDR-H3 loop critical for binding. Followed by, we issued a natural language command: 💡**Rosalind, redesign the CDR-H1, H2, and H3 loops of Chain B. Optimize for binding affinity to Chain A. Generate 100 variants and filter for high structural confidence.** Instead of sampling from a fixed library (as in phage display), Rosalind initiates a diffusion-based design workflow built on the BoltzGen model for backbone and sequence generation, with Boltz-2 used downstream for co-folding and structural validation of the designed complex. Diffusion models work by progressively adding noise to atomic coordinates and then learning to denoise them into valid protein structures conditioned on the target antigen. This lets the model explore loop geometries that natural evolution may never have produced. ![The left panel shows the structural mapping of selected residues targeted for redesign, while the right panel provides a sequence-level comparison specifying the precise mutation sites and the corresponding amino acid changes between the original and redesigned antibody](images/5e27a8afef3857265633f9cf3fa0dde08e044a20-1000x504.png) ![The Picture panel provides a sequence-level comparison specifying the precise mutation sites and the corresponding amino acid changes between the original and redesigned antibody](images/963453f4f001746f01b7cd9f9d1ff0f2bf5bcd19-691x739.png) Within minutes, Rosalind generated the variants. Crucially, it did not just output sequences; it co-folded them. Using the Boltz-2 engine, she predicted the 3D structure of each new antibody variant docked against the PD-1 antigen. She also discarded unstable "hallucinations" (low pLDDT) and non-binders. She then presented the top candidate based on the **Complex iPDE** metric, generating the visualization dashboard attached in this report. ![Image](images/0086383c5c8114a83162ea967dd5c1e0127751a7-1000x676.png) ## The Metric Engine And Results Generating a sequence is easy; validating it is hard. LiteFold provides an automated battery of confidence metrics that serve as a proxy for wet-lab success. For this experiment, we focused on four key indicators: - **pLDDT (Predicted Local Distance Difference Test):** A measure of local structural confidence. - **pTM (Predicted Template Modeling):** A measure of global topological accuracy. - **Clash Score:** A check for physical realism (steric overlaps). - **Complex iPDE (Interface Predicted Distance Error):** The gold standard for binder ranking. To evaluate the success of the experiment, we compared the LiteFold-generated variant against a baseline run of the unmutated, natural antibody (Keytruda). The data, visualized in the attached analytical dashboards, reveals a striking narrative. ![Image](images/b1e772567ffaf7590e278893431ae18b8e57d1d0-781x398.png) The most significant finding in this dataset is the **Complex iPDE (Interface Predicted Distance Error)**. - **What it means:** iPDE is the model's predicted positional error at the binding interface, expressed in Å. Values below 1.0 Å are commonly used as a strong prioritization signal in computational design pipelines, indicating that the structure model is highly confident in the predicted interface geometry. This is a confidence score, not a measurement of binding affinity. - **The Result:** The natural antibody baseline run returned an iPDE of **2.42 Å**. While this is acceptable, it suggests some degree of predicted flexibility or uncertainty in the model's handling of the wild-type loop. - **The Design:** The LiteFold-generated variant achieved an iPDE of 0.612 Å **Analysis:** It is important to be precise about what this means. iPDE is the model's predicted positional error at the interface, that is, a confidence score the model assigns to its own prediction. A lower value means the model is more certain about where the interface atoms sit, not that the molecule binds better in reality. The honest reading is that the diffusion model produced a sequence whose predicted complex geometry the structure model resolves with high internal confidence. This is a strong prioritization signal for downstream evaluation, and it makes this variant a high-priority candidate for wet lab testing. It is not, by itself, evidence of improved binding. ## Molecular Dynamics Simulation (MD) Setup To validate the structural stability and behavior of the designed biologic, molecular dynamics (MD) simulations were performed. MD simulations model the motion of atoms over time using physical laws, allowing us to observe how the protein behaves in a realistic environment . The designed protein structure obtained from LiteFold Rosalind was used as the starting structure. The system was prepared by adding hydrogen atoms Assigning appropriate force field parameters Solvating the system in a water box Neutralizing with ions. This setup ensures that the system mimics physiological conditions before running the simulation. The molecular dynamics simulation was carried out in multiple stages to ensure system stability: 1. Energy Minimization The system was minimized to remove steric clashes and unfavorable interactions. 1. Equilibration PhaseNVT ensemble (constant volume and temperature) NPT ensemble (constant pressure and temperature)This step stabilizes temperature and pressure of the system. 1. Production Run A full MD simulation was performed for 100 ns (10 simulations ran for 10ns each), generating trajectory data representing atomic motion over time. During simulation, parameters such as temperature, pressure, and energy were monitored to ensure system equilibrium. ## Visualizing the structures ![Image 1: The design The visualization of the mutated design highlights the redesigned CDR regions (visible in yellow/orange against the blue framework). The high pLDDT (0.96) indicates that these loops are predicted to be structured and stable. Often in de novo or redesigned structures, loops are "hallucinated" as floppy or disordered. Here, Rosalind engineered a structured loop that likely adopts a rigid conformation upon binding, minimizing the entropic cost of interaction.](images/758676357ab055738e66822ee3273dc5d2ee7a3d-470x443.png) ![Image 2 (The Baseline): The baseline structure, while robust, shows a lower global pTM (0.733). This might reflect the natural flexibility of the PD-1/PD-L1 interaction axis. PD-1 is known to have flexible loops (C'D and FG loops) that undergo conformational selection. The fact that the AI Design stabilized this interaction (higher pTM and lower iPDE) suggests the generated antibody might possess "lock-and-key" characteristics that are highly desirable for therapeutic potency.](images/f542140087e63c690fb941dadfe82a14957a7307-1000x438.png) **** ## MD Simulation Results and Analysis The MD simulation results were analyzed to evaluate structural stability, flexibility, and interaction patterns of the designed biologic. > Please note that at LiteFold, we perform composable Molecular Dynamics. This means in the above figure you are seeing one 10ns simulations. Similarly we performed 10 more 10ns simulation where each new simulation started from the last checkpoint. This design decision massively helps in parallelizing different MD jobs and also run as many nano seconds as possible. **RMSD (Root Mean Square Deviation)** RMSD was calculated to assess structural stability over time. A stable RMSD indicates that the protein maintains its conformation during simulation. ![Image](images/f1306c8ea456006e98ce6d4aae3dc1cc3a798f72-892x362.png) RMSD stabilized over the trajectory, indicating that the designed antibody maintains its overall fold without major structural deviation across the 100 ns aggregate sampling window. **RMSF (Root Mean Square Fluctuation)** RMSF analysis shows residue-level flexibility. ![Image](images/94f95da12c9035ee211fca3496f76f115d1d46b4-849x343.png) We observed higher fluctuations observed in loop regions. However the core residues remained stable. **Radius of Gyration (Rg)** Rg measures compactness of the protein. ![Image](images/adb57a8510d58d2ccccb6fad1bff5bc5aab5b3c7-868x350.png) Stable Rg indicates compact and well-folded structure ## **Why This Matters for Drug Discovery** The success of this internal LiteFold experiment has broad implications for the biopharma industry. ### **Speed and Efficiency** Traditional antibody engineering humanization and affinity maturation can take months of phage display cycles. In this experiment, Rosalind generated and evaluated a high-probability candidate in minutes. The Complex iPDE of 0.612 Å demonstrates the platform's ability to generate high-confidence computational candidates, potentially allowing researchers to prioritize molecules for wet lab validation and reduce experimental screening burden. ### **The Rise of "Bio-Betters"** A major commercial application of this technology is the creation of "bio-betters" and biosimilars. As patents on major biologics expire, there is strong commercial interest in molecules that engage the same target with different sequences (to navigate IP) and ideally with improved properties such as stability or formulation. This case study demonstrates LiteFold's ability to generate sequence-diverse CDR variants against a known epitope. The redesigned loops are chemically distinct from the Keytruda CDRs, which is the right starting point for that kind of program. Whether the resulting molecule is functionally equivalent to the originator is an empirical question for the wet lab. ### **Democratizing High-Performance Computing** Perhaps the most important takeaway is the accessibility. Running a co-folding model like Boltz-2 typically requires significant computational infrastructure. LiteFold packages this into a web-accessible platform. A researcher can upload a PDB, ask Rosalind to "optimize the interface," and receive a detailed report with industry-standard metrics like pLDDT and iPDE automatically calculated. ### Role of MD in This Project Molecular dynamics simulation plays a critical role in validating computational protein design.While LiteFold Rosalind generates structurally optimized proteins, MD simulation ensures that: - The structure is stable over time - The protein behaves realistically in a simulated biological environment - No hidden instabilities or structural breakdown occur MD analysis also provides insights into molecular interactions, flexibility, and thermodynamic stability, which are essential for biologic design and therapeutic applications .Thus, MD acts as a bridge between **design (AI-based prediction)** and **real-world biological feasibility**. ## **Technical Deep Dive: The Metrics of Success** To fully appreciate the results, it is worth expanding on the specific metrics provided by the LiteFold engine in this experiment. ### **The Confidence Score (pLDDT)** The Design achieved a **Mean pLDDT of 0.96**. In the AlphaFold/Boltz scale, anything above 0.90 is considered "high accuracy," comparable to crystal structures. This tells us that the side-chains in the core of the antibody and at the interface are packed tightly. A common failure mode in AI design is "molten globule" structures proteins that look right from a distance but lack internal packing. The 0.96 score confirms that the Rosalind-designed antibody has a solid, drug-like hydrophobic core. ### **The Reality behind clash scores** Both the Design (5.41) and the Baseline (4.88) show moderate clash scores. In raw PDB files from generative models, this is standard. Diffusion models approximate the atom positions from a noise distribution. These scores are easily resolved by a standard energy minimization (relaxation) step using a force field like Amber or Rosetta. The fact that the Design's clash score is comparable to the Baseline indicates that the mutations did not introduce any severe steric conflicts that would render the molecule unfolded. ### **Ligand ipTM vs. Protein ipTM** The dashboard shows Ligand ipTM: 0.000 and Protein ipTM: 0.911. This is the expected behavior for this system. PD-1 is a protein, so the protein-protein ipTM is the relevant score and is high. The Ligand ipTM is 0.000 because there is no small molecule ligand in the complex; the field is simply inactive, not a meaningful zero score. The Protein ipTM of 0.911 indicates strong predicted compatibility at the protein-protein interface.**Conclusion: The LiteFold Advantage** This case study serves as a proof-of-concept demonstration of the LiteFold platform's capabilities. By applying our computational workflow to a well-characterized therapeutic antibody, we generated variants that achieve superior predicted structural confidence metrics compared to baseline computational models. While these represent computational predictions requiring experimental validation, the results demonstrate the platform's potential to accelerate antibody discovery timelines. The Complex iPDE of 0.612 Å is the headline figure. It indicates that the structure model is highly confident in the predicted geometry of the designed interface, which makes this variant a high-priority candidate for downstream evaluation. Whether the AI-generated CDR loops actually engage the PD-1 epitope as predicted is a question for the wet lab. As we advance LiteFold's development, we continue integrating enhanced molecular dynamics simulation and expanded computational chemistry capabilities into Rosalind. This case study demonstrates the platform's current ability to rapidly generate and evaluate antibody variants, positioning LiteFold as a valuable tool in modern computational antibody discovery pipelines that complement experimental approaches. **Summary of Key Findings:** - **Platform Demonstration**: Rosalind successfully generated novel CDR loop variants for the Pembrolizumab framework, showcasing the computational design workflow. - **Improved Computational Metrics**: The AI-generated variant achieved superior predicted Interface Distance Error compared to baseline (0.612 Å vs 2.42 Å), indicating higher model confidence. - **Strong Predictive Scores**: Global structure prediction (pTM) and local confidence (pLDDT) metrics were exceptionally high for the generated variant. - **Experimental Testing Priority**: The computational metrics suggest this variant represents a high-priority candidate for wet lab validation studies. - **Structural Stability Prediction**: MD simulations indicate the designed variant maintains structural integrity and exhibits realistic dynamic behavior over time. - **Platform Validation**: Results demonstrate LiteFold's capability to rapidly generate computationally optimized antibody candidates suitable for experimental pipelines. The MD simulation confirms that the designed biologic is structurally stable, maintains its conformation, and exhibits realistic dynamic behavior over time. This supports the reliability of the protein design generated using LiteFold Rosalind. ## **Appendix: Technical Data & Visual Evidence** This appendix contains the raw data tables extracted from the experiment, serving as the evidential basis for the analysis above. ## **Appendix A: The Design (Mutated CDRs)** *The LiteFold-generated variant with redesigned CDR loops.* ![Image](images/943df0310e99743073d753777a83ed1b65db074c-1000x499.png) **Visual Description:** The 3D visualization displays the antibody Fab fragment (blue) bound to the PD-1 antigen. The redesigned CDR loop is highlighted in **yellow/orange**. The structure appears compact with no visible backbone breaks. ![Image](images/6b6ff213849980a4b41d63a400db35e0cb40725c-1000x617.png) ## **Appendix B: The Baseline (Wild-Type Pembrolizumab)** *The unmutated, natural antibody control run.* ![Image](images/28a77031f42cc6f40f991df3e7c1879f5e81d949-1000x498.png) **Visual Description:** The 3D visualization shows the standard wild-type complex. The coloring is uniform (blue), representing the native sequence without modification highlights. ![Image](images/029e6fb7ddc2c5e52c4f8dbd00b4417baa638efa-1000x687.png) --- # Ensemble Docking vs Static docking. When Does Protein Flexibility Matter? URL: https://www.lite.bio/blogs/ensemble-docking-vs-static-docking-when-does-protein-flexibility-matter Author: Aditi Sinha Date: 2026-01-05 If one were to trust the diagrams found in introductory biology textbooks, molecular recognition would appear to be a serene, orderly, and deterministic affair. The "Lock and Key" model, proposed by Emil Fischer in 1894, depicts the protein as a rigid, Pac-Man-like entity with a mouth the active site gaping open in a fixed, immutable geometry. The ligand, shaped conveniently like a specific wedge of cheese or a geometrically perfect key, floats passively through the cytosol until it slots perfectly into the gap. *Click.* Biochemistry happens. This model is elegant. It is intuitive. And, for the doctoral student staring at a docking failure rate of 90% on a Friday evening, it is an absolute fabrication. The persistence of this analogy has created a "square peg in a round hole" cognitive dissonance in the field of Computer-Aided Drug Design (CADD). The reality of the microscopic world is not static; it is a chaotic storm of thermal fluctuations. Proteins are not locks; they are more akin to Jell-O molds vibrating in a washing machine. Ligands are not keys; they are flexible chains of atoms desperately seeking an energetic minimum while being bombarded by solvent molecules. When a researcher attempts to perform rigid-receptor docking forcing a static ligand into a static crystal structure they are essentially trying to park a Cadillac in a garage that is currently breathing, shifting, and occasionally collapsing upon itself. ![Fig.1| Molecular Recognition models](images/bc12b0cedd8b12b77406390c7e341075d34a10ef-1000x498.png) ### 1. The Cultural Despair of the Docking Community The psychological toll of this thermodynamic reality is well-documented in the digital confessionals of the scientific community. On forums like Reddit, the frustration manifests in memes and threads titled "Molecular docking struggles," where bioinformaticians vent about the absurdity of their results. Users share the existential dread of obtaining binding affinity scores of +500 kcal/mol because a single side-chain rotamer in the receptor clashed with the ligand, resulting in a physical impossibility that the software scores as "highly unfavorable" rather than "physically impossible". These cultural artifacts the memes about "forbidden" bond angles and the despair over software errors mask a serious scientific bottleneck. The humor is a coping mechanism for the limitations of the "Lock and Key" paradigm. The inability of static docking to account for protein flexibility results in false negatives: missed drug candidates that could have cured a disease but were rejected because the crystal structure's active site was closed by 1.5 Ångströms at the moment of crystallization. ### 1.2 The Transition to Rigor However, the "funny" side of docking failures quickly turns serious when the stakes are human health. The exclusion of protein flexibility is not just a source of frustration; it is the primary source of error in Structure-Based Drug Design (SBDD). Rigid docking efforts typically show performance rates between 50% and 75%, while methods that account for flexibility can enhance pose prediction accuracy to 80–95%. This Blog transitions from the "static fallacy" to the rigorous, high-performance computing solutions that address it. We explore the paradigm shift from static docking to **Ensemble Docking**, a methodology that acknowledges the chaotic, wiggling nature of proteins. We will dissect the integration of Molecular Dynamics (MD) simulations to generate "snapshots" of protein motion, analyze when this computational expense is justified, and provide exhaustive case studies from the cryptic pockets of IL-2 to the shapeshifting active sites of kinases demonstrating that in the world of molecular docking, flexibility is not just a feature, it is the function. ![Fig.2| From static to Flexible docking](images/8e9185056fa6a8a4fd80fb6ac4176e17e2dedfa4-1000x508.png) ## 2. The Physics of Wiggling: Thermodynamics of Molecular Recognition To understand why ensemble docking is necessary, we must first dismantle the physics of the static model. Proteins are thermodynamic ensembles, not statues. They exist as a population of conformations, navigating a complex energy landscape. ### 2.1 The Models of Binding The limitations of the "Lock and Key" model led to two sophisticated successors that inform how we approach flexible docking. #### 2.1.1 The Induced Fit Model Proposed by Daniel Koshland, this suggests the "Hand in Glove" analogy. The ligand binds to a ground-state conformation, and the interaction energy drives the protein into a new, bioactive conformation. Modeling this computationally is difficult because it implies the protein's shape change is dependent on the specific ligand, requiring expensive "Induced Fit Docking" (IFD) protocols that refine side-chains *after* the ligand is placed. #### 2.1.2 The Conformational Selection Model Ensemble docking is theoretically grounded in **Conformational Selection**. This posits that the protein naturally samples a variety of conformations in solution, even without a ligand. The "bioactive" conformation is simply one of these pre existing high-energy states (a "snapshot") within the equilibrium ensemble. The ligand selectively binds to and stabilizes this specific conformation. If the protein visits the open state naturally, we can capture it using Molecular Dynamics (MD) simulations. ### 2.2 Scales of Flexibility Protein motions occur across vast temporal and spatial scales. Static docking fails catastrophically with Loop and Domain motions, which ensemble docking is designed to capture. - **Side-Chain Rotation:** Rotation of amino acid side chains (picoseconds). - **Loop Displacement:** Movement of flexible surface loops (nanoseconds). - **Domain Motion:** Hinge bending between protein domains (microseconds). ![Fig.3| Scales of protein flexibility](images/2a2d62f0e8c5ea25cca5243f78748452d243e87a-1000x478.png) ### 2.3 The "Unhappy Valley" of Scoring Scoring failures often peak when the RMSD (Root Mean Square Deviation) between the docked pose and the native structure is between 1.5 and 2.0 Å. This "unhappy valley" represents poses that are geometrically close to the correct answer but are penalized by the scoring function due to minor clashes that a flexible receptor would easily accommodate. By ignoring these "wiggles," static docking ignores the entropic component of binding. ## 3. The Methodology of Ensemble Docking Ensemble docking is a practical compromise between the extreme cost of fully flexible docking and the inaccuracy of rigid docking. Instead of making the protein flexible *during* the docking, we make the protein flexible *before* the docking by generating a discrete set of representative conformations. ### Stage 1: Ensemble Generation via Molecular Dynamics While NMR ensembles or multiple crystal structures can be used, Molecular Dynamics (MD) simulations are the gold standard for exploring the conformational space. MD simulations generate a "trajectory" a high frame rate movie of the protein wiggling in explicit solvent. This captures the crucial role of solvent in stabilizing transient conformations. For rare events, such as the opening of cryptic pockets, techniques like **Accelerated MD (aMD)** or **Metadynamics** are used to force the protein to explore high-energy states. ### Stage 2: Clustering and Representative Selection A raw trajectory contains too much redundancy; docking to 6,000 structures is computationally prohibitive. **Clustering Algorithms** reduce the ensemble to a manageable number of representative structures (typically 3 to 20). - **RMSD Clustering:** Groups structures based on backbone deviation. - **K-Medoids:** Selects the most "central" actual snapshot from the trajectory, avoiding artificial average structures. In a Lysozyme case study, a 100 ns simulation yielded 15 clusters, but the top 4 clusters accounted for 90% of the population, allowing researchers to focus their docking efforts efficiently. ### Stage 3: Cross-Docking and Ranking The ligand is docked into each representative structure independently. The final ranking often uses the "Best Score Strategy," taking the single best score across all conformers. This mimics Conformational Selection: the ligand "finds" the best fitting shape. ![Fig.4| various stages in ensemble docking](images/68b568f7b69dad4c97bc21a9d68d67d151557287-1000x546.png) ## 4. Molecular Dynamics: The Engine of Flexibility The validity of ensemble docking rests on the quality of the MD simulation. If the simulation does not sample the relevant "open" state, the docking will fail. ### 4.1 MD Post-Processing: "Dynamic Docking" Validation MD is not just for *pre-docking* generation; it is also used *post-docking* to refine poses. This "dynamic docking" validation involves running a short simulation (e.g., 5-100 ns) on the static docking pose. - **The Logic:** If a ligand is a true binder, it should be stable. If it is a decoy, it will drift away or exhibit high RMSD fluctuations. - **Results:** This method has been shown to improve ROC AUC (Receiver Operating Characteristic) scores by over 22% compared to docking scores alone, effectively filtering out unstable false positives. ## 5. Case Study I: Lysozyme and Flavokawain B To illustrate the power of ensemble docking, we examine the interaction between **Lysozyme (LYZ)** and the ligand **Flavokawain B (FB)**. Static docking of FB to the crystal structure of lysozyme yields mediocre binding energies. Researchers performed a 100 ns MD simulation of Lysozyme in water, clustered the trajectory, and docked FB to the top 4 representative structures. **Results:** - Cluster 1 (Dominant state): -28.41 kJ/mol. - **Cluster 2 (14% of population): -29.37 kJ/mol (Best Binder).** The ensemble approach identified **Cluster 2, **a conformation representing only a minority of the simulation time as the optimal binding state. This conformation allowed for specific interactions with residues **Glu-35, Trp-108, and Arg-114** that were not accessible in the dominant crystal like conformation. ## 6. Case Study II: Hunting Ghosts: Cryptic Pockets in IL-2 The most dramatic application of ensemble docking is the discovery of **Cryptic Binding Sites **pockets that do not exist in the ligand-free crystal structure but open transiently due to protein dynamics. ### 6.1 The Interleukin-2 (IL-2) Problem IL-2 is a critical cytokine with a "flat" interface considered undruggable. Researchers at D. E. Shaw Research used unbiased, long-timescale MD simulations to study the binding of inhibitor **SP4206**. ### 6.2 The Solution The simulations revealed that the "flat" surface of IL-2 breathes. A groove opens transiently even without the ligand. When SP4206 is present, its hydrophobic **dichlorobenzene group** slides into this transient groove, preventing it from closing. The protein then clamps down, locking the ligand in place. A static docking attempt on the apo IL-2 structure would have failed 100% of the time because the pocket simply *did not exist* in the input coordinates. This demonstrates that for cryptic pockets, ensemble docking is the only viable approach. ![Fig.5| Cryptic Pockets in IL-2](images/ae851d7a1836b176551f3752bea617ba47cfb819-1000x486.png) ## 7. Case Study III: Disorder and Mutual Induced Fit (Bcl-xL) **Bcl-xL**, a cancer drug target, binds to the **Bim** protein, which contains an **Intrinsically Disordered Region (IDR)**. Docking a flexible snake (Bim) into a flexible pocket (Bcl-xL) is the ultimate nightmare for rigid algorithms. Simulations using **McMD (Multicanonical MD)** revealed a mechanism of "Mutual Induced Fit". As they approach, *both* molecules change shape: Bcl-xL opens a cryptic pocket, and Bim folds into an α-helix *upon* binding. Ensemble docking using snapshots from the McMD trajectory successfully identified the intermediate "open" states, explaining why inhibitors like ABT-737 bind with high affinity by mimicking the hydrophobic residues of the Bim helix. ![Fig.6| Understanding mutual flexibility is key. Inhibitors can be designed to mimic the transient, folded states of disordered proteins](images/a5be991f6d504fa188c2d0907936779b3e36cd9b-1000x546.png) ## 8.The LiteFold Solution: A Unified Infrastructure for Ensemble Docking The transition from static to ensemble docking has historically been hindered by high technical barriers. Traditional workflows are often fragmented, requiring researchers to juggle disparate command-line tools for simulation, clustering, and docking, all while managing heavy computational resources. **LiteFold** addresses these challenges by providing a unified, cloud-native infrastructure specifically designed to streamline the complexities of dynamic drug discovery. ### 8.1 Removing Infrastructure Friction Conducting ensemble docking typically requires access to High-Performance Computing (HPC) clusters to run computationally intensive MD simulations. LiteFold abstracts this complexity entirely. By offering a browser-based "Physics Engine Infrastructure," it allows researchers to launch simulations ranging from nanoseconds to microseconds without configuring complex environments or managing GPUs. This democratization of compute power ensures that the "wiggling" of proteins is accessible to all researchers, not just those with dedicated supercomputing access. ### 8.2 Seamless Integration of Dynamics and Docking A critical bottleneck in ensemble docking is the handover between the MD simulation and the docking protocol. Traditionally, this involves manual extraction of frames, file conversion, and scripting to batch-process docking jobs. LiteFold integrates these steps into a cohesive pipeline. - **Automated Ensemble Generation:** Users can run MD simulations via the **Dynamo** module directly in the browser. The platform enables the inspection of trajectories and the selection of representative conformations without the need to download gigabytes of trajectory files. - **Unified Workflow:** Once representative structures (snapshots) are identified, they can be immediately utilized in the docking workflow. This integration reduces the "feedback latency" between observing a protein's motion and testing its druggability. ### 8.3 Advanced Analysis Capabilities LiteFold moves beyond simple pose generation by incorporating analysis tools that are essential for evaluating ensemble results. The platform’s infrastructure supports the handling of large-scale experimental datasets and annotations, ensuring that the ensembles generated are biologically relevant. By centralizing the storage and analysis of simulations, LiteFold allows for a more rigorous assessment of ligand stability and binding thermodynamics, directly addressing the "Unhappy Valley" of scoring where static methods fail. ## 9. Bridging the Gap: From Static AI Predictions to Dynamic Ensembles While AI-driven structure prediction has revolutionized structural biology, it often yields static snapshots that bias towards the most stable state, frequently missing the high-energy "open" conformations required for ligand binding. LiteFold serves as the bridge between these static AI predictions and the dynamic reality required for successful drug design. ### 9.1 Breathing Life into Static Models AI models excels at predicting the ground-state structure of proteins from sequence. However, these models often fail to capture cryptic pockets or transient conformational changes. LiteFold leverages these static predictions as a starting point, using its integrated physics engine to "breath life" into the structures. By running MD simulations on AI-predicted models, LiteFold generates a conformational ensemble that explores metastable states, such as the DFG-out conformation in kinases, which are often missed by static prediction alone. ### 9.2 Interactive Feedback Loops One of the most powerful features of LiteFold is the reduction of feedback loops in the design process. In a traditional workflow, modifying a ligand to fit a new protein conformation might take days of set-up and calculation. LiteFold’s **DeNovo** and interactive design modules allow researchers to edit molecules and immediately observe changes in predicted binding metrics against the generated ensemble. - **Real-Time Optimization:** Researchers can use fragment growing or the molecule editor to refine compounds within the context of the dynamic pocket. - **Dynamic Validation:** As new molecules are designed, their binding affinity is recomputed against the ensemble, providing instant insight into how chemical modifications influence binding to flexible targets. ### 9.3 The Convergence of AI and Physics LiteFold represents a convergence of generative AI and physics-based simulation. It uses neural networks to predict initial structures and pockets, and then applies rigorous physics (MD) to validate and explore the conformational landscape. This hybrid approach ensures that drug discovery is not limited by the static nature of initial predictions but is enhanced by the rigorous thermodynamic sampling provided by the platform's infrastructure. ## 10. Conclusion: The Democratization of Flexibility The "Lock and Key" model, while a foundational concept in biochemistry, has historically constrained drug discovery by promoting a static view of molecular interactions. We now understand that proteins are dynamic entities that dance, shift, and reshape themselves in response to their environment and binding partners. Rigid docking is akin to examining a single frame of a complex film; it captures a moment but misses the narrative. Ensemble docking provides the necessary context, capturing the protein in its various states of motion. However, the complexity of generating and managing these ensembles has often restricted this powerful technique to computational specialists with access to massive infrastructure. **LiteFold** fundamentally changes this landscape. By creating an integrated, cloud-native workspace that seamlessly combines Molecular Dynamics with docking and design, LiteFold democratizes access to protein flexibility. It removes the barriers of hardware configuration and fragmented software, allowing researchers to focus purely on the science. Whether identifying cryptic pockets in "undruggable" targets or refining lead compounds against a shifting active site, LiteFold provides the unified infrastructure necessary to navigate the chaotic, wiggly reality of the molecular world. The future of drug discovery is dynamic, and with platforms like LiteFold, that future is now accessible. ## References Amaro, R. E., Baudry, J., Chodera, J., Demir, Ö., McCammon, J. A., Miao, Y., & Smith, J. C. (2018). Ensemble Docking in Drug Discovery. *Biophysical Journal*, *114*(10), 2271–2278. https://doi.org/10.1016/j.bpj.2018.02.038 cnapolitan2. (2021, September 27). *Docking and scoring - Schrödinger*. Schrödinger. https://www.schrodinger.com/life-science/learn/white-papers/docking-and-scoring/ Damilola. (2023, February 16). *The Docking Method Showdown: Rigid Receptor Docking vs. Induced Fit Docking vs. QPLD*. Medium. https://medium.com/@dbodun56/the-docking-method-showdown-rigid-receptor-docking-vs-induced-fit-docking-vs-qpld-8251d8927c7f Gathiaka, S., Liu, S., Chiu, M., Yang, H., Stuckey, J. A., Kang, Y. N., Delproposto, J., Kubish, G., Dunbar, J. B., Carlson, H. A., Burley, S. K., Walters, W. P., Amaro, R. E., Feher, V. A., & Gilson, M. K. (2016). D3R grand challenge 2015: Evaluation of protein–ligand pose and affinity predictions. *Journal of Computer-Aided Molecular Design*, *30*(9), 651–668. https://doi.org/10.1007/s10822-016-9946-8 Lexa, K. W., & Carlson, H. A. (2012). Protein flexibility in docking and surface mapping. *Quarterly Reviews of Biophysics*, *45*(3), 301–343. https://doi.org/10.1017/s0033583512000066 Ricci-Lopez, J., Aguila, S. A., Gilson, M. K., & Brizuela, C. A. (2021). Improving Structure-Based Virtual Screening with Ensemble Docking and Machine Learning. *Journal of Chemical Information and Modeling*, *61*(11), 5362–5376. https://doi.org/10.1021/acs.jcim.1c00511 Tripathi, A., & Bankaitis, V. A. (2018). Molecular Docking: From Lock and Key to Combination Lock. *Journal of Molecular Medicine and Clinical Applications*, *2*(1). https://doi.org/10.16966/2575-0305.106 --- # The Generative Geometric Turn in AI Drug Discovery URL: https://www.lite.bio/blogs/the-stochastic-architect-diffusion-models-in-structure-based-drug-design-and-the-boltzgen-paradigm Author: Aditi Sinha Date: 2025-12-28 The pharmaceutical industry stands at a critical juncture, often described through the lens of "Eroom's Law" the observation that drug discovery is becoming slower and exponentially more expensive over time, despite aggregate improvements in technology. The process of identifying a therapeutic candidate, optimizing its properties, and guiding it through clinical trials is historically a venture of immense attrition, costing upwards of $2.6 billion per successful launch. Central to this inefficiency is the "lock and key" problem: finding a small molecule (ligand) that binds with high affinity and specificity to a biological target (protein), typically by fitting into a complex, three-dimensional binding pocket. For decades, this challenge was addressed through High-Throughput Screening (HTS) of physical libraries or, more recently, virtual screening using classical docking algorithms like AutoDock Vina or Glide. While these methods have yielded successes, they are fundamentally limited by their reliance on static representations, rugged energy landscapes, and the sheer vastness of chemical space estimated at over 1060 pharmacologically active compounds. A profound shift is now underway, driven by the integration of geometric deep learning and generative artificial intelligence. This "Geometric Turn" acknowledges that molecules are not merely strings of text (SMILES) or 2D graphs, but dynamic physical objects embedded in continuous Euclidean space. The most potent engine of this revolution is the **Diffusion Model**. Originally popularized in computer vision for generating photorealistic images from noise, diffusion models have been mathematically adapted to the non-Euclidean manifolds of molecular geometry. By learning to reverse the thermodynamic process of entropy turning structure into noise and back again these models are enabling the *de novo* generation of binders that are physically plausible, synthetically accessible, and tailored to novel targets that have historically been deemed "undruggable". This Blog provides an exhaustive technical analysis of the application of diffusion models to Small-Molecule Generation and Binding Pose Prediction. We will dissect the theoretical underpinnings of SE(3)-equivariant networks, analyze the landscape of current state-of-the-art architectures (from DiffDock to TargetDiff), and finally, explore the next frontier in unified biomolecular design: the **BoltzGen**. As part of the LiteFold ecosystem, BoltzGen represents a paradigm shift from isolated task specific models to a unified all atom generative framework capable of simultaneous folding and binder design, democratizing access to the kind of computational infrastructure that was once the exclusive preserve of big pharma. ![Fig.1| From Static representations to dynamic 3D diffusion models for De novo Drug design, The goal is to accelerate discovery of novel theurapeutics for undruggable targets using Generative AI in 3D space.](images/b85bec6e7e212f3edd0ba8adcde4dff56aa9c00d-1000x415.png) ### Theoretical Foundations: Thermodynamics and Geometric Deep Learning To understand the efficacy of diffusion models in chemistry, one must first grasp their statistical mechanical roots. Unlike Generative Adversarial Networks (GANs), which rely on unstable adversarial training, or Variational Autoencoders (VAEs), which frequently suffer from posterior collapse, diffusion models are grounded in a stable, iterative denoising process inspired by non-equilibrium thermodynamics. The framework operates through two dual processes: **1. The Forward Process (Diffusion):** This is a fixed Markovian chain that progressively destroys structure. Given a data distribution X0 (e.g., a valid crystal structure of a ligand), Gaussian noise is added over discrete time steps t=0…..T. As t to T, the molecular geometry dissolves into an isotropic Gaussian distribution N(0,1). This process is chemically equivalent to raising the temperature of a system until it becomes a high-entropy gas. **2. The Reverse Process (Generative Denoising):** The generative model learns to reverse this entropy. Starting from pure noise sampled from a standard Gaussian, a neural network approximates the reverse transition Pθ(xt-1|xt), effectively "cooling" the system to recover a low-energy, chemically valid state. In the context of Structure-Based Drug Design (SBDD), this is often formulated as **Score-Based Generative Modeling (SGM)**. Instead of trying to learn the intractable probability density function of all possible molecules, the network learns the "score function"—the gradient of the log-density of the data (∇xlog pt(x)). Intuitively, this vector field points the "atoms" from their noisy positions toward high-density regions of chemical validity. ![Fig.2| Thermodynamic and Geometric deep learning in Diffusion models](images/4c977b60dff54c540888d511dbf121eff5b3292c-1000x468.png) ### The Non-Negotiable Requirement: SE(3) Equivariance Applying neural networks to 3D chemistry introduces a constraint that does not exist in text or image generation: the laws of physics must be invariant to the observer's frame of reference. A molecule's internal energy, bond lengths, and binding affinity do not change if the molecule is rotated 90 degrees or translated 5 angstroms to the left. Standard neural architectures (CNNs, RNNs) are not naturally invariant to 3D rotations. If a standard network is trained on a molecule, and that molecule is rotated, the network might perceive it as a completely different object. This necessitates **SE(3) Equivariance** - **SE(3) Group:** The Special Euclidean group in 3 dimensions, comprising all continuous rigid body translations and rotations. - **Equivariance vs. Invariance:** - *Invariant:* *f(R.x) = f(x) *Global properties like "predicted toxicity" or "solubility" should be invariant; rotating the molecule shouldn't change the prediction. - *Equivariant:* *f(R.x) = R. f(x)*. Generative tasks that predict atomic coordinates or forces must be equivariant. If the input molecule rotates, the predicted update vectors must rotate in exact unison. The first deployment of diffusion models for 3D molecular generation, the **E(3) Equivariant Diffusion Model (EDM)**, utilized E(n)-equivariant Graph Neural Networks (EGNNs). These networks operate on relative distances rather than absolute coordinates, ensuring that the generative process respects the symmetry of physical space. Without this inductive bias, models require massive data augmentation (showing the network every possible rotation of a molecule) and still fail to generalize to novel orientations. ![Fig.3| The non- negotiable requirement: SE(3) Equivariance in molecular modeling](images/f08573e0ca524d4df4f2754f6151ffee42e2ae64-1000x501.png) ### Diffusion on Non-Euclidean Manifolds Molecules are constrained systems. Atoms are not free-floating points; they are tethered by covalent bonds with specific lengths and rigid angles. Modeling a molecule purely in Euclidean space (R3) often leads to "broken" geometry carbon rings that are not flat, or bonds that stretch physically impossible distances. Advanced diffusion approaches, particularly in molecular docking (e.g., DiffDock), operate on the **product manifold** of the ligand's degrees of freedom: M= R3 x SO(3) x Tm Where: - R3 represents global translation. - SO(3) represents global rotation (Special Orthogonal group). - Tm represents the $m$ rotatable torsion angles (modeled as a hypertorus). By defining the diffusion process directly on this manifold, the model ensures that the local geometry (bond lengths/angles) remains rigid and chemically valid, while the system explores the conformational space (folding/twisting) and the pose space (position/orientation) stochastically. ## The Landscape of Small-Molecule Generation The application of diffusion models to small molecules is broadly divided into two domains: **Unconditional Generation** (exploring chemical space for novel scaffolds) and **Conditional Generation** (designing ligands to fit a specific protein pocket SBDD). ### Unconditional Generation: From Clouds to Graphs #### a. EDM: The E(3) Equivariant Diffusion Model The **E(3) Equivariant Diffusion Model (EDM)** was a foundational architecture that treated molecules as point clouds of atoms in continuous 3D space. Using an EGNN to estimate noise, EDM demonstrated that diffusion models could generate thermodynamically stable conformations better than previous VAE or Flow based methods. However, because it treated atoms essentially as a "gas" that condenses into a molecule, it sometimes struggled with precise integer bond orders and ring planarity without post-hoc refinement. #### b. MolDiff: Enforcing Chemical Validity To address the limitations of pure point-cloud generation, **MolDiff** introduced a bond-aware diffusion framework. MolDiff explicitly models the interrelationships between atoms and bonds during the generative process. By incorporating bond attributes (single, double, aromatic) into the graph diffusion, it constrains the generation to chemically valid graphs. This approach significantly outperforms EDM in terms of "validity" (percentage of generated structures that respect valency rules) and "uniqueness," ensuring the model isn't just memorizing the training set. #### c. 3D-EDiffMG: Scaffold-Constrained Optimization In practical drug discovery, chemists rarely start from scratch. They often engage in **Lead Optimization**, where a core scaffold (e.g., a benzodiazepine ring) is known to bind, and the goal is to modify side chains to improve solubility or reduce toxicity. **3D-EDiffMG** utilizes a dual-encoder architecture (Dual-SWLEE) that can encode strong interactions (covalent bonds) and weak interactions (non-bonded forces) separately. This allows for masked diffusion or "inpainting," where the core scaffold is fixed, and the diffusion model generates optimal R-groups in the context of the scaffold's 3D geometry. ![Fig.4| An overview of evolving diffusion models for 3D molecular generation, progressing from basic point clouds to chemically valid graphs and targeted scaffold optimization.](images/719a123c2c902e85596dc79990b3210fc192268b-1000x503.png) ### Structure-Based Drug Design (SBDD) The "Holy Grail" of generative chemistry is SBDD: given a protein pocket *P*, generate a ligand *L* that maximizes the probability *p(L|P)*. #### a. TargetDiff: The Baseline for SBDD **TargetDiff** represents the first successful application of equivariant diffusion to non-autoregressive SBDD. It defines a joint diffusion process over atom types and coordinates, conditioned on the protein pocket's residues. By modeling the interaction as a distribution, TargetDiff captures the multi-modal nature of binding—there isn't just *one* molecule that fits a pocket, but a diverse family of potential binders. However, early evaluations showed that while TargetDiff generated high-affinity ligands (high Vina scores), it often produced molecules with **atomic collisions** (clashes) with the protein, as the vanilla diffusion loss didn't penalize van der Waals violations heavily enough. #### b. NucleusDiff: Solving the Collision Problem Atomic collisions render a docked pose physically impossible (infinite energy). **NucleusDiff** addresses this by introducing explicit geometric constraints into the diffusion process. It places auxiliary "mesh points" on the spherical van der Waals surface of each atom. The loss function includes a regularization term that enforces minimum distances between these mesh points and the protein atoms. - **Result:** NucleusDiff reduces the atomic collision rate by up to **100%** compared to TargetDiff. - **Affinity:** It simultaneously improves the average Vina score by **22.16%** on the CrossDocked2020 benchmark, proving that physical plausibility correlates with predicted affinity. #### BINDDM: Hierarchical Subcomplex Attention Proteins are large macromolecules; generating a small ligand while attending to thousands of protein atoms creates signal-to-noise issues. **BINDDM** (Binding-Adaptive Diffusion Model) employs a hierarchical approach. It dynamically extracts a "subcomplex" the subset of residues most critical for binding and focuses the SE(3)-equivariant network's attention there. This "zoom-in" mechanism allows BINDDM to generate ligands with more realistic 3D structures and higher binding affinities (up to -5.92 kcal/mol improvement in Vina score) by effectively ignoring the irrelevant parts of the protein surface. ![Fig.5| Graphical Representation for SBDD](images/d50847aaf743a0bfa8e7f5f40507bdf835e184cc-1000x503.png) ## The Revolution in Molecular Docking: From Search to Generation While SBDD focuses on generating *new* molecules, **Molecular Docking** focuses on predicting the binding pose of *existing* molecules. This is the workhorse of virtual screening. ### a. The Limitations of Traditional Algorithms Traditional docking tools like **AutoDock Vina**, **Glide**, and **Surflex-Dock** treat docking as an optimization problem on a rugged energy landscape. - **Mechanism:** They use a scoring function (physics-based or empirical) and a search algorithm (Genetic Algorithm, Monte Carlo, or Simulated Annealing) to find the global minimum. - **Failures:** These search algorithms frequently get trapped in local minima. They are computationally expensive (minutes per ligand), scaling poorly to billion-compound libraries. Furthermore, they perform poorly on "blind docking" (where the pocket location is unknown) because the search space is too vast. ### b. DiffDock: A Generative Paradigm Shift **DiffDock** (Diffusion Docking) reframes the problem entirely. It does not "search" for a minimum energy; it "generates" the binding pose distribution. - **Architecture:** DiffDock operates on the product manifold of translations, rotations, and torsions. - **Process:** It samples random poses (noise) and iteratively refines them using a learned score model that "pushes" the ligand toward the pocket. - **Confidence Model:** Crucially, DiffDock includes a confidence model that ranks the generated poses. Since diffusion is stochastic, it can generate multiple hypotheses; the confidence model predicts which one is correct. **Benchmarking Dominance:** On the PDBBind benchmark, DiffDock achieves a **38% Top-1 success rate** (RMSD < 2Å), significantly outperforming traditional docking (23%) and previous deep learning regression models like EquiBind (20%). Even more impressively, its inference time is orders of magnitude faster than exhaustive search methods, making it feasible for library-scale screening. ### c. DiffDock-PP and Protein-Protein Interactions The principles of DiffDock have been extended to **DiffDock-PP** for rigid protein-protein docking. Protein-protein interactions (PPIs) are notoriously difficult due to the large surface areas and subtle energetics. DiffDock-PP frames PPI docking as learning a probability distribution over the SE(3) transformation that aligns the ligand protein to the receptor. It runs **5 to 60 times faster** than traditional methods on GPU hardware, opening the door for interactome-scale mapping. ## Challenges and Frontiers ### Synthetic Accessibility (SA) A recurring criticism of generative AI in chemistry is the "hallucination" of unsynthesizable molecules. A model might propose a structure that binds perfectly *in silico *maximizing Vina scores but contains highly strained rings, impossible bond angles, or functional groups that would be chemically unstable or impossible to synthesize in a lab. - **The Metric:** Synthetic Accessibility (SA) score is a heuristic ranging from 1 (easy) to 10 (hard), penalizing complexity, chiral centers, and fused rings. - **Integration:** Modern pipelines like LiteFold integrate SA scoring directly into the generation loop. By using classifier guidance or filtering, the diffusion process can be steered away from high-SA regions of chemical space, ensuring that the "fail fast" philosophy applies to synthesis planning as well as binding. ### The "Static Receptor" Fallacy vs. Co-Folding Most current docking and SBDD models assume a "rigid receptor" the protein is treated as a static statue. In biological reality, proteins are breathing, dynamic entities. Binding often involves **induced fit** (the pocket reshapes to accommodate the ligand) or **conformational selection**. - **The Problem:** Rigid docking fails when the holo (bound) structure of the protein differs significantly from the apo (unbound) structure. - **The Solution:** The field is moving toward **Co-Folding** models (like AlphaFold Multimer and Boltz-1), which predict the structure of the protein and ligand *simultaneously* from sequence. This allows the model to predict the induced fit changes, offering a higher fidelity representation of the interaction. ## The LiteFold Ecosystem and the BoltzGen Pipeline The fragmentation of the current AI drug discovery landscape where a researcher needs one tool for folding, another for docking, and a third for design creates immense friction. **LiteFold** aims to unify this stack, positioning itself as the "Infrastructure for Drug Discovery." By providing high-performance compute, secure data handling, and browser-based interfaces for complex MD and docking simulations, LiteFold democratizes access to these advanced capabilities. ![Fig.6| The boltzgen Pipeline in Litefold Ecosystem](images/696d8a02a62e7f3d0c161b883bc387516dc74e20-1000x457.png) The crown jewel of this ecosystem is the upcoming **BoltzGen pipeline**, a unified all-atom generative framework that integrates the state of the art **Boltz-1** biomolecular interaction model. ### BoltzGen: Unifying Design and Structure Prediction **BoltzGen** is not merely a docking tool; it is a generative engine that unifies **design** and **structure prediction**. Developed in collaboration with MIT’s Jameel Clinic and validated by major biotech partners, it represents a departure from the "pipeline of distinct models" approach. #### Architectural Innovation: Geometry-Based Residues Unlike language models that treat amino acids as text tokens, BoltzGen employs a **purely geometry-based representation** of designed residue types. - The model predicts the geometric cluster of atoms that constitutes a residue (e.g., the specific cloud of atoms that makes an Alanine). - This allows the model to train on structure prediction (folding) and design (generation) simultaneously. The gradient flow is continuous, enabling the model to "reason" about steric clashes and packing density in real time during generation. - This unification allows BoltzGen to match the folding performance of state of the art predictive models (like Boltz-1) while retaining the creativity of a generative model. ### Unprecedented Performance on Novel Targets The true test of any AI model in biology is **generalization**. Models often "memorize" the PDB, performing well on familiar targets but failing on novel ones. - **The Benchmark:** BoltzGen was tested on **9 novel targets** that have **<30% sequence identity** to any complex in the PDB (effectively "aliens" to the model). - **The Result:** BoltzGen successfully designed nanomolar (nM) binders for **66%** of these targets (6 out of 9). - **Efficiency:** This was achieved by testing fewer than **15 designs per target** in the wet lab. This efficiency, a hit rate of >5% on the very first batch of <15 designs is orders of magnitude better than traditional HTS (hit rates often <0.1% of thousands of compounds) and earlier AI generation methods. It suggests that BoltzGen has learned a generalized "physics of binding" rather than just memorizing specific interaction motifs. ### Multi-Modal Capabilities BoltzGen is modality-agnostic. Through its Design Specification Language, users on LiteFold can define constraints for: - **Proteins & Nanobodies:** Designing large biologic binders for immunotherapy. - **Peptides:** Designing cyclic peptides or macrocycles for "undruggable" flat interfaces (e.g., PPIs). - **Small Molecules:** While primarily a protein design model, the pipeline supports small molecule contexts, allowing for the design of protein binders *to* small molecules (e.g., biosensors) or vice-versa. ### Operational Excellence: LMI4Boltz and Hardware Optimization Deploying these massive all-atom models requires significant computational resources. The standard Boltz-2 model is VRAM-hungry, limiting inference on consumer hardware. The LiteFold implementation leverages **LMI4Boltz** (Low Memory Inference for Boltz), a set of optimizations including: - **Tensor Offloading:** Moving large, infrequently used tensors (like MSA representations) to host memory. - **In-Place Operations:** Using FlashAttention and in-place updates to avoid memory spikes. - **Result:** These optimizations increase the token limit by **66.7%**, allowing for the modeling of large complexes (>2,660 tokens) on standard 24GB VRAM GPUs (like the A10G or 3090/4090), making the pipeline accessible to a broader range of researchers via the LiteFold cloud. ### Comparative Benchmarking: Boltz-1 vs. AlphaFold3 As the foundation of BoltzGen, the **Boltz-1** model's accuracy is paramount. In comparative benchmarks using **ABCFold** standards: - **Accuracy:** Boltz-1 achieves **AlphaFold3-level accuracy** in predicting 3D structures of biomolecular complexes. - **CASP15 Performance:** On the critical LDDT-PLI metric (protein-ligand interactions), Boltz-1 achieves a score of **65%**, matching AlphaFold3 and significantly outperforming Chai-1 (40%). - **Open Source:** Unlike AlphaFold3, which is restricted, Boltz-1 is fully open-source (MIT License), allowing LiteFold to integrate it deeply into the generation pipeline without API restrictions or data privacy concerns. | Model | Core Mechanism | Input Data | Output | Key Innovation | | --- | --- | --- | --- | --- | | EDM | E(3)-Equivariant Diffusion | Point Cloud | Atom Coordinates | First stable 3D molecular diffusion | | MolDiff | Bond-Aware Diffusion | Graph + 3D | Graph + Coordinates | Enforces chemical validity & bond orders | | TargetDiff | Conditional SE(3) Diffusion | Protein Pocket + Ligand | Ligand Structure | Jointly models atom types & coords in pocket | | NucleusDiff | Constrained Diffusion | Protein Pocket + Ligand | Ligand Structure | Mesh-based collision constraints (<1% clashes) | | DiffDock | Manifold Diffusion | Protein (Blind) + Ligand | Binding Pose | Diffusion on R3 x SO(3) x Tm | | BoltzGen | Unified All-Atom Diffusion | Target + Constraints | Binder (Protein/Mol) | Simultaneous Design & Folding | **** ## Conclusion: The Fail-Fast Future The integration of **BoltzGen** into the **LiteFold** platform signals the maturation of generative biology. We are moving away from the era of "computer-aided drug design" (where computers acted as assistants to human intuition) to "computer-driven drug design" (where algorithms act as the primary architects). By combining the stochastic exploration of **Diffusion Models**, the rigorous physical constraints of **SE(3) Equivariance**, and the validation power of **Co-Folding** within a unified pipeline, we can now "fail fast and fail cheap" *in silico*. Instead of synthesizing 10,000 compounds to find one hit, researchers can generate 100 high-confidence designs, validate them computationally with Boltz-2, and proceed to the wet lab with a focused set of 15 candidates that have a >60% probability of success. This is not just an incremental improvement in efficiency; it is a fundamental restructuring of the economics of drug discovery. For the LiteFold community, the release of the BoltzGen pipeline is the key to unlocking this potential, transforming the browser into a portal for de novo biological engineering. --- # Molecular Docking vs. QSAR: How Smart Computing Shapes ADMET Decisions URL: https://www.lite.bio/blogs/molecular-docking-vs-qsar Author: Aditi Sinha Date: 2025-12-18 ### Why ADMET Needs a Better Plan Most drug projects begin with bright hopes and a pile of molecules that look good on paper. Then reality walks in. A compound that seemed like a star in early screens may vanish in the gut, stick to plasma proteins, clog a liver enzyme, or cause heart issues no team wants to explain in a meeting. Anyone who has worked in discovery knows this moment. The drug was fine until ADMET said no. ADMET profiling has moved from a late filter to a front-line requirement in modern drug design. Poor PK and unexpected toxicity were once responsible for nearly half of all failures in the clinic. This pushed research teams toward early in silico testing, allowing them to flag the weak links before money and months are spent on long wet-lab cycles. Yet even with this shift, experimental ADMET work remains slow and costly when dealing with wide chemical space. This is where strong computational screening becomes essential. The focus of this work is to inspect two major approaches used in current pipelines: QSAR, which studies patterns in chemical structure, and molecular docking, which studies how a molecule might fit in a protein pocket. These tools are often spoken of as rivals, but in practice they answer different questions and fail in different ways. Used with care, they fill the gaps in each other’s blind spots. The goal here is to break down these methods, highlight their uses across common ADMET endpoints, and prepare the ground for a full benchmarking study. This includes decisions on datasets, proper validation, and the design of mixed models that combine QSAR features with docking outputs. The final aim is to give a clear path for building a reliable ADMET module that can support real project needs. ![Fig.1 From QSAR hits to ADMET reality checks: why early predictions must face real drug behavior.](images/0b9d493db4a64e28f7bd814f0c9214975f5977ca-1000x546.png) ## **2. Theoretical Paradigms in Computational Modeling** To design a valid set of experiments for an ADMET module, one must first deconstruct the theoretical axioms that govern QSAR and molecular docking. The fundamental distinction lies in their approach to biological reality: QSAR is an inductive process relying on statistical inference from known data, while molecular docking is a deductive process simulating physical interactions based on first principles and empirical scoring. ### **2.1 Quantitative Structure-Activity Relationships (QSAR)** QSAR is predicated on the central axiom of chemoinformatics: *structural similarity implies functional similarity*. This principle asserts that the biological activity or physicochemical property of a molecule is a deterministic mathematical function of its molecular structure. The objective of QSAR modeling is to approximate this function (***f***) by mapping a vector of calculated molecular descriptors (***D***) to a biological endpoint ***Pi= f(Di1, Di2,...,Din) + ϵ*** Where ***Pi*** is the property of molecule ***i***, ***D*** represents the vector of descriptors, and ***ϵ*** represents the error term. **The Descriptor Landscape** The fidelity of a QSAR model is intrinsically limited by the information content of its descriptors. In the context of ADMET prediction, descriptors are generally categorized into hierarchies of increasing complexity: • **1D Descriptors:** Scalar values representing bulk properties, such as Molecular Weight (MW), atom counts, and total charge. These are computationally trivial to generate but lack topological context. • **2D Descriptors:** Topological indices that encode the connectivity of the molecule. This includes graph invariants (e.g., Wiener index, Balaban index) and molecular fingerprints (e.g., MACCS keys, ECFP4). Extended-Connectivity Fingerprints (ECFPs) are particularly dominant in ADMET modeling because they capture circular substructural environments, which often correspond to "toxicophores" or metabolic liabilities. • **3D Descriptors:** These encode the spatial arrangement of atoms, capturing stereochemistry, molecular fields, and potential pharmacophoric points (e.g., CoMFA, CoMSIA fields). While theoretically richer, 3D descriptors introduce dependency on conformation generation, significantly increasing computational cost and noise if the bioactive conformation is unknown. ![Fig.2 Basic molecular numbers to full 3D features, better descriptors raise QSAR accuracy but also cost more compute](images/494a356250fc38f75ddca990d81163b682edec6e-1000x546.png) **Statistical and Machine Learning Engines** The mathematical "engine" driving QSAR has evolved substantially. Early approaches relied on Multiple Linear Regression (MLR) and Partial Least Squares (PLS), which assume linearity between descriptors and activity. While interpretable, these methods often fail to capture the complex, non-linear landscapes of toxicity and metabolism. Modern ADMET modules predominantly utilize non-linear Machine Learning (ML) algorithms: • **Random Forest (RF) and Gradient Boosting (XGBoost):** These ensemble methods are currently the "gold standard" baselines for ADMET benchmarking. They are robust to noise, handle high-dimensional feature spaces (like fingerprints) effectively, and provide measures of feature importance. • **Support Vector Machines (SVM):** Effective for defining decision boundaries in high-dimensional spaces, particularly for binary classification tasks like "Toxic/Non-Toxic". • **Deep Learning (GNNs):** Graph Neural Networks (GNNs) and Message Passing Neural Networks (MPNNs), such as ChemProp, represent the frontier. Instead of using pre-calculated descriptors, these models learn optimal representations directly from the molecular graph during training. ### **2.2 Molecular Docking** Molecular docking is a structure-based simulation technique that predicts the preferred orientation (pose) and binding affinity of a ligand to a macromolecular target. In the context of ADMET, docking is restricted to endpoints mediated by specific proteins, such as metabolic enzymes (e.g., CYP3A4, CYP2D6), transporters (e.g., P-gp/MDR1), and specific toxicity targets (e.g., hERG, Androgen Receptor). **The Mechanics** The docking process involves two coupled components: a. **Search Algorithm:** This component explores the conformational space of the ligand (and occasionally the receptor). It navigates degrees of freedom including translation, orientation, and torsion angles of rotatable bonds. Common algorithms include: ◦ *Genetic Algorithms (Lamarckian GA):* Used by AutoDock, this method evolves a population of ligand conformations over generations to minimize energy. ◦ *Systematic Search:* Used by Glide, this method exhaustively explores conformational space using hierarchical filters. ◦ *Monte Carlo Simulated Annealing:* Used by Rosetta and Vina, this method makes random perturbations and accepts/rejects them based on the Metropolis criterion. b. **Scoring Function:** This component evaluates the generated poses to estimate the free energy of binding (Δ*G bind*). The accuracy of docking is entirely dependent on the scoring function's ability to discriminate between the correct bioactive pose and decoys. ◦ *Empirical Scoring Functions:* Calibrated against experimental affinity data (e.g., PDBbind). They sum contributions from terms like hydrogen bonding, lipophilic contact, and entropy penalties (e.g., GlideScore, AutoDock Vina). ◦ *Knowledge-Based Potentials:* Derived from statistical analysis of interatomic contact frequencies in large crystal structure databases (e.g., DrugScore). ![Fig.3 Molecular docking tests how a ligand fits a protein, scores the fit, and separates useful poses from false ones.](images/748c6c0f63cfcb1a900b03eeb35af355a8829ecd-1000x546.png) **The Challenge of Flexibility** A critical limitation of docking in ADMET is protein flexibility. While standard docking treats the receptor as a rigid body, ADMET targets like CYP450s and P-glycoprotein exhibit significant plasticity (Induced Fit). A ligand might bind to an "open" conformation of CYP3A4 that is structurally distinct from the "closed" crystal structure. Rigid docking often produces false negatives in these scenarios, necessitating advanced techniques like Induced-Fit Docking (IFD) or Molecular Dynamics (MD) refinement, which drastically increase computational cost. ## **3. Comparative Analysis by ADMET Endpoint** To build an effective ADMET module, one cannot simply choose "QSAR" or "Docking" universally. The optimal strategy is endpoint-dependent. The following analysis evaluates the performance and suitability of each method across the major categories of the ADMET spectrum.(*ADMET Property Prediction through Combinations of Molecular Fingerprints*, 2023) ### **3.1 Absorption and Distribution (A/D)** **Dominant Paradigm:** QSAR (Ligand-Based) Properties such as Aqueous Solubility (LogS), Lipophilicity (LogP), Caco-2 Permeability, and Blood-Brain Barrier (BBB) penetration are largely governed by the global physicochemical nature of the molecule rather than a specific lock-and-key interaction with a single protein pocket.(*ADMET Benchmark Group*, 2025) - **Solubility & Lipophilicity:** These are thermodynamic equilibrium properties determined by solvation energy and crystal lattice energy. Molecular docking has no theoretical basis here, as there is no "solubility receptor." QSAR models, particularly those using descriptors like LogP, Molecular Weight, and Topological Polar Surface Area (TPSA), are the industry standard. Deep learning models (e.g., ChemProp) trained on datasets like AqSolDB (approx. 10,000 compounds) have achieved Mean Absolute Errors (MAE) rivaling experimental variance. - **Permeability (Caco-2/PAMPA):** While transporters play a role in intestinal absorption, passive diffusion is often the rate-limiting step for the majority of drug-like space. QSAR models utilizing TPSA and hydrogen bond counts are highly predictive. Docking is only relevant if one specifically suspects carrier-mediated transport (e.g., PepT1) is the primary absorption mechanism, which is a minority case. - **Blood-Brain Barrier (BBB):** Penetration is a function of lipid solubility, size, and P-gp efflux liability. While docking to P-gp can inform the efflux component (discussed below), the overall BBB status is best predicted by classification QSAR models. Benchmarks on the TDC BBB dataset (~2,000 compounds) show that classical Random Forest models using 2D fingerprints often perform as well as complex GNNs, providing a robust baseline for any new module.(Stratiichuk et al., 2025) ![Fig.4 QSAR links basic chemical properties to ADMET outcomes without using protein structure.](images/ab3358e65cf6e60dd256f5d8bc8c7e36d470f116-511x410.png) ### **3.2 Metabolism (M): The Cytochrome P450 Battleground** **Status:** Synergistic Application (Hybrid Models) The metabolism of xenobiotics is primarily mediated by the Cytochrome P450 (CYP) superfamily, with isoforms 3A4, 2D6, 2C9, 1A2, and 2C19 accounting for >75% of drug metabolism. This domain represents the most fertile ground for comparing and combining docking and QSAR.(Palestro et al., 2014) ### **The Case for QSAR in Metabolism** For high-throughput profiling, QSAR is the workhorse. Large public datasets (e.g., PubChem, ChEMBL) contain bioactivity data for thousands of compounds against major CYP isoforms. - **Classification:** Predicting if a molecule is an Inhibitor or Non-inhibitor. QSAR models trained on these binary labels are fast and effective. For example, classifiers for CYP2C9 and CYP2D6 inhibition in the TDC benchmark group utilize ~12,000 data points each, allowing Random Forest and GNN models to learn the "chemical signature" of inhibition without needing to solve the binding pose. - **Strengths:** Handles the "promiscuity" of CYPs well by learning features of diverse substrates. Computationally inexpensive, enabling the screening of millions of compounds.(Jain, 2025) ### **The Case for Docking in Metabolism** QSAR models generally fail to predict the *Site of Metabolism* (SOM) i.e., *where* on the molecule the oxidation will occur. This is where docking excels. - **SOM Prediction:** By docking a ligand into the heme-containing active site of a CYP isoform (e.g., CYP2D6 PDB 3QM4 or CYP3A4 PDB 5VCC), one can measure the distance between the heme iron and potential metabolic sites (e.g., methyl groups, aromatic rings). The geometry of the pose provides a mechanistic hypothesis for the metabolic product. - **Isoform Selectivity:** Docking can rationalize why a drug is metabolized by CYP2D6 (narrow, acidic pocket) versus CYP3A4 (large, flexible, hydrophobic pocket). This structural insight is invaluable for lead optimization when chemists attempt to "dial out" metabolic liability by modifying specific steric features. **Failure Modes:** The large, plastic active site of CYP3A4 is a notorious challenge for rigid docking. Standard AutoDock Vina or Glide protocols often fail to capture the binding of large ligands that induce significant conformational changes in the protein. In such cases, QSAR models often outperform docking in pure affinity prediction because they implicitly "learn" the flexibility from the training data.(Trott & Olson, 2009) ![Fig.5 comparative analysis by ADMET endpoint](images/6f45cd953f56c1019fa46e774dca31d933d1a85c-504x406.png) ### **3.3 Transporters: The P-glycoprotein Challenge** **Status:** QSAR Dominance due to Structural Complexity P-glycoprotein (P-gp/MDR1) is a broad-spectrum efflux pump responsible for multidrug resistance. It has a massive, flexible binding cavity (~6000 ų) that can accommodate multiple ligands simultaneously, making it a nightmare for standard docking protocols.(Koirala et al., 2025) - **Docking Limitations:** The "polyspecificity" of P-gp means standard scoring functions struggle to discriminate between binders and non-binders, as the binding energy landscape is flat and defined by vague hydrophobic interactions rather than specific hydrogen bonds. While flexible docking protocols (e.g., using induced-fit or ensemble docking) have shown some utility, they are computationally too expensive for routine screening modules. - **QSAR Utility:** Due to these structural difficulties, ligand-based QSAR classifiers (Substrate vs. Non-substrate; Inhibitor vs. Non-inhibitor) remain the industry standard. Dataset sizes for P-gp in TDC are modest (~1,200 compounds), but sufficient for Random Forest models to achieve usable accuracy (AUROC > 0.85). ### **3.4 Toxicity: hERG and Beyond** **Status:** High-Value Target for Hybrid Modeling **hERG Channel Inhibition:** Blockage of the hERG potassium channel is a primary cause of drug-induced QT prolongation and fatal arrhythmias. It is a regulatory hard-stop. - *QSAR:* The primary screen. Models trained on electrophysiology patch-clamp data (IC50) are standard. However, simple QSAR sometimes struggles with "activity cliffs"—where a small structural change causes a massive safety shift.(Creanza et al., 2021) - *Docking:* Recent cryo-EM structures of hERG (e.g., PDB 5VA1) have revolutionized this field. Docking reveals that many blockers bind within the central pore via pi-stacking interactions with residues like Tyr652 and Phe656.(Maroua Fattouche et al., 2024) - *Hybrid Advantage:* Benchmarking studies have demonstrated that integrating docking scores (specifically the interaction energy with key pore residues) as features *into* a QSAR model (Structure-Based QSAR) significantly improves predictive accuracy over 2D descriptors alone. This "hybrid" approach captures both the general chemical properties (QSAR) and the specific steric fit into the channel pore (Docking). | ADMET Endpoint | Primary Method | Secondary Method | Key Limitation of Secondary | Recommended Workflow | | --- | --- | --- | --- | --- | | Solubility (LogS) | QSAR | None | No specific target | QSAR (GNN or RF) | | Permeability (Caco-2) | QSAR | Docking | Only for specific carriers | QSAR (RF with TPSA/MW) | | Metabolism (CYP450) | QSAR | Docking | Protein flexibility (Induced fit) | QSAR for screening; Docking for SOM | | Clearance (CLint) | QSAR | PBPK | Requires Vmax/Km inputs | QSAR -> PBPK Integration | | hERG Toxicity | QSAR | Docking | Scoring accuracy | Hybrid (QSAR + Docking Features) | | P-gp Efflux | QSAR | Docking | Huge, flexible binding site | QSAR (Classification) | | Mutagenicity (Ames) | QSAR | None | Target complexity (DNA/Enzymes) | QSAR (Structural Alerts) | Table. 1 Summary of Methodological Applicability ## 4. Case Study: Navigating the hERG Activity Cliff of Terfenadine vs. Fexofenadine To illustrate the practical friction between QSAR and structural modeling, we examined one of the most famous "activity cliffs" in pharmaceutical history: the relationship between the withdrawn antihistamine **Terfenadine** and its safe metabolite, **Fexofenadine**. Terfenadine was withdrawn from the market in 1997 due to fatal arrhythmias caused by potent hERG channel blockade. Fexofenadine, which differs structurally by only a single carboxylic acid group on the terminal benzene ring, exhibits a 200-fold reduction in hERG affinity and is safe for daily clinical use. We subjected both molecules to the **ADMET-AI** QSAR model (Swanson et al., 2024) and compared these predictions against experimental IC50 data and structural literature. ### 4.1. The QSAR Blind Spot: Sensitivity vs. Magnitude The ADMET-AI model successfully identified the trend but failed to capture the magnitude of the safety margin. - **Terfenadine:** Predicted hERG probability of **97.7%** (High Risk). - **Fexofenadine:** Predicted hERG probability of **68.2%** (High Risk). While the model correctly ranked Terfenadine as the more toxic entity (a 29.5% probability gap), it still classified Fexofenadine as a "high risk" compound. In a binary screening funnel, **both molecules might have been discarded.** This stands in stark contrast to experimental reality. Electrophysiology data (Kamiya et al., 2008; Scherer et al., 2002) establishes Terfenadine’s IC50 at **15–56 nM**, whereas Fexofenadine’s IC50 shifts dramatically to **~4,590 nM**. **Why did QSAR struggle?** The failure is rooted in the descriptors. The molecules are physicochemical twins: they share a ~30 Da molecular weight difference and very similar lipophilicity profiles (LogP 6.45 vs 5.51). Because the ADMET-AI model relies on learned patterns from 2D and 3D descriptors, it views these two molecules as structurally synonymous. It sees the "toxicophore" (the central piperidine/phenyl motif) in both, but cannot fully weigh the subtle electronic effect of the carboxylic acid that neutralizes the toxicity. ### 4.2. The Structural Imperative: The Tetramer Trap If QSAR lacks the resolution to distinguish these molecules, molecular docking should theoretically fill the gap. However, this case study highlights a critical dependency in structure-based design: **Biological Assembly.** Preliminary docking of both ligands into a **single hERG subunit** (monomer) failed to reproduce the experimental 200-fold affinity gap. The scoring functions could not differentiate the binding energies significantly. This aligns with findings by **Vaz et al. (2011)** and **Wu et al. (2021)**, which demonstrate that accurate hERG predictions require the full **tetrameric pore assembly**. The toxicity of Terfenadine arises from specific pi-stacking interactions with **Phe656** and **Tyr652** within the central pore. The addition of the carboxylic acid in Fexofenadine creates an electrostatic penalty and steric clash *only* when the channel is modeled in its complete tetrameric state. ### Summary This case study perfectly encapsulates the "Hybrid" argument. - **QSAR** correctly flagged the scaffold as dangerous but lacked the nuance to "clear" the safe metabolite. - **Docking** holds the potential to explain the mechanism (the carboxylic acid clash), but only if the simulation environment (the tetramer) accurately reflects biological reality. Relying on either method in isolation would have resulted in a false positive (killing a safe drug via QSAR) or a potential false negative (missing the toxicity via monomer-based docking). ![Fig. 6 HeRg activity Cliff case study](images/787327217323f5492738d25c1c3ef53ea4f02d69-1000x483.png) ## **Conclusions and Strategic Recommendations** For the development of a robust internal ADMET module, the "Docking vs. QSAR" debate should be reframed as a "Tiered Integration" strategy. The evidence suggests neither tool is sufficient alone, but together they form a comprehensive filter. ### **Strategic Recommendations** 1. **Backbone with QSAR:** The primary engine of the module should be QSAR. It is the only method scalable to large libraries and applicable to all endpoints (solubility, permeability, global toxicity). 1. **Use Random Forest as the Benchmark:** Do not assume Deep Learning is necessary. Start with Random Forest (ECFP4). Only deploy ChemProp/GNNs if they demonstrate a statistically significant improvement (>0.02 AUROC) on **Scaffold Splits**. Complexity without gain is technical debt. 1. **Deploy Hybrid Models for High-Risk Targets:** For critical safety endpoints like hERG and major metabolic liabilities like CYP3A4, implement a hybrid pipeline. Dock the compounds, extract the scores, and use them as features. This adds mechanistic resilience to the statistical predictions. 1. **Rigorous Validation:** Reject any model validated solely on random splits. Implement scaffold splitting and use imbalance-aware metrics (AUPRC, MCC) to ensure the module provides value in real-world discovery campaigns. ## References *ADMET Benchmark Group*. (2025). TDC. [https://tdcommons.ai/benchmark/admet_group/overview/](https://tdcommons.ai/benchmark/admet_group/overview/) *ADMET property prediction through combinations of molecular fingerprints*. (2023). Ar5iv. [https://ar5iv.labs.arxiv.org/html/2310.00174](https://ar5iv.labs.arxiv.org/html/2310.00174) Creanza, T. M., Delre, P., Ancona, N., Lentini, G., Saviano, M., & Mangiatordi, G. F. (2021a). Structure-Based Prediction of hERG-Related Cardiotoxicity: A Benchmark Study. *Journal of Chemical Information and Modeling*, *61*(9), 4758–4770. [https://doi.org/10.1021/acs.jcim.1c00744](https://doi.org/10.1021/acs.jcim.1c00744) Creanza, T. M., Delre, P., Ancona, N., Lentini, G., Saviano, M., & Mangiatordi, G. F. (2021b). Structure-Based Prediction of hERG-Related Cardiotoxicity: A Benchmark Study. *Journal of Chemical Information and Modeling*, *61*(9), 4758–4770. [https://doi.org/10.1021/acs.jcim.1c00744](https://doi.org/10.1021/acs.jcim.1c00744) Farid Amellal, & Billette, J. (1996). Selective Functional Properties of Dual Atrioventricular Nodal Inputs. *Circulation*, *94*(4), 824–832. [https://doi.org/10.1161/01.cir.94.4.824](https://doi.org/10.1161/01.cir.94.4.824) Jain, A. (2025, January 5). *Undersampling, Oversampling and SMOTE, Ensemble Method and Cost Sensitive Learning techniques for…*. Medium. [https://medium.com/@abhishekjainindore24/undersampling-oversampling-and-smote-ensemble-mehtod-and-cost-sensitive-learning-techniques-for-08efb557ec68](https://medium.com/@abhishekjainindore24/undersampling-oversampling-and-smote-ensemble-mehtod-and-cost-sensitive-learning-techniques-for-08efb557ec68) Kamiya, K., Niwa, R., Morishima, M., Haruo Honjo, & Sanguinetti, M. C. (2008). Molecular Determinants of hERG Channel Block by Terfenadine and Cisapride. *Journal of Pharmacological Sciences*, *108*(3), 301–307. [https://doi.org/10.1254/jphs.08102fp](https://doi.org/10.1254/jphs.08102fp) Knape, K., Linder, T., Wolschann, P., Beyer, A., & Stary-Weinzinger, A. (2011). In silico Analysis of Conformational Changes Induced by Mutation of Aromatic Binding Residues: Consequences for Drug Binding in the hERG K+ Channel. *PLoS ONE*, *6*(12), e28778. [https://doi.org/10.1371/journal.pone.0028778](https://doi.org/10.1371/journal.pone.0028778) Koirala, M., Yan, L., Mohamed, Z., & DiPaola, M. (2025). AI-Integrated QSAR Modeling for Enhanced Drug Discovery: From Classical Approaches to Deep Learning and Structural Insight. *International Journal of Molecular Sciences*, *26*(19), 9384. [https://doi.org/10.3390/ijms26199384](https://doi.org/10.3390/ijms26199384) Maroua Fattouche, Salah Belaidi, Oussama Abchir, Walid Al-Shaar, Younes, K., Muneerah Mogren Al-Mogren, Samir Chtita, Soualmia, F., & Majdi Hochlaf. (2024). ANN-QSAR, Molecular Docking, ADMET Predictions, and Molecular Dynamics Studies of Isothiazole Derivatives to Design New and Selective Inhibitors of HCV Polymerase NS5B. *Pharmaceuticals*, *17*(12), 1712–1712. [https://doi.org/10.3390/ph17121712](https://doi.org/10.3390/ph17121712) Palestro, P. H., Gavernet, L., Estiu, G. L., & Bruno, L. E. (2014). Docking Applied to the Prediction of the Affinity of Compounds to P-Glycoprotein. *BioMed Research International*, *2014*, 1–10. [https://doi.org/10.1155/2014/358425](https://doi.org/10.1155/2014/358425) Scherer, C. R., Lerche, C., Decher, N., Dennis, A. T., Maier, P., Ficker, E., Busch, A. E., Bernd Wollnik, & Steinmeyer, K. (2002). The antihistamine fexofenadine does not affect *I*Kr currents in a case report of drug‐induced cardiac arrhythmia. *British Journal of Pharmacology*, *137*(6), 892–900. [https://doi.org/10.1038/sj.bjp.0704873](https://doi.org/10.1038/sj.bjp.0704873) Stratiichuk, R., Shevchuk, N., Kyrylenko, R., Vozniak, V., Koleiev, I., Voitsitskyi, T., Husak, V., Ostrovsky, Z., Khropachov, I., Starosyla, S., Yesylevsky, S., & Nafiev, A. (2025). *Improving ADMET prediction with descriptor augmentation of Mol2Vec embeddings*. [https://doi.org/10.1101/2025.07.14.664363](https://doi.org/10.1101/2025.07.14.664363) Swanson, K., Walther, P., Leitz, J., Mukherjee, S., Wu, J. C., Shivnaraine, R. V., & Zou, J. (2024). ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries. *Bioinformatics*, *40*(7). [https://doi.org/10.1093/bioinformatics/btae416](https://doi.org/10.1093/bioinformatics/btae416) Trott, O., & Olson, A. J. (2009). AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. *Journal of Computational Chemistry*, *31*(2), 455–461. [https://doi.org/10.1002/jcc.21334](https://doi.org/10.1002/jcc.21334) --- # Structural Plasticity in the Mutome: Mechanisms of Binding Pocket Alteration and Therapeutic Intervention URL: https://www.lite.bio/blogs/structural-plasticity-in-the-mutome-mechanisms-of-binding-pocket-alteration-and-therapeutic-intervention Author: Aditi Sinha Date: 2025-12-10 ### **1. From Handshakes to Dynamic Landscapes** Imagine a handshake. It seems like a simple gesture, two hands clasping. But consider the nuance. If your hand is rigid like stone, the handshake fails. If it is too limp, the connection is weak. A perfect handshake requires the hand to conform, to adjust its pressure and shape in response to the other person. It is a dynamic, mutual adaptation. For over a century, biology textbooks taught us that proteins and drugs interact like a "Lock and Key". In this rigid, mechanical analogy, the protein (the lock) sits frozen in a single shape, waiting for a perfectly carved ligand (the key) to slide in. While this model, championed by Emil Fischer in 1894, gave us the basic concept of specificity, it is, in the modern view, functionally dead. Proteins are not locks. They are breathing, vibrating, shifting molecular machines. They are closer to the "handshake" or a "glove fitting a hand". They exist in a constantly shifting landscape of shapes an ensemble of conformations. In the high-stakes world of oncology, this dynamic nature is where the battle is fought. Cancer is often described as a disease of the genome, but functionally, it is a disease of the proteome's structure. A single mutation in a DNA sequence translates to a swapped amino acid in a protein. This tiny change can act like a piece of grit in a gearbox or a wedge in a door. It can lock a protein in an "always-on" handshake, signaling cells to divide uncontrollably. It can reshape a binding pocket just enough that a life-saving drug can no longer fit, or it can create a brand-new, secret cavity a "cryptic pocket" that we never knew existed.(*Structural Biochemistry/Protein Function/Lock and Key - Wikibooks, Open Books for an Open World*, 2022) This Blog explores the structural biology of the "Mutome" the universe of mutant proteins driving cancer. We will journey from the atomic physics of binding pockets to the clinical realities of drug resistance. We will dismantle the mechanisms of notorious killers like EGFR, KRAS, and BRAF, and see how computational wizardry is helping us find pockets that don't exist until we look for them. This is the story of how shape dictates destiny in cancer biology. ![Fig.1 Beyond the Lock and Key. Visualizing the paradigm shift](images/87b9e7f98fc1e860025b39fe1b41dfa3d8946b99-1000x546.png) ### **2. The Biophysics of Binding: Beyond the Lock and Key** To understand how cancer mutations wreak havoc, we must first establish the physical rules of how drugs bind to proteins. The interaction is governed by thermodynamics, specifically the Gibbs free energy of binding (ΔG), which dictates the affinity of a drug for its target.(Tripathi & Bankaitis, 2018) **2.1 The Thermodynamics of Specificity** The binding event is a delicate balance between enthalpy (ΔH) and entropy (ΔS). **Enthalpy (ΔH):** This represents the heat energy released or absorbed. It is driven by specific interactions: hydrogen bonds (like Velcro hooks), salt bridges (magnetic attraction between charges), and van der Waals forces (shape complementarity). A mutation that removes a hydrogen bond donor, for example, changing a Threonine to an Alanine, directly penalizes enthalpy. **Entropy (ΔS):** This reflects the disorder of the system. When a drug binds, it loses its freedom to tumble in solution (unfavorable entropy). However, binding often displaces highly ordered "unhappy" water molecules trapped in the pocket, releasing them into the bulk solvent (favorable entropy). This "hydrophobic effect" is a primary driver of drug binding. ### **2.2 The Solvation Landscape** Water is not just a background solvent; it is a structural participant. In many binding pockets, water molecules bridge the interaction between protein and ligand. A mutation that changes the polarity of a pocket say, replacing a hydrophobic Leucine with a polar Arginine completely reorganizes these water networks. The drug must now pay a higher energetic penalty to strip these water molecules away (desolvation penalty) before it can bind. This subtle shifting of "water architecture" is often how resistance mutations work without causing obvious steric clashes.(*RLO: Lock and Key Hypothesis*, 2025) ### **2.3 Models of Binding: Induced Fit vs. Conformational Selection** The "Lock and Key" model fails because it assumes rigidity. Current biophysics relies on two more sophisticated models : 1. **Induced Fit:** The protein is flexible but exists in a ground state. As the ligand approaches, its electrostatic and chemical field induces a conformational change in the protein, molding the pocket around the ligand. This is the "hand in glove" model. 1. **Conformational Selection:** The protein naturally fluctuates between many shapes (an ensemble) even without the ligand. The ligand selectively binds to and stabilizes one of these pre-existing, high-energy conformations. This model is crucial for understanding allosteric inhibitors, which bind to shapes that might only exist for a microsecond Cancer mutations perturb these dynamics. A mutation might destabilize the "inactive" shape of a kinase, shifting the entire population toward the "active" shape. If a drug prefers the inactive shape (as many do), it suddenly finds no target to bind to, even if the binding pocket itself looks unchanged in a static crystal structure. ### **3. Mechanisms of Mutation-Induced Pocket Alteration** Somatic missense mutations single amino acid substitutions are the architects of structural resistance. They alter binding pockets through distinct physical mechanisms. ### **3.1 Steric Hindrance: The "Gatekeeper" Phenomenon** The most direct mechanism is the introduction of bulk. Inhibitor binding pockets often contain a deep hydrophobic cleft. A residue located at the entrance or "gate" of this cleft controls access. - **Mechanism:** If a small residue (like Threonine) is mutated to a bulky one (like Methionine or Isoleucine), the side chain physically protrudes into the space occupied by the drug. This is the "Gatekeeper" mutation seen across the kinome (e.g., T790M in EGFR, T315I in ABL). - **Nuance:** It is rarely *just* steric. As we will see with EGFR, bulky residues also enhance van der Waals interactions with the natural substrate (ATP), effectively increasing the "fuel" affinity and allowing the enzyme to outcompete the inhibitor. ### **3.2 Electrostatic Remodeling** Replacing a neutral residue with a charged one (or vice versa) alters the electrostatic potential surface of the pocket. In the ABL kinase, resistance mutations often occur not in the deep pocket but on the P-loop (phosphate-binding loop) which clamps down on the drug. A mutation here might alter the charge distribution, disrupting the electrostatic steering that guides the drug into the pocket. - **pH Dependence:** Some mutations alter the pKa of surrounding residues, making binding sensitive to the cellular pH environment, a factor often overlooked in standard assays. ### **4. Case Study: The EGFR Gatekeeper Saga** The Epidermal Growth Factor Receptor (EGFR) is the "poster child" for structural oncology. Its journey from a targetable driver to a resistant mutant and back again illustrates the arms race between drug design and protein evolution. **4.1 The Structural Basis of Activation (L858R and Del19)** In healthy cells, EGFR exists in an equilibrium favoring an autoinhibited tethered conformation. Ligand binding (like EGF) induces dimerization and activation. • **L858R:** This mutation, located in the activation loop, substitutes a hydrophobic Leucine with a large, positively charged Arginine. In the Wild Type (WT) protein, Leucine 858 packs into a hydrophobic cluster that stabilizes the inactive helical conformation. The Arginine disrupts this cluster, preventing the inactive state from forming. Consequently, the kinase snaps into the active conformation an asymmetric dimer even without a ligand. This "always on" state drives the cancer. ![Fig.2 Structural Basis of EGFR Activation: WT vs L858R](images/64d00def064c3111c54ca0ad2b0b785bb84ae80d-1000x491.png) **4.2 The First Generation: Competitive Inhibition** First generation drugs like **Gefitinib** and **Erlotinib** are reversible, ATP competitive inhibitors. They work because the L858R mutation, while activating the kinase, also lowers its affinity for ATP compared to the drug (specifically, it has a lower Km for the drug relative to ATP than WT). This "therapeutic window" allows the drug to shut down the mutant receptor while sparing the wild type receptor in healthy tissues.(Zhu et al., 2018) ![Fig.3 Drug Selectivity affecting the Mutant and the Wild type](images/290fa993ace30eaf78ecd2827cc569a854231248-1000x546.png) **4.3 The T790M Resistance Mechanism:** **A Biophysical Debate** Inevitably, tumors develop resistance. In ~50% of cases, this is due to the T790M mutation. Threonine 790 is the "gatekeeper" residue located deep in the ATP-binding cleft. • *The Steric Hypothesis*: Initially, it was believed that the Methionine side chain was simply too big, physically blocking Gefitinib binding. • *The Affinity Hypothesis (The Real Culprit)*: Detailed kinetic and structural studies revealed a more sophisticated mechanism. Methionine does *not* completely block Gefitinib binding; the drug can still fit (albeit poorly). However, the T790M mutation drastically increases the receptor's affinity for ATP. The Methionine side chain creates a perfect hydrophobic environment for the adenine ring of ATP, locking it in tighter. In the cell, where ATP concentration is high (millimolar range), the drug can no longer compete. T790M restores the ATP affinity that the original L858R mutation had lost.(Spellmon et al., 2017) ![Fig.4 The Steric Hypothesis vs The affinity Hypothesis](images/c4fe9ea375aa31796a2fa729576400acaeab982b-1000x476.png) **4.4 Third-Generation Covalent Solutions** To defeat T790M, chemists changed the rules. Instead of competing reversibly, they designed **Osimertinib** (and others like Rociletinib). These drugs carry a "warhead" an acrylamide group. They bind to the pocket and then undergo a Michael addition reaction with a specific residue: **Cysteine 797 (C797)**. • **The Structural Trick:** Once the covalent bond forms, the affinity is effectively infinite. The drug cannot wash off. Even if T790M makes the pocket love ATP, the covalent drug permanently occupies the site, shutting down the kinase.(Maloney et al., 2021) **4.5 The C797S Counter-Move and the "Hydrophobic Clamp"** The tumor eventually counters with **C797S** (Cysteine to Serine). - **Loss of Handle:** Serine lacks the thiol group needed for the covalent bond. Osimertinib becomes a reversible inhibitor again, and because of T790M, it loses to ATP.(Yun et al., 2008) - **Structural Remodeling:** But it goes deeper. Structural analysis of C797S mutants revealed the importance of a **"Hydrophobic Clamp"** (residues Leu718 and Val726). The mutation alters the flexibility of this region. New "fourth-generation" allosteric inhibitors (like EAI045) attempt to bypass the ATP site entirely by binding to a separate allosteric pocket created by the displacement of the C-helix, effectively clamping the kinase jaw shut from the outside(Kar et al., 2010) | Stage | Mutation | Structural Effect | Drug Class | Resistance Mechanism | | --- | --- | --- | --- | --- | | Oncogenesis | L858R / Del19 | Destabilizes inactive state; promotes active asymmetric dimer. | 1st Gen (Gefitinib) | High affinity for active mutant state. | | Acquired Resistance | T790M | Increases affinity for ATP; restores Km for ATP to WT levels; steric clash. | 2nd Gen (Afatinib) / 3rd Gen (Osimertinib) | T790M outcompetes reversible drugs; 3rd Gen uses covalent bond to C797. | | Late Resistance | C797S | Removes nucleophilic thiol (SH to OH); prevents covalent bonding. | 4th Gen (EAI045) | Drug reverts to reversible binding and is outcompeted by ATP. | | Allosteric Escape | L718Q / G724S | Alters "Hydrophobic Clamp" geometry; steric clash with drug core. | Allosteric Inhibitors | Changes in pocket shape preventing binding. | Table: Examples of few more mutation and their drug class ### **5. Computational Frontiers: Hunting the Invisible** How do we find pockets that only exist for a microsecond? We use computational microscopes. ### **5.1 Molecular Dynamics (MD) and "Mixed Solvents"** Static crystallography is blind to cryptic pockets. MD simulations act as a movie. **Mixed Solvent MD,** Researchers flood the virtual simulation box not just with water, but with small organic probes (benzene, acetonitrile). These probes wiggle into transient crevices on the protein surface(Dmitri Beglov et al., 2018). If a probe lingers in a spot (high residence time), it indicates a potential "cryptic" binding site that can be wedged open by a drug.(Li et al., 2014) ### **5.2 Markov State Models (MSMs)** Proteins transition between thousands of micro-states. MSMs map these states into a network.(Creative Biostructure, 2024) **Mapping the Path:** By simulating thousands of short trajectories and stitching them together, MSMs can predict the pathway a protein takes to open a cryptic pocket. This revealed, for instance, that the "undruggable" p53 Y220C mutant possesses a transiently open trench that could be targeted to stabilize the protein.(Oleinikovas et al., 2016) ### **5.3 AI and Graph Neural Networks (PocketMiner)** MD is computationally expensive (months of supercomputer time). New AI tools like **PocketMiner** treat the protein structure as a geometric graph. **Pattern Recognition:** Trained on thousands of MD trajectories, the AI learns to recognize the subtle geometric "wobble" of residues that precedes the opening of a pocket. It can scan the entire human proteome in days, predicting which "smooth" proteins might actually have hidden pockets waiting to be drugged.(Meller et al., 2023) ![Fig.5 Hunting invisible cryptic pockets Through various methods](images/b52af77f051da7ecac4bf4a07c68b1e7e2524dec-1000x465.png) ### **6. Conclusion:** The battle against cancer mutations is no longer just about finding a molecule that fits a hole. It is about understanding the fourth dimension of biology: **time and motion**. - **Enthalpy** dictates the strength of the grip. - **Entropy** dictates the cost of freezing the motion. - **Water** dictates the hidden energy penalties. - **Kinetics** dictates how long the drug holds on. We have moved from the "Lock and Key" to the "Dynamic Ensemble." We are designing drugs that act as molecular glue, allosteric wedges, and covalent traps. We are using supercomputers to watch proteins breathe and AI to predict their next move. The Mutome is complex and adaptive, but by understanding the deep physics of its structural plasticity, we are slowly learning how to lead the dance. ### References **References** Creative Biostructure. (2024, August 18). *Protein-Ligand Interaction*. [Creative-Biostructure.com](http://Creative-Biostructure.com); Creative Biostructure. [https://www.creative-biostructure.com/proteinligand-interation.htm](https://www.creative-biostructure.com/proteinligand-interation.htm) Dmitri Beglov, Hall, D. R., Wakefield, A. E., Luo, L., Allen, K. N., Dima Kozakov, Whitty, A., & Vajda, S. (2018). Exploring the structural origins of cryptic sites on proteins. *Proceedings of the National Academy of* *Sciences*, *115*(15), E3416–E3425. [https://doi.org/10.1073/pnas.1711490115](https://doi.org/10.1073/pnas.1711490115) Kar, G., Keskin, O., Gursoy, A., & Nussinov, R. (2010). Allostery and population shift in drug discovery. *Current Opinion in Pharmacology*, *10*(6), 715–722. [https://doi.org/10.1016/j.coph.2010.09.002](https://doi.org/10.1016/j.coph.2010.09.002) Li, M., Petukh, M., Alexov, E., & Panchenko, A. R. (2014). Predicting the Impact of Missense Mutations on Protein–Protein Binding Affinity. *Journal of Chemical Theory and Computation*, *10*(4), 1770–1780. [https://doi.org/10.1021/ct401022c](https://doi.org/10.1021/ct401022c) Maloney, R. C., Zhang, M., Jang, H., & Nussinov, R. (2021). The mechanism of activation of monomeric B-Raf V600E. *Computational and Structural Biotechnology Journal*, *19*, 3349–3363. [https://doi.org/10.1016/j.csbj.2021.06.007](https://doi.org/10.1016/j.csbj.2021.06.007) Meller, A., Ward, M., Borowsky, J., Meghana Kshirsagar, Lotthammer, J. M., Oviedo, F., Ferres, J. L., & Bowman, G. R. (2023). Predicting locations of cryptic pockets from single protein structures using the PocketMiner graph neural network. *Nature Communications*, *14*(1), 1177–1177. [https://doi.org/10.1038/s41467-023-36699-3](https://doi.org/10.1038/s41467-023-36699-3) Oleinikovas, V., Saladino, G., Cossins, B. P., & Gervasio, F. L. (2016). Understanding Cryptic Pocket Formation in Protein Targets by Enhanced Sampling Simulations. *Journal of the American Chemical Society*, *138*(43), 14257–14263. [https://doi.org/10.1021/jacs.6b05425](https://doi.org/10.1021/jacs.6b05425) *RLO: Lock and Key Hypothesis*. (2025). [Nottingham.ac.uk](http://Nottingham.ac.uk). [https://www.nottingham.ac.uk/helmopen/rlos/pharmacology/pharmacodynamics/lock_and_key/2.html](https://www.nottingham.ac.uk/helmopen/rlos/pharmacology/pharmacodynamics/lock_and_key/2.html) Spellmon, N., Li, C., & Yang, Z. (2017). Allosterically targeting EGFR drug-resistance gatekeeper mutations. *Journal of Thoracic Disease*, *9*(7), 1756–1758. [https://doi.org/10.21037/jtd.2017.06.43](https://doi.org/10.21037/jtd.2017.06.43) *Structural Biochemistry/Protein function/Lock and Key - Wikibooks, open books for an open world*. (2022). [Wikibooks.org](http://Wikibooks.org). [https://en.wikibooks.org/wiki/Structural_Biochemistry/Protein_function/Lock_and_Key](https://en.wikibooks.org/wiki/Structural_Biochemistry/Protein_function/Lock_and_Key) Tripathi, A., & Bankaitis, V. A. (2018). Molecular Docking: From Lock and Key to Combination Lock. *Journal of Molecular Medicine and Clinical Applications*, *2*(1). [https://doi.org/10.16966/2575-0305.106](https://doi.org/10.16966/2575-0305.106) Yun, C.-H., Mengwasser, K. E., Toms, A. V., Woo, M. S., Greulich, H., Wong, K.-K., Meyerson, M., & Eck, M. J. (2008). The T790M mutation in EGFR kinase causes drug resistance by increasing the affinity for ATP. *Proceedings of the National Academy of Sciences*, *105*(6), 2070–2075. [https://doi.org/10.1073/pnas.0709662105](https://doi.org/10.1073/pnas.0709662105) Zhu, S.-J., Zhao, P., Yang, J., Ma, R., Yan, X.-E., Yang, S.-Y., Yang, J.-W., & Yun, C.-H. (2018). Structural insights into drug development strategy targeting EGFR T790M/C797S. *Oncotarget*, *9*(17), 13652–13665. [https://doi.org/10.18632/oncotarget.24113](https://doi.org/10.18632/oncotarget.24113) --- # Small Molecule vs Peptide Competition for the Same Pocket URL: https://www.lite.bio/blogs/small-molecule-vs-peptide-competition-for-the-same-pocket Author: Aditi Sinha Date: 2025-11-29 In the high stakes world of modern drug discovery, finding a binding pocket on a disease causing protein is only half the battle. The real challenge and the subject of intense biophysical debate is deciding what kind of "key" should fit that lock. For most of the 20th century, the pharmaceutical industry placed its bets on **Small Molecules**: rigid, synthetically crafted chemical structures (typically under 500 Daltons) designed to fit into deep enzymatic pockets with sub angstrom precision. These are the "lock picks" of biology. They are precise, efficient, and easily swallowed in a pill. But biology is rarely static. As we move toward the frontier of "undruggable" targets specifically **Protein-Protein Interactions (PPIs)** the terrain changes. We are no longer looking at deep, stable caverns, but at vast, flat, and shifting landscapes. Here, a new competitor dominates: **Peptides**. This is not merely a difference in size; it is a clash of biophysical philosophies. Small molecules rely on **Enthalpy** (the energy of static attraction), while peptides are masters of **Entropy** (the energy of disorder and dynamics). In this post, we will dissect the competition between these two modalities. We will move beyond the simple "lock and key" analogy to explore the deep physics of shape complementarity, the chaotic war for water displacement, and the kinetic race to the pocket. To understand why small molecules and peptides behave so differently, you first have to look at the battlefield. The structural "terrain" of a binding site dictates the rules of engagement. ## **The Small Molecule: The "Lock Pick"** Small molecules are designed for deep concavity. They thrive in active sites the catalytic clefts of enzymes like kinases or proteases. In these deep pockets, a small molecule can be surrounded on almost all sides by protein residues. This high degree of enclosure allows the molecule to maximize **Van der Waals interactions** (packing forces) despite having a small mass. Because they have a limited surface area (typically burying only 300–500 Å ), small molecules cannot afford "wasted" space. They must achieve what medicinal chemists call **Ligand Efficiency (LE)**. Every atom must contribute to the binding energy. They rely on "hotspots" specific residues (often Tryptophan, Tyrosine, or Arginine) that act as anchors. A small molecule must position a hydrogen bond donor or acceptor with sub-angstrom precision to match these anchors. If the pocket shifts by even 0.5 Å, the "lock pick" may fail to turn. ## **The Peptide: The "Velcro Handshake"** Peptides, by contrast, are evolved to dominate open plains. Protein-Protein Interaction (PPI) interfaces are typically large (1,000–2,000 Å), flat, and hydrophobic. A small molecule trying to bind here is like a climber trying to scale a glass wall there are no deep crevices to grab.(Vikram Gaikwad et al., 2025) Peptides solve this through **Distributed Affinity**. Instead of relying on a single deep anchor, they spread their interactions across a massive surface area. They act like "molecular velcro." Even if one section of the interface releases, the rest holds firm. This allows peptides to bridge discontinuous sub-pockets that are far apart. For example, the p53 tumor suppressor peptide interacts with the MDM2 protein by inserting three specific residues (Phe19, Trp23, Leu26) into shallow clefts spread across a long helix. A small molecule has to be artificially stretched or "linked" to cover that same distance, often breaking the rules of oral bioavailability (Rule of Five) in the process. (Atanu Maity et al., 2020). ![Fig.1 Interaction of peptide and small molecule with the same MDM2 protein.](images/3b495defc2ae8eda984da2ae2b77dfa5471985e4-1000x546.png) # The Water Game: "Unhappy" Water vs. The Hydrophobic Effect Water is not just a passive background solvent; it is an active, aggressive competitor for the binding site. Before a drug can bind to a protein, it must strip away the water molecules clinging to the protein surface. This "desolvation" cost is the thermodynamic gatekeeper of binding. ## Small Molecules: The Sniper Approach In deep binding pockets, water molecules often get trapped. A water molecule in the bulk solvent is "happy", it can tumble freely and form 3 to 4 hydrogen bonds with its neighbors. But when a water molecule is trapped in a hydrophobic protein cavity, it loses this freedom. It cannot rotate (entropic penalty) and may not find enough partners to hydrogen bond with (enthalpic penalty). These are termed "Unhappy Waters" or high-energy hydration sites.(DeAngelo et al., 2025) Small molecules exploit this misery. A central strategy in modern drug design is to place a hydrophobic group (like a methyl or chloro group) exactly where an "unhappy" water molecule sits. By displacing this frustrated water molecule and releasing it back into the bulk solvent, the drug liberates energy. - The Energy Payoff: Releasing a single "unhappy" water molecule can yield between -2 to -6 kcal/mol of free energy. - Precision: Small molecules act like snipers, targeting these specific high-energy waters to generate affinity without needing a massive surface area.(Raddi & Voelz, 2023) ![Fig.2 Small molecules sniper approach and unhappy water displacement](images/2b174c2ccacaead7cda809db6fa04ad43d64ae28-1000x508.png) ## Peptides: The Bulk Approach Peptides play a different game. Because they cover such a vast surface area, they don't just displace one or two waters; they displace a whole network. This is the Hydrophobic Effect operating on a macro scale. When a peptide binds to a flat hydrophobic interface, it releases dozens of water molecules from the protein surface. While these surface waters are not as "unhappy" as the deep cavity waters (the energy gain per molecule is lower, perhaps -0.5 to -1.0 kcal/mol), the quantity is what matters. The release of 20 or 30 water molecules generates a massive surge of solvent entropy ( ΔS{solv}. This entropic boost is often the primary engine that drives peptide binding, compensating for the fact that their individual hydrogen bonds might not be as optimized as those of a synthetic small molecule. # The Kinetic Race: How They Enter the Pocket Thermodynamics tells us *if* a binding event will happen; kinetics tells us *how fast*. Recent research has shown that the speed of entry (K{on}) and the duration of stay (Residence Time) are critical for drug efficacy. Here, the flexible nature of peptides gives them a surprising advantage. ## The "Fly-Casting" Mechanism You might assume that a floppy, disordered peptide would bind slower than a rigid, compact small molecule. The opposite is often true. Peptides, particularly intrinsically disordered ones (IDPs), utilize a mechanism known as Fly-Casting. Imagine trying to catch a ball (the protein) with your hand (the small molecule) versus catching it with a large net (the peptide). 1. **The Net:** In solution, the disordered peptide has a large "capture radius." It occupies a greater volume of space than a compact molecule. 1. **The Hook:** One end of the peptide (the "anchor" residue) makes the initial contact with the protein. 1. **The Reel-In:** Once anchored, the rest of the peptide folds down onto the protein surface in a "dock-and-coalesce" maneuver. This allows peptides to have association rates (K{on}) that are 1.5 to 2.5 times faster than the diffusion limit of a rigid sphere.They don't have to wait for a perfect collision; they snag the protein and reel themselves in.(“Overcoming the Shortcomings of Peptide-Based Therapeutics,” 2022) ## Small Molecules: The Diffusion Limit Small molecules rely on Brownian diffusion. They zip around the solution and must hit the binding pocketw with the correct orientation. If the pocket is "cryptic" (meaning it is closed most of the time and only flickers open occasionally), the small molecule is at the mercy of the protein's internal dynamics.(Wang et al., 2021) It must wait for the protein to "breathe" open—a mechanism called Conformational Selection. This can severely limit how fast a small molecule can bind, making it difficult to target proteins with short half lives.(Wang et al., 2023) ![Fig.3 The kinetic Race of peptide vs small molecules](images/d20460154af6b02235b883843bcc2b549a1f0ba5-1000x508.png) # The Thermodynamic Bill: The "Entropy Tax" Everything in biophysics comes with a price tag. The currency is Gibbs Free Energy (ΔG = ΔH − TΔS). While peptides have advantages in water release and kinetics, they pay a massive tax that small molecules avoid. ## The Conformational Entropy Penalty A small molecule is usually rigid. It looks the same in solution as it does bound to the protein. It is "pre-organized." A peptide, however, is flexible. In solution, it is a chaotic ensemble of thousands of different shapes. To bind, it must freeze into a specific shape (like an alpha-helix). This loss of freedom is the Conformational Entropy Penalty (Δ*Sconf*). 1. **The Cost:** For a standard peptide, freezing its backbone into a helix costs a massive amount of energy (10–20 kcal/mol of unfavorable entropy). 1. **The Compensation:** To overcome this "entropy tax," the peptide must generate enormous binding energy (Enthalpy) and release huge amounts of water (Solvent Entropy). This is why natural peptides often have only micromolar affinity. They spend most of their binding energy just paying off the debt of folding. Small molecules, having pre-paid this cost during synthesis (by being rigid rings), can achieve nanomolar affinity with far fewer contacts.(Greives & Zhou, 2014) ## Entropy-Enthalpy Compensation This leads to a phenomenon known as Entropy-Enthalpy Compensation. If medicinal chemists try to make a peptide tighter and more rigid (to improve Enthalpy), they often find that the Entropic gain decreases. However, peptides exhibit a unique property called Entropy-Enthalpy Transduction. Because they are large and coupled to the protein's own movements, a strong interaction at one end of the peptide can loosen the protein structure at the other end, releasing entropy elsewhere. This allows peptides to "cheat" the compensation rule in ways rigid small molecules cannot. ![Fig. 4 The balancing of entropy and enthalpy](images/36cc31a44fa1a2c6d72c4e7b96cfc003a162862f-1000x496.png) # Structural Plasticity: The Resistance Problem One of the most compelling arguments for peptide-based therapeutics (or "peptidomimetics") is their resilience against drug resistance. ## The Brittle Small Molecule Small molecules are often described as "brittle." Because they rely on precise, lock and key complementarity, a single point mutation in the protein can destroy binding. - Case Study: Bcl-2 and Venetoclax. Venetoclax is a powerful small molecule drug that binds to the Bcl-2 protein to treat leukemia. However, patients can develop a resistance mutation, G101V. This mutation replaces a tiny Glycine residue with a bulky Valine. - The Clash: The rigid Venetoclax molecule cannot accommodate this extra bulk. It clashes with the Valine, and binding affinity drops over 100-fold, rendering the drug ineffective. ## The Plastic Peptode The native ligand for Bcl-2 is the BIM BH3 peptide. Remarkably, the BIM peptide binds to the mutant G101V protein almost as tightly as it binds to the wild type. The peptide is plastic. When it encounters the bulky Valine mutation, the peptide backbone slightly shifts or "unwinds" by a fraction of an Angstrom. It routes its side chains around the obstacle. The entropic cost of this small adjustment is negligible compared to the total binding energy. This "structural plasticity" suggests that peptides (or flexible, peptide like molecules) are more robust long-term solutions for targets prone to mutation, such as viral proteins or cancer targets.(Vogt & Cera, 2012) # The Future: Hybrids and Chimeras The binary distinction between "small molecule" and "peptide" is rapidly fading. The pharmaceutical industry is now chasing Chimeras. A molecules that attempt to capture the best of both worlds. ## Stapled Peptides To solve the "Entropy Tax" problem of peptides, chemists use Stapling. By chemically locking a peptide into an alpha-helix *before* it enters the body, they pre-pay the entropy cost. Stapled peptides (like ATSP-7041 targeting MDM2) bind tighter than native peptides because they don't lose as much entropy upon binding. They also resist degradation by enzymes and can sometimes penetrate cells. ## PROTACs (Proteolysis Targeting Chimeras) These are large, distinct molecules connected by a flexible linker. They behave like peptides in their kinetics. The flexible linker allows the molecule to search for its target using a mechanism similar to fly-casting, and they induce the formation of a ternary complex (Target-PROTAC-Ligase) that mimics a Protein-Protein Interaction. | Feature | Small Molecule | Peptide | | --- | --- | --- | | Analogy | The Lock Pick | The Velcro Handshake | | Primary Target | Deep, rigid pockets (Enzymes) | Flat, broad interfaces (PPIs) | | Water Strategy | "Sniper": Displace high-energy trapped waters | "Bulk": Massive surface water release | | Binding Entry | Diffusion & Conformational Selection | Fly-Casting & Induced Fit | | Thermodynamics | Enthalpy-driven (Pre-organized) | Entropy-driven (Solvent release vs. Folding cost) | | Resistance | Brittle (Mutation kills binding) | Plastic (Adapts to mutations) | | Key Weakness | Undruggable flat surfaces | Membrane permeability & Metabolic stability | Table: Overall comparison between small molecules and peptide # Conclusion The competition for the pocket is not just about which molecule fits best; it is about which molecule manages the chaos of biology most effectively. Small molecules remain the champions of the deep pocket being efficient, precise, and lethal to enzymes. But as we tackle the complex web of Protein-Protein Interactions, the peptide’s ability to hug the surface, release bulk water, and adapt to mutations offers a distinct advantage. The future of drug discovery likely lies in the middle ground: rigidifying peptides to lower their entropy tax (stapling) or making small molecules large enough to mimic the "handshake" of a protein (macrocycles). Understanding these biophysical rules from the "unhappy" water molecule to the "fly-casting" entry is the key to unlocking the next generation of therapeutics. # References Atanu Maity, Choudhury, A. R., Chakrabarti, R., Atanu Maity, Choudhury, A. R., & Chakrabarti, R. (2020). Effect of Stapling on the Thermodynamics of mdm2-p53 Binding. *BioRxiv (Cold Spring Harbor Laboratory)*. [https://doi.org/10.1101/2020.12.28.424518](https://doi.org/10.1101/2020.12.28.424518) DeAngelo, T. M., Utsarga Adhikary, Korshavn, K. J., Seo, H.-S., Brotzen-Smith, C. R., Camara, C. M., Sirano Dhe-Paganon, Bird, G. H., Wales, T. E., & Walensky, L. D. (2025). Structural insights into chemoresistance mutants of BCL-2 and their targeting by stapled BAD BH3 helices. *Nature Communications*, *16*(1), 8623–8623. [https://doi.org/10.1038/s41467-025-63657-y](https://doi.org/10.1038/s41467-025-63657-y) Gaikwad, V., Choudhury, A. R., & Chakrabarti, R. (2025). Microscopic Insights into the Solvation of Stapled Peptides–A Case Study of p53-MDM2. *The Journal of Physical Chemistry B*, *129*(29), 7499–7510. [https://doi.org/10.1021/acs.jpcb.5c03100](https://doi.org/10.1021/acs.jpcb.5c03100) Greives, N., & Zhou, H.-X. (2014). Both protein dynamics and ligand concentration can shift the binding mechanism between conformational selection and induced fit. *Proceedings of the National Academy of Sciences*, *111*(28), 10197–10202. [https://doi.org/10.1073/pnas.1407545111](https://doi.org/10.1073/pnas.1407545111) Overcoming the Shortcomings of Peptide-Based Therapeutics. (2022). *Future Drug Discovery*. [https://doi.org/10.4155//fdd-2022-0005](https://doi.org/10.4155//fdd-2022-0005) Vikram Gaikwad, Choudhury, A. R., & Chakrabarti, R. (2025). Microscopic Insights into the Solvation of Stapled Peptides–A Case Study of p53-MDM2. *The Journal of Physical Chemistry B*, *129*(29), 7499–7510. [https://doi.org/10.1021/acs.jpcb.5c03100](https://doi.org/10.1021/acs.jpcb.5c03100) Vogt, A. D., & Cera, E. D. (2012). Conformational Selection or Induced Fit? A Critical Appraisal of the Kinetic Mechanism. *Biochemistry*, *51*(30), 5894–5902. [https://doi.org/10.1021/bi3006913](https://doi.org/10.1021/bi3006913) Wang, J., Do, H. N., Koirala, K., & Miao, Y. (2023). Predicting Biomolecular Binding Kinetics: A Review. *Journal of Chemical Theory and Computation*, *19*(8), 2135–2148. [https://doi.org/10.1021/acs.jctc.2c01085](https://doi.org/10.1021/acs.jctc.2c01085) Wang, X., Ni, D., Liu, Y., & Lu, S. (2021). Rational Design of Peptide-Based Inhibitors Disrupting Protein-Protein Interactions. *Frontiers in Chemistry*, *9*, 682675–682675. [https://doi.org/10.3389/fchem.2021.682675](https://doi.org/10.3389/fchem.2021.682675) --- # Generative models in designing novel scaffolds URL: https://www.lite.bio/blogs/generative-models-in-designing-novel-scaffolds Author: Aditi Sinha Date: 2025-11-21 Drug discovery in past was like hunting for a rare spice in an endless pantry. Scientists would sift through mountains of existing molecules a “virtual screening” of billions of compounds hoping that one might bind to a disease target. This process has often been compared to finding a needle in a haystack. But what if, instead of rummaging through that pantry, we could *magically cook up *brand-new molecular recipes on demand? Thanks to advances in AI, that fantasy is becoming reality. Modern generative models allow researchers to design molecules from scratch, a shift from selection to creation in drug discovery. In practical terms, the old approach meant testing large libraries of known molecules (or docking them computationally) to see which ones might work. For example, traditional highthroughput screening might evaluate billions of compounds to find a handful of leads, at enormous cost and time. In contrast, virtual generation uses neural networks to propose entirely new compounds that have never been synthesized before. Instead of only picking from pre-made parts, AI lets us imagine and build novel “lego blocks” of chemistry that fit our target. This new paradigm aims to explore chemical space far beyond what any existing database holds. In fact, recent reports show generative models discovering molecules with *better* predicted binding and drug-like properties than any hit found by a brute force screen. In one case, an AI called IDOLpro used guided diffusion models to create ligands with 10–20% higher binding scores than the best hits from an exhaustive virtual screen, while also improving synthetic accessibility. And it can even include ADME/Tox (absorption, distribution, metabolism, excretion, toxicity) predictors in its scoring, optimizing multiple constraints at once. These successes hint at a future where virtual generation yields rich leads faster and smarter than ever. ![Fig.1 Uses Of AI in Drug Discovery](images/e3119339a334a9d495b1964d04284a8e3041cd49-860x392.png) ## Generative Models: VAEs, GANs and Transformers How do these AI wizards actually conjure new molecules? At the heart of virtual generation are deep generative models. Three popular classes are Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Transformer based models. Each has a different “creative engine,” but all share the goal of learning patterns from known chemicals and using them to make new ones. ![Fig.2 Various models involved for Drug discovery](images/d2eeec5f09643b67c6340dcb6d2b2c2100c22875-1000x498.png) ### Variational Autoencoders (VAEs) A VAE learns to compress and then reconstruct molecules. In practice, a molecule is often represented as a SMILES string (a text notation of its structure) or as a graph of atoms and bonds. The VAE has two parts: an *encoder* that turns the input (SMILES or graph) into a point in a continuous “latent” space, and a *decoder* that tries to map back from that latent vector to a valid molecule. By training on a dataset of existing compounds, the VAE learns a smooth map: similar molecules end up near each other in latent space. Once trained, we can sample or optimize points in this latent space and decode them to generate new SMILES strings. This lets us search for molecules with desired features. For example, Gomez-Bombarelli et al. used a VAE to map discrete SMILES into a low-dimensional continuous space, enabling efficient optimization for properties like drug likeness. In other words, VAEs let us treat molecules somewhat like malleable clay in a continuous space, rather than as isolated strings. There are also graph-based VAEs. One example is the Junction-Tree VAE (JT-VAE), which generates molecules in two phases(A N M Nafiz Abeer et al., 2024). First it builds a tree of chemical substructures (think functional groups) that outline a scaffold. Then it pieces together the atomic graph according to that tree. The result is often a valid, novel molecule. This two-stage approach helps ensure chemical validity: the model first plans a roughly “legal” structure (the tree of known fragments) and then fills in details. ![Fig.3 Working abstract of VAE models for Drug discovery](images/a7ebabf87c08da6f00b393b82f934904d532fc26-1000x546.png) ### Generative Adversarial Networks (GANs) Think of a GAN as two neural networks playing a game of cat and mouse. One network (the *generator*) creates candidate molecules, while the other (the *discriminator*) tries to tell generated molecules apart from real ones it saw in training. Over many rounds, the generator learns to produce molecules that look real enough to fool the discriminator. In effect, GANs learn the distribution of the training data without explicitly modeling it. For molecules, this can mean generating new SMILES or even molecular graphs. Some models (like MolGAN) directly generate graph adjacency matrices of small molecules. Others incorporate reinforcement learning: they add a reward signal for desirable properties. For instance, the ORGAN model is a GAN-like framework where molecules scoring well (e.g. good binding or drug-likeness) give higher reward. This reward is incorporated into the training so that the generator gradually focuses on producing high-scoring molecules(Abeer et al., 2024). In short, GANs are another way to learn to write new SMILES strings (or build graphs) by competing against a critic. ![Fig. 4 Working abstract of GAN models for Drug discovery](images/6bee2765b3f35f518bb4690230b8b647a1523bf3-1000x546.png) ### Transformer-Based Models The recent NLP revolution with Transformers has spilled into chemistry. Here, SMILES strings are treated literally as a chemical language. A Transformer model (like a molecule-version of GPT or BERT) is trained on massive databases of SMILES. By learning the “grammar” of chemistry, it can then predict missing parts of a string or generate entire new strings. Since Transformers attend to all parts of the sequence, they can capture complex dependencies (like ring closures, branching) in SMILES. One can simply ask a trained “SMILES-GPT” to continue a prefix or sample new molecules from its learned distribution. Researchers have indeed built “chemical language models” that generate valid, novel SMILES with desired patterns (Bran & Schwaller, 2023). The power of this approach is that any improvement in language modeling (like the GPT-4 scale models) could translate into better molecule generation. Transformers have already been used for retrosynthesis, property prediction, and now generation all by exploiting analogies between chemical and natural language. In short, these models read and write chemistry in the language of SMILES. (Advanced variants even treat molecules as graphs with attention mechanisms or use hybrid graph-transformer models like the TGVAE, but at core it’s “make SMILES with language AI.”) ![Fig. 5 Fundamentals for Transformer based models](images/60d58e6a622cdd3720e6478d0f7f15be40c608ad-1000x546.png) | Model | Input format | Key components | Training Objective | Strengths | Challenges | | --- | --- | --- | --- | --- | --- | | Variational Autoencoder | SMILES / Graph | Encoder, Decoder, Latent Space | Reconstruct input while regularizing latent space | Smooth latent space, useful for optimization and interpolation | May generate invalid molecules, hard to control output distribution | | Generative Adversarial Network | SMILES / Graph | Generator, Discriminator | Fool the discriminator | High-quality outputs, captures complex data distributions | Training instability, mode collapse | | Transformer-based | SMILES | Attention Layers, Token Embeddings | Learn contextual representations to predict next tokens | Handles long-range dependencies, scalable with large datasets | Requires large datasets, computationally intensive | *Table: Comparison Between various models discussed previously* Each of these generative approaches has its strengths. VAEs provide a smooth latent space useful for interpolation and systematic exploration of novel scaffolds. GANs focus on generating realistic examples and can achieve sharp, detailed designs (though at the risk of reduced diversity or mode collapse). Transformers, meanwhile, excel at learning complex sequential or graph patterns and can scale with data (especially when leveraging large pretrained models). In practice, one might compare them by metrics like validity (percentage of chemically valid molecules generated), uniqueness (how many outputs are distinct), and novelty (how different they are from training data), or by downstream performance after property optimization. Notably, many real world tools use combinations of these methods. For example, the open-source **REINVENT 4** framework uses recurrent neural networks *and* Transformers to drive molecular generation, wrapped in reinforcement learning loops (Loeffler et al., 2024). It can propose R-group substitutions, design libraries, or hop scaffolds, all through generative AI ![Fig.6 Information flow in Reinvent 4](images/9d19dc5c5ae08b3dccdd23dc5d9cf215c4b56607-1000x428.png) # Balancing Act: Multi-Objective Molecular Design Designing a molecule isn’t just about making it bind tightly to the target. A good candidate must also be *realizable* and *drug like*. This means it should be synthesizable in the lab, and have acceptable ADMET properties (like solubility and low toxicity). Thankfully, modern generative models are getting better at juggling multiple criteria at once. The key is multi-objective optimization. In practice, this means the AI isn’t just aiming for one score, but balancing several. For instance, some approaches build a single “score” that combines binding affinity prediction, synthetic accessibility, and drug-likeness. Others use Pareto-front ideas or guided retraining to emphasize trade offs. A recent example is the IDOLpro model (from the Chemical Science group). It uses a diffusion based generator whose latent variables are steered by differentiable scoring functions for multiple targets. In tests, IDOLpro created molecules with 10–20% higher *predicted* binding affinities than the best competing generative method, while also improving synthetic accessibility. In other words, it found molecules that not only fit the target pocket tightly but were also easier to make. The model even beat an exhaustive virtual screen: none of the molecules in the huge database had both the high binding and good synthesis scores that IDOLpro’s generated compounds did. List of some of the typical constraints that these AI models consider: - **Target Affinity:** Models often include a docking or physics-based score (like AutoDock Vina or a ML binding predictor) so that generated molecules should snugly fit the protein pocket. - **Synthetic Accessibility:** Neural networks can incorporate scores that estimate how hard it is to synthesize the molecule (e.g. SA score or retrosynthesis-derived metrics). Generators are then biased toward structures with higher feasibility. - **ADMET Properties:** Predictors for absorption, solubility, metabolic stability, toxicity, etc., can be used as additional scoring functions. By optimizing for these, the model avoids crazy chemical features that would make a drug impossible. By tweaking the training or the scoring during generation, AI can propose molecules that hit all these checkpoints. For example, some tools let you “plug in” any differentiable function (including ADME/Tox predictors) as an optimization target(Kadan et al., 2025). The end result is a kind of controlled creativity: the model is no longer just freewheeling random chemist, but a chemist with a rulebook. This ability to satisfy multiple constraints is what makes generative design so exciting to pharma: one can, in principle, design a compound that not only binds its target strongly but also *already* looks like a viable drug candidate in terms of chemistry and safety. ## Diffusion Models: A New Direction for Pocket-Aware Molecule Generation While VAEs, GANs, and Transformers have pushed molecular generation forward, **diffusion models** are quickly becoming serious contenders, especially for structure-based drug design. Inspired by image generation tools like DALL·E and Stable Diffusion, these models work in reverse: they start with random noise and slowly “denoise” it to form meaningful outputs. In chemistry, that output can be a valid, 3D molecular structure. In ligand design, diffusion models are particularly powerful because they can **build molecules atom-by-atom or fragment-by-fragment inside the 3D context of a protein pocket**. That makes them extremely pocket-aware, which is a major advantage over models that generate SMILES strings independently of structural context. **Pocket2Mol** is one of the first 3D diffusion based frameworks built for generating ligands directly within a protein binding pocket. It takes as input a 3D representation of the pocket (derived from a protein-ligand complex or structure prediction) and uses a two-stage generation process: 1. **Molecular topology prediction** – the model first generates a graph structure that defines which atoms are connected and how. 1. **3D coordinate generation** – then it produces the 3D spatial arrangement of the atoms so that the final molecule fits snugly into the pocket. Pocket2Mol was trained on real protein-ligand complexes from the PDBbind dataset, learning how ligands typically conform to their pockets. A key strength is its ability to produce ligands that are *physically plausible* right out of the model, requiring less post-generation docking or clean-up. It also tends to propose **diverse scaffolds**, not just variations on known templates, because it learns directly from structure rather than relying on SMILES data. Pocket2Mol's innovation is in treating molecular generation as a **geometrically conditioned graph construction problem**—which is ideal for tasks like fragment growing, scaffold hopping, or structure-based screening when the binding site is known.(Yu et al., 2022) ![Fig.7 The Generation Procedure of Pocket2Mol. For each panel, the left part is protein and the right part is the sampled molecular fragment.](images/2e601b456995cbc9991d3213dba7e87fc0b8a5ce-680x444.png) ## PocketXMol **PocketXMol** builds on the idea of pocket-conditioned generation using a more advanced **3D denoising diffusion probabilistic model (DDPM)**. Where Pocket2Mol uses a two-stage system, PocketXMol operates in a fully unified framework: - It takes a 3D binding pocket as input (represented as a voxel grid or mesh) - It incorporates **spatial, electrostatic, and residue-type features** into its conditioning - Then, through a denoising diffusion process, it gradually constructs ligand atoms and bonds in 3D space One of the standout features of PocketXMol is its **multi-step ligand optimization **the model doesn’t just generate one shot; it can iteratively refine molecules during generation to improve fit, binding potential, and drug-likeness. It also supports **multi-objective optimization** natively, factoring in synthetic accessibility and ADMET predictions during generation (Orvieto et al., 2023). Compared to traditional pipelines, PocketXMol eliminates the need for post-generation docking. Since it builds ligands directly within the 3D pocket during training, its outputs tend to be highly accurate in both shape complementarity and interaction hotspots. ![Fig.8 Schematic representation of the sampling and Training process of PocketXMol.](images/7fa1cdb95c75177ca839015e5f81a01fa90597bd-701x551.png) # LiteFold unifies all of it for Rapid Ligand Generation Denovo Drug Design using LiteFold Platform These ideas aren’t just academic. Several tools and platforms are bringing generative design into the hands of chemists and biologists. LiteFold’s “DeNovo” module lets users specify a protein pocket (from experiment or a predicted structure) and then churns out hundreds of candidate ligands in minutes. Under the hood, LiteFold uses diffusion based neural network models to predict molecules that fit the pocket. Users can then visualize each new molecule docked in 3D, tweak them, and see metrics like drug-likeness or synthetic accessibility in real time. In short, LiteFold automates and accelerates the structure-based design cycle. ![Fig.9 Infromational flow of the denovo Module for generative ligand design through Litefold](images/a979c79ae0d70df3fef3878ee736c42cfb5b7a17-1000x433.png) LiteFold’s design engine tightly couples multiple generative models in one workflow (see the provided system-flow diagram). In practice, the platform first encodes the target protein pocket and then simultaneously leverages diffusion models to propose new ligands The model learns a continuous latent embedding of molecules (and protein induced conformations) to enable smooth sampling and interpolation of chemical space. In effect it “polices” the diffusion model outputs, nudging generated candidates toward realistic, valid chemistries. In this way the three architectures complement each other. We have explained this in more details in our blogpost here: # Conclusion The shift from screening to generation is reshaping how we think about drug discovery. Rather than passively searching existing libraries, scientists can now actively explore uncharted chemical terrain. This AI-powered creativity is, in effect, turning the drug lab into a playground of molecules. As with any emerging technology, there are challenges: ensuring generated molecules are truly novel (not just close copies of training data), validating AI predictions experimentally, and integrating human intuition with machine imagination. Yet the progress is undeniable. Modern generative models can already propose *hundreds* of candidate structures that are both potent and practical. For pharmaceutical researchers, this means a new toolkit: instead of brute-forcing screens, teams can seed the process with a target structure and ask AI to dream up leads. For scientifically curious readers, it’s an exciting proof that concepts from language and vision (like Transformers and diffusion) are finding home in chemistry. In the coming years, we may see even more synergy. for example, combining generative chemistry models with AI-predicted protein structures or real-time bioassay data. The key message is that AI no longer just reads the periodic table, it’s learning to write it. In short, drug discovery is becoming more like engineering than guesswork. With VAEs, GANs, and Transformers as our new molecular apprentices, we have the potential to design tailor-made therapies at a scale and speed previously impossible. We’ve gone from scouring the shelves of known compounds to playing in a generative sandbox of possibilities. The future drug is no longer just waiting to be found it can now be imagined, built, and tested by machines (and people) together. # References A N M Nafiz Abeer, Urban, N. M., Weil, M. R., Alexander, F. J., & Yoon, B.-J. (2024). Multi-objective latent space optimization of generative molecular design models. *Patterns*, *5*(10), 101042–101042. [https://doi.org/10.1016/j.patter.2024.101042](https://doi.org/10.1016/j.patter.2024.101042) Abeer, A. N. M. N., Urban, N. M., Weil, M. R., Alexander, F. J., & Yoon, B.-J. (2024). Multi-objective latent space optimization of generative molecular design models. *Patterns*, *5*(10), 101042. [https://doi.org/10.1016/j.patter.2024.101042](https://doi.org/10.1016/j.patter.2024.101042) Bran, A. M., & Schwaller, P. (2023). *Transformers and Large Language Models for Chemistry and Drug Discovery*. [ArXiv.org](http://ArXiv.org). [https://arxiv.org/abs/2310.06083#:~:text=,data%2C like spectra from analytical](https://arxiv.org/abs/2310.06083#:~:text=,data%2C%20like%20spectra%20from%20analytical) Kadan, A., Ryczko, K., Lloyd, E., Roitberg, A., & Yamazaki, T. (2025). Guided multi-objective generative AI to enhance structure-based drug design. *Chemical Science*, *16*(29), 13196–13210. [https://doi.org/10.1039/d5sc01778e](https://doi.org/10.1039/d5sc01778e) Loeffler, H. H., He, J., Alessandro Tibo, Janet, J. P., Alexey Voronov, Mervin, L. H., & Engkvist, O. (2024a). Reinvent 4: Modern AI–driven generative molecule design. *Journal of Cheminformatics*, *16*(1), 20–20. [https://doi.org/10.1186/s13321-024-00812-5](https://doi.org/10.1186/s13321-024-00812-5) --- # When Physics meet AI URL: https://www.lite.bio/blogs/when-physics-meet-ai Author: Aditi Sinha Date: 2025-11-14 Docking scores offer a quick first estimate, but often miss the underlying thermodynamics that drive real binding. Chemists have been chasing the dream of predicting binding affinity with accuracy for years, yet the gap between computer scores and lab results keeps showing up like that one stubborn error bar that won’t go away. The problem? Traditional docking is like judging a relationship by the first handshake. It captures the pose, but not the full thermodynamics that determine real binding affinity. Real binding depends on many hidden players: water molecules, entropy, conformational changes, and sometimes even long-range electrostatics. All that complexity makes accurate prediction tough when using only simplified scoring functions. That’s where physics and AI come together, one grounded in first principles, the other driven by pattern and inference. Physics brings the rules, thermodynamics, molecular motion, water effects. while AI brings intuition, pattern recognition, and a bit of chaos magic. Together, they’re turning guesswork into something close to real chemistry. Techniques like **Free Energy Perturbation (FEP)** and **Absolute Binding Free Energy (AB-FEP)** have long been the “serious calculators” of the bunch, but they were slow and complex. Now, with **machine learning–boosted scoring** and **generative models** stepping in, we’re speeding up the process without losing the science. Machine learning can help predict which molecular poses are worth simulating, approximate energy landscapes faster, and even learn corrections from past free energy errors. Instead of running millions of calculations, AI can help the system “guess smarter." And as generative models start suggesting molecules that already “fit” a binding pocket, these free energy methods become the filter that decides which ones truly make sense in the physical world. Together, they’re building a feedback loop AI creates, physics validates, and both learn. Now that we know the story; now it’s time to open the lab notebook and see what’s really going on inside these models. ![Fig.1 Molecular dating app: the overconfident ligand](images/070f788c2af78e61a0c163dc26b567f8c4a940aa-853x869.png) # Why Quick Scores Are Fails In the early stages of drug discovery, we rely heavily on molecular docking. Think of docking as a nightclub bouncer: it quickly screens a massive line of potential compounds (virtual screening) and gives a lightning-fast "yes/no" based on whether the molecule physically fits the binding pocket. It's superb for speed and initial structural insights into the *mode* of interaction between the ligand and the receptor. The catch is that the bouncer only looks at the size and shape, not the chemistry. Docking relies on highly simplified, empirical scoring functions designed for throughput, not thermodynamic rigor. These functions intentionally adopt a series of simplifications to reduce complexity, which means they cannot reliably account for the subtle, yet critical, energy changes that occur when a molecule binds, such as changes in solvation (how water molecules rearrange), or entropic effects (how the molecules' flexibility changes). The resulting predicted score is therefore a highly unreliable stand in for the molecule's true value that is the quantitative binding free energy, or ΔG . This fundamental inability to accurately predict the binding strength means drug pipelines are often littered with "false positives" molecules that look good on paper but fail in the lab. Addressing this scoring flaw is currently the most critical priority for industrial drug design, as it limits the ultimate success of virtual screening. ## The Mirror Twin Test The moment you challenge docking with a subtle problem, its facade crumbles. Take **enantiomers**: mirror-image molecules that are chemically identical but might fit a protein's chiral environment in wildly different ways. Predicting which twin binds better requires capturing every nuance of molecular movement, solvent interaction, and subtle electrical forces—the very things simplified scoring skips over. When docking software (like Glide, considering its extra precision (XP), standard precision (SP), and high-throughput virtual screening (HTVS) modes, as well as AutoDock Vina) was put through the "Mirror Twin Test" against 141 pairs of enantiomers with biological activities reported across seven targets, it repeatedly failed. (Rawal & Braga, 2024) The conclusion? If the software correctly identified the stronger binder, it was purely due to chance. This isn't a simple bug; it's confirmation that the underlying mathematical *model* cannot capture true thermodynamic reality. When you need precision, you must turn to a model built on physics. ![Figure 2: Generated by AI](images/b6ba20ce736fe29e3a405ae7bf8ec29ea727be40-1000x1272.png) # Free Energy's High Cost of Precision The industry's answer to the prediction crisis is Free Energy Perturbation (FEP). FEP is the forensic accountant of molecular biology, recognized as the gold standard for high-accuracy binding affinity prediction, especially when refining a promising drug candidate (hit-to-lead and lead optimization phases). FEP derives the true ΔG from the laws of statistical mechanics by simulating an "alchemical transformation", a theoretical pathway where one ligand is smoothly mutated into another (Relative FEP) or slowly vanishes from the protein pocket (Absolute FEP, or AFEP). This process is theoretically sound and precise. ![Fig. 3 Schematic Overview of Active Learning Enhanced Sampling Techniques for Library Screening with FEP](images/73f68cfde42260ddd145e56b795f0f23eea4fb13-787x594.png) ## The Triple Barrier to Entry: Why FEP Stayed Exclusive In theory, FEP sounds perfect. In practice, it was the molecular equivalent of a VIP computing club. The FEP formula can be applied without calculating intermediate states (the "end-state-only" approach) only if there is a very high degree of phase-space overlap between the initial (λ=0) and final (λ=1) states. Since this condition is rarely met in complex molecular binding, FEP must painstakingly sample dozens of intermediate states (λ windows) to ensure the full energetic path is correctly mapped. (Braga & Rawal, 2025). This intensive requirement created a "Triple Barrier": 1. The Time Sink: Sampling intermediate states demands vast computational power. Even small changes in a molecule require substantial resources, historically limiting FEP to small, exclusive groups of highly congeneric molecules that are chemically very similar. 1. The Expert Tax: FEP calculations are highly advanced and require deep, specialized computational expertise for setup, execution, and thorough analysis. You couldn't just hand it to an intern; this requirement acted as a significant barrier to widespread adoption. 1. The New Discovery Bottleneck (AFEP): Absolute FEP (AFEP) has the theoretical power to find entirely new, *de novo* hits because it doesn't require a structural reference molecule. However, its inherently slow throughput historically made it impractical for real-world screening campaigns. ![Fig. 4 machine learning (ML), especially active learning (AL) and deep learning (DL), can enhance the efficiency, accessibility, accuracy, and precision of FEP workflows.](images/c66cfc1f716ee508c447988ec32f438b7b1e78a6-499x500.png) The core issue was simple: FEP’s power was crippled by inefficient sampling, the computer was effectively wandering aimlessly through the molecular landscape. The mission for AI became clear: Stop the wandering. Provide the GPS. Democratize the gold standard, making it fast and accessible for everyone in drug design. # AI-Driven Sampling Acceleration Machine learning now provides the algorithmic intelligence to tackle FEP's high computational cost head-on, deploying "smart sampling" to accelerate the exploration of the Free Energy Surface (FES) and the underlying molecular kinetics. ## ML for Optimal Path Finding Traditional molecular dynamics (MD) simulations take too long waiting for a molecule to spontaneously find the complex, low-energy path to its binding site. AI fixes this by defining superior Collective Variables (CVs). These are custom, machine learned coordinates that simplify the high dimensional binding event into a single, optimized path, forcing the simulation to find the "express lane." The COMet-Path hybrid method exemplifies this, combining an enhanced sampling technique (funnel metadynamics) with a machine-learned optimal association CV. The initial step requires optimizing the coefficients of this pathlike variable, which must be performed on at least one converged free energy landscape, typically obtained with another set of variables. Once optimized, this approach ensures the free energy profile converges significantly faster, particularly for complex ligands in exposed and large binding cavities. This powerful combination delivers a potent blend of accuracy and speed for calculating Absolute Binding Free Energies (ABFE), successfully overcoming the major speed limitation of traditional AFEP methods with minimal additional computational overhead. COMet-Path can provide a satisfactorily accurate estimation of the ABFE for noncongeneric ligands, increasing the success rate compared to using funnel metadynamics alone. (Crivelli-Decker et al., 2024) ![Fig.5 Double-decoupling alchemical protocol.](images/857882ba73ef06027b630d89f148beae76f8be8e-500x319.png) ## Learning the Point of No Return: Transition State Mapping For methods that trace the physical path of binding, such as Transition Path Sampling (TPS) and Steered Molecular Dynamics (MD), knowing the precise "Point of No Return" the transition state, or committor function is crucial for efficiency. ML algorithms are trained to learn this function directly. The Automated Iterative Molecular Dynamics (AIMMD) method, for instance, iteratively combines TPS with a neural network that estimates the committor function. This intelligence then guides intelligent "shooting" from the transition state to accelerate the exploration of the binding mechanism. Variational approaches are also proposed that minimize the total squared displacement over equilibrium trajectories, bypassing the need to estimate committor values explicitly. A major bonus here is that scientists gain not just the final ΔG number, but rich details on the *physical mechanism* itself, including binding intermediates, transient pockets, and molecular mechanisms, critical insights that the non-physical, alchemical paths of standard FEP mask entirely. ## The AI Art Generator for Molecules: Phase Space Exploration Generative models, such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs) and Diffusion models, are deployed to enhance molecular structure sampling and implicit density estimation. These deep-learning methods represent a considerable effort focused on enhanced samplings for extracting the free energy surface and kinetics. A VAE-based network can learn simplified, low-dimensional maps (embeddings) of time lagged protein conformations, helping to reveal the slow dynamics of stochastic protein motions crucial for understanding induced fit binding and accurate affinity prediction. Specifically, Wasserstein GANs minimize the difference between the true distribution of molecular structures and the synthetic ones generated. This process, known as implicit density estimation, ensures that the generated samples are highly representative of valid molecular structures and helps avoid issues common in other generative models, such as the *posterior collapse problem* often seen in VAEs. (Donald & Willem Jespers, 2025) This capability guarantees efficiency by ensuring computational cycles aren't wasted exploring invalid or non-representative conformations, thereby accelerating the exploration of phase space. Furthermore, the capability to perform efficient *conditional generation* of molecules, estimating the distribution of a structure given certain desired chemical properties is significantly enhanced through these generative models. Classical molecular dynamics simulations rely on simplified energy rules (force fields). While fast, these rules are often parameterized based on limited training data and sometimes fail to capture the subtle, critical electronic, polarization, and non-covalent forces crucial for high-precision binding affinity. This is where AI delivers a crucial upgrade in physical fidelity. ## Neural Network Potentials Neural Network Potentials (NNPs) are sophisticated ML models trained on the "master blueprint" of high-level Quantum Mechanical (QM) calculations. They are trained on extensive datasets derived from these QM calculations and learn to represent the complex potential energy surface, achieving near-QM accuracy but running at a fraction of the cost compared to performing full QM dynamics. Integrating these NNPs into FEP directly elevates the quality of the underlying physics, leading to higher fidelity ΔG predictions. These enhancements mark clear progress toward more practical ML-augmented force fields for FEP workflows. ![Fig.6 Schematic representation of Molecular dynamics simulation using NNP](images/e6d2bfb80fbfb7494fe8a25f990fe8a8b3a21c01-992x508.png) ## Trajectory Tipping: Solving the Speed vs. Accuracy Trade-off Running an entire simulation entirely with a highly accurate NNP is still significantly more computationally expensive compared to using a cheap, traditional classical force field. This threatens to reintroduce the old dilemma: accuracy versus throughput. The brilliant solution is the reweighting strategy, a clever operational hack developed by researchers like [Tkaczyk et al](https://www.semanticscholar.org/paper/Reweighting-from-Molecular-Mechanics-Force-Fields-Tkaczyk-Karwounopoulos/149f0e354cd11d6865d048988faac37680c26d61). Trajectories are first run cheaply using the simple molecular mechanics (MM) force field. Then, the expensive, accurate NNP (such as ANI-2x) is used to *reweight* those results, applying a high-precision correction after the fact. This strategy retains the accuracy benefits associated with the neural network potential without committing to the full computational grind for the entire simulation. a powerful and practical compromise that makes high-accuracy, ML-augmented force fields viable for high-throughput FEP workflows. (Zeng et al., 2022) # The Hybrid AI-FEP Pipeline The combination of physics-based rigor and AI-driven efficiency has operationalized FEP, transforming it from a bespoke lab curiosity into a scalable, industrial-grade predictive platform that actively guides lead optimization in pharmaceutical research. ## Active Learning for Intelligent Screening FEP, even accelerated, is pricier than docking. Active Learning (AL) algorithms act as the shrewd stock picker, guiding molecule selection to dramatically reduce the total number of necessary FEP calculations during virtual screening. The process is self-correcting, AL algorithms are trained and informed using small, high-quality virtual activity data sets generated by FEP itself. This ability to generate highly valuable virtual activity data from FEP to train AL algorithms is a crucial factor in overcoming the typical data paucity challenge faced by ML in drug discovery. This establishes a virtuous cycle where FEP provides the ultimate ground truth, which in turn trains the AL to prioritize only the most promising candidates for the next expensive FEP calculation, effectively replacing exhaustive searching with intelligent, iterative optimization. ![Fig.7 The ultimate synergy](images/0590a9cdac83458b19c2b4c4e52ee85d192791d1-500x163.png) Physics Augments AI. Machine Learning models often suffer from poor performance due to Limited SAR data. By harnessing Free Energy Perturbation (FEP) to create highly reliable, synthetic data points, we overcome this bottleneck. This 'Data Augmentation' transforms a failing model (red curve) into a highly accurate predictor (blue curve), proving that physics is the key to unlocking AI's full potential in drug design ## Eliminating the Expert Barrier Historically, preparing the initial protein-ligand complex structure, the "assembly required" phase was manual, error prone, and demanded deep expertise. This complexity was a major roadblock to FEP adoption. Deep Learning (DL) methods have now completely automated this crucial step. DL based co-folding tools, such as AlphaFold variants, NeuralPLexer, and DragonFold, automatically generate accurate complex structures, which are essential starting points for FEP. This is revolutionary because it bypasses the need for traditional docking or manual complex preparation steps, which drastically lowers the operational complexity and the reliance on specialized experts. This key automation step is crucial, making FEP a robust, accessible tool for high throughput industrial pipelines. ## The Report Card: Validating Predictive Power This hybrid AI-FEP system has moved past academic papers and into the boardroom, proven by hard, quantitative metrics that dictate drug development decisions. Industrial case studies, like those from [Bayer Pharma AG](https://pubmed.ncbi.nlm.nih.gov/40958764/),(Donald & Willem Jespers, 2025) confirm that modern free energy calculations have achieved the necessary speed and usability to effectively drive real drug discovery projects, running hundreds of compounds across more than ten different targets (including both retrospective and prospective cases). Quantitatively, the system delivers a retrospective analysis of 147 calculations showed a gold-standard Mean Unsigned Error (MUE) 0.94 kcal/mol and a high correlation R-value of 0.77. This is a critical milestone, as an MUE below 1.0 kcal/mol is the consensus threshold for a genuinely useful, predictive tool in drug discovery. Furthermore, when FEP was applied prospectively to optimize a lead compound series, it successfully guided optimization efforts that led to a tenfold improvement in activity for lead molecules like imidazotriazine I, without increasing atom count or overall size. These strong quantitative metrics confirm FEP’s graduation to a trusted, reliable decision-making technology for pharmaceutical optimization.(van Pinxteren & Jespers, 2025) | Metrics | Molecular Docking (VS) | Traditional FEP/MD | AI-Augmented FEP Hybrid | | --- | --- | --- | --- | | Primary Output | Docking Score (Simplified ΔG) | Relative/Absolute Binding Free Energy ΔG | Highly Accurate ΔG (Accelerated) | | Physical Rigor | Low (Empirical/Knowledge-Based) | High (Statistical Mechanics) | Highest (ML-NNP Enhanced) | | Accuracy (MUE) | Poor (Prediction due to chance) | High (approx 1-2 kcal/mol) | Predictive (approx 0.94 kcal/mol) | | Computational Cost | Very Low (High Throughput) | Extremely High (Limited Throughput) | Medium-High (Accelerated via ML/AL) | | Key AI/ML Role | Minimal | Enhanced Sampling (CVs, Committor) | Force Fields (NNP), Protocol Optimization (AL) | | Typical Use Case | Hit Identification (Large Screens) | Lead Optimization (Small, Congeneric Sets) | Hit-to-Lead and Lead Optimization (Broad Scope) | Table. 1 Comparative Metrics: Docking vs. Traditional FEP vs. Hybrid AI-FEP The successful strategy emphasizes a hybrid approach that integrates human expertise with sophisticated ML tools, accelerating and democratizing FEP-based drug discovery.8 AI serves not to replace the fundamental chemical and physical models, but to optimize and enable them. | FEP Challenge Addressed | AI/ML Technique Deployed | Mechanism of Enhancement | | --- | --- | --- | | Slow Convergence/Sampling | Machine-Learned Collective Variables (COMet-Path) | Defines optimal, low-dimensional reaction coordinate to speed up free energy surface exploration. | | Transition State Identification | Committor Function NN/TPS (AIMMD) | Learns the probability of reaching the bound state, promoting smarter sampling from transition regions | | Classical Force Field Inaccuracy | Neural Network Potentials (NNPs) | Trained on QM data to provide higher physical fidelity for energy calculations. | | NNP High Computational Cost | Reweighting Strategy (ANI-2x) | Applies the accurate NNP corrections to trajectories generated by cheaper classical force fields. | | Setup Complexity (Initial Pose) | DL Co-folding (AlphaFold variants) | Automates the generation of accurate protein-ligand complex starting structures. | | Operational Cost/Efficiency | Active Learning (AL) | Guides molecule selection, minimizing the total number of expensive FEP calculations needed. | Table 2: AI’s Role in Enhancing FEP Workflows # The Next Horizon: Quantum Computing Meets Free Energy Looking ahead, the ultimate evolution involves integrating **quantum computing** into this AI-physics framework. Quantum computers hold the potential to tackle the deepest, most intractable problems in quantum chemistry those related to complex electronic interactions and large-scale reaction dynamics that even the most powerful classical supercomputers cannot solve. Current research is focused on developing **qubit-efficient variational quantum algorithms (VQAs)** that are resource aware on current quantum hardware. This involves minimizing qubit requirements while maintaining quantum advantages, particularly when solving high-dimensional, combinatorial problems like optimizing coordination in multi-agent systems, a task analogous to optimizing a molecular system. The future remains strictly hybrid: smartly delegating tasks between classical computation and quantum accelerators to solve the immense, high-dimensional, combinatorial optimization problems inherent in molecular systems. This strategic delegation is crucial because it leverages the advantage of quantum acceleration without overburdening the limited qubit capacity of current devices. Mid-term goals include optimizing quantum circuits specifically for agent interactions and designing more efficient integration strategies that balance quantum and classical computation to expand scalability. Ultimately, the long-term vision involves using fully fault-tolerant quantum computers (FTQCs) to enable entirely quantum driven agent interactions, decision making, and optimization, finally surpassing all known classical computational limits in molecular simulation and enabling breakthroughs in drug design that are currently impossible. # Conclusion The convergence of statistical physics (FEP) and Artificial Intelligence is nothing short of a paradigm shift, a true computational revolution. The quick and dirty scores of simplified docking have been definitively eclipsed by a highly predictive platform that operates with near-quantum accuracy and industrial scale throughput. By providing the molecular GPS for sampling, automating the painful setup phase, simplifying the system through coarse graining, and elevating force fields to quantum fidelity, AI has solved FEP’s three major historical roadblocks. The result is a validated predictive engine with an error rate below the crucial 1.0 kcal/mol threshold. This hybrid approach has not just accelerated the drug discovery timeline; it has profoundly raised the bar for computational rigor in the pharmaceutical industry, ushering in an era where speed and precision no longer have to be mutually exclusive. The future of drug discovery is not just faster or more accurate, it is fundamentally more intelligent, ensuring that speed and precision are, at long last, inseparable allies in the quest for human health. # References Cournia, Z., Allen, B., & Sherman, W. (2017). Relative Binding Free Energy Calculations in Drug Discovery: Recent Advances and Practical Considerations. *Journal of Chemical Information and Modeling*, *57*(12), 2911–2937. [https://doi.org/10.1021/acs.jcim.7b00564](https://doi.org/10.1021/acs.jcim.7b00564) ‌ Hong, S. H., Ryu, S., Lim, J., & Kim, W. Y. (2019). Molecular Generative Model Based on an Adversarially Regularized Autoencoder. *Journal of Chemical Information and Modeling*, *60*(1), 29–36. [https://doi.org/10.1021/acs.jcim.9b00694](https://doi.org/10.1021/acs.jcim.9b00694) ‌ Burger, P. B., Hu, X., Balabin, I., Muller, M., Stanley, M., Joubert, F., & Kaiser, T. M. (2024). FEP Augmentation as a Means to Solve Data Paucity Problems for Machine Learning in Chemical Biology. *Journal of Chemical Information and Modeling*, *64*(9), 3812–3825. [https://doi.org/10.1021/acs.jcim.4c00071](https://doi.org/10.1021/acs.jcim.4c00071) ‌ York, D. M. (2023). Modern Alchemical Free Energy Methods for Drug Discovery Explained. *ACS Physical Chemistry Au*, *3*(6), 478–491. [https://doi.org/10.1021/acsphyschemau.3c00033](https://doi.org/10.1021/acsphyschemau.3c00033) ‌ van Pinxteren, D. J. M., & Jespers, W. (2025). Integrating Machine Learning into Free Energy Perturbation Workflows. *Journal of Chemical Information and Modeling*, *65*(19), 9856–9864. [https://doi.org/10.1021/acs.jcim.5c01449](https://doi.org/10.1021/acs.jcim.5c01449) ‌ Crivelli-Decker, J. E., Beckwith, Z., Tom, G., Le, L., Khuttan, S., Salomon-Ferrer, R., Beall, J., Gómez-Bombarelli, R., & Bortolato, A. (2024). Machine Learning Guided AQFEP: A Fast and Efficient Absolute Free Energy Perturbation Solution for Virtual Screening. *Journal of Chemical Theory and Computation*. [https://doi.org/10.1021/acs.jctc.4c00399](https://doi.org/10.1021/acs.jctc.4c00399) ‌ Braga, D. M., & Rawal, B. (2025). Harnessing AI and Quantum Computing for Revolutionizing Drug Discovery and Approval Processes: Case Example for Collagen Toxicity. *JMIR Bioinformatics and Biotechnology*, *6*, e69800–e69800. [https://doi.org/10.2196/69800](https://doi.org/10.2196/69800) ‌ Ramírez, D., & Caballero, J. (2016). Is It Reliable to Use Common Molecular Docking Methods for Comparing the Binding Affinities of Enantiomer Pairs for Their Protein Target? *International Journal of Molecular Sciences*, *17*(4), 525–525. [https://doi.org/10.3390/ijms17040525](https://doi.org/10.3390/ijms17040525) ‌ Matsumura, N., Yoshimoto, Y., Yamazaki, T., Amano, T., Noda, T., Ebata, N., Kasano, T., & Sakai, Y. (2025). Generator of Neural Network Potential for Molecular Dynamics: Constructing Robust and Accurate Potentials with Active Learning for Nanosecond-Scale Simulations. *Journal of Chemical Theory and Computation*, *21*(8), 3832–3846. [https://doi.org/10.1021/acs.jctc.4c01613](https://doi.org/10.1021/acs.jctc.4c01613) --- # The overlooked role of intrinsic water in protein–ligand binding URL: https://www.lite.bio/blogs/intrinsic-water Author: Aditi Sinha Date: 2025-10-19 Water is essential for life, that’s something we all know. But beyond keeping us alive, water also shapes how molecules recognize, interact, and bind to each other. In the world of proteins and ligands, it’s far more than just a background solvent. Inside a protein’s binding pocket, water molecules don’t simply float around they take part in the action. They help drugs attach correctly, stabilize the structure, and even influence how strong or weak the binding will be. In modern drug design, scientists now realize that success isn’t just about pushing water out of the binding site. It’s about understanding how each tiny water molecule behaves, which ones to keep, which to replace, and how they quietly hold everything together. So, how exactly does water manage this backstage role of keeping protein–ligand complexes stable, flexible, and functional? Let’s look closer at the science behind this unsung hero. ![Fig.1 Water in the Pocket: Friend, foe, or just a mediator in drug binding?](images/9ce8598b9c3d3226f3fb2c46d5245f848b8f134c-1000x1000.png) # Proteins Live in a Watery World Inside every cell, proteins exist in a watery soup. But this is no ordinary soup, it’s one where each water molecule has a job to do. Some waters move around freely like audience members milling about, while others are fixed in place like security guards near the stage. These fixed ones are vital, they keep the protein’s shape intact, help residues stay where they should, and maintain that delicate 3D fold that makes everything work. Without water, proteins wouldn’t just act up they’d lose their shape completely and collapse into a floppy, useless structure. ![Fig 2: The red dots are water. See how the amino acid is interacting with it.](images/107a46964ca405c8ec8513ccd01cf044fd8d1805-1000x614.png) Some water molecules inside a protein are permanent residents, not just passing guests. These are called structural waters, and they act like tiny screws and bolts that keep everything together. They help amino acids hold their positions and stabilize the protein’s inner architecture. When a ligand binds, some of these water molecules may leave, but others stay behind to keep the structure balanced. It’s a bit like rearranging your room, you move a few things, but the walls better stay in place, or the whole setup collapses. ## The Water Bridge Effect Here’s where water gets creative. Not every amino acid residue can directly connect to the ligand, sometimes the geometry just doesn’t work out. That’s where bridging waters come into play. These molecules form hydrogen bond networks between the ligand and nearby residues, acting like microscopic matchmakers. Instead of forcing a direct link, water steps in, holds both hands, and says, “Relax, I got this.” This bridging effect adds flexibility and makes the complex more resistant to breaking apart. Even if one bond weakens, the water bridge can help maintain the communication between protein and ligand. Think of it as adding an extra Wi-Fi repeater, the signal stays strong even if the distance increases. A single water bridge can sometimes stabilize a structure that would otherwise fall apart, especially when the ligand is bulky or irregular. In some enzyme systems, these bridging waters are so important that removing them ruins the entire binding process. ![Fig 3: Water bridge effect between 2 distant proteins](images/7ed36325785c05497ee2a2e94c4daa1d96bd1cbf-1000x667.png) # Water and Binding Energy: The Balancing Game When a ligand binds to a protein, some water molecules are displaced. At first glance, it might seem that removing water always strengthens binding, but the reality is more nuanced. The energy released by displacing water must balance the energy gained from the new protein-ligand interactions. Structural studies, especially crystallography, often show water molecules conserved across multiple protein structures, highlighting their crucial roles in stability and function. Binding itself is governed by the binding free energy (ΔG = ΔH − TΔS), where ΔH represents the enthalpic contribution (energy change) and TΔS represents the entropic contribution (disorder). Removing highly organized, “unhappy” water molecules those that are energetically constrained can increase entropy and lower ΔG, thereby favoring binding. Conversely, displacing “happy,” well-connected water molecules can weaken binding if the ligand cannot compensate with stronger direct interactions or new hydrogen bonds. In the realm of molecular docking and simulation studies, water can literally make or break the binding energy, acting as either a stabilizing ally or a destabilizing obstacle. ![Fig.4. multifaceted role of water in protein-ligand binding](images/cf94b7ac4b5a0630af9650664f65d6596b792552-1000x855.png) When a water molecule forms favorable hydrogen bonds, it strengthens the interaction, but if it is trapped in an energetically unfavorable position, it can weaken binding instead. This delicate balance is reminiscent of arranging a team: having the right players in the right positions makes all the difference. Hydrophobic (water-repelling) and hydrophilic (water-loving) interactions must also stay balanced. When a ligand enters a binding pocket, hydrophobic regions tend to push water away, while hydrophilic regions attract it in. This tug-of-war between forces helps determine the final binding affinity, a measure of how tightly a ligand binds to the protein. Interestingly, many computational studies have shown that explicitly including a few key water molecules gives much more accurate results than removing them entirely. So, when modeling protein–ligand systems, it’s wise not to discard all water molecules; some of them serve as the MVPs of stability, influencing both the energy landscape and the geometry of binding. | Action | Energy Effect (ΔG = ΔH − TΔS) | Impact on Binding Strength | Drug Design Strategy | | --- | --- | --- | --- | | Pushing out Unhappy Water | Large increase in Disorder (+TΔS). Energy cost (ΔH) is paid off. | Significant Strength Increase | Design the drug to take the exact spot of highly organized, unstable water. | | Keeping Water (Bridging) | Favorable drop in Energy (-ΔH) due to strong connections. Disorder (TΔS) is penalized (water is stuck) . | Enhanced Precision and Stability | Design the drug to use the structured water molecule as a connecting point (Retention Strategy). | | Pushing out Happy Water | Small change in Disorder (low gain). Energy cost (ΔH) may be unfavorable. | Strength Penalty (Weaker Binding) | Avoid pushing out water that is highly stable and well-connected, unless the drug forms a much stronger direct connection. | *Table.1 Summary of thermodynamic effects of water displacement on ligand binding and drug design strategies.* Water molecules are dynamic participants in the binding site. They often slip in and out, forming temporary hydrogen bonds and adapting to the protein’s subtle movements. This process resembles a well-choreographed dance, where water moves in rhythm with the protein’s flexing structure, stabilizing newly formed conformations. Some water molecules are so consistently present that they are treated as semi-permanent structural components, while others act like fleeting extras, appearing briefly but performing critical roles in folding, orientation, and proper ligand positioning. Beyond structural stabilization, water helps proteins recover from slight unfolding or shape fluctuations. By cushioning the interior, it allows atoms to “breathe” naturally rather than sticking unnaturally, preventing local collapse and maintaining the overall protein architecture. Moreover, water can actively participate in chemical reactions. In enzymes, individual water molecules may act as reactants, breaking chemical bonds or assisting in proton transfer, while in receptor-ligand interactions, water can guide the ligand into its proper orientation before binding occurs. Classic examples include **HIV protease** and **carbonic anhydrase**, where well-positioned water molecules are critical for both structural integrity and enzymatic function. Removing these waters can dramatically reduce efficiency and alter the protein’s behavior. In essence, water is not merely a passive background solvent in the biochemical theater; it is an active member of the cast, influencing binding energies, reaction pathways, folding dynamics, and the delicate choreography of molecular interactions. By understanding the subtle roles of water, researchers can gain deeper insights into protein-ligand recognition, drug design, and enzymatic mechanisms, appreciating that each water molecule whether fleeting or permanent can be essential to the molecular story. ![Fig.5 Hydration sites identified using the water density distribution, and displayed as purple spheres. in HIV protease complex](images/7050b89f7940292836a931e152a286e760d80e57-1000x707.png) ## Water's Outer Ring and Binding Speed Water's influence is not limited to just the surface where the protein and drug touch (the first hydration shell). Research increasingly confirms that the second ring of water molecules can also be critical for how strongly a drug binds. For computer simulations (like MM/PBSA) to accurately predict binding, they must fully consider the effects of this second ring of water to match experimental data reliably. Moreover, water affects not only the final binding energy ΔG but also how fast the drug attaches (kinetics). The necessary rearrangement of the water network during binding is a crucial step that affects the energy barrier of the reaction. Detailed modeling that included corrections for water free energy helped researchers understand why the drug xk263 binds 1,000 times faster to HIV-1 protease (HIVp) than the drug ritonavir. This shows that the structure of the water around the binding site greatly influences the speed of binding, making water essential for both strong and fast drugs. ![Fig.6 First hydration shells](images/ef45185ae68ad51124503670656ef22e9bcc4e2e-926x780.png) # Structural Synergy: Water-Mediated Bridges Water bridges act as flexible connectors when direct protein-ligand bonds aren’t feasible. Research shows that over 85% of protein-ligand complexes in the Protein Data Bank (PDB) contain one or more bridging waters, averaging 3–4 per complex. These waters stabilize complexes, enhance binding precision, and sometimes dictate target specificity. ## Real-World Examples The stabilizing power of water bridges is well-known in many drug targets: 1. **CDK2 Inhibitors:** Studies on CDK2 complexes show that a water-mediated hydrogen bond between the drug and the backbone of residue Glu81 is critical for holding the drug in place . Statistical surveys show that water-connected hydrogen bonds involving an oxygen atom on the drug are about twice as common as those involving a nitrogen atom. 1. **Thrombin Inhibitors:** In one striking example, adding a simple hydrogen-donating ammonium group to a potent thrombin inhibitor led to a massive increase in binding strength (over 500-fold) . X-ray analysis showed this huge boost was because the ammonium group formed a **charge-assisted hydrogen bond** with the protein and surrounding water, proving how kept water can amplify favorable charge interactions . These examples confirm that water bridges are not just random parts of the structure, but essential pieces that dictate the precise fit and strength needed for a drug to bind strongly. ![Fig.7. Chemical structure of a series inhibitors of human factor Xa in the X-ray cocrystal structure of human factor Xa ](images/4c1e0fec93b277668ea7740fd78edd63ad240954-661x583.png) | Target System | Water's Role in Complex | Key Structural/Energy Finding | | --- | --- | --- | | Factor Xa (FXa) Inhibitor | Pushing out highly structured water. | Removing thermodynamically costly water significantly increased binding strength (ΔG = -1.92 kcal/mol) | | CDK2 Inhibitors | Water-connected hydrogen bonds (N-H···O). | Water acts as a bridge between the drug and key protein residues (like Glu81 backbone), locking the drug's shape | | Thrombin Inhibitors | Charge-assisted water bridge. | Adding an ammonium group led to a >500-fold strength increase via water-mediated, charge-assisted connections . | | Bosutinib/Kinase Targets | Using a fixed water network. | The drug recognizes its target by engaging a pair of conserved water molecules; the protein's "gatekeeper" residue controls access to them. | | HIV-1 Protease | Strongly fixed water molecule. | The calculation must account for the large energy cost of keeping this water molecule highly organized . | *Table.2 Examples of how water molecules influence ligand binding across different drug target systems.* # Hydration Hot Spots: Mapping Critical Water Sites Hydration hot spots are tiny regions within protein binding pockets where water molecules have an outsized impact on binding energy. Mapping these spots helps chemists make smart choices: should a water molecule be **pushed out** to increase disorder (and potentially strengthen binding), or **kept** to maintain structural stability? Understanding hydration hot spots gives drug designers two main strategies: 1. **The Push-Out Strategy:** This involves carefully designing a part of the drug molecule to sit exactly where an unstable, highly organized water molecule is located. pushing out these highly restricted molecules leads to a substantial gain in disorder (entropy), which lowers ΔG and makes the drug bind stronger. Drugs designed to expel this high-density, unhappy water have been shown to provide the greatest increase in binding strength. 1. **The Keep Strategy:** If a hot spot corresponds to a water molecule that is very strongly connected and contributes favorably to the system's energy (e.g., highly structured water that forms multiple hydrogen bonds), the best approach is to design the drug to keep this organized water molecule and use it as part of the overall connection network. This captures the structural stability provided by the water bridge. Water molecules in binding pockets aren’t always happy campers. The inner layer of water is often more crowded than the surrounding solvent, creating a physical barrier that any incoming drug must overcome. When a ligand binds, it may **displace** some of these waters. While it might seem that removing water always strengthens binding, the reality is subtler. Some waters contribute essential hydrogen bonds, and kicking them out can reduce stability. Advanced docking algorithms now often predict which waters are “happy” to leave and which prefer to stay. Crystallography frequently shows structural waters tucked neatly inside binding pockets their repeated presence across multiple protein structures underscores their importance. MD and Monte Carlo simulations reveal how water dynamically rearranges during binding, forming temporary bridges, stabilizing charged regions, and adjusting networks to accommodate the ligand. Some of these key waters are so predictable that modern drug-design software includes them automatically during virtual screening. Ignoring them can lead to false predictions or missed opportunities for effective binding. ![fig.8 Molecular mechanism of water in ligand binding](images/6cc50fc0c87762bc5ac1c4e01b0c36f690b21580-1000x667.png) ## Crowded Water at the Binding Site The physical nature of the water shell places specific limits on drug binding. Looking closely at the protein-water surface shows that water molecules in the inner layer, even near nonpolar parts, are, on average, more crowded than water in the main solvent body. This crowding, which is about 6% denser for globular proteins, means that putting a drug into the binding site requires overcoming a significant barrier related to the water molecules being physically squeezed. This squeezing challenge contributes directly to the energy penalty during computational calculations. Therefore, to get the maximum benefit from pushing out organized water, designers must create drugs that minimize the necessary squeezing penalty while achieving the largest possible gain in disorder. # The Computational Challenge: Modeling Water Accurately ## The Limits of Simple Models Historically, computer-aided drug design used quick, simple scoring methods, often relying on implicit solvent models (simple water models). These models are fast but cannot accurately capture the specific, local effects of individual, organized water molecules or properly account for the subtle disorder changes that define the hot spots. Because these simple models fail to capture the molecular detail of the water environment, the estimated binding energies are often inaccurate. ## The Problem of Slow Water Movement and Hybrid Solutions While detailed MD simulations are necessary, they have a major limitation: slow water movement. When parts of the binding site are deeply buried or hard for the main solvent to reach, the time needed for water molecules to move in or out to reach a stable balance (equilibration) can be much longer than the time the simulation can practically run. This slow sampling of water movement is particularly problematic when calculating binding energies. For example, if a drug is modified to be smaller, the pocket might accommodate an extra water molecule. If the simulation doesn't run long enough to see that water move in, the calculated energy will be wrong. To fix this, hybrid Monte Carlo/Molecular Dynamics (MC/MD) methods have been developed . MC/MD speeds up the process of water reaching balance between the bulk solvent and buried pockets by including random moves (Monte Carlo) that allow water to easily enter or leave confined spots in a way that is still thermodynamically correct . Using MC/MD dramatically improves the accuracy of calculated results and reduces the difference between calculating the change forward and backward, confirming that getting the water balance right is essential for accurately calculating the final binding energy (Δ*G*). ## Sensitivity to Water Models The specific explicit water model chosen (like TIP3P or TIP5P) is another critical factor affecting the accuracy of calculated energies. Different popular models show variations in physical properties, such as the calculated water density and how much steric squeezing occurs. For instance, the steric compression penalty, measured by the van der Waals decoupling free energy, was found to be noticeably higher for TIP3P and TIP4P models compared to TIP5P. Furthermore, simulations of complexes, like the CD44–HA complex, showed that the TIP5P water model gave the lowest structural difference (RMSD) from the experimental starting structure compared to other models . This variation means the choice of water model is important it is a source of thermodynamic uncertainty and requires researchers to choose models that accurately reflect the known physical properties of the solvent . # Moving forward from here The strong evidence showing water's dual role as both a driver of disorder (entropy) through the strategic removal of unhappy molecules and as an energy stabilizer (enthalpy) through water-mediated connections makes its complete inclusion mandatory for modern drug development . Overcoming the technical difficulties in accurately modeling this complex system is quickly turning water from a frustrating variable into a powerful tool for fine-tuning how drugs interact with their targets. Computer methods continue to advance to include water more explicitly and intelligently. Recent approaches, such as GraphWater-Net, use network structures to map protein atoms, drug atoms, and the complex network of water molecules and their interactions . By using these networks to extract interaction details, this model significantly improves the prediction of drug-protein binding strength, outperforming previous advanced methods by a margin of 0.022 to 0.129 in correlation . To be truly successful, therapeutic development needs to fully consider water effects not just in the final binding strength (Δ*G*), but also in the speed of binding (kinetics). As computing power increases and hybrid simulation techniques become standard, researchers are gaining a clearer atomic-level view of the water environment. This allows for the smart design of drugs that maximize the thermodynamic benefits of water rearrangement while minimizing physical and entropic costs. This complete approach ensures that water, the most abundant molecule in biological systems, is finally recognized and utilized as a decisive partner in molecular recognition. ![Fig.9. Four methods explored in this work for dealing with water molecules in the binding site.](images/03aa79f0ad41fc2f5f0ad6fa5b9d689315579a3e-1000x410.png) # Conclusion For drug developers, water is both a friend and a challenge. If you include too much of it in a model, the system gets noisy and slow. If you ignore it, your predictions may go completely off track. A well-designed ligand often fits into a binding site in a way that either makes smart use of bridging waters or replaces them efficiently. Some drugs even *depend* on water bridges for optimal binding. Pharmaceutical chemists now pay close attention to these “hydration sites” specific points where water molecules consistently help in maintaining stability. Understanding which waters to keep and which to replace can make the difference between a weak binder and a powerful therapeutic. You could say that water is the **silent support staff** that keeps the stars the protein and the ligand looking good on stage. Without it, the structure falls flat, and the whole biochemical show loses balance. So the next time you see a molecular model and think, “Those water molecules are just background,” remember some of them might be the real heroes holding the story together. # **References** - Michel, J., Tirado-Rives, J., & Jorgensen, W. L. (2009). *Prediction of the water content in protein binding sites.* **Chemical Reviews**, 109(9), 3509–3529. - Huggins, D. J., et al. (2011). *Role of water in biological recognition.* **Journal of Medicinal Chemistry**, 54(21), 7105–7123. - Young, T., Abel, R., Kim, B., Berne, B. J., & Friesner, R. A. (2007). *Motifs for molecular recognition exploiting hydrophobic enclosure in protein–ligand binding.* **PNAS**, 104(3), 808–813. - Ball, P. (2008). *Water as an active constituent in cell biology.* **Chemical Reviews**, 108(1), 74–108. - Ross, G. A., Bodnarchuk, M. S., & Essex, J. W. (2015). *Water sites, networks, and free energies with grand canonical Monte Carlo.* **JACS**, 137(47), 14930–14943. - Aldeghi, M., Bluck, J. P., & Biggin, P. C. (2022). *Computational mapping of hydration sites in proteins.* **WIREs Computational Molecular Science**, 12(1), e1525. - Abel, R., Young, T., Farid, R., Berne, B. J., & Friesner, R. A. (2008). *Role of the active-site solvent in the thermodynamics of factor Xa ligand binding.* **Chemical Reviews**, 108(9), 3716–3756. - Poornima, C. S., & Dean, P. M. (1995). *Hydration in drug design. 1. Multiple hydrogen-bonding features of water molecules in mediating protein–ligand interactions.* **Journal of Computer-Aided Molecular Design**, 9(6), 500–512. --- # Fail Fast, Fail Cheap: In-Silico Toxicology Pipelines for Early Drug Candidate URL: https://www.lite.bio/blogs/tox Author: Nabajit Borah Date: 2025-09-24 Drug discovery is basically a casino where the house almost always wins. Around more than 90% of drug projects never make it to patients, and the price of failure climbs steeply the further along you go. Flop early in discovery? That’s about a million dollars down the drain. Flop in late-stage clinical trials? That’s a jaw-dropping $2.6 billion burned. One of the biggest culprits behind these late-stage wipe outs is **toxicity, drugs that look promising but turn out to be harmful once tested in humans.** Toxicity accounts for nearly a third of candidate terminations. For a well-funded pharma giant, that’s painful. For a startup? It’s lethal. ![Figure 1: Drug Toxicity and its affects](images/464c0f3bd57f6723f47512187866083f53e64b63-1000x844.png) # Why All Toxicities Aren’t the Same “Toxicity” sounds like a single villain, but in drug development it shows up wearing different masks. If you want to catch it early, you need to know which version you’re fighting. Broadly, toxic effects fall into three buckets: - **On-target toxicity** – the drug is hitting exactly what it was designed to, but too hard. The biology is right, the dose is wrong. Blood pressure drugs that drop pressure so low you faint are a classic example. It’s the pharmacological equivalent of enjoying loud music until the speakers blow. - **Off-target toxicity** – here the drug starts flirting with proteins it was never meant to bind. Kinase inhibitors are notorious: one was designed for a single kinase, but ended up interacting with dozens. Sometimes this broad activity helps (blocking cancer escape routes), other times it’s deadly (triggering arrhythmias through hERG channel binding). - **Chemical-based toxicity** – some compounds don’t even need proteins to wreak havoc. They’re chemically unstable or reactive, forming toxic metabolites or generating free radicals. Acetaminophen overdose frying the liver is a textbook case. Understanding which bucket a toxic effect falls into isn’t trivia, it shapes how you predict it, test it, and engineer around it. ![Figure 2: Three major type of drug toxicity](images/364501fecd86ffbeb4606296460bb7b0135b9441-1000x605.png) # The Death of the Magic Bullet For most of the last century, drug discovery worshipped the **magic bullet** idea: a single, exquisitely selective molecule that hits one target and leaves everything else alone. Clean, precise, safe. Reality check: molecules are promiscuous. On average, a small-molecule drug binds between **six and eleven different proteins**. This “polypharmacology” used to be considered sloppy chemistry. Now it’s seen as the natural state of drug action—and sometimes, a feature. **Which is why the mindset has shifted from snipers to shotguns**. The new ambition is a **magic shotgun**: a drug rationally designed to hit multiple targets in a coordinated way. In diseases like cancer, Alzheimer’s, or metabolic syndrome, nudging several nodes in a pathway often works better than hammering one in isolation. And from a practical angle, a single multi-target drug can also reduce pill burden, simplify pharmacokinetics, and cut down drug–drug interactions compared to combination therapy. This is where computational toxicology earns its keep. Predictive models spit out results like *“Molecule X binds Target Y”.* On their own, those predictions are meaningless. The value comes from how you interpret them. - If **Target Y** is the hERG potassium channel (the classic culprit behind drug-induced arrhythmias), then binding is a **red flag:** a signal your compound could be cardiotoxic. - If **Target Y** is a receptor in an inflammatory pathway linked to another disease, the *exact same prediction* is a **green light:** an unexpected chance to repurpose the compound for a new indication. That’s the trick: the computational output is neutral. It’s the **biological context,** what that protein does in health and disease, that flips it from liability to opportunity. ![Figure 3: The old vs The New ways to drug discovery](images/f21590b2d7c3009dc288174b026b64a193aab97a-1000x430.png) This is why drug repurposing has shifted from serendipity to strategy. Aspirin started as a painkiller and became a cornerstone of cardiovascular prevention. Today, instead of waiting for lucky accidents, we use predictive models to systematically chart a molecule’s binding “constellation,” then sort the stars into two camps: hazards that demand caution, and potential guides toward new therapeutic directions. # A tour of computational predictive methods Broadly, there are two methods, both perspectives are essential, and the best pipelines use them together. ## **Ligand-Based Methods***: look at the molecule itself (the “key”)* These methods work under a simple but powerful rule of thumb: *similar molecules tend to act in similar ways*. If you don’t have the 3D structure of the protein, you let the drug’s structure do the talking. ### **A: Quantitative Structure–Activity Relationship (QSAR)** The process involves translating the 2D or 3D chemical structure of a molecule into a set of numerical features. You take features of the molecule, like molecular weight, lipophilicity (*logP*), or how many hydrogen bonds it can form and convert them into numbers called *descriptors.* A statistical model, such as a regression or classification algorithm, is then trained on a dataset of molecules with known activities (e.g., toxic or non-toxic) to find a mathematical equation that links the descriptors to the activity. The resulting model, in its simplest form, looks like this Activity=f(descriptors)+error where ***f*** can be any regression or classification function, such as a linear model, nonlinear regression, or machine learning model and **error** represents the difference between the observed experimental biological activity and the activity predicted by the model Feed in enough data about known molecules (toxic or safe), and the model learns patterns. Once trained, it can spit out predictions for new compounds. Like predicting house prices, the quality of the answer depends entirely on the quality of your dataset. ### **B: Pharmacophore Modeling** While QSAR considers the molecule as a whole, pharmacophore modeling focuses on identifying the essential 3D arrangement of features required for a drug to interact with a target or the “essential features” that matter for binding: hydrogen bond donors/acceptors, hydrophobic spots, aromatic rings, charges. It’s a kind of 3D blueprint of what’s needed for activity. You can then use this blueprint as a search query, scanning vast compound libraries for molecules that *fit the pattern, *even if they look nothing like the original. ## Structure-Based Methods: *look at the protein it might fit into (the “lock”).* If you’ve got a high-resolution protein structure (from X-ray crystallography or cryo-EM), you can get a lot more specific. ### A: Molecular Docking As explained earlier in[ my blog on docking](__GHOST_URL__/molecular-docking-in-drug-discovery/), the software tries to “fit” a drug molecule into the binding pocket of the protein. It tests millions of poses, scoring each one to estimate how well it sticks. The result is not just a binding score but a visual snapshot of how the drug might interact at the atomic level. ### B: Reverse Docking ("Target Fishing") Flip the game: instead of many drugs against one protein, test one drug against thousands of proteins. The goal is to spot all the possible targets,both intended and unintended. This is how you can flag potential off-target toxicities or repurpose existing drugs. One striking example: Researchers used reverse docking to show that PFAS, environmental contaminants, bind strongly to the folic acid receptor, which is crucial for brain development. This binding disrupts folate uptake, explaining the link between PFAS exposure and neurodevelopmental problems. Table 1: Summarizes computational predictive methods | Method Category | Core Principle | What You Need (Data) | Example Techniques | Pros | Cons | | --- | --- | --- | --- | --- | --- | | Ligand-Based | Similar molecules have similar activities. | A dataset of molecules with known activity data. | QSAR, Pharmacophore Modeling, 2D/3D Similarity | Doesn't require a 3D target structure. Fast for large libraries. | Less mechanistic insight. Can struggle with novel chemical scaffolds. | | Structure-Based | A drug's activity depends on its 3D fit to a target. | A 3D structure of the target protein(s) | Molecular Docking, Reverse Docking | Provides detailed mechanistic insight into binding. | Requires a known 3D target structure. Computationally intensive | No single method rules them all. Instead, computational scientists build workflows that: - Use **ligand-based methods** when speed is needed or when protein structures are missing. - Use **structure-based methods** when detailed mechanistic insight is critical. - Integrate both into **layered pipelines**, cross-validating predictions for reliability. But here’s the non-negotiable truth: your model is only as good as your input data. QSAR collapses if the training dataset is biased or messy. Docking results are meaningless if the protein structure is poor quality. A huge part of the field isn’t the glamorous modeling, but the grunt work of collecting, cleaning, and curating solid data. Garbage in, garbage out, it’s the universal law of computational toxicology. # The AI Revolution in Toxicology QSAR and docking have been around since your professors were grad students, but in the last decade, toxicology got rocket fuel: artificial intelligence (AI). What used to be linear regressions and rule-based “if this, then that” systems is now a playground of neural nets, graph models, and enough acronyms to make you more confuse. At its core, AI changed toxicology in two big ways: ## Old School: Expert Systems Before AI hype, software relied on wisdom from chemists into codified ruled based systems. These tools run on curated libraries of structural alerts (a.k.a. toxicophores): chemical substructures repeatedly associated with toxic outcomes. Some examples includes: - **[Derek Nexus (Lhasa Ltd.)](https://optibrium.com/products/stardrop/modules/derek-nexus/)**– Probably the most widely used. It flags toxicity risks across endpoints like mutagenicity, carcinogenicity, skin sensitization, and more. It’s trusted in pharma pipelines and even considered by regulators. - **[Toxtree (Open source, JRC)](https://toxtree.sourceforge.net/)** – A free alternative that encodes decision trees of toxicophores. Great for academic use or early-stage projects. - **[HazardExpert / OncoLogic](https://compudrug.com/hazardexpertpro)** – Older U.S. EPA and commercial tools that applied similar alert logic, especially around carcinogenicity. Now the question is, what counts as a “structural alert”? - **Nitroaromatic groups** – associated with mutagenicity (due to metabolic activation into reactive intermediates). - **Epoxides** – reactive three-membered rings that can covalently bind DNA or proteins, leading to genotoxicity. - **Anilines** – linked to methemoglobinemia (blood toxicity). - **Michael acceptors** (α,β-unsaturated carbonyls) – electrophilic hotspots prone to react with nucleophiles in proteins → toxicity. - **Aromatic amines** – long-known red flags for carcinogenic risk. These systems flag and provide reasoning: *“Compound contains an aromatic amine substructure. Literature links similar compounds to hepatotoxicity due to reactive metabolite formation.”* That transparency is their superpower. Table 2: Strengths vs Weaknesses of ruled-based systems | Aspect | Strengths | Weaknesses | | --- | --- | --- | | Coverage | Spans multiple toxicity endpoints with minimal input requirements | Over-pessimistic predictions leading to false positives | | Speed | Rapid analysis using only SMILES strings | Endpoint bias (missing complex toxicities like hERG cardiotoxicity) | | Validation | Well-established usage in regulatory settings | Difficulty handling novel chemical structures | | Integration | Easy to incorporate into existing screening workflows | Limited ability to consider complex biological contexts | | Interpretability | Every alert has a documented rationale with traceable evidence | Static knowledge base limited to known mechanisms | ## New School: Machine Learning Systems If expert systems are *rulebooks written by chemists*, machine learning (ML) systems are *data-hungry interns that figure out the rules on their own*. They rely on models that chew through **huge datasets of molecules with measured properties,** toxicity endpoints, ADME (Absorption, Distribution, Metabolism, Excretion) profiles, etc. and learn the patterns automatically. Popular Platforms and Tools: - **[ADMET Predictor (Simulations Plus)](https://www.simulations-plus.com/software/admetpredictor/)****:** One of the longest-standing commercial platforms. It bundles models for solubility, metabolism, toxicity, and pharmacokinetics. Widely used in pharma pipelines. - **[ADMET-AI (recent, open-access)](https://admet.ai.greenstonebio.com/)****:** Web platform that runs ML/DL models for a wide range of ADMET endpoints. Think of it as a “plug-and-play” toxicology predictor. - **[DeepTox (NIH Tox21 Challenge winner, 2014)](https://www.frontiersin.org/journals/environmental-science/articles/10.3389/fenvs.2015.00080/full)****:** A pioneering deep learning framework trained on >10k compounds across 12 toxicity endpoints. It showed that neural networks could outperform traditional QSAR methods. - **[ProTox-III](https://tox.charite.de/protox3/)****:** A webserver trained on massive datasets, predicts acute toxicity (LD50), hepatotoxicity, carcinogenicity, and more, with probability scores. - **[ChemProp / ChemBERTa](https://arxiv.org/abs/2010.09885)****:** Academic ML toolkits for molecular property prediction, often used in research settings. These aren’t toxicology-specific but have been adapted to endpoints like hERG inhibition or mutagenicity. ![Figure 4: Old vs New ways in Drug Toxicology](images/80ed714ee883794264148d4b72fed18f18a639ac-947x750.png) ## The ML Evolution: From Classic Models to Deep Learning Let's take a quick journey through the evolution of machine learning in toxicology. ### First Wave: Classic ML Workhorses Before the deep learning hype train arrived, toxicologists were already getting impressive results with these tried-and-true approaches: ![Image](images/f5f204cbb8d0ab9cd9266d9e0817e4e7b7c4c944-1000x667.png) - **Random Forests (RF):** By combining multiple decision trees, they've proven remarkably effective for QSAR-type predictions while gracefully handling noisy datasets. - **Support Vector Machines (SVMs):** Excel at binary classification (toxic vs. non-toxic). These were the go-to tools for early hERG cardiotoxicity screening, drawing clean boundaries between "safe" and "risky" compounds. - **Artificial Neural Networks (ANNs):** The precursors to today's deep learning revolution. Even these simpler architectures showed impressive flexibility, though they demanded substantial training data. What all these models had in common: they relied heavily on **hand-crafted molecular descriptors,** carefully calculated properties like logP (lipophilicity), polar surface area, or molecular fingerprints such as ECFP4 that capture structural patterns. ### Second Wave: The Deep Learning Revolution The game-changer? Modern deep learning architectures that bypass those human-designed descriptors entirely, learning directly from the raw molecular structure: ![Image](images/0a8acbcd459560579d40741a4954da57a70031c3-1000x1000.png) - **Graph Neural Networks (GNNs):** These treat molecules as they truly are graphs where atoms are nodes and bonds are edges. What makes GNNs revolutionary is their ability to discover complex structural patterns without being explicitly programmed. They can recognize dangerous motifs like "electron-rich aromatic ring adjacent to a nitro group" without a toxicologist spelling it out. They've become workhorses for predicting mutagenicity, cardiotoxicity, and liver damage. GNNs mimic how information flows through a molecule, with atoms "talking" to their neighbors through bonds. This captures the subtle electronic and structural dependencies that drive toxicity. - **Transformer-based Models:** The same technology powering modern chatbots has been adapted for molecules. Models like ChemBERTa and MolBERT treat chemical structures (represented as SMILES strings) like sentences, learning the "grammar" of toxic compounds. These have achieved remarkable performance across the ADMET spectrum. Table 3: Strengths vs Weaknesses of ML/DL based methods | Aspect | Strengths | Weaknesses | | --- | --- | --- | | Interpretability | Advanced techniques like attention maps and SHAP values provide some insight | Often function as "black boxes" with limited explanation of predictions | | Coverage | Can handle diverse chemical spaces and novel structural patterns | Heavily dependent on training data distribution and quality | | Speed | Fast inference once trained; can process thousands of compounds quickly | Initial training can be computationally expensive and time-consuming | | Validation | Often outperforms traditional methods in benchmark datasets | Regulatory acceptance still evolving; concerns about reproducibility | | Integration | Modern APIs and containerization make deployment flexible | May require specialized infrastructure or expertise to implement | | Data Requirements | Can extract patterns from complex, heterogeneous datasets | Performance degrades significantly with insufficient training examples | | Adaptability | Can be retrained or fine-tuned as new data becomes available | May struggle with domain shift (e.g., novel chemical classes) | ## The Full ADMET Gauntlet Predicting off-target toxicity is crucial, but it’s only one piece of the survival game a drug has to play. To make it from bench to bedside, a molecule has to run the full **ADMET gauntlet**: - **Absorption** – Can the drug actually get into the body (say, across the gut lining)? - **Distribution** – Once inside, does it reach the right tissues at the right concentration? - **Metabolism** – Does the liver shred it to pieces before it has a chance to work? - **Excretion** – How efficiently does the body get rid of it, and through which routes? - **Toxicity** – Does it cause collateral damage along the way? You can have a perfectly targeted drug, but if it has poor solubility the project fails. This is why modern computational platforms predict the complete ADMET profile simultaneously, offering a comprehensive assessment of "drug-likeness." Multi-task learning models use a single molecular representation to predict multiple properties at once, rather than creating separate models for each characteristic. This approach recognizes the interconnected nature of these properties: - Lipophilicity influences both absorption and metabolic stability - Plasma protein binding affects distribution and clearance - Metabolic processes directly impact toxicity profiles By training models to predict multiple properties simultaneously, we capture biological interdependencies that create **more generalizable models** representing how molecules behave as complete systems. This represents an evolution from isolated property analysis to understanding drugs as dynamic entities in biological systems. ![Figure 5: ADMET Prediction](images/5511585aaa1e7ebbfb48791ca791840a8d85187c-1000x703.png) # Today's Toxicology Prediction The real magic happens when these computational methods join forces in integrated screening workflows. Rather than betting on a single algorithm, modern pipelines create a consensus safety profile by combining multiple prediction strategies. Take the Off-Target Safety Assessment (OTSA) framework - it's basically computational toxicology's version of a Swiss Army knife. It merges 2D similarity searches, QSAR models, and 3D pocket analysis into one automated system. Its secret weapon? It doesn't just analyze the drug itself but also its *metabolites*. Smart move, since those breakdown products often cause the actual toxicity. Here's a quick rundown of the tools toxicologists are using right now: | Tool | Approach | Cost | What It Does Best | | --- | --- | --- | --- | | ADMET Predictor | AI/ML + QSAR | Commercial | All-in-one package with comprehensive risk scoring | | Derek Nexus | Rule-Based | Commercial | Spots toxic fragments and explainswhythey're dangerous | | ProTox-II | ML + Fragment-based | Free (Web) | Predicts organ-specific toxicity and mechanistic pathways | | pkCSM | Graph-based ML | Free (Web) | Models complete ADMET profiles using molecular signatures | ### Industry‑standard tools at a glance - **TEST (EPA Toxicity Estimation Software Tool):**Ensemble QSAR. Combines multiple statistical models built on curated datasets and molecular descriptors, then uses a consensus prediction.Strengths: transparent descriptor-based rationale and EPA familiarity. Best for early hazard screening and benchmarking when you want quick, explainable calls. - **OECD QSAR Toolbox:**Read‑across and chemical category formation. Uses profilers and mechanistic knowledge to group similar substances, infer properties, and document applicability domains and mechanistic hypotheses.Strengths: regulatory‑accepted workflows and rich provenance. Best for dossiers, justification of predictions, and mechanistic reasoning. - **Leadscope**Knowledge‑based structural alerts plus statistical modeling over large, curated toxicology corpora.Strengths: broad endpoint coverage and documentation supporting regulatory review. Best for hazard identification and safety assessment with traceable evidence. ### When to use what (quick guide) - Need explainable, regulator‑friendly rationale fast → Derek Nexus, TEST, Leadscope, OECD QSAR Toolbox. - Need broad ADMET coverage across endpoints at once → ADMET Predictor, pkCSM, ProTox‑II. - Need mechanistic or target‑specific insight → Docking or reverse docking, then cross‑check with rule‑based alerts and ML. - Preparing a dossier or justification package → OECD QSAR Toolbox and Leadscope for provenance plus consensus with ML outputs. ## Beyond Molecules: The Systems View Finding that your compound binds to an off-target protein is just the beginning. The million-dollar question is: "Will this actually hurt someone?" Enter **systems toxicology** - where we model how a molecular hiccup cascades through biological networks. This approach connects the dots between a single binding event and clinical adverse reactions by tracing the ripple effects through genes, proteins, and pathways. Real-world example: You predict your drug binds to a certain protein. Great, but is that protein even expressed in liver tissue where toxicity might occur? Systems approaches layer in this crucial biological context. Despite all this progress, some tough challenges remain: - **The Black Box:** AI is powerful but often can't explain its predictions. Regulators (rightfully) demand to know *why* a model flags something as toxic. - **Garbage In, Garbage Out:** Even the fanciest AI can't overcome flawed training data. The field desperately needs standardized, high-quality datasets. - **Regulatory Hurdles:** Getting the FDA to accept computer predictions as primary evidence remains an uphill battle requiring rock-solid validation and transparency. The future is clear: we're headed toward integrated platforms that combine *in silico* predictions with lab data in AI-powered frameworks that speak the language regulators understand. # Conclusion Drug development is a high-stakes game where toxicity can sink even the most promising candidates. The computational tools we've discussed aren't just nice-to-have - they're essential for designing safer drugs more efficiently. We've evolved from the naive "magic bullet" idea to understanding that drugs interact with multiple targets throughout the body. Our computational arsenal now lets us predict these complex interaction profiles before spending a dime on synthesis. While challenges in data quality and model interpretability persist, the trajectory is promising. By catching problems early through computational prediction, we're steering drug discovery toward molecules with better odds of success - ultimately delivering safer medicines to the patients who need them. --- # Molecular Simulations: The Fun Way to Predict Binding Affinity URL: https://www.lite.bio/blogs/mds Author: Nabajit Borah Date: 2025-09-15 Picture a thriller where the hero tracks down the villain using perfect surveillance footage, kicks down the warehouse door, and finds nothing. The target vanished hours ago. This is exactly what happened when our docking algorithm ranked millions of "perfect" binders in an afternoon. The wet lab results? Most were phantom leads, impressive on paper but in reality, most of them were a dud firecracker. Just like static surveillance can't capture real movement, molecular docking gives us beautiful snapshots that miss the dynamic reality of molecular life. This disconnect between computational prediction and experimental truth is the central challenge of structure-based drug discovery. Docking, as I explained in my [previous article on molecular docking](__GHOST_URL__/molecular-docking-in-drug-discovery/), is an efficient starting point. But it’s still a snapshot: proteins are treated as mostly rigid, water is simplified, and entropy is largely ignored. ## Beyond Static Snapshots Proteins are not frozen sculptures; they flex and shift in solution, revealing hidden pockets or collapsing apparent ones. Docking can’t capture this motion or the entropy cost of locking a flexible ligand into place, which often explains why “perfect” binders fail experimentally. Molecular dynamics (MD) simulations step in here, running proteins and ligands in motion over nanoseconds to microseconds, letting us see if a pose is truly stable. Coupled with methods like free energy perturbation (FEP), MD brings us closer to realistic binding affinities than docking ever could. ![Figure 1: The gap between computational prediction and experimental reality. (AI Generated)](images/2f7caa0fde74835bd3157a536839674b5db5c0a0-1000x1000.png) These limitations of docking are not just theoretical nitpicks. Large benchmarking efforts, like the D3R Grand Challenges [1], repeatedly show how docking struggles in practice. Even when a ligand’s pose looks right with less than 2 Å deviation from the crystal structure, the scoring functions often misjudge the actual binding strength. Flexible receptors, solvent effects, and entropy penalties are simplified or skipped altogether, so docking scores drift away from experimental affinities. The result: plausible poses but poor predictions. ## Simulating Life in Silico Docking is a snapshot; **molecular dynamics is the movie**. In MD, the computer doesn’t just line up shapes and call it a match. It actually solves **Newton’s equations of motion for every atom,** protein, ligand, and even the surrounding water. The result is a frame-by-frame movie of how the system behaves in time. Side chains wiggle, water molecules drift in and out, and the binding pocket, sometimes opening wider, sometimes shut immediately. Instead of a rigid lock-and-key guess, MD shows the interaction as it really is: alive, flexible, and constantly shifting. This is powerful for two reasons: 1. **Refinement:** It tells you if a docking pose is stable or falls apart once the protein undergoes conformational change. 1. **Thermodynamics:** From the trajectory, you can estimate **binding free energy** , essentially, how favorable it is for the ligand to stay bound versus floating away. These calculations are extremely tedious often needing long simulations (hundreds of nanoseconds, sometimes microseconds) to see meaningful motions, especially for flexible targets thus using a lot of computational recourses. ![Figure 2: Simulate atom movements. Track shape changes. Measure interactions [2\]](images/e9ce67b662a658fd2e18da6e987e771408eb7fae-1000x481.png) ## Force Fields and why it matters so much If molecular dynamics is a movie, then the **force field is the script.** Atoms don’t know physics on their own, computers need a set of rules that say how strongly bonds stretch, how angles bend, how charges attract and how van der Waals forces push and pull. To simply put, a force field is set of equations and parameters that translate Newton’s laws into numbers the simulation can run with. Different MD software packages implement these rules in slightly different ways. Some focus on biomolecular accuracy, others on speed, GPU performance, or flexibility for custom setups. Table 1 (below): Some of the most widely used MD engines handle force fields and related features: | MD Engine | Best For | Force Field Specialties | Performance Considerations | Accessibility | | --- | --- | --- | --- | --- | | AMBER | Proteins, nucleic acids, small molecules | AMBER (ff14SB, ff19SB) for proteins; GAFF/GAFF2 for drug-like molecules | Good CPU scaling, improving GPU support | Free for academics, paid for commercial use | | CHARMM | Membrane proteins, lipid bilayers, carbohydrates | CHARMM36 for proteins and lipids; CGenFF for drug-like molecules | Extensive tools for setup and analysis | Academic source code access, commercial licenses required | | GROMACS | Large biomolecular systems, membrane proteins | Compatible with most force fields Optimized for AMBER and CHARMM | Excellent CPU scaling, highly efficient code | Open source, free for all uses | | OpenMM | Method development, custom protocols | Compatible with standard force fields and Excellent for custom force fields | Best GPU performance, flexible programming model | Open source, Python API makes it accessible | | OPLS | Drug discovery, organic molecules | OPLS-AA and OPLS-2005/OPLS3e for proteins and drug-like molecules | Optimized for Schrödinger software suite | Commercial license through Schrödinger | ## The role of Enhanced Sampling and Binding Free Energies MD is great for short clips, but it is painfully slow. Most simulations runs are in microseconds, while the real biology happens on the millisecond-to-second scale. It’s like trying to watch a 2-hour film by capturing a few seconds of footage, you miss the big plot twists. That’s where Enhanced Sampling Methods come in: they tweak the rules of the movie so we can fast-forward through the boring parts and jump to the rare, high-impact scenes. Just like there’s no single shortcut to skipping boring scenes in a movie, there isn’t one trick to speeding up MD. Some popular approaches include: - **Replica Exchange MD (REMD):**In normal MD, the system is stuck because at room temperature it doesn't have the energy to hop over barriers. REMD solves this by running clones of the same system at different temperatures. High-temp replicas shake harder, cross barriers more easily, and then occasionally trade places with the low-temp replicas.Result: the cold replicas get access to high-energy states they'd never reach on their own. - **Metadynamics:**Proteins love to sit in the same comfy valley. Metadynamics continuously adds small "energy hills" (Gaussians) to discourage the system from revisiting the same states, flattening the landscape so it explores new conformations. - **Umbrella Sampling:**Some processes, like pulling a ligand out of a protein pocket, are so energetically steep that doing it in one go doesn't work, the ligand just falls back in. Umbrella sampling breaks the journey into small, manageable steps, each stabilized with a gentle restraint (a spring). - **Other tricks (ABF, Funnel MD, etc.):**These are more specialized, but they all share the same idea: add just enough bias to help the system cross barriers, then carefully remove the bias in analysis to get the *true* physics back. ![Figure 3: Enhanced Sampling Methods](images/2e736d24dac3202039c63fa59f8ae5ad5258a48e-965x529.png) Once we’ve explored these states and confirmed that the ligand actually stays bound, the next big question is: ***“How strong is that binding?”*** This is where **binding free energy (ΔG_bind)** comes in. It’s the number that tells us, quantitatively, how tightly a ligand grips its protein, directly related to the experimental dissociation constant (Kd). Computing it is hard because you need to capture both enthalpy (the energy of interactions) and entropy (the freedom molecules lose when they stick together). Over the years, computational chemists have built a toolbox of methods that trade accuracy for speed. - **Docking scores** are fast, great for triaging huge libraries, but not real free energies. - **Molecular Mechanics Poisson-Boltzmann Surface Area (MM-PBSA)** takes MD snapshots, calculates energies for bound vs. unbound states, and adds implicit solvation. It’s cheap and sometimes gets trends right, but absolute errors can be brutal. - **Linear Interaction Energy and Extended LIE (LIE/ELIE)** estimate binding free energy by focusing on the main interactions between a ligand and its protein, electrostatics and van der Waals forces. Combining this energies and a few coefficients (learned from experimental binding data), reasonably increases prediction. - **Free Energy Perturbation (FEP )** is the heavyweight champion. It literally “alchemically” morphs one ligand into another inside the protein pocket across a series of simulations. Done carefully, it reaches near-experimental accuracy. The catch: it eats a lot of compute and requires expertise. | Method | Principle | Strengths | Weaknesses | | --- | --- | --- | --- | | Docking Score | Empirical / force-field scoring of poses | Very fast, screens millions | Not a true free energy; ignores entropy | | MM-PBSA | Energy + solvation (PB/GB) from MD snapshots | Cheap, widely used | High errors (>6 kcal/mol), sensitive to setup | | LIE / ELIE | Empirical scaling of interaction energies | Faster than FEP, better than MM-PBSA | Needs training, poor transferability | | FEP | Alchemical transformation of ligands | Near-experimental accuracy | Heavy compute, expert setup | Table 2: Overview of Computational Methods for Estimating Binding Free Energy. ## The Limits of Classical Force Fields For decades, simulations have leaned on **classical force fields** like AMBER, CHARMM, and OPLS. They’ve been workhorses, aiding in countless discoveries and drug design projects. But they’re not flawless. But simplicity comes at a cost. The biggest issue? **Fixed charges.** Real molecules aren’t static; their electron clouds shift, polarize, and sometimes even transfer charge between atoms. Classical force fields freeze those charges in place, which makes them less accurate for charged or highly polar systems. Another problem is **parameter transferability.** A force field calibrated for one class of molecules may break down for another, forcing researchers into tedious re-parameterization before they can even hit “run.” ![Figure 4: Conceptual visualization of classical force field limitations (AI Generated)](images/415c1050c6b302104614f2cbad32b271875d6a5b-939x430.png) ## Neural Network Force Fields: The New Frontier This is where machine learning steps in. Neural network force fields don’t start with a fixed equation, they learn the energy landscape directly from high-level quantum mechanical (QM) data, calculations which capture how electrons really behave. Because they learn directly from QM, these models can “feel” things classical models can’t. For example: - **Polarization:** In reality, the electron cloud around an atom distorts when another charge comes nearby. Classical force fields freeze charges in place, but neural nets can respond dynamically. - **Charge transfer:** Sometimes electrons actually shift from one molecule to another (common in binding pockets). Classical models can’t handle that, but neural nets can. - **Many-body effects:** Instead of treating each interaction pair by pair (like atom A pulling on atom B), neural nets can capture the collective influence of multiple atoms acting at once. The difference isn’t just theoretical. In one benchmark, the neural network force field Espaloma beat a widely used classical force field (OpenFF 1.2.0) for the Tyk2 kinase system, cutting prediction error from 1.10 kcal/mol down to 0.73 kcal/mol. In drug discovery, where fractions of a kcal/mol can reorder your entire hit list, that’s a big deal [3]. ![Figure 5: Espaloma small molecule parameters can be used for accurate protein-ligand alchemical free energy calculations](images/066da3f1c40787375b61caa979f53560e667d825-850x488.png) ## How an MD Simulation Actually Run? Picture MD as building a movie studio. You set the stage with physical laws, then capture the dynamic dance of atoms. While traditional MD requires arcane command-line knowledge, modern platforms like [Litefold.ai](http://litefold.ai/) simplify this complexity into intuitive interfaces. 1. **System Setup; The Script & Cast: **We begin by preparing the protein and ligand, fixing missing atoms, assigning charges, and choosing a force field. A bad setup here is like miscasting the lead actor, the whole show falls apart. 1. **Environment Creation; Setting the Stage:** Molecules don’t live in a vacuum. We place them in a water box, add ions, and mimic physiological salt levels. This ensures the interactions unfold in something resembling real biology. 1. **Energy Minimization; Safety Check:** Before rolling the cameras, we need to relax any tense spots in our molecular system. This step removes any unrealistic high-energy positions that could cause the simulation to crash. Otherwise, the simulation could “explode” from bad starting positions. 1. **Equilibration; The Rehearsal:** Now we gently warm the system to body temperature (300K), holding the main structure steady while water and side chains find their natural rhythm. This step ensures everything is stable before the real run. 1. **Production Run; The Filming:** This is the actual simulation. The restraints come off, the system runs freely, and we record atomic motion over nanoseconds to microseconds. These snapshots become our molecular movie. 1. **Analysis: The Editing Room: **Once the movie’s shot, we dig into the frames: tracking RMSD, hydrogen bonds, conformational changes, or calculating binding free energies. This is where raw motion turns into scientific insight. 1. **Validation & Enhancement; Quality Control:** Finally, we check consistency by running repeats, ensuring the results are statistically solid. If needed, we apply enhanced sampling to capture rare but important events. ![Figure 6: Conceptual visualization of a typical MD simulation run](images/3dd9a373c6f62d4fc5b7f8ab638d48ecd2cdbe53-1000x1000.png) ## Where MD Still Falls Short MD gives us incredible atomic detail, yet there are walls it can’t break. There are several reasons for that. For example: **1. The Timescale Problem**: Simulations run for nanoseconds to microseconds, but biology moves on milliseconds to seconds. It’s like trying to understand a symphony by sampling a few microseconds of each note. Enhanced sampling helps, but we’re still orders of magnitude away from the true pace of life. **2. The Accuracy Ceiling**: Even the best free energy methods top out at ~0.5–1.0 kcal/mol accuracy. Sounds tiny, but in binding affinity that’s a 3–5x error. In drug discovery, where the difference between nanomolar and micromolar binding can decide success or failure, that margin reshuffles your hit list. **3. The Transferability Gap**: No force field or neural net works everywhere. Something tuned for soluble proteins may flop on membrane proteins; trained on kinases, it may stumble on GPCRs. There’s no universal recipe, every new target demands painful re-validation. **4. The Sampling Myth**: Even with enhanced methods, you never know if you’ve really seen everything. A hidden allosteric pocket might only open at the 10-millisecond timescale, and your microsecond run won’t catch it. Missing those rare states can mean missing the mechanism of resistance or selectivity. **5. The Computational Reality Check**: Yes, AI force fields and FEP can approach experimental accuracy. But they eat compute for breakfast and demand expertise. Academic labs and small biotech are stuck between costly “rigorous” methods and cheap-but-dodgy approximations like MM-PBSA. **6. The Validation Bottleneck**: The real choke point isn’t silicon, it’s biology. You can run a thousand MD jobs in a week, but testing ten compounds in the lab takes months. **The bottom line?** It’s brilliant for understanding mechanisms, generating hypotheses, and prioritizing experiments. Just don’t confuse computational confidence with biological truth. The wet lab still has the final word. ## Conclusion Molecular dynamics is not just an add-on to docking. It is the bridge between static prediction and biological reality. By letting proteins and ligands “breathe” in silico, MD helps separate true binders from false positives, refines docking poses, and allows a first look at the thermodynamic cost of binding. It does not eliminate the need for experiments, but it dramatically narrows the search space and gives wet-lab teams higher-quality hypotheses to test. The future of MD lies at the intersection of physics and machine learning. Neural network force fields, GPU acceleration, and better enhanced-sampling protocols are already pushing simulations toward biological timescales and near-experimental accuracy. Combined with smart triaging (docking, FEP, AI-driven pose scoring), MD will remain one of the most valuable tools for turning computational predictions into actionable leads. The goal is not just to simulate reality but to guide it, prioritizing the right molecules, reducing wasted syntheses, and accelerating the path from idea to drug. ## References [1] [D3R Grand Challenge Overview – ](https://pmc.ncbi.nlm.nih.gov/articles/PMC6472484/)*[Community-wide evaluation of computational drug design methods.](https://pmc.ncbi.nlm.nih.gov/articles/PMC6472484/)* [2] [Chemspace – ](__GHOST_URL__/molecular-docking-in-drug-discovery/)*[Molecular Dynamics Simulations: Concepts and Applications](__GHOST_URL__/molecular-docking-in-drug-discovery/)* [3] [ResearchGate – ](https://www.researchgate.net/figure/Espaloma-small-molecule-parameters-can-be-used-for-accurate-protein-ligand-alchemical_fig6_363403092)*[Espaloma small-molecule parameters for accurate protein-ligand alchemical free energy calculations.](https://www.researchgate.net/figure/Espaloma-small-molecule-parameters-can-be-used-for-accurate-protein-ligand-alchemical_fig6_363403092)* --- # Molecular Docking in Drug Discovery URL: https://www.lite.bio/blogs/molecular-docking-in-drug-discovery Author: Nabajit Borah Date: 2025-09-08 Imagine spending over a decade and billions of dollars chasing a single medicine, only to see most candidates fail before they ever reach a patient’s hands. That’s the reality of drug development today. On average, it takes 12 to 15 years and billion of dollars to bring a new drug from the lab to the pharmacy shelf. Out of the thousands of compounds discovered, maybe one will make it all the way through. What if a computational method could accelerate this timeline from years to weeks? Molecular Docking has revolutionized how we approach protein-ligand interactions in structure-based drug design. This computational tool has become fundamental in modern drug discovery pipelines, enabling the virtual screening of millions of compounds before expensive experimental validation. As pharmaceutical R&D spending reached $260 billion in 2023, docking serves as a cost-effective filter in the early stages of drug development. # What is Molecular Docking Picture a scientist staring at a 3D model of a protein on their computer screen. Somewhere on that protein lies a tiny pocket, just the right size for a molecule that could become tomorrow’s life-saving drug. But which molecule will fit? Testing them one by one in the lab would take years. ![Figure 1: Moving beyond manual testing to quickly identify life-saving molecules that fit just right](images/017276413d522e00a9d86ecba68bb9ab9183b5f2-1000x1000.png) This is where molecular docking comes in. It is a computational method that seeks to predict how two or more molecules will bind to form a stable complex. In the most common scenario, a small molecule, referred to as the "ligand," is docked into the binding site of a large biomolecule, the "receptor," which is typically a protein or a nucleic acid. Molecular Docking heavily relies on structure-based drug design and therefore needs high resolution experimental structures obtained from techniques like X-ray crystallography, NMR spectroscopy, and cryo-electron microscopy. This process has two primary and interconnected goals: 1. **Predicting the binding pose:** This is the geometric challenge of docking. The objective is to determine the most favorable 3D orientation of the ligand within the receptor's binding site. But here’s the tricky part: the ligand doesn’t just drop in like a rigid Lego block. It can flex around its bonds, changing shape until it finds a position where everything clicks into place like hydrogen bonds and hydrophobic contacts stabilizing the complex. 1. **Affinity Prediction (Scoring):** This is the energetic challenge. The goal here is to estimate the strength of the interaction between the ligand and the receptor, quantified as the binding free energy. This is represented by a "docking score," a numerical value calculated by a scoring function. By convention, a more negative (or lower) score indicates a stronger, more stable predicted interaction, suggesting a higher binding affinity. # The Docking Algorithm If molecular docking were a video game, the docking algorithm would be the “engine” running under the hood. It decides how the pieces move around and how the game keeps score. Every docking program relies on two things: a way to search (to figure out how a ligand might fit into a receptor) and a way to score (to judge which fits are actually good). ## Sampling (The Search Algorithm) The search algorithm is responsible for exploring the vast "conformational space" of the system. This space includes every possible way a ligand can sit inside the receptor, plus all the shapes it can adopt as its bonds rotate. The combinations are essentially astronomical, which makes brute-force searching impossible. To tackle this, docking programs use smart search algorithms that balance speed with accuracy. Instead of checking every possibility, they explore the most promising regions of the space. Common strategies include: 1. **Genetic Algorithms:** Used by programs like AutoDock and GOLD, these algorithms mimic principles of biological evolution. The idea borrows from evolution: start with a “population” of random ligand poses, then let them evolve. Poses undergo mutation (small random tweaks), crossover (mixing features of two good poses), and selection (keeping the strongest fits). With each generation, weaker poses drop out while stronger ones survive, and therefore, in the end an optimal binding pose. 1. **Monte Carlo Methods:** It involves making random changes to the ligand's position, orientation, or conformation. Each new pose is evaluated, and it is accepted or rejected based on a criterion that favors lower-energy states. 1. **Systematic Searches**: The algorithm doesn’t rely on randomness. Instead, it tries to cover the conformational space in an orderly way. A common trick is to split the ligand into smaller fragments and place them one by one into the binding site, gradually building up the full molecule. This helps manage complexity, but the number of possibilities still grows fast, therefore, pruning strategies are used to cut off bad fits early. ![Figure 2: Different types of Search Algorithms used in Molecular Docking](images/6d33faa0e8dd6587ce32819fff094ad69989644e-1000x608.png) ## Scoring Function It is the mathematical engine that estimates how strongly a ligand binds in a given pose. It plays two key roles. During the search, it gives quick feedback whether the current pose is better than the last pose and, in the end, it ranks all the generated poses to highlight the most likely binding mode. Without scoring, docking would generate endless possibilities. Scoring functions are usually approximations and are usually of three types: 1. **Physics-Based (Force-Field):** They apply laws of physics to estimate binding strength which relies on several factors such as energies from van der Waals forces, electrostatics, and molecular force fields such as CHARMM or AMBER. Accurate but painfully slow. 1. **Empirical:** Based on experimental data, these functions use a weighted sum of simple terms such as hydrogen bonds, hydrophobic contacts, and penalties for rotational flexibility. They’re faster than full physics and tuned to match real-world results. 1. **Knowledge-Based:** Instead of physics or experiments, these functions learn from statistics. By analyzing thousands of known protein–ligand structures, they assign “potentials” to atomic contacts: the more often an interaction shows up in nature, the more favorable it’s assumed to be. ![Figure 3: Scoring functions and it’s types](images/a55ac9160baee0e7cc75b3fa0842f97d098d7972-1000x667.png) Docking lives on the trade-off between accuracy and speed. The search has to test millions of ligand poses, which means the scoring function used during sampling must be fast but approximate. This sacrifices physical realism, so the final ranking can’t rely on those scores alone. To handle this, most workflows use a two-stage process: a rapid initial search with simplified scoring to generate candidate poses, followed by more rigorous re-scoring of the top hits. This balance between efficiency and accuracy defines both the strength and the limitation of docking. # Types of Docking We will briefly differentiate the types of docking in the following basis: 1. **On the basis of flexibility of the interacting molecules** 1. **On the basis of interacting partners** 1. **On the basis of Binding Site knowledge** Let's understand each of them in proper details: ## On the basis of flexibility of the interacting molecules Under this we got Rigid Docking, Flexible docking, Flexible-Ligand / Rigid-Receptor Docking. ### Rigid Docking Historically, the concept of docking was rooted in the "lock-and-key" model. Protein and ligand as rigid shapes that fit together perfectly. But biological molecules aren’t rigid shapes, they are “dynamic”; they wiggle, bend and adapt to the surrounding environment. The search algorithm only explores the six degrees of translational and rotational freedom to find the best geometric fit. This approach is computationally very fast but is biologically unrealistic for most systems. Primarily used in initial screenings. ![Figure 4: Lock and Key model vs Induced-Fit model where P is the protein (receptor) and L is the ligand](images/f12b91dae08af84aaf396b840165df7bfa28a3c0-679x511.png) ### Flexible Docking (Flexible Receptor) The more biologically accurate “induced-fit” or “glove” model recognizes that proteins and ligands aren’t rigid shapes. When they interact, both partners can shift, bend, or twist their conformations to achieve a tighter, more favorable fit. This flexibility is what actually happens in living systems, making induced fit a much closer reflection of reality than the older lock-and-key view. This process is computationally expensive, so it's usually saved for the very end of a docking study to refine the best-looking candidate structures. ### Flexible-Ligand / Rigid-Receptor Docking: This is the most common and widely used approach in drug discovery, especially for large-scale virtual screening. The ligand can rotate around its bonds and adopt different shapes, but the protein is treated as rigid. This makes the model more realistic than rigid–rigid docking while keeping the computational cost reasonable. It is widely used by docking tools like AutoDock Vina and Glide. ## On the basis of interacting partners Under this we have Protein-Ligand Docking, Protein-Protein Docking, Protein-Nucleic Acid Docking. ### Protein – Ligand Docking: The bread and butter of drug discovery. Here the goal is to predict how a small, drug-like molecule fits into a protein pocket. It is central to drug discovery because it supports virtual screening (computer testing of millions of molecules to find promising ones, called “hits”) and lead optimization (improving those hits by tweaking their structure so they bind better and act more like real drugs). ![Figure 5: Protein - Ligand (Green) Docking](images/61efdd8e6128a3c157f3f91b494ea092dccd9177-875x666.png) ### Protein – Protein Docking Proteins often team up to carry out signaling and cellular functions. Predicting how two proteins fit together is harder than protein – ligand docking as interfaces are larger, flatter, and more flexible, and the binding energies are subtle. Tools like RosettaDock and UDock2 specialize in this tough challenge. ![Figure 6: Protein - Protein Docking](images/88b868fbb394ee173ad1694483c5a8567af4b528-498x289.png) ### Protein – Nucleic Acid Docking Unlike proteins, Nucleic Acid docking is hard due to their charged phosphate backbone and groove structures mean you need special handling unlike amino acids chains (proteins). Many important processes such as replication, transcription, and translation relies on it and therefore requires dedicated tools like NPDock and HDOCK. ![Figure 7: Protein (Blue) - Nucleic Acid Docking](images/e2c297ad8b4dec6f0fd283b1750eec77d072dbb3-544x430.png) ## On the basis of Binding Site knowledge Docking approaches can be grouped by how much we know about the protein’s binding site in advance (using prior experimental results). Under this, we got: Targeted / Focused Docking, Blind Docking. ### Targeted/ Focused Docking Think of this as searching with a flashlight. You already know where the action is, so you shine directly on that spot. Here, the ligand is only tested in a specific region of the protein, usually a well-known active site or binding pocket. Because the search is focused, it’s much faster and easier on computational resources. This is why targeted docking is the go-to method for most drug discovery projects where the binding site is already mapped out. ### Blind Docking Now imagine switching off the flashlight and exploring the entire room in the dark. In blind docking, the ligand is allowed to scan the whole surface of the protein without assumptions about where it might bind. It’s computationally more expensive, but it can uncover hidden or unexpected binding sites, especially *allosteric sites*, which are alternative pockets that can regulate protein function. ![Figure 8: Focused vs Blind Docking](images/6e8a9ad43443b0d819a3f215d728b447a9de4a83-850x512.png) Figure 8 explains clearly, in focused (targeted) docking, the ligand is only tested against a single, predefined binding pocket, making the process fast and precise. In blind docking, the ligand explores the entire protein surface, which is slower but powerful for uncovering unknown or allosteric sites that could become new drug targets. The selection of a docking method is therefore a critical step in the research process, guided by the specific scientific question. A researcher aiming to quickly screen a million-compound library for potential starting points might choose a fast, flexible-ligand/rigid-receptor approach. Whereas a medicinal chemist seeking to understand the precise binding mode of a high-affinity lead compound would opt for a more computationally intensive flexible docking. # The workflow of Molecular Docking While the theory of docking is complex, the practical workflow has become increasingly accessible thanks to a variety of powerful software tools. Figure 9 sums it up neatly. ![Figure 9 : Docking Pipeline](images/b393257101190c6766c0953e81d9b615693f8c00-809x399.png) 1. **Protein and Ligand Selection**You start by picking your players: the protein (target) and the small molecule (ligand). Think of it as choosing the lock and the potential key. Structural files from a public repository like Protein Data Bank (PDB) are downloaded. 1. **Protein - Ligand Preparation**Before docking, both molecules need a cleanup from non-protein atoms like zinc and copper which act as catalysts in biological reactions. Hydrogens are added, protonation states are adjusted on the basis of amino acids states, and the correct molecular form (tautomer, protomer, stereoisomer) is chosen. It’s like sharpening the key so it actually fits. 1. **Binding Site Definition/Grid Box**Where on the protein should the ligand try to fit? If you already know the active site, great, you define that cavity and carry out targeted docking. If not, computational tools can scan the entire protein surface to spot potential “pockets” using a three-dimensional grid of points which are used for energy calculation. The box must be large enough to allow the ligand to move and rotate freely within the binding site but not so large as to waste computational effort on irrelevant space. 1. **Structural Water**Water molecules often disturb the docking process, forming hydrogen bonds between protein and ligand. Therefore, they are usually removed. 1. **Molecular Docking**This is the actual search. Algorithms explore different orientations and conformations of the ligand inside the binding site. Then, scoring functions rank which poses are most likely to be biologically relevant. 1. **Evaluation**Finally, you check: Do the predicted interactions (like hydrogen bonds) make chemical sense? If experimental data exists, does the docked pose reproduce reality? This step validates whether your “key” actually works in the “lock.” On top of that, you also look at the docking score, the more negative the value, the stronger (and usually better) the predicted binding. | Software | Availability | Core Algorithm | Primary Application | Key Feature | | --- | --- | --- | --- | --- | | AutoDock / Vina | Free (Academic) | Lamarckian Genetic Algorithm / Iterated Local Search | Academic research, general-purpose docking | Widely used, well-documented, Vina is known for its speed and ease of use. | | Glide | Commercial | Systematic Search + Optimization | High-throughput virtual screening (HTVS) in industry | Highly regarded for its accuracy and robust performance in ranking compounds. | | GOLD | Commercial | Genetic Algorithm | Lead optimization, handling of protein flexibility | Features multiple scoring functions and advanced options for modeling side-chain flexibility and bridging waters | | DOCK6 | Academic License | Anchor-and-Grow Incremental Construction | Academic virtual screening, fragment-based design | A long-standing program from UCSF with a focus on shape complementarity. | | FlexX | Commercial | Fragment-Based Incremental Construction | Fast virtual screening, scaffold hopping | Builds ligands piece-by-piece into the active site based on interaction patterns. | | HDOCK | Free (Web Server) | Hybrid Template-based + FFT-based Free Docking | Protein-protein and protein-nucleic acid docking | Specialized for Nucleic Acid- Protein complexes, can use sequences as input. | Table 1: Summary of Commonly Used Docking Platforms ## A Critical Look at the Limitations of Molecular Docking Docking generates educated guesses about how molecules might interact, trusting the results blindly would result in a lot of wasted time and money chasing false leads. It's a powerful tool for narrowing down millions of possibilities, but it needs to be validated by real-world experiments. ### Scoring Function Problem The biggest source of error? The scoring function, the formula docking programs use to rank which molecules “fit” best. These functions have to be fast, therefore, they trade off speed with accuracy. They often miss out: - **Entropy calculations**: Biological molecules are dynamic, free to wiggle, rotate, bend but when it is docked, it is forced into a single pose. Losing that flexibility costs energy, because the system becomes more ordered. Docking usually can’t fully capture these changes. - **Solvation:** Proteins and ligands are surrounded by water in real life. Before docking, one of the steps is to remove water. Making bonds involves displacing water. This is not free energy-wise: **displacing water costs energy, and forming new bonds might not fully compensate**. Docking scoring functions simplify this, often treating solvent as a background, which can affect the true binding energy of docked complex. ### Protein Flexibility Most standard docking protocols treat the receptor as a single, rigid structure to keep the computational costs low. Proteins are not static. Upon binding a ligand, a protein's active site often undergoes conformational changes, from small side -chain rotations to large-scale domain movements to accommodate the ligand. ### Docking Preparation The principle of "garbage in, garbage out" (GIGO) is paramount in molecular docking. Common mistakes involved: - **Wrong Protonation States:** Charges matter. The protonation states of acidic and basic groups on both the ligand and the protein are highly dependent on pH. Docking the wrong form of a molecule leads to wrong predictions due to fundamentally flawed electrostatic calculations. - **Wrong Tautomer:** Many organic molecules can exist in multiple tautomeric forms, which are structural isomers that readily interconvert. ****These forms can have different shapes and hydrogen bonding capabilities; docking the wrong tautomer can give poor results. - **Low-Resolution Structures:** If your protein structure is fuzzy, docking results will be unreliable. The accuracy of a docking result is limited by the quality of the receptor structure used. # Conclusion Molecular docking might not be perfect, but it’s still one of the most powerful tools in drug discovery. While docking scores shouldn't be taken as absolute truths, they provide valuable initial guidance for prioritizing compounds and understanding potential binding modes. The key lies in recognizing docking as the first step in a validation pipeline, not the final answer. To further validate molecular docking predictions, several experimental and computational approaches can be employed. **Biochemical assays** such as enzyme inhibition studies, binding affinity measurements such as **Surface Plasmon Resonance (SRC) and Isothermal Titration Calorimetry (ITC**). Molecular dynamics simulations let you watch the docked molecule and protein in motion, checking if the predicted pose is stable. Free energy perturbation (FEP) calculations offer more accurate binding affinity predictions by accounting for entropic effects that standard docking scoring functions miss. And structure-activity relationship (SAR) studies with related compounds help confirm whether the predicted binding mode explains real-world trends in potency and selectivity. --- # Structure Based Drug Design just got easier than ever URL: https://www.lite.bio/blogs/litefold-denovo Author: Cory Kornowicz Date: 2025-09-05 We present LiteFold *DeNovo, *our second flagship feature, a highly efficient structure based drug design workflow to accelerate your lead candidate research to the next level. In this blog, we will discuss the features that are available today, how to get started, and the new features that are coming very soon. ## What is *Structure Based Drug Design* If you are fairly new to this field and want to learn more about the basics of Structure Based Drug Design (SBDD), this section is for you. For decades, drug discovery was akin to searching for a needle in a haystack only without knowing what the needle looked like or even if there was one to find. Scientists would test thousands of compounds against disease targets, hoping something would work, often with little understanding of why successful drugs succeeded or failed drugs didn't. This trial-and-error approach was expensive, time-consuming, and frequently led to dead ends after years of research. Structure Based Drug Design (SBDD) transformed this process by *turning on the lights*. To put it simply, SBDD is the process of creating new molecules using the three dimensional structure of the binding pocket of the protein target. The binding pocket can be thought of as a lock that a molecule can fit snugly into. By using the three dimensional structure of the interior of the lock, molecules can be crafted to specifically match that complementary shape to perfectly fit that lock. The goal is to always maximally fit a particular lock while minimizing the amount of other locks that molecule can fit into – fitting into multiple binding pockets is a common cause of *off-target *effects and can lead to unintended side effects. ![Image generated by ChatGPT (this image is just for intuition purpose only).](images/388dc9d711053fa28a7cabd58150df67fa9d2a23-1000x667.png) ## Why *Structure Based Drug Design* This structure-guided approach offers dramatic advantages over traditional methods. Rather than synthesizing and testing thousands of random compounds, a process which can take years and cost millions, SBDD allows researchers to focus on molecules most likely to succeed. Traditional high-throughput screening might test a billion compounds to find a few promising leads, while SBDD can identify strong candidates from much smaller, rationally designed libraries. This means fewer failed experiments, reduced costs, and faster paths to potential therapies. The power of SBDD has been proven repeatedly in breakthrough medications.[ HIV protease inhibitors](https://en.wikipedia.org/wiki/Discovery_and_development_of_HIV-protease_inhibitors) like saquinavir, one of the first major SBDD successes, were designed by analyzing the precise structure of the HIV enzyme. [Cancer drugs like imatinib (Gleevec)](https://www.rroij.com/open-access/structurebased-drug-design-of-kinase-inhibitors-integrating-organic-chemistry-and-therapeutic-applications.pdf) were crafted to fit perfectly into specific kinase binding sites. More recently, COVID-19 antivirals like [nirmatrelvir (Paxlovid)](https://pdfs.semanticscholar.org/f74e/079ba427018b162244aaba2eb5d352a84020.pdf) were developed using SBDD principles, demonstrating how structural knowledge can accelerate drug development even under urgent timelines. Beyond initial discovery, structural knowledge enables systematic optimization that's impossible with traditional approaches. When researchers see exactly how a molecule sits in its target binding site–which atoms interact where, which parts contribute to binding–they can methodically modify the drug to enhance potency, improve selectivity, or reduce side effects. It's like being able to file down specific parts of the key while watching exactly how it fits into the lock, rather than guessing which modifications might work. This rational optimization process can transform a weak initial hit into a potent, selective drug through cycles of structural analysis and informed design. The success of SBDD raises an exciting possibility: if structural knowledge is so powerful, what if we could rapidly generate not just a few optimized molecules, but hundreds? LiteFold makes this possible by using neural network models to predict molecular structures specifically designed to fit into target pockets—completing predictions in minutes rather than hours or days. This rapid structure-based drug design approach allows our platform to generate hundreds of molecule candidates in a single workflow, enabling rational design of molecules with higher specificity, higher potency, and fewer off-target effects. ![Image](images/df6c507b2b9c9e4ff486a5916ec887cc9ca49d7e-1000x454.png) **Example of structure-based molecular design:** A computationally generated molecule (center) sits precisely within its target protein binding pocket (highlighted in purple box), demonstrating how structural knowledge enables the rational design of drug candidates with optimal fit and specificity. ## How to use LiteFold *DeNovo* in your drug design workflow To get started with LiteFold *DeNovo*, navigate to the *DeNovo s*ection inside the *Lab *tab. ![Image](images/1ca61d6331847700eb26c11178dc7e8ca14c43d7-1000x532.png) You can upload a PDB file of your own or use a protein structure previously predicted through our Structure Prediction pipeline. The platform will first guide you to choose a target pocket, which can be entered manually or our AI pocket prediction service will find and detect pockets that could be targeted for you. ![Image](images/c108f2ff42a2b14f0670d319339c23e73793f8af-1000x402.png) Now you are ready to generate molecules for your target pocket! While there is no limit to how many molecules you can generate, our free tier has a cap of 100 molecules. ![Image](images/251848ce8adbb6d520103bfcf0fc97a86c3b17a8-1000x536.png) Once the molecules are generated, you are presented with an intuitive, integrated workspace that transforms complex drug discovery into an accessible, visual experience. View your candidates in an interactive 3D window to see exactly how each molecule fits within the target binding pocket, then seamlessly switch to the built-in molecule editor to refine promising structures with just a few clicks. All critical metrics are displayed at a glance, allowing you to instantly identify the most promising candidates. This streamlined workflow accelerates your path from initial concept to lead compounds, condensing what traditionally takes weeks of computational work into a single, efficient session. With features like molecular dynamics simulation (coming soon!), you can validate and optimize your designs without switching between multiple software tools or waiting for external analysis. ![Image](images/11e427c42344dda8d9354afb5030f0f3ed533ab1-1000x540.png) ## Fragment Growing & Molecule Editing LiteFold *DeNovo* offers two ways to quickly edit the generated molecules: fragment growing and the molecule editor. Fragment growing allows you to take an already existing molecule and grow or add new pieces within the context of the binding pocket. The molecule editor allows you to make fine tuned adjustments by hand to change the structure of the molecule and then quickly visualize the changes. Any new changes to a molecule recomputes its metrics and binding affinity score so you can see if your edits increase the potency and drug-likeliness properties. ## Supported Metrics Each molecule generated has a few properties computed automatically: QED, Synthetic Accessibility, and Vina docking scores. The QED (Quantitative Estimate of Drug-likeness) score ranges from 0 to 1, with higher values indicating that a molecule possesses favorable drug-like properties such as appropriate molecular weight, lipophilicity, and other physicochemical characteristics that correlate with successful oral drugs. The Synthetic Accessibility (SA) score predicts how challenging it would be to synthesize the molecule in the laboratory, with lower scores indicating easier synthetic routes—a crucial factor for determining whether promising candidates can be practically manufactured. Finally, the Vina docking scores estimate the binding affinity between each generated molecule and the target protein, with more negative values suggesting stronger binding interactions and potentially higher potency. Together, these three metrics provide an immediate assessment of each candidate's drug-likeness, synthetic feasibility, and predicted efficacy, allowing researchers to quickly prioritize the most promising molecules for further development. ## What's Next In our next release, LiteFold will unveil a new and improved Rosalind and offer the ability for users to run molecular dynamic simulations! The *Simulation* module will open the door for users to run different variations of molecular dynamics, spanning anything between long duration 1µs simulations and short 10ns stability measuring simulations (accelerated with metadynamics). This means you can actually watch how your molecules behave and how the protein reacts dynamically. Rosalind's new upgrades will allow her to operate as a *medicinal chemist*. This includes editing molecules using her built-in chemical intuition, accessing the tools provided through LiteFold's Lab, and envisioning new mechanistic hypotheses for how compounds interact with their targets. She can iterate on designs, suggest improvements, and help you explore chemical space more efficiently than ever before. Lastly, we are working on a High Throughput Virtual Screening workflow with Structure Based Virtual Screening (SBVS) and Ligand Based Virtual Screening (LBVS) options. This will bridge our technology stack with current state of the art extensive library filtering pipelines and improve upon them with user directed custom scoring functions; including pharmacophore, shape & electrostatic, and distance based similarity scores. We can't wait to see the life-changing medicines you'll discover using these tools! --- # Edit, predict, evaluate your proteins structures in bulk with LiteFold URL: https://www.lite.bio/blogs/bulk-predictions Author: Anindyadeep Date: 2025-08-07 We just launched something new at LiteFold, a **structure prediction editor**. Yes, an editor, not just a prediction tool. Let me explain. After AlphaFold 2 and 3 came out of DeepMind, the open-source community didn’t sit back. They’ve been actively building strong alternatives, not just replicas, but next-gen structure predictors. Some of the standout ones include [OpenFold from AQ Laboratory](https://github.com/aqlaboratory/openfold), [ProteinX by ByteDance](https://github.com/bytedance/Protenix), [Chai-1/2 by Chai Discovery](https://github.com/chaidiscovery/chai-lab), and [Boltz 1/2 by MIT](https://github.com/jwohlwend/boltz). What makes these exciting is that they're not just doing plain protein folding anymore they’re moving towards generalized biomolecular structure prediction. That includes [antibody complexes](https://www.chaidiscovery.com/), protein-RNA structures, ligands the whole range. If you're skeptical or wondering how Boltz-2 stacks up against AlphaFold, take a close look at the latest benchmarks. The results speak for themselves. | Metric | AlphaFold 2 | AlphaFold 3 | Boltz-2 | | --- | --- | --- | --- | | Antibody–Antigen DockQ > 0.23 | 27% (AF2-Multimer v2.3) | 85% | 53% | | Protein–Protein DockQ > 0.23 | ~60% | 83% | ~67% | | Protein–Ligand Pose Accuracy (<2Å) | Not supported | ~66% (PoseBusters) | 62.5% (FEP+ subset) | | Protein–RNA LDDT | ~60 | 78 | 74 | | RMSF Pearson (ATLAS) | - | - | 0.82 | | RMSF Spearman (mdCATH) | - | - | 0.77 | | Binding Affinity Pearson (FEP+) | - | - | 0.62–0.66 | | Binding Affinity R² (FEP+) | - | - | ~0.55 | | SKEMPI-2 ΔΔG Pearson | - | 0.86 | – | | Structure-only Inference Time | ~30–90 min | ~5–20 min | ~5–15 sec | | Binding Affinity Inference Time | Not supported | Not supported | ~5–15 sec | From the above table, you can see that Boltz-2 performs competitively with AlphaFold 3 across several benchmarks and outperforms it in inference speed and flexibility. For instance below images shows the comparisons between Boltz-2 model and others in CASP-16 and other affinity related benchmarks. ![Image](images/ce7aabc40e13dd6a8ffd4d375268329a2434bc88-1000x552.png) Figure 1: Comparison of Boltz-2 with different benchmark including CASP-16. Source: [Boltz-2 Github](https://github.com/jwohlwend/boltz) ![Image](images/c10249a33e32b3a2265747ba9f0a542ce20a3f39-1000x886.png) Figure 2: Comparison of Boltz-2 with other models in docking / interactions related benchmark. Source: [Boltz-2 GitHub](https://github.com/jwohlwend/boltz)💡While Boltz‑2 is still catching up with AlphaFold‑3 in some areas, it offers a strong balance of speed and base-level accuracy, both critical when the goal is to predict and filter thousands of backbone structures at scale. For high-throughput workflows, this makes Boltz‑2 a highly practical choice. We also plan to integrate additional models in the near future to further enhance coverage and performance. While AlphaFold 3 is not [technically Open Source](https://github.com/google-deepmind/alphafold3/blob/main/LICENSE), Boltz-2 brings a lot of that capability in an open, fast, and extendible way. That’s why we chose it. We’re excited to share that LiteFold’s structure prediction editor is powered by Boltz-2 under the hood. Try out today We are on a mission to make molecular and structural biology related experiments easier than ever. Whether you are doing research on protein design, drug design or want to run and organize your experiments, LiteFold helps to manage that with ease. Try out, it's free. [ Try out LiteFold ](https://app.litefold.in) ## What's this Structure Prediction Editor? The most common input for structure prediction is still FASTA files. You can define multiple chains: proteins, DNA/RNA, small molecules (via SMILES or CCD codes) and even point to precomputed MSAs. For most basic use cases, that’s enough. But if you're using the Boltz‑2 model, you're leaving a lot of power on the table by sticking to FASTA. Here's what you can not do in FASTA: 1. **No Modified Residues:** You can’t specify non-standard amino acids or modified nucleotides. This means you’re limited to plain sequences, which doesn’t reflect real biological scenarios where modifications are common (e.g., phosphoserine, methylated cytosine). 1. **No Covalent Bonds:** You can’t define explicit bonds between atoms, for example, a covalent link between a ligand and a residue. That’s essential if you’re modeling inhibitors or covalently bound cofactors. 1. **No Pocket Constraints:** You can’t specify where a ligand should bind, e.g., “this SMILES should interact with residues 35, 89, and 110 in chain A.” Boltz allows adding physical constraints to guide sampling around pockets. FASTA can’t capture this. 1. **No Affinity Prediction Setup:** You can’t flag which ligand you want the model to compute binding affinity for. That means Boltz won’t run its affinity head, and you miss out on one of its best features: rapid, structure-informed affinity estimates. ## Sign up for LiteFold LiteFold blogs are here to educate the community about AI for drug discovery related topics / structural and molecular biology. No fuss, just pure content for the love of the scientific community. Subscribe Email sent! Check your inbox to complete your signup. No spam. Unsubscribe anytime. Above customizations are something you can not do with FASTA file. However the Boltz community let's you to do this using YAML files. YAML files are very similar in nature like JSON files. However we understand that there is another learning curve to understand on how to use the YAML files to the fullest. So that's why we have created a simple editor, where you start with uploading all your FASTA files normally under the folders tab. After finishing the editor go to structure prediction (under the Labs tab) and then select the folder you are interested to predict the structures. You will be seeing a simple input box, here simply write what you want to edit. In the below example we wrote, add an ATP in chain A and calculate the binding affinity. ![Image](images/09b5bdb8af1a6777fd2d6d52c82c17460eee9308-834x540.gif) However you can write anything which you want to add and our AI will add that in the YAML configuration. In case our AI understand that it is not feasible, it will not do it. 😰Consider a scenario, where you have 100 fasta files, and you want to make edits to each of them. Think how inaccurate and in-efficient and time wasting would it be, to edit each of them, copy nucleotide sequences all manually. Whereas you can use our editor and do it simply on the fly without all the manual hassel. We are living in a times, where we can introduce edits to bio-molecular structures and how they should behave. Anyways, once done then, we can start doing structure predictions. ## Do not wait until predictions are finished But wait, again consider the scenario, where you have 1000 + structure to predict. Normally what happens is that you need to wait for each of them to finish and then start doing the downstream task. But this is an utter waste of time. What if, you could asses structure predictions in real time? With LiteFold's bulk structure predictions, we exactly do that. All you need to do is press the structure prediction button, and the predictions will start. Here is a examples of predicting 25+ structures in around 3-4 minutes. We also show how you filter it, see the structures etc. 0:00 /1:26 1× This video is intended to show a live demostration of bulk structure prediction using LiteFold.You can skip parts or speed it up to 2x to watch the full video faster. ### What happens under the hood If you are interested to learn how this has been operating under the hood, then feel free to read this section else you can skip this part. So under the hood, we have deployed Boltz-2 on [Modal](https://modal.com/). Modal is a serverless GPU provider. We perform efficient and scalable GPU inference there. So, once you start doing structure predictions, then we handle parallelization in both ways. One is the batched inference in the same GPU. We do dynamic batching based on the residue sequence length. And accordingly, we handle batch inference. Not only that, we also handle multi GPU parallelization, where we distribute, multiple FASTA files batches in multiple GPUs. So, to put it simply, suppose you are predicting 100 fasta files. Assume, we do predictions in a batch of 5 files. So this means, we have 20 such batches. So each device (containing one GPU) will be computing 5 files and we will be having 5 gpus. So in one go we are predicting 25 FASTA files, and in just 4 passes, we will complete our overall predictions. Visually how this looks like: ![Image](images/232302ab666857d2a41ad4c1060af8ef37ab984d-1000x533.png) Figure 5: Inside how parallel FASTA computation happens Inside each GPU, we compute structure prediction, affinity prediction (if mentioned in the input to compute) and also compute different types of metrics. Metrics are super important to asses and filter the quality of the structure. Let's discuss in the next section. ## Asses quality right after predictions Once structure predictions are finished, researchers anyways tend to either write scripts or use CLIs to understand, filter and asses structure quality. However in LiteFold, we have metrics already built in. Let's discuss each of the metrics which are available in LiteFold. ### Predicted Local Distance Difference Test (pLDDT) pLDDT estimates the confidence in the position of each residue in the predicted structure. It ranges from 0 to 100 (or 0 to 1 in normalized form), where higher values indicate higher confidence. For a given *i, *pLDDT estimates the expected distance deviation from the true structure over a small local window. ![Image](images/040cd3b528a243e5c212e51e66d75d79d817ec78-712x322.png) This helps visualize local model confidence per residue. Scores > 70 are generally considered trustworthy, while scores < 50 indicate low confidence. In LiteFold platform, we also color the residue structure based on the pLDDT score. ![Image](images/c34ff15b2425c542072d65b10718ef0f585fb666-423x472.png) Figure 7: Protein structure with varying level of pLDDT scores As you can see in the above picture, we have residues colored in different shades as follows: 1. > 90 (dark blue): very high confidence. 1. 70-90 (yellow - lightblue): confident 1. < 70 (orange - red): Low confidence (might be disordered or un-certain) These are super important, because sometimes, these folds could be associated with different binding sites / regions. So, if the confidence is not good, then the downstream follows could be subjected to change which will save a good time, because in that case you can reject the structure. ### Ramachandran Plot Validation Ramachandran plots show the distribution of backbone dihedral angles (ϕ, ψ) for residues, used to assess stereochemical quality. To put it simply, the plot helps us to visualize the allowed conformation of polypeptide backbones in the proteins. It helps to asses the quality of the protein structures and understand the preferred angles of rotations in the backbone. Proteins are made up of amino acid and each amino acid are linked by peptide bonds. So in a polypeptide, each amino acid has a: 1. A phi (ϕ) torison angle - rotation around N-Cα bond 1. A psi (ψ) torsion angle - rotation around the Cα–C bond ![Image](images/13cee812efb3da31c00094bce216cb972fc0e4c6-484x338.png) Figure: Phi and Psi angles in amino acid of a protein Each residue (except glycine and proline) in a protein has a specific (ϕ, ψ) combination. When you plot these, you get a map of allowed and disallowed regions based on steric hinderance. ![Image](images/9e7a098f2c472fed9841b9979d0d499e19ba2997-1000x589.png) Figure: Example Ramachandran Plot. [Source](https://www.peptideweb.com/ramachandran-plot-as-a-tool-for-peptide-and-protein-structures-quality-determination) So inside LiteFold, mathematically, for each residue *ri *we compute: ![Image](images/cc3d1498826ae74d440b377d6969ec9e1929d620-1000x142.png) Outliers are defined as residues where (ϕi​,ψi​) lie outside all allowed regions. The final metric is: ![Image](images/224fadb1adf96b49430a735205d4f33a70303142-1000x244.png) This percentage provides a quantitative measure of stereochemical quality — lower values indicate better structural geometry. ### Clash Score We compute the clash score to quantify steric overlaps (clashes) between atoms in a predicted protein structure. This follows [MolProbity standards](https://pmc.ncbi.nlm.nih.gov/articles/PMC5734394/), where a clash is defined as two non-bonded atoms being closer than the sum of their van der Waals (vdW) radii minus a threshold of 0.4 Å. Small allowances are made for potential hydrogen bonds. So, let ![Image](images/7668bc674fd73318c375da73f70aa0eb8f37a85a-555x307.png) This gives the number of steric clashes per 1000 atoms — lower values indicate more physically realistic and better-packed structures. ### Binding Affinity Prediction This metric is a model based metric. This means we use another deep learning model, in our case it is Boltz-2 Affinity to predict this score. This predicts how strongly the ligand binds to the protein expressed as log(IC50) or binding probability. The Boltz-2 Affinity model has two heads, the regression head predicts the log(IC5o [μM]), which represent the log scale inhibitory concentration (IC5o) in micromolar units. Lower values indicates stronger binding (nanomolar to picomolar) while higher values indicate weak or no binding. The binary classifier head predicts the probability that a ligand is a binder. This is useful for hit indentification where the goal is to distinguish binders from non-binders. So, now we can interpret binding strength in energetic terms, we can approximate Gibbs free energy of binding ΔG from the predicted log(IC₅₀) using: ![Image](images/35f6230c32aebfd6a26014409b61c1ea89d2f61e-810x224.png) We can even compute approximate IC₅₀ and ΔG values based on the predicted log(IC₅₀) values (denoted here as y) as shown in the table below. | Predicted y | IC₅₀ (μM) | Approx. ΔG (kcal/mol) | Binding Strength | | --- | --- | --- | --- | | -3 | 0.001 nM | ≈12.6 | Very strong binder | | 0 | 1 μM | ≈8.2 | Moderate binder | | 2 | 100 μM | ≈5.5 | Weak binder / decoy | Predicted yy IC₅₀ (μM) Approx. ΔG (kcal/mol) Binding Strength -3 0.001 nM ≈12.6 Very strong binder 0 1 μM ≈8.2 Moderate binder 2 100 μM ≈5.5 Weak binder / decoy ## Conclusion & What's coming next So, that's all about LiteFold's structure prediction module. We discussed about how you can predict structures with Boltz, can edit protein residues easily, predict structures at scale. We also take a peak at different metrics, we support, what they do and their significance. ### So what's coming next In the coming versions, we will be adding support for AlphaFold-2 model as well. AlphaFold-3 might be coming later, because that requires a different integration. We will also bring more metrics so more granular filtering and evaluating structure quality. In the end, we are making the AI-powered lab for early stage drug discovery and research. So, in the structure prediction in that respect, we will be bringing things like relaxing the structure, using molecular dynamics to see the stability of the predicted complex and many more. However we are not just limiting ourselves to structure predictions. More sophisticated pipelines like generating denovo molecules, molecular docking, molecular dynamics are all coming soon. Not to forget, we are also building Rosalind, our intelligent AI Co-Research assistant, which will help you to build multiple experiments faster than ever, so that you can get the results much faster and focus on doing the core research. We are super excited. Until next time. Try out today We are on a mission to make molecular and structural biology related experiments easier than ever. Whether you are doing research on protein design, drug design or want to run and organize your experiments, LiteFold helps to manage that with ease. Do try out, it's free. [ LiteFold Platform ](https://app.litefold.in) --- # Structural biology and AlphaFold URL: https://www.lite.bio/blogs/structural-biology-and-alphafold Author: Anindyadeep Date: 2025-07-30 Structural biology delves into fundamental biological structures such as proteins, DNA, and RNA, examining their behaviors, formations, and conformations. The integration of structural biology with AI has opened numerous avenues, including predicting 3D structures of proteins and simulating interactions, thereby advancing fields like [omics biology](https://en.wikipedia.org/wiki/Omics) in general and medicine 3.0, artificial drug discovery in particular. Among these challenges, protein folding—understanding how proteins fold and predicting the folding patterns of unknown proteins—stands out as a very fundamental problem. Accurate predictions can accelerate biological research by minimizing wet lab tests and escalating the usage of dry and soft research tools. DeepMind's AlphaFold models, particularly AlphaFold 2, have made groundbreaking strides in this area, achieving accuracies exceeding 90% in protein structure predictions. This achievement was recognized with the [2024 Nobel Prize in Chemistry awarded to Demis Hassabis](https://www.nobelprize.org/prizes/chemistry/2024/hassabis/facts/) and John Jumper of DeepMind, alongside David Baker of the University of Washington, for their pioneering work in AI-driven protein structure prediction and computational protein design. This blog post aims to help fellow readers understand the *significance* of secondary and tertiary structure of protein and folding in general. How does a protein fold? How does it even form in the first place? What’s the relationship between structures like DNA, RNA, and proteins—and why is this trio one of the *most* fundamental building blocks of what we call life? Tighten your seatbelts, because we’re about to uncover the answers to these questions one layer at a time. ## How do Proteins form Well, lot of us tech bros only care about proteins to gain good muscle lol. But if you think about it, every living organism right from tiniest bacteria to humans is fundamentally run by some macro and supra molecules like DNA, DNP (Deoxy Ribo Nucleo protein), RNP (Ribo Nucleo Protein), Enzymes etc. They do everything (like literally everything) from providing structure to cells, catalyzing reactions, transmitting signals, defending our body against invaders (like diverse pathogen, likeVirus) etc. In the end each of the proteins and how they work is just some set of massive chains of biochemical reactions. However they are fascinating. Now the question comes how and from where does protein originates? In one word, It all starts with DNA (the instruction manual of life). Let’s brush up our high school knowledge about cells and take a roller coaster ride right from cells to DNA. Try out today We are on a mission to make molecular and structural biology related experiments easier than ever. Whether you are doing research on protein design, drug design or want to run and organize your experiments, LiteFold helps to manage that with ease. Do try out, it's free. [ LiteFold Platform ](https://app.litefold.in) ## A small rollercoaster ride of Analogies Since I expect most readers here have very little background in biology, let’s start from the basics. You probably know this: every living organism is made up of **cells**. Now, inside a cell, we have the **nucleus**. Inside the nucleus, there are tiny openings called **nuclear pores**—think of them as selective gates that allow specific molecules to enter and exit. Within the nucleus, we find **chromosomes**. Chromosomes are essentially highly organized, tightly coiled structures made up of **DNA** wrapped around **histone proteins** (imagine threads wrapped around tiny spools). **DNA**, or deoxyribonucleic acid, is this long, twisted molecule (a double helix, if you remember) that carries all the genetic instructions needed to build and operate an organism. Specific sections of the DNA that actually do something meaningful are called **genes** (we’ll come back to this later). Within the nucleus, we find chromosomes. Chromosomes are essentially highly organized, tightly coiled structures made up of DNA wrapped around histone proteins (imagine threads wrapped around tiny spools). Fig 1: Image generated by ChatGPT-4o ## Sign up for LiteFold LiteFold blogs are here to educate the community about AI for drug discovery related topics / structural and molecular biology. No fuss, just pure content for the love of the scientific community. Subscribe Email sent! Check your inbox to complete your signup. No spam. Unsubscribe anytime. DNA, or deoxyribonucleic acid, is this long, twisted molecule (a double helix, if you remember) that carries all the genetic instructions needed to build and operate an organism. Specific sections of the DNA that actually *do* something meaningful are called genes (we’ll come back to this later). Now, programmatically, you can think of DNA as the main gateway to a massive, ancient codebase. Some parts of this codebase make perfect sense, while other parts are just commented-out, outdated, or gibberish lines of code that don’t seem useful at all. The meaningful, functional parts of the codebase are what we call **genes**. And here's another fun analogy: think of genes like **classes** in object-oriented programming. A class acts as a blueprint for creating objects; similarly, genes are blueprints for making proteins. Now, just like how we instantiate an object from a class, we *instantiate* something called **mRNA** from a gene. And just like how you use objects to perform different functions in a program, mRNA is used as the working copy to **build proteins**. The process of creating mRNA from DNA is called **transcription**, and the process of creating proteins from mRNA is called **translation**. ## Time to dive deeper to Transcription Now let’s dive deeper into how proteins form and fold. As we know, DNA is the starting point for creating proteins. DNA is massive in length, but most of it “does not make sense” in terms of coding for proteins. The sections of DNA that contain useful information are called genes (see the image below). Figure 2: Image generated from AI Above is a simplified structure of DNA. The four letters you see here—A, T, C, G—are called nucleotides. A stands for Adenine, T for Thymine, C for Cytosine, and G for Guanine. These are complex chemical structures where A always pairs with T (through two hydrogen bonds), and C always pairs with G (through three hydrogen bonds). The 3’ and 5’ marks indicate the direction of the strands. You can think of DNA like a coiled zip, where one strand (from 3' to 5') is the sense strand, and the other (from 5' to 3') is the antiparallel strand (behave as anitsense strand). These are useful as you will see in the next section. We begin with the process of transcription, which takes place inside the nucleus. During transcription, a protein called RNA polymerase binds to the DNA and unwinds it, causing the double helix to "unzip." RNA polymerase then starts “parsing” the DNA from the 3' to 5' direction. In the process, it synthesizes a single-stranded structure known as mRNA (messenger RNA). Fig 3: Animation of transcription. Source: Reddit post We begin with the process of transcription, which takes place inside the nucleus. During transcription, a protein called RNA polymerase binds to the DNA and unwinds it, causing the double helix to "unzip." RNA polymerase then starts “parsing” the DNA from the 3' to 5' direction. In the process, it synthesizes a single-stranded structure known as mRNA (messenger RNA). This mRNA is complementary to the DNA template strand and is synthesized in the 5' to 3' direction. As you can see in the gif above, the purple-yellow double coiled structure is DNA and a single coiled structure is the mRNA. The big-blob like structure is our RNA polymerase. One fun fact, it’s interesting on how RNA polymerase finds where to start and where to end. A fun fact: It's interesting how RNA polymerase knows where to start and stop. If you're familiar with tokenization in LLMs, it's similar to how we have special tokens like ` and `. In DNA, specific sequences called promoters signal where RNA polymerase should begin, and terminators mark the end, guiding the process just like special tokens in LLMs. mRNA or messenger RNA is a coiled single stranded structure. Just like DNA, RNA consists mainly of the four letters A, G, C, and U. Notice that instead of T (Thymine), we have U (Uracil), which is a key differentiator between RNA and DNA in terms of chemical structure. Similar to mRNA, another type of RNA called **tRNA** (transfer RNA) is also formed during transcription, and we'll explore its function later. The point is, each type of RNA has a unique structure and shape that enables it to perform specific functions. Before protein synthesis begins, there's an interesting step right after RNA is formed. The early transcribed RNAs, like mRNA and tRNA, are called nascent RNA, meaning they are unstable. When RNA polymerase parses the DNA, it doesn't distinguish between gene sequences and non-coding (junk) DNA. It simply makes a single pass. As a result, the mRNA that forms is nascent and includes both functional and non-functional parts. To fix this, a protein called spliceosome comes in and removes the non-coding regions (introns) of the mRNA. After this splicing, we have a fully functional mRNA (exons) that’s ready for protein synthesis. Figure 4: Simplified transcription Before moving forward, please note that: structures like mRNA or tRNA are not really a single thread like structure. Rather they are complex molecular structure which are itself 3D in shape (unlike all other molecules). However there shape matters a lot. For instance, here is how different RNA looks like. Figure 5: Different type of RNA shapes and structure. Source: [ResearchGate](https://www.researchgate.net/figure/Diversity-of-RNA-Types-and-Their-Functionalities-in-a-Cell-This-comprehensive-diagram_fig1_377437774) These are just some quite a handful of RNAs I have shown. However I assume you got the idea here. These shapes are very much significant which we will uncover more in the next section. Great job if you have understood till here now. If not, then I highly encourage you to [watch this 2 mins video](https://www.youtube.com/watch?v=gG7uCskUOrA) which simulates the process of transcription and translation. More on translation on the next section. ## Translation: Journey from RNA to Protein Once the RNA molecules are synthesized and processed, they need to exit the nucleus to reach the cytoplasm where protein synthesis occurs. This transport happens through specialized channels in the nuclear envelope called nuclear pores. Only properly processed RNA, especially mature mRNA, is allowed to pass through these pores, ensuring that only functional transcripts enter the cytoplasm. Inside the cytoplasm, the mature mRNA associates with Ribosomes (a Ribosome is composed of complex molecular structures like Ribosomal RNA and proteins). Pay attention to figure 5 from our previous section. Notice how mRNA looks like a single strand, tRNA looks like a hair pin (or clover leaf), rRNA looks like a blob where you can fit something. These are very important. Ribosome (made up with these rRNA) tries to “read” the sequences of the mRNA in set of three nucleotides called codons. Figure 6: Codons in mRNA Each codon will code for one amino acid (proteins are just a chain of amino acid). So picture this, the mRNA binds with the ribosome which is “parsing” mRNA in the set of 3 sequences each. On the other side tRNA brings the amino acid and sits on each codon. tRNA has an anticodon region which sits on top of mRNA codons (an anticodon is a codon but with complementary sequence. For example; if AUG is a codon then UAC is the anticodon). Figure 7: Simplified view of translation. So like this with each pass of “parsing” the mRNA sequence (set of 3), each amino acid gets linked with each other form a chain. And this will not stop till the stop codon is reached. Again you can think it like when a LLM reaches the stop token, then it no more generates a new token. Below is a simple GIF that shows how the process of translation takes place. Figure 8: Process of transcription visualized. [Resource](https://plantlet.org/translation-mrna-to-protein/) Awesome, that’s pretty much it. We now know how exactly transcription followed by translation works. As you can see that, chain of amino acids gets linked with each other like beads of necklace. However once the process of translation stops, the chain of amino acid goes out of the Ribosome. The initial nascent protein undergoes into lots of conformational changes which turns it into a three dimensional structure. We will discuss more about it in the next section. 🪴The entire central dogma of the life irrespective of hierarchy is regulated a lot more number of factors other than translation and transcription. However that is out of the scope of the current blog. ## Folding We are almost there. The linear nascent chain of amino acid formed just after translation is not functional yet. This current structure state is called primary structure of a protein. Beads of amino acid linked with peptide bonds. The folding starts to happen in many cases mostly driven by chemical properties like[ van-der-wals forces](https://en.wikipedia.org/wiki/Van_der_Waals_force), electrostatic forces etc. As soon as amino acid starts to fold it forms repeated patterns stabilized by hydrogen bonds. There are two such types of structure: 1. **Alpha helices (α-helices)**: Coiled structures like a spring 1. **Beta sheets (β-sheets)**: Flattened, sheet-like structures formed when segments of the chain align side by side. Figure 9: Process of folding visualized. These local structures gives the protein initial stability and helps to form the overall shape. This structure state is called secondary structure. Finally these secondary structure further folds into 3D globular shape due other various types of interactions like Hydrophobic interactions, Hydrogen bonds, Ionic bond etc. In the end this series of chemical interactions happens to make the large molecules undergo to a state of equilibrium. Very simply in AlphaFold, we will be given the series of amino acid (kinda like the primary structure) and we need to predict the tertiary structure out from it. ## Structure is everything In the end, **structure is everything**. It is the fundamental factor that determines how a complex system looks and functions. Think about it — **coal and diamond** are made of the exact same element, **carbon**, yet they have completely different physical properties. One is brittle, the other is incredibly hard. The difference lies purely in **how the carbon atoms are arranged structurally**. Similarly, in biology, biomolecules like **DNA, RNA, enzymes**, or any kind of **protein** must have the correct shape to function properly and interact with other molecules. For example, **enzymes** are often called **biochemical catalysts** because they speed up chemical reactions by lowering activation energy. But how exactly do they do it? The answer again lies in structure: every enzyme has a **specific three-dimensional shape** that allows it to **bind precisely** to its target molecule (substrate), facilitating the reaction. The substrate (i.e. our target molecule) fits into the active site like a key fits into a lock (also called as lock-and-key model). Once bound, the enzyme and the substrate undergoes a chemical reaction making a new product. ## AlphaFold2 As mentioned previously, for the very first time in history, the Nobel Prize has been awarded to an AI-native solution in the field of chemistry. That’s a huge deal. The prize was awarded for solving a decades-old problem, predicting the 3D structure of a protein from its amino acid sequence, a task known as the protein folding problem. The name of this groundbreaking project? AlphaFold2 by Google DeepMind. In simple terms, Given a sequence of amino acids (i.e., the primary structure), can we predict how it folds into its final 3D shape (i.e., the tertiary structure)? And not just "kind of guess it", but predict it with near-experimental accuracy — something biologists have dreamt of for decades. AlphaFold2 treats this as a geometry prediction problem guided by biological constraints. It leverages Transformers (yes, the same ones behind LLMs) to understand spatial relationships between amino acids and predicts pairwise distances and angles, iteratively refining them through a mechanism called recycling. It doesn’t simulate the physics directly like old-school methods; instead, it learns structure implicitly from huge datasets of known protein structures and alignments. At its core, AlphaFold2 is a highly intricate attention-based architecture that encodes both sequence and spatial relationships — making it a perfect blend of bio + geometry + deep learning. We will learn more about AlphaFold2 details in the coming blog posts. ## Conclusion Congratulations , you now have enough domain knowledge to not be afraid of tackling the protein folding problem. In our next blog post, we’re going to take a brief look at the current state of protein folding powered by deep learning. We’ll dive into some key bio-computational literature that will help us navigate more complex ideas with confidence. Along the way, we’ll explore different types of architectures that have contributed to breakthroughs in protein structure prediction.