Proteins Aren't Sentences: Why Bigger Protein Models Don't Win
Anindyadeep, Nabajit Borah ·
A protein language model (PLM) turns an amino-acid sequence into vectors that can be reused for downstream protein tasks: stability, localization, binding, mutation fitness, evolutionary structure, and more. In the frozen-embedding setting, the pretrained PLM is not fine-tuned. We extract a sequence embedding from it, train a small supervised probe on top, and ask how much useful biological signal was already present in the representation.
Several protein language models have been developed in recent years. Among the earliest, Meta (formerly Facebook) FAIR introduced the ESM-2 family of models. Since then, a growing number of PLMs, including: ESM2, ProGen, and DPLM, have demonstrated strong performance across various protein tasks.
Each of these models learns its representations in a fundamentally different way, and that difference propagates into downstream research and the decisions built on it. To make the comparison concrete, we benchmark the following models:
| Model | Parameters | Objective family | Readout direction | Frozen embedding interface |
|---|---|---|---|---|
| ESM-C 300M | 300M | masked language model | bidirectional | mean-pooled encoder hidden state |
| DPLM 150M | 150M | discrete diffusion | bidirectional | teacher-forced hidden-state pooling |
| ESM-2 150M | 150M | masked language model | bidirectional | mean-pooled encoder hidden state |
| ProGen3 219M | 219M | autoregressive language model | left-to-right | teacher-forced hidden-state pooling |
| ProGen2-small | 151M | autoregressive language model | left-to-right | teacher-forced hidden-state pooling |
BenchPLM
Protein sequences read a lot like text: residues behave like words and local motifs like phrases. That resemblance is what led researchers to borrow architectures from natural language processing.
That observation led researchers to borrow ideas from natural language processing and develop different architectures for learning representations of proteins. How a model reads a protein determines what its embeddings can represent. Masked models, causal models, diffusion models all see sequence context differently, so they inevitably learn different representations. However those variations naturally leads to the following questions:
-
Which protein language models learn the best representations, and why?
-
Which protein language model should be used for downstream tasks?
-
What kind of representations have these models learned, and how much evolutionary knowledge do they capture?
To answer these, we introduce PLMBench, a collection of 9 tasks designed to systematically evaluate protein language model representations.
| Task | Split size (train / validation / test) | Primary metric | What it probes |
|---|---|---|---|
| Thermostability Regression | 5310 / 706 / 706 | Spearman rho | Sequence-level stability signal. |
| DeepLoc multiclass localization | 10414 / 1368 / 1368 | Accuracy | Subcellular localization across multiple classes. |
| DeepLoc binary localization | 6707 / 698 / 807 | Accuracy | A simpler binary localization ablation. |
| Metal ion binding | 5797 / 719 / 719 | Accuracy | Compact function and binding-site signal. |
| FLIP2 Alpha Amylase | 2574 / 644 / 488 | Spearman rho | Mutation-fitness ranking under a one-to-many split. |
| FLIP2 Hydrophobic Core | 9974 / 2493 / 12468 | Spearman rho | Mutation-fitness ranking under a larger low-to-high shift. |
Let’s understand the tasks in more details. In case if you are more interested to know the results, you can skip this section.
-
Thermostability Regression: Given a protein sequence, predict its thermostability score. We use Spearman's ρ, which measures how well the model ranks more-stable proteins above less-stable ones.
-
DeepLoc Tasks: Two classification tasks. In multi-class classification, we predict the subcellular localization of a protein (e.g., nucleus, cytoplasm, mitochondrion, membrane, extracellular). In binary classification, we predict whether a protein is membrane-bound or soluble.
-
Metal Ion Binding: Predict whether a given protein sequence binds metal ions. We do not predict the specific metal or binding site.
-
FLIP2 Alpha Amylase: Given mutant alpha-amylase sequences, the model ranks variants from worse to better function. We use Spearman's ρ, where higher values indicate better ranking accuracy.
-
FLIP2 Hydrophobic Core: Similar to above, but using mutant sequences from the hydrophobic core dataset. The task is to predict mutation fitness or structural packing quality. As a regression task, we again use Spearman's ρ, where higher values indicate stronger predictive ranking.
Every model was evaluated on the same six tasks with the same train/validation/splits, the same embedding policy and the same shallow probe family. The protocol was:
-
Freeze the pre-trained PLM backbone
-
Extract sequence embeddings with sliding window mean pooling
-
Train a probe on the training split.
-
Select probe hyperparameters on the validation split.
-
Report the held-out test metric.
One probe family is held fixed across all models, so every score reflects the representation and that probe jointly. We keep the probe shallow to stay close to the raw embedding, but absolute numbers would shift under a different probe; the cross-model ranking is the durable signal, not the decimal.
Interpretability
We also ran three diagnostic analyses. The first two are mechanistic probes; the third is a representation-geometry check. They are useful for interpretation, but they are not included in the benchmark score.
| Diagnostic | Models covered | Question |
|---|---|---|
| Evolutionary layer localization | ESM-2 scales, DPLM 150M, ProGen2-small | At which layer is evolutionary distance most readable? |
| Context occlusion | ESM-2 150M, DPLM 150M, ProGen2-small | Does a residue-level score depend on left or right context? |
| Embedding PCA and same-family retrieval | ESM-2 150M, DPLM 150M, ProGen2-small; larger cached check with ESM-2 650M, ESM-2 3B, ProGen2-medium | Do frozen embeddings form clean ortholog-family neighborhoods? |
ESM-C and DPLM Leads Overall
The overall result is close at the top. ESM-C 300M and DPLM 150M both have a mean task rank of 1.83 across the six tasks. ESM-C wins more individual tasks (3 versus 2), while DPLM has the marginally higher average normalized primary score (0.938 versus 0.924). With six tasks and a single probe seed, a gap this small is inside the noise floor; the honest reading is that ESM-C and DPLM are tied at the top, not that one edges the other.
ESM-2 150M is the next strongest model. Both ProGen checkpoints fall below the bidirectional models in this frozen-probe setting. ESM-C is also the largest backbone here, so this aggregate alone cannot separate architecture from scale. The same-size comparison does: ESM-2 and DPLM at 150M beat both ProGen checkpoints of equal or larger size, which isolates the effect we return to in the conclusion.
The per-task view explains why the aggregate is close:
| Task | Winner | Winning score | Runner up |
|---|---|---|---|
| Thermostability | ESM-C 300M | 0.664 Spearman | ESM-2 150M at 0.645 |
| DeepLoc multiclass | ESM-C 300M | 0.814 accuracy | DPLM 150M at 0.812 |
| DeepLoc binary | ESM-2 150M | 0.914 accuracy | DPLM 150M at 0.912 |
| Metal ion binding | DPLM 150M | 0.711 accuracy | ESM-2 150M at 0.702 |
| FLIP2 Alpha Amylase | ESM-C 300M | 0.676 Spearman | DPLM 150M at 0.613 |
| FLIP2 Hydrophobic Core | DPLM 150M | 0.371 Spearman | ESM-C 300M at 0.355 |
ProGen3 remains usable on some stability and classification tasks, but it does not win any task in this suite. ProGen2-small is weaker across the benchmark, with the clearest gap on FLIP2 Alpha Amylase: 0.105 Spearman, compared with 0.676 for ESM-C and 0.613 for DPLM.
The pattern is clear: on downstream tasks, models that condition each residue on the full sequence outperform models restricted to a left-to-right prefix. This is not a verdict on the ProGen family.
ProGen models were trained to generate viable protein sequences. However the pattern is biologically plausible. Many protein properties depend on residues that are distant in sequence but coupled through structure, family constraints, or functional motifs. A frozen embedding that can integrate both upstream and downstream context gives a shallow probe more of the relevant signal.
Evolutionary Understanding
When Protein Language Models were trained, researchers saw that the models were learning evolutionary relationship. This is one reason ESMFold dropped the MSA: the language model was meant to stand in for it. The payoff is generalization to sequences with few or no homologs. In this section, we tried to quantify how much evolutionary understanding does different protein language models carries.
TimeTree
TimeTree is a database of species divergence times. Given two species, it tells you how many millions of years ago they last shared a common ancestor. So for three proteins drawn from species X, Y, and Z, TimeTree can say whether X branched off closer to Y or to Z.
This probe asks a single question: at what layer depth does a protein language model start encoding evolutionary history rather than surface sequence? For each model we sample a set of layers, pull the hidden states, mean-pool over residues to get one vector per protein, and check whether cosine distances between those vectors reproduce TimeTree's branching order.
The test set is built from identity-discordant triplets. Each triplet has three proteins:
-
A, the anchor
-
B, the history-correct neighbor (TimeTree says A and B diverged more recently)
-
C, the identity-misleading decoy (raw sequence identity says A looks more like C)
The two signals are deliberately put in conflict. By sequence identity alone, A resembles C more than B. By evolutionary history, A is closer to B. A model that only tracks surface similarity will be pulled toward C and fail. A model that has internalized deeper phylogenetic structure will still place B nearer.
Triplet accuracy is just the fraction of triplets where the embedding agrees with history:
distance(A, B) < distance(A, C)
Chance is 50%. Scores above that mean the layer carries evolutionary signal beyond raw identity; scores at or below mean it does not.
The same result is easier to inspect as a heatmap. ESM-2 and DPLM become more readable with depth, while ProGen2-small does not show the same late-layer gain.
Not only accuracy, we also ask another question: Across many protein pairs, do embedding distances increase as TimeTree divergence times increase? Now consider this table:
| Model | Objective | Direction | Best layer | Best triplet acc | Best Spearman rho |
|---|---|---|---|---|---|
| ESM-2 8M | Masked LM | Bidirectional | 6 / 6 | 0.602 | 0.374 |
| ESM-2 35M | Masked LM | Bidirectional | 12 / 12 | 0.630 | 0.435 |
| ESM-2 150M | Masked LM | Bidirectional | 30 / 30 | 0.609 | 0.412 |
| DPLM 150M | Discrete diffusion | Bidirectional | 30 / 30 | 0.627 | 0.468 |
| ProGen2-small | Causal LM | Left-to-right | 0 / 12 | 0.460 | 0.158 |
Scale is not the axis here. ESM-2 35M matches 150M on this triplet probe, and DPLM 150M holds the strongest TimeTree-distance correlation despite no size advantage. What moves the signal is how the model reads context, not how many parameters it has.
For ESM-2 and DPLM, the signal gets stronger in deeper layers. That suggests the model gradually builds a more evolution-aware representation as sequence context is processed. For ProGen2, the best TimeTree signal was at the input/early layer, and deeper causal layers did not improve it, which supports the blog’s argument that left-to-right models are weaker at forming bidirectional evolutionary geometry.
Context Occlusion
To test whether bidirectional models genuinely use context from both directions, we run a simple occlusion experiment inspired by input-perturbation methods common in neural network interpretability.
The setup is straightforward. Pick a target residue and mask it. Record the model's log-probability for the correct amino acid at that position. Then mask one additional context residue at a known offset from the target and record the log-probability again. The difference between the two measurements tells us how much that context position contributed to the prediction. A large change means the context residue mattered; a small or zero change means the model was not relying on it.
For ESM-2 and DPLM, both masked language models, we mask the target and measure perturbations symmetrically on both sides. For ProGen2, a causal language model, we score the target from its prefix only, since by construction it never sees residues to the right.
The binned heatmap reports mean absolute log-probability change by relative position. Two patterns stand out.
-
First, all three models show strongest sensitivity in the immediate neighborhood of the target (offsets -20 to +20), with effects decaying at greater distances. This is expected: local sequence context carries the strongest signal for residue identity.
-
Second, and more telling, is the asymmetry. ESM-2 and DPLM show measurable sensitivity on both sides of the target, with the right-side effect at +20 (0.050 for ESM-2, 0.044 for DPLM) comparable to or exceeding the left-side effect at -20. ProGen2-small shows a sharp spike at -20 (0.104) and exactly zero for every right-side bin. The right side is blank because those residues simply do not exist in the causal prefix used to score the target.
The bar plot compresses the full heatmap into a single left-versus-right summary. For each model, we sum the absolute effects from all left-side bins and all right-side bins, then report the right-context fraction.
-
ESM-2 150M: left 0.021, right 0.026, right fraction 0.552
-
DPLM 150M: left 0.018, right 0.022, right fraction 0.556
-
ProGen2-small: left 0.031, right 0.000, right fraction 0.000
ESM-2 and DPLM split their context dependence nearly evenly across both sides, with a slight lean toward the right. ProGen2's entire context mass sits on the left, exactly as a causal architecture dictates.
Proteins are not sentences. In natural language, left-to-right context is a reasonable inductive bias because meaning largely builds sequentially. In proteins, the constraints on a residue's identity are fundamentally non-sequential. A residue may be constrained by a disulfide partner dozens of positions downstream, by residues it packs against in the folded structure, by compensatory mutations elsewhere in the sequence, or by a distant active-site motif it co-evolves with. None of these relationships respect sequence order.
The occlusion results confirm that ESM-2 and DPLM capture this: their representations at any given position are shaped by the full sequence context. ProGen2 behaves correctly for its architecture, but its representations at each position are blind to everything that follows. This asymmetry likely explains part of the performance gap we observe in downstream tasks that require whole-sequence understanding.
Do Ortholog Families Stay Together?
As a final representation-level check, we reused the same TimeTree panel containing 1,024 proteins from 128 ortholog families, with exactly eight sequences per family. For ESM-2 150M, DPLM 150M, and ProGen2-small, we used the cached final mean-pooled embeddings from the phylogenetic-geometry runs.
The visualization below fits PCA separately for each model after L2-normalizing the embeddings, matching the cosine geometry used in the TimeTree benchmark. Colored points mark eight evenly spaced ortholog families; gray points are the remaining families. The plotted colored points use a tiny deterministic display jitter so duplicated or near-identical PCA coordinates remain visible. The panels are qualitative and independently scaled, so the quantitative readout is the bar plot underneath: same-family retrieval across all 128 families.
Consider this table:
| Source | Same-family NN@1 | Same-family P@7 |
|---|---|---|
| Position identity | 0.661 | 0.431 |
| 3-mer Jaccard | 0.850 | 0.787 |
| ESM-2 150M | 0.855 | 0.798 |
| DPLM 150M | 0.859 | 0.820 |
| ProGen2-small | 0.641 | 0.397 |
DPLM and ESM-2 produce the cleanest same-family neighborhoods, with DPLM slightly ahead by precision@7. ProGen2-small still retrieves a same-family nearest neighbor well above random chance, but its local neighborhoods decay quickly: across the seven nearest non-self neighbors, fewer than half are from the same ortholog family.
We also repeated the same diagnostic for the larger cached checkpoints available on this panel: ESM-2 650M, ESM-2 3B, and ProGen2-medium as well.
The larger ESM-2 checkpoints improve family-neighborhood purity on this diagnostic, with ESM-2 3B reaching 0.875 P@7. ProGen2-medium does not show the same scaling behavior here: it is below the 3-mer baseline and below ProGen2-small by P@7. That result is local to this ortholog-family retrieval probe, but it reinforces the broader pattern that the causal ProGen embeddings are less useful as frozen neighborhood representations in these experiments.
The Curse of Sparse Data
Most of the real world protein engineering datasets are small. A useful embedding should not only perform well with thousand of labels; it should also be sample-efficient when labels are scarce. Measured variants could be in the terms of 20,100 or 700. To benchmark this, the setup is simple.
-
We take a frozen PLM Embedding.
-
Train a small regression probe on K-labeled examples.
-
Predict fitness on held-out variants.
-
Repeat this on different label budgets.
-
Plot it as the curve and compare the differences.
| Task | Training budgets | Test size used | What it stresses |
|---|---|---|---|
| FLIP2 GB1 one-vs-rest | 8 / 16 / 24 / 28 | 4096 | Extremely low-data fitness transfer. |
| FLIP2 Alpha Amylase close-to-far | 64 / 128 / 256 / 512 / 1024 / 1782 | 1924 | Extrapolation from closer mutations to farther mutations. |
| FLIP2 Rhodopsin by-wild-type | 32 / 64 / 128 / 256 / 512 / 700 | 184 | Membrane-protein fitness transfer. |
So this means, if a model gets a good Spearman score with only 16 or 64 labels, that means its embedding is sample-efficient: the probe does not need much supervision because the representation already carries useful signal.
The sparse-transfer result largely agrees with the full frozen benchmark. ESM-C has the best mean normalized curve score (0.771) and wins two of the three sparse tasks. ESM-2 is second by normalized curve score (0.691). DPLM is third overall (0.657) but is the strongest model on Rhodopsin at full budget.
| Task | Best frozen model | Full-budget score | Interpretation |
|---|---|---|---|
| GB1 one-vs-rest | ESM-C 300M | 0.535 Spearman | ESM-C separates early, especially after 16 labels. |
| Alpha Amylase close-to-far | ESM-C 300M | 0.299 Spearman | ESM-C is strongest under this mutation-distance shift. |
| Rhodopsin by-wild-type | DPLM 150M | 0.688 Spearman | DPLM is the clearest winner on this membrane-protein fitness task. |
The ProGen models do not close the frozen-transfer gap under sparse labels. ProGen3 improves on Rhodopsin as the budget grows, reaching 0.472 at 700 examples, but remains behind DPLM and ESM-2. ProGen2-small remains weak across the sparse suite. The limited-data conclusion is therefore aligned with our initial study: ESM-C is the strongest broad frozen representation, while DPLM deserves special attention on Rhodopsin-like fitness transfer.
Conclusion
Across every probe in this study, the most consistent predictor of frozen-representation quality was not parameter count or pretraining scale. It was the conditioning set the model reads. Bidirectional encoders, ESM-2 and ESM-C as masked models and DPLM as a diffusion model, beat the causal ProGen checkpoints on every task that rewards whole-sequence understanding: downstream probes, evolutionary geometry, and same-family retrieval alike. The reason is visible if we write down what each model computes. A causal model factorizes the sequence left to right:
$ p(x) = \prod_{i=1}^{L} p!\left(x_i \mid x_{<i}\right) $
so its representation for residue i depends only on the prefix x₍<ᵢ₎. A masked or diffusion encoder reconstructs each residue from both sides:
$ p(x_i \mid x_{\setminus i}), \qquad x_{\setminus i} = (x_1, \dots, x_{i-1}, x_{i+1}, \dots, x_L) $
so its embedding for residue i is conditioned on the entire rest of the chain. The occlusion experiment makes this concrete and measurable: ESM-2 and DPLM split their context dependence almost evenly across both sides of the target, with a right-context fraction near 0.55, while ProGen2 places exactly zero mass to the right. Protein constraints are non-sequential. A residue is shaped by disulfide partners, packing neighbors, and coevolving active-site motifs that sit anywhere in the chain. The terms a causal model drops from x₍<ᵢ₎ are precisely the structurally informative ones, and the weaker fitness regressions, the softer evolutionary signal, and the faster neighborhood decay all follow from that single fact. Scale does not change this; it moves you within it. ESM-C (300M) tops both the broad and the sparse benchmarks, but the clean isolation of the effect is the same-size comparison: 150M bidirectional models decisively outperform causal models of equal or larger size. Within a good architecture, scale helps, and ESM-2 3B reaches 0.875 family-retrieval precision. It cannot rescue a poor one, and ProGen2-medium falls below even the 3-mer baseline. On the hardest evolutionary triplet probe, scale is not reliable even inside ESM-2, where 35M matches 150M. Architecture sets the ceiling; size and data are how you climb toward it. For a frozen pipeline the practical reading is direct. Default to a bidirectional backbone: ESM-C as the broad workhorse, DPLM where membrane-protein or far-mutation fitness transfer matters, as on Rhodopsin. Do not use causal embeddings as frozen features. And consistent with Ko, Parkinson and Wang (2026) and the NbBench results, no single model wins everywhere, so fusing complementary bidirectional representations is the natural next step. Bigger models can become better models, but only within the limit set by how they read a protein sequence.
References
-
Rao RM, Meier J, Sercu T, et al. Transformer protein language models are unsupervised structure learners. Presented at: International Conference on Learning Representations (ICLR); 2021. OpenReview: https://openreview.net/forum?id=fylclEqgvgd (https://openreview.net/forum?id=fylclEqgvgd) arXiv: https://arxiv.org/abs/2007.06225 (https://arxiv.org/abs/2007.06225)
-
Raffel C, Shazeer N, Roberts A, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research. 2020;21(140):1-67. arXiv: https://arxiv.org/abs/1910.10683 (https://arxiv.org/abs/1910.10683)
-
Wang X, Zheng Z, Ye F, Xue D, Huang S, Gu Q. Diffusion language models are versatile protein learners. Proceedings of the 41st International Conference on Machine Learning (ICML). PMLR. 2024;235:52309-52333. arXiv: https://arxiv.org/abs/2402.18567 (https://arxiv.org/abs/2402.18567)
-
Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123-1130. DOI: https://doi.org/10.1126/science.ade2574 (https://doi.org/10.1126/science.ade2574) bioRxiv: https://www.biorxiv.org/content/10.1101/2022.07.20.500902v2 (https://www.biorxiv.org/content/10.1101/2022.07.20.500902v2)
-
Ko YS, Parkinson J, Wang W. Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations. Briefings in Bioinformatics. 2026;27(1):bbag014. DOI: https://doi.org/10.1093/bib/bbag014 (https://doi.org/10.1093/bib/bbag014)
-
Zhang Y, Tsuda K. NbBench: benchmarking language models for comprehensive nanobody tasks. Machine Learning: Science and Technology. 2025;6:040502. DOI: https://doi.org/10.1088/2632-2153/ae20ec (https://doi.org/10.1088/2632-2153/ae20ec) arXiv: https://arxiv.org/abs/2505.02022 (https://arxiv.org/abs/2505.02022)