Proteins Aren't Sentences: Why Bigger Protein Models Don't Win

Anindyadeep, Nabajit Borah ·

A protein language model (PLM) turns an amino-acid sequence into vectors that can be reused for downstream protein tasks: stability, localization, binding, mutation fitness, evolutionary structure, and more. In the frozen-embedding setting, the pretrained PLM is not fine-tuned. We extract a sequence embedding from it, train a small supervised probe on top, and ask how much useful biological signal was already present in the representation.

Several protein language models have been developed in recent years. Among the earliest, Meta (formerly Facebook) FAIR introduced the ESM-2 family of models. Since then, a growing number of PLMs, including: ESM2, ProGen, and DPLM, have demonstrated strong performance across various protein tasks.

Each of these models learns its representations in a fundamentally different way, and that difference propagates into downstream research and the decisions built on it. To make the comparison concrete, we benchmark the following models:

Model Parameters Objective family Readout direction Frozen embedding interface
ESM-C 300M 300M masked language model bidirectional mean-pooled encoder hidden state
DPLM 150M 150M discrete diffusion bidirectional teacher-forced hidden-state pooling
ESM-2 150M 150M masked language model bidirectional mean-pooled encoder hidden state
ProGen3 219M 219M autoregressive language model left-to-right teacher-forced hidden-state pooling
ProGen2-small 151M autoregressive language model left-to-right teacher-forced hidden-state pooling

BenchPLM

Protein sequences read a lot like text: residues behave like words and local motifs like phrases. That resemblance is what led researchers to borrow architectures from natural language processing.

That observation led researchers to borrow ideas from natural language processing and develop different architectures for learning representations of proteins. How a model reads a protein determines what its embeddings can represent. Masked models, causal models, diffusion models all see sequence context differently, so they inevitably learn different representations. However those variations naturally leads to the following questions:

  • Which protein language models learn the best representations, and why?

  • Which protein language model should be used for downstream tasks?

  • What kind of representations have these models learned, and how much evolutionary knowledge do they capture?

To answer these, we introduce PLMBench, a collection of 9 tasks designed to systematically evaluate protein language model representations.

Task Split size (train / validation / test) Primary metric What it probes
Thermostability Regression 5310 / 706 / 706 Spearman rho Sequence-level stability signal.
DeepLoc multiclass localization 10414 / 1368 / 1368 Accuracy Subcellular localization across multiple classes.
DeepLoc binary localization 6707 / 698 / 807 Accuracy A simpler binary localization ablation.
Metal ion binding 5797 / 719 / 719 Accuracy Compact function and binding-site signal.
FLIP2 Alpha Amylase 2574 / 644 / 488 Spearman rho Mutation-fitness ranking under a one-to-many split.
FLIP2 Hydrophobic Core 9974 / 2493 / 12468 Spearman rho Mutation-fitness ranking under a larger low-to-high shift.

Let’s understand the tasks in more details. In case if you are more interested to know the results, you can skip this section.

  1. Thermostability Regression: Given a protein sequence, predict its thermostability score. We use Spearman's ρ, which measures how well the model ranks more-stable proteins above less-stable ones.

  2. DeepLoc Tasks: Two classification tasks. In multi-class classification, we predict the subcellular localization of a protein (e.g., nucleus, cytoplasm, mitochondrion, membrane, extracellular). In binary classification, we predict whether a protein is membrane-bound or soluble.

  3. Metal Ion Binding: Predict whether a given protein sequence binds metal ions. We do not predict the specific metal or binding site.

  4. FLIP2 Alpha Amylase: Given mutant alpha-amylase sequences, the model ranks variants from worse to better function. We use Spearman's ρ, where higher values indicate better ranking accuracy.

  5. FLIP2 Hydrophobic Core: Similar to above, but using mutant sequences from the hydrophobic core dataset. The task is to predict mutation fitness or structural packing quality. As a regression task, we again use Spearman's ρ, where higher values indicate stronger predictive ranking.

Every model was evaluated on the same six tasks with the same train/validation/splits, the same embedding policy and the same shallow probe family. The protocol was:

  1. Freeze the pre-trained PLM backbone

  2. Extract sequence embeddings with sliding window mean pooling

  3. Train a probe on the training split.

  4. Select probe hyperparameters on the validation split.

  5. Report the held-out test metric.

One probe family is held fixed across all models, so every score reflects the representation and that probe jointly. We keep the probe shallow to stay close to the raw embedding, but absolute numbers would shift under a different probe; the cross-model ranking is the durable signal, not the decimal.

Interpretability

We also ran three diagnostic analyses. The first two are mechanistic probes; the third is a representation-geometry check. They are useful for interpretation, but they are not included in the benchmark score.

Diagnostic Models covered Question
Evolutionary layer localization ESM-2 scales, DPLM 150M, ProGen2-small At which layer is evolutionary distance most readable?
Context occlusion ESM-2 150M, DPLM 150M, ProGen2-small Does a residue-level score depend on left or right context?
Embedding PCA and same-family retrieval ESM-2 150M, DPLM 150M, ProGen2-small; larger cached check with ESM-2 650M, ESM-2 3B, ProGen2-medium Do frozen embeddings form clean ortholog-family neighborhoods?

ESM-C and DPLM Leads Overall

The overall result is close at the top. ESM-C 300M and DPLM 150M both have a mean task rank of 1.83 across the six tasks. ESM-C wins more individual tasks (3 versus 2), while DPLM has the marginally higher average normalized primary score (0.938 versus 0.924). With six tasks and a single probe seed, a gap this small is inside the noise floor; the honest reading is that ESM-C and DPLM are tied at the top, not that one edges the other.

ESM-2 150M is the next strongest model. Both ProGen checkpoints fall below the bidirectional models in this frozen-probe setting. ESM-C is also the largest backbone here, so this aggregate alone cannot separate architecture from scale. The same-size comparison does: ESM-2 and DPLM at 150M beat both ProGen checkpoints of equal or larger size, which isolates the effect we return to in the conclusion.

The per-task view explains why the aggregate is close:

Task Winner Winning score Runner up
Thermostability ESM-C 300M 0.664 Spearman ESM-2 150M at 0.645
DeepLoc multiclass ESM-C 300M 0.814 accuracy DPLM 150M at 0.812
DeepLoc binary ESM-2 150M 0.914 accuracy DPLM 150M at 0.912
Metal ion binding DPLM 150M 0.711 accuracy ESM-2 150M at 0.702
FLIP2 Alpha Amylase ESM-C 300M 0.676 Spearman DPLM 150M at 0.613
FLIP2 Hydrophobic Core DPLM 150M 0.371 Spearman ESM-C 300M at 0.355

ProGen3 remains usable on some stability and classification tasks, but it does not win any task in this suite. ProGen2-small is weaker across the benchmark, with the clearest gap on FLIP2 Alpha Amylase: 0.105 Spearman, compared with 0.676 for ESM-C and 0.613 for DPLM.

The pattern is clear: on downstream tasks, models that condition each residue on the full sequence outperform models restricted to a left-to-right prefix. This is not a verdict on the ProGen family.

ProGen models were trained to generate viable protein sequences. However the pattern is biologically plausible. Many protein properties depend on residues that are distant in sequence but coupled through structure, family constraints, or functional motifs. A frozen embedding that can integrate both upstream and downstream context gives a shallow probe more of the relevant signal.

Evolutionary Understanding

When Protein Language Models were trained, researchers saw that the models were learning evolutionary relationship. This is one reason ESMFold dropped the MSA: the language model was meant to stand in for it. The payoff is generalization to sequences with few or no homologs. In this section, we tried to quantify how much evolutionary understanding does different protein language models carries.

TimeTree

TimeTree is a database of species divergence times. Given two species, it tells you how many millions of years ago they last shared a common ancestor. So for three proteins drawn from species X, Y, and Z, TimeTree can say whether X branched off closer to Y or to Z.

This probe asks a single question: at what layer depth does a protein language model start encoding evolutionary history rather than surface sequence? For each model we sample a set of layers, pull the hidden states, mean-pool over residues to get one vector per protein, and check whether cosine distances between those vectors reproduce TimeTree's branching order.

The test set is built from identity-discordant triplets. Each triplet has three proteins:

  • A, the anchor

  • B, the history-correct neighbor (TimeTree says A and B diverged more recently)

  • C, the identity-misleading decoy (raw sequence identity says A looks more like C)

The two signals are deliberately put in conflict. By sequence identity alone, A resembles C more than B. By evolutionary history, A is closer to B. A model that only tracks surface similarity will be pulled toward C and fail. A model that has internalized deeper phylogenetic structure will still place B nearer.

Triplet accuracy is just the fraction of triplets where the embedding agrees with history:

distance(A, B) < distance(A, C)

Chance is 50%. Scores above that mean the layer carries evolutionary signal beyond raw identity; scores at or below mean it does not.

The same result is easier to inspect as a heatmap. ESM-2 and DPLM become more readable with depth, while ProGen2-small does not show the same late-layer gain.

Not only accuracy, we also ask another question: Across many protein pairs, do embedding distances increase as TimeTree divergence times increase? Now consider this table:

Model Objective Direction Best layer Best triplet acc Best Spearman rho
ESM-2 8M Masked LM Bidirectional 6 / 6 0.602 0.374
ESM-2 35M Masked LM Bidirectional 12 / 12 0.630 0.435
ESM-2 150M Masked LM Bidirectional 30 / 30 0.609 0.412
DPLM 150M Discrete diffusion Bidirectional 30 / 30 0.627 0.468
ProGen2-small Causal LM Left-to-right 0 / 12 0.460 0.158

Scale is not the axis here. ESM-2 35M matches 150M on this triplet probe, and DPLM 150M holds the strongest TimeTree-distance correlation despite no size advantage. What moves the signal is how the model reads context, not how many parameters it has.

For ESM-2 and DPLM, the signal gets stronger in deeper layers. That suggests the model gradually builds a more evolution-aware representation as sequence context is processed. For ProGen2, the best TimeTree signal was at the input/early layer, and deeper causal layers did not improve it, which supports the blog’s argument that left-to-right models are weaker at forming bidirectional evolutionary geometry.

Context Occlusion

To test whether bidirectional models genuinely use context from both directions, we run a simple occlusion experiment inspired by input-perturbation methods common in neural network interpretability.

The setup is straightforward. Pick a target residue and mask it. Record the model's log-probability for the correct amino acid at that position. Then mask one additional context residue at a known offset from the target and record the log-probability again. The difference between the two measurements tells us how much that context position contributed to the prediction. A large change means the context residue mattered; a small or zero change means the model was not relying on it.

For ESM-2 and DPLM, both masked language models, we mask the target and measure perturbations symmetrically on both sides. For ProGen2, a causal language model, we score the target from its prefix only, since by construction it never sees residues to the right.

The binned heatmap reports mean absolute log-probability change by relative position. Two patterns stand out.

  1. First, all three models show strongest sensitivity in the immediate neighborhood of the target (offsets -20 to +20), with effects decaying at greater distances. This is expected: local sequence context carries the strongest signal for residue identity.

  2. Second, and more telling, is the asymmetry. ESM-2 and DPLM show measurable sensitivity on both sides of the target, with the right-side effect at +20 (0.050 for ESM-2, 0.044 for DPLM) comparable to or exceeding the left-side effect at -20. ProGen2-small shows a sharp spike at -20 (0.104) and exactly zero for every right-side bin. The right side is blank because those residues simply do not exist in the causal prefix used to score the target.

The bar plot compresses the full heatmap into a single left-versus-right summary. For each model, we sum the absolute effects from all left-side bins and all right-side bins, then report the right-context fraction.

  • ESM-2 150M: left 0.021, right 0.026, right fraction 0.552

  • DPLM 150M: left 0.018, right 0.022, right fraction 0.556

  • ProGen2-small: left 0.031, right 0.000, right fraction 0.000

ESM-2 and DPLM split their context dependence nearly evenly across both sides, with a slight lean toward the right. ProGen2's entire context mass sits on the left, exactly as a causal architecture dictates.

Proteins are not sentences. In natural language, left-to-right context is a reasonable inductive bias because meaning largely builds sequentially. In proteins, the constraints on a residue's identity are fundamentally non-sequential. A residue may be constrained by a disulfide partner dozens of positions downstream, by residues it packs against in the folded structure, by compensatory mutations elsewhere in the sequence, or by a distant active-site motif it co-evolves with. None of these relationships respect sequence order.

The occlusion results confirm that ESM-2 and DPLM capture this: their representations at any given position are shaped by the full sequence context. ProGen2 behaves correctly for its architecture, but its representations at each position are blind to everything that follows. This asymmetry likely explains part of the performance gap we observe in downstream tasks that require whole-sequence understanding.

Do Ortholog Families Stay Together?

As a final representation-level check, we reused the same TimeTree panel containing 1,024 proteins from 128 ortholog families, with exactly eight sequences per family. For ESM-2 150M, DPLM 150M, and ProGen2-small, we used the cached final mean-pooled embeddings from the phylogenetic-geometry runs.

The visualization below fits PCA separately for each model after L2-normalizing the embeddings, matching the cosine geometry used in the TimeTree benchmark. Colored points mark eight evenly spaced ortholog families; gray points are the remaining families. The plotted colored points use a tiny deterministic display jitter so duplicated or near-identical PCA coordinates remain visible. The panels are qualitative and independently scaled, so the quantitative readout is the bar plot underneath: same-family retrieval across all 128 families.

Consider this table:

Source Same-family NN@1 Same-family P@7
Position identity 0.661 0.431
3-mer Jaccard 0.850 0.787
ESM-2 150M 0.855 0.798
DPLM 150M 0.859 0.820
ProGen2-small 0.641 0.397

DPLM and ESM-2 produce the cleanest same-family neighborhoods, with DPLM slightly ahead by precision@7. ProGen2-small still retrieves a same-family nearest neighbor well above random chance, but its local neighborhoods decay quickly: across the seven nearest non-self neighbors, fewer than half are from the same ortholog family.

We also repeated the same diagnostic for the larger cached checkpoints available on this panel: ESM-2 650M, ESM-2 3B, and ProGen2-medium as well.

The larger ESM-2 checkpoints improve family-neighborhood purity on this diagnostic, with ESM-2 3B reaching 0.875 P@7. ProGen2-medium does not show the same scaling behavior here: it is below the 3-mer baseline and below ProGen2-small by P@7. That result is local to this ortholog-family retrieval probe, but it reinforces the broader pattern that the causal ProGen embeddings are less useful as frozen neighborhood representations in these experiments.

The Curse of Sparse Data

Most of the real world protein engineering datasets are small. A useful embedding should not only perform well with thousand of labels; it should also be sample-efficient when labels are scarce. Measured variants could be in the terms of 20,100 or 700. To benchmark this, the setup is simple.

  1. We take a frozen PLM Embedding.

  2. Train a small regression probe on K-labeled examples.

  3. Predict fitness on held-out variants.

  4. Repeat this on different label budgets.

  5. Plot it as the curve and compare the differences.

Task Training budgets Test size used What it stresses
FLIP2 GB1 one-vs-rest 8 / 16 / 24 / 28 4096 Extremely low-data fitness transfer.
FLIP2 Alpha Amylase close-to-far 64 / 128 / 256 / 512 / 1024 / 1782 1924 Extrapolation from closer mutations to farther mutations.
FLIP2 Rhodopsin by-wild-type 32 / 64 / 128 / 256 / 512 / 700 184 Membrane-protein fitness transfer.

So this means, if a model gets a good Spearman score with only 16 or 64 labels, that means its embedding is sample-efficient: the probe does not need much supervision because the representation already carries useful signal.

The sparse-transfer result largely agrees with the full frozen benchmark. ESM-C has the best mean normalized curve score (0.771) and wins two of the three sparse tasks. ESM-2 is second by normalized curve score (0.691). DPLM is third overall (0.657) but is the strongest model on Rhodopsin at full budget.

Task Best frozen model Full-budget score Interpretation
GB1 one-vs-rest ESM-C 300M 0.535 Spearman ESM-C separates early, especially after 16 labels.
Alpha Amylase close-to-far ESM-C 300M 0.299 Spearman ESM-C is strongest under this mutation-distance shift.
Rhodopsin by-wild-type DPLM 150M 0.688 Spearman DPLM is the clearest winner on this membrane-protein fitness task.

The ProGen models do not close the frozen-transfer gap under sparse labels. ProGen3 improves on Rhodopsin as the budget grows, reaching 0.472 at 700 examples, but remains behind DPLM and ESM-2. ProGen2-small remains weak across the sparse suite. The limited-data conclusion is therefore aligned with our initial study: ESM-C is the strongest broad frozen representation, while DPLM deserves special attention on Rhodopsin-like fitness transfer.

Conclusion

Across every probe in this study, the most consistent predictor of frozen-representation quality was not parameter count or pretraining scale. It was the conditioning set the model reads. Bidirectional encoders, ESM-2 and ESM-C as masked models and DPLM as a diffusion model, beat the causal ProGen checkpoints on every task that rewards whole-sequence understanding: downstream probes, evolutionary geometry, and same-family retrieval alike. The reason is visible if we write down what each model computes. A causal model factorizes the sequence left to right:

$ p(x) = \prod_{i=1}^{L} p!\left(x_i \mid x_{<i}\right) $

so its representation for residue i depends only on the prefix x₍<ᵢ₎. A masked or diffusion encoder reconstructs each residue from both sides:

$ p(x_i \mid x_{\setminus i}), \qquad x_{\setminus i} = (x_1, \dots, x_{i-1}, x_{i+1}, \dots, x_L) $

so its embedding for residue i is conditioned on the entire rest of the chain. The occlusion experiment makes this concrete and measurable: ESM-2 and DPLM split their context dependence almost evenly across both sides of the target, with a right-context fraction near 0.55, while ProGen2 places exactly zero mass to the right. Protein constraints are non-sequential. A residue is shaped by disulfide partners, packing neighbors, and coevolving active-site motifs that sit anywhere in the chain. The terms a causal model drops from x₍<ᵢ₎ are precisely the structurally informative ones, and the weaker fitness regressions, the softer evolutionary signal, and the faster neighborhood decay all follow from that single fact. Scale does not change this; it moves you within it. ESM-C (300M) tops both the broad and the sparse benchmarks, but the clean isolation of the effect is the same-size comparison: 150M bidirectional models decisively outperform causal models of equal or larger size. Within a good architecture, scale helps, and ESM-2 3B reaches 0.875 family-retrieval precision. It cannot rescue a poor one, and ProGen2-medium falls below even the 3-mer baseline. On the hardest evolutionary triplet probe, scale is not reliable even inside ESM-2, where 35M matches 150M. Architecture sets the ceiling; size and data are how you climb toward it. For a frozen pipeline the practical reading is direct. Default to a bidirectional backbone: ESM-C as the broad workhorse, DPLM where membrane-protein or far-mutation fitness transfer matters, as on Rhodopsin. Do not use causal embeddings as frozen features. And consistent with Ko, Parkinson and Wang (2026) and the NbBench results, no single model wins everywhere, so fusing complementary bidirectional representations is the natural next step. Bigger models can become better models, but only within the limit set by how they read a protein sequence.

References

  1. Rao RM, Meier J, Sercu T, et al. Transformer protein language models are unsupervised structure learners. Presented at: International Conference on Learning Representations (ICLR); 2021. OpenReview: https://openreview.net/forum?id=fylclEqgvgd (https://openreview.net/forum?id=fylclEqgvgd) arXiv: https://arxiv.org/abs/2007.06225 (https://arxiv.org/abs/2007.06225)

  2. Raffel C, Shazeer N, Roberts A, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research. 2020;21(140):1-67. arXiv: https://arxiv.org/abs/1910.10683 (https://arxiv.org/abs/1910.10683)

  3. Wang X, Zheng Z, Ye F, Xue D, Huang S, Gu Q. Diffusion language models are versatile protein learners. Proceedings of the 41st International Conference on Machine Learning (ICML). PMLR. 2024;235:52309-52333. arXiv: https://arxiv.org/abs/2402.18567 (https://arxiv.org/abs/2402.18567)

  4. Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123-1130. DOI: https://doi.org/10.1126/science.ade2574 (https://doi.org/10.1126/science.ade2574) bioRxiv: https://www.biorxiv.org/content/10.1101/2022.07.20.500902v2 (https://www.biorxiv.org/content/10.1101/2022.07.20.500902v2)

  5. Ko YS, Parkinson J, Wang W. Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations. Briefings in Bioinformatics. 2026;27(1):bbag014. DOI: https://doi.org/10.1093/bib/bbag014 (https://doi.org/10.1093/bib/bbag014)

  6. Zhang Y, Tsuda K. NbBench: benchmarking language models for comprehensive nanobody tasks. Machine Learning: Science and Technology. 2025;6:040502. DOI: https://doi.org/10.1088/2632-2153/ae20ec (https://doi.org/10.1088/2632-2153/ae20ec) arXiv: https://arxiv.org/abs/2505.02022 (https://arxiv.org/abs/2505.02022)