AminoWeb: Crystallizing the Web for the Finest Protein Data at Scale
LiteFold Research ·
AminoWeb at a glance. Metagenomic and predicted-structure corpora dominate the footprint, while functional and task datasets are small but high-value.
Protein machine learning is no longer limited by model architecture alone. It is limited, just as often, by whether the data means what we think it means.
Today, LiteFold is open-sourcing AminoWeb, a 29-dataset collection of cleaned, standardized, ML-ready protein data on Hugging Face. The release totals approximately 8 TB of Parquet-backed data and spans the major data regimes used in modern protein science: sequence corpora, experimental and predicted structures, functional annotations, evolutionary alignments, variant-effect assays, stability measurements, binding datasets, and specialized peptide or ligand resources.
Proteins are not documents. A sequence can be redundant with a thousand near-neighbors and still fail in a new organism. A structure can have excellent local geometry and still be the wrong biological assembly. A mutation can look beneficial in one assay and deleterious in another because the assays measure different phenotypes. AminoWeb is built around those biological constraints.
Every dataset in the collection is converted into a typed, sharded Parquet format; records are deduplicated where deduplication is scientifically appropriate; splits are designed to reduce homology leakage; and the processing choices are documented so that users can decide whether a filter is right for their question.
Explore AminoWeb: huggingface.co/LiteFold
Why we need a collection like AminoWeb
Public protein datasets are abundant, but abundance is not the same thing as readiness.
Datasets like UniProt is not for training loops. PDB entries are stored as mmCIF structures with metadata that require domain-specific parsing. Metagenomic proteins live across FTP mirrors and project-specific releases. Deep mutational scanning studies are often attached to papers as supplementary files with assay-specific score conventions. Stability and binding datasets use different units, mutation notations, and sign conventions. Hence this data fragmentation leads us to the following consequences.
- It wastes time: A lab that wants to train a protein language model, evaluate a variant-effect predictor, or build a structure-conditioned design model often has to spend weeks reconstructing the same parsers and filters before asking a biological question.
- It weakens comparisons: A model trained on one version of UniProt and tested on a random PDB split is not directly comparable to a model trained on another release and tested on a homology-aware split. Without shared data provenance, "state of the art" can become a statement about preprocessing rather than biology.
AminoWeb is our attempt to make the curation layer explicit.
What is in AminoWeb
AminoWeb: 29 datasets, ~7.56 TB of curated protein data
AminoWeb groups 29 datasets into four broad biological use cases.
| Category | Examples | Biological role |
|---|---|---|
| Foundation sequence and evolutionary data | UniProtKB, UniRef50, UniRef90, BFD, MGnify, NCBI, OpenProteinSet, Evolutionary | Learn protein grammar, family structure, taxonomic diversity, and co-evolutionary signal |
| Structure and complex data | PDB, AlphaFoldDB, ESMAtlas, CATH, PDB-CCD, protenix-data | Learn 3D geometry, assemblies, ligand context, structural classes, and predicted-structure confidence |
| Function and ontology data | GO, GOA, InterPro, Pfam, STRING, IEDB, Human Protein Atlas | Connect sequences and structures to molecular function, cellular context, immune recognition, and interaction networks |
| Experimental task datasets | ProteinGym, FLIP2, MegaScale-Tsuboyama2023, FireProtDB, SKEMPI2, CycPepMPDB, DisProt, SwissSidechain | Fine-tune and evaluate on mutation effects, stability, binding, disorder, permeability, and noncanonical chemistry |
How we curated the data
Different data types fail in different ways, so we filter them differently.
Sequences
We drop fragments, extreme-length outliers, internal stops, entries with high ambiguous-residue content, and records the upstream database already flags as incomplete. Canonical and noncanonical residues are validated separately. For UniProtKB, Swiss-Prot (manually curated) and TrEMBL (auto-annotated) stay in separate columns rather than flattened together. Same treatment for Pfam A vs B and InterPro member databases. Taxonomic source is a column too, so users can sample intentionally instead of inheriting upstream organism bias.
Functional annotations
Over 95% of GO annotations are IEA, inferred electronically. Training on IEA-heavy data risks circular annotation propagation. We expose the evidence code per record and ship two precomputed views: a high-confidence view (EXP, IDA, IPI, IMP, IGI, IEP, plus curator-reviewed IBA), and a full-coverage view that retains IEA.
Structure
We preserve method, resolution, assembly metadata, chain composition, ligand context, and confidence fields. Predicted structures from AlphaFoldDB and ESMAtlas are not interchangeable with experimental PDB. pLDDT and PAE report prediction confidence, not experimental accuracy, and the systematic biases go further than calibration: AFDB extends low-pLDDT regions through what should be labeled disorder, ESMAtlas degrades on orphan metagenomic sequences, and neither captures bound states or conformational ensembles. pLDDT, PAE summaries, AFDB model-version tags, and a disorder-content estimate are all surfaced as columns.
We also don't deduplicate PDB by sequence alone. The same protein can matter biologically across ligand-bound states, conformations, and oligomeric assemblies. Sequence-cluster columns are exposed so users can split however the task demands.
Mutation-effect data
Raw scores are not comparable across assays. ProteinGym mixes DMS and clinical annotations across many proteins. MegaScale measures folding stability, not organismal fitness. SKEMPI2 reports binding-affinity changes for protein-protein complexes. FireProtDB and MegaScale both concern stability but differ in scale and assay design. The schema keeps these distinctions instead of hiding them behind one "fitness" column.
The curation pipeline
Every dataset moves through the same pipeline, regardless of upstream format.
- Start with the raw primary sources: UniProt, RCSB/wwPDB, EBI (AlphaFoldDB, MGnify), NCBI, and original benchmark releases. Source and release metadata stored per record.
- Parse to typed Parquet: FASTA, mmCIF, OBO, XML, Stockholm, A3M, CSV, TSV, JSON, and paper supplementary files all materialized into structured records. No archival format parsing inside training loops.
- Normalize IDs without erasing provenance: Every record keeps its native ID alongside normalized cross-references (UniProt accessions, PDB IDs, taxonomy, source tags, release dates).
- Domain-specific filters: Sequence, structure, annotation, and assay datasets fail differently. Fragments, low-resolution structures, IEA annotations, unparsable mutation strings, and ambiguous sign conventions are not the same problem.
- Careful deduplication: Sequence corpora use exact-match and identity clustering. Structural datasets keep clustering columns instead of collapsing entries, since sequence-level dedup would strip meaningful conformational and ligand-state information.
- Homology-aware splits: MMseqs2 easy-cluster at 30% identity, 80% coverage for sequence-only tasks; CATH topology-level splits for fold-generalization; 40% sequence-identity entity clusters for PDB complexes; original published splits for ProteinGym, FLIP, and other benchmarks. Same-protein variants are split at the protein level. Exact parameters published per dataset card.
- Streamable storage: All large datasets sharded for streaming through Hugging Face datasets, Arrow, Polars, Dask, and pandas.
- Release-date stamps: Every record carries an upstream release date (UniProt release tag, PDB deposition date, AlphaFoldDB model version). Lets users build date-aware splits matching CASP, CAMEO, or custom cutoffs, avoiding one of the most common silent leakage failures.
Case study 1: stability is a measured thermodynamic property, not a generic score
MegaScale-Tsuboyama2023 is one of the clearest examples of why biological curation matters. The original Tsuboyama et al. study introduced cDNA display proteolysis and used it to measure proteolytic resistance as a stability proxy at very large scale across natural and designed small protein domains, roughly 40 to 72 residues in length. That scope is important: the dataset captures folding stability for small globular domains assayed by proteolysis, not for full-length therapeutic proteins, membrane proteins, or large multidomain enzymes.
The dashboard analysis covers 540,759 single-substitution measurements across 376 domains. For clarity in AminoWeb visualizations, we report ΔΔG using ΔG of unfolding (positive when the folded state is favored), so:
ΔΔG = ΔG_unfold(WT) − ΔG_unfold(mutant), in kcal/mol
Under this convention, positive ΔΔG means the mutant is less stable than wild type; negative ΔΔG means the mutant is stabilizing. The Tsuboyama source release reports ΔG of folding (negative when folded is favored), so we apply an explicit sign flip during ingestion. The convention is encoded in the column docstring of the processed Parquet so downstream joins with FireProtDB or SKEMPI2 do not silently invert outcomes.
Most single substitutions are mildly destabilizing, which is exactly what a biochemist would expect. The distribution also gives a useful sanity check for the pipeline: introduced proline has a median ΔΔG of about +1.99 kcal/mol, and roughly 92% of introduced-proline substitutions are destabilizing in this processed convention.
The value of this dataset is not only the global histogram. It is the ability to look at a protein domain position by position and ask which substitutions are tolerated.
This 4G3O domain heatmap shows the core idea: each column is a position, each row is an introduced amino acid, and each cell is the mean ΔΔG for that substitution. Red bands identify substitutions and positions that tend to destabilize the fold; blue cells mark substitutions that appear stabilizing under the assay. Missing cysteine coverage is not filled in by interpolation.
Aggregating across domains recovers familiar substitution chemistry. Introducing proline is broadly disruptive; substitutions involving bulky or charged residues show context-dependent effects; and self-substitutions sit near zero, providing a measured noise floor and per-domain reproducibility estimate rather than a structural artifact.
This is the kind of figure AminoWeb is designed to make routine. A user should not need to rediscover the sign convention, parse wide stability tables, and collapse duplicate measurements before getting to the biological question.
Case study 2: variant-effect assays are local maps, not one universal score
ProteinGym is a benchmark suite for protein fitness prediction and design. It brings together deep mutational scanning assays and clinical variant annotations, including substitutions and indels. In the AminoWeb processing snapshot analyzed here, the normalized table contains 2,931,539 rows across 281 assays. Broken down by variant type: roughly 0.82M clean single substitutions, 0.55M multi-substitution rows in the 2 to 10 range, 0.13M higher-order combinations, plus indel and clinical records that require separate evaluation protocols. About 1.04 million rows carry a usable score under our parser; the remainder are wild-type controls, unparsable mutation strings, or records held back for benchmark-defined splits. Users should pick the variant-type subset that matches the model class they are evaluating.
But ProteinGym is not a single homogeneous experiment. A viral replication assay, an enzyme-activity screen, a binding assay, and a clinical label set should not be placed on the same raw numerical axis.
This distribution is a useful warning. Single-substitution heatmaps are only one slice of ProteinGym. Multi-mutants, indels, missing mutation strings, and clinical records require separate treatment. For many biological questions, those records are the interesting part; for a position-by-amino-acid heatmap, they are outside the valid input format.
When a clean single-substitution assay is available, the map is powerful, but a single heatmap can hide what assay-specific scales actually mean. The strongest illustration is to plot two different assays for the same protein side by side. For proteins in ProteinGym that have both an abundance readout and an activity readout, the within-assay z-scored landscapes diverge at many positions: a variant that is well tolerated for abundance can be strongly depleted for activity. The MTHFR map below uses within-assay robust z-score (red below median, teal above) and is shown as a single example. The substantive point is that a teal cell in MTHFR is not biologically equivalent to a teal cell in a different assay, even if the colors match.
This is the level at which ProteinGym is most biologically honest: as a collection of assay-specific landscapes that can be normalized carefully, evaluated consistently, and compared with their limitations in view.
Case study 3: structure datasets need complexes, chemistry, and evolutionary context
The protenix-data portion of AminoWeb covers processed structural-training artifacts at large scale: PDB-derived complex features, MSA files, entity clusters, chemical-component context, and parsed structure summaries. The local dashboard artifacts analyzed here report approximately 1.05 TiB of data across 931,270 files, with 170,123 unique PDB IDs, 157,865 MSA sequence records, 32,448 chemical-component rows, and 369,163 entity clusters.
The first biological point is simple: these are not mostly isolated single chains.
Protein-ligand structures dominate, followed by protein-only multimers, protein-ion complexes, monomers, glycans, and protein-nucleotide complexes. The categories are heuristic because ligand identifiers in PDB-style data mix true ligands, cofactors, crystallization additives, glycans, metals, and modified residues. That caveat is not a nuisance; it is part of the biology.
The structure dashboard also includes exploratory feature clusters derived from metadata and parsed structural features. These clusters are useful for atlas building, but they are not CATH classes, Pfam families, or a biological taxonomy.
The value of the cluster map is pragmatic. It separates regimes such as small high-resolution protein-ligand structures, monomers, ion-rich complexes, glycan-containing entries, protein-nucleotide structures, and large assemblies. For model training, those regimes matter because they stress different parts of the representation: local chemistry, interfaces, chain pairing, token budget, and assembly complexity.
Evolutionary evidence is similarly heterogeneous. MSA file size and homolog count are useful diagnostics, but they are not independent proof that a sequence has rich evolutionary signal. The Protenix MSA landscape separates broad homolog clouds, source-specific homolog sets, and paired/concatenated signal patterns.
This matters for biologists because MSA quality is not a backend detail. Co-evolutionary signal can help structure prediction, but a deep alignment dominated by near-duplicates or source-specific homologs is not the same thing as a diverse, well-paired family alignment. To make the distinction filterable rather than visual, every MSA-bearing record in AminoWeb carries an N_eff column computed at 80% sequence identity (the AlphaFold convention), alongside raw depth, source breakdown, and pairing strategy for complexes. Users can filter for N_eff above a chosen threshold instead of trusting raw depth.
Redundancy control is another place where the structural data layer must be explicit. The 40% entity-cluster histogram shows that most clusters are singletons, while a smaller number contain many related entities. We chose 40% identity to match the upstream Protenix release for direct comparability, but this threshold sits inside the twilight zone for sequence-based homology detection, which means "singleton" at 40% does not imply structural independence. For stricter splits we additionally expose 30% identity and CATH-topology-level cluster columns. Users running fold-generalization benchmarks should prefer those.
Those clusters help users reason about leakage, split design, and overrepresented structural families. They do not solve every leakage problem automatically, but they make the problem visible.
Finally, structural biology is chemical biology. Models trained only on polymer coordinates miss a large part of what makes proteins function in cells and assays.
The processed component counts mix two very different categories: chemistry that determines function (heme, ADP, FAD, ATP, NAD, structural zinc, catalytic magnesium, biologically attached N-glycans) and chemistry that reflects how the structure was solved (chloride and sodium from buffer, magnesium from crystallization conditions, crystallographic additives). AminoWeb annotates each component with the BioLiP "biologically relevant ligand" flag and the PDB Chemical Component Dictionary subset class, so users can filter to functional ligands rather than reading the raw frequency chart as ligand pharmacology.
How to use AminoWeb
The simplest path is Hugging Face datasets:
from datasets import load_dataset
# Stream UniProtKB without downloading the full dataset locally.
uniprot = load_dataset("LiteFold/UniProtKB", split="train", streaming=True)
for record in uniprot.take(5):
print(record["accession"], len(record["sequence"]))
For task datasets:
from datasets import load_dataset
pg = load_dataset("LiteFold/ProteinGym", split="train")
mega = load_dataset("LiteFold/MegaScale-Tsuboyama2023", split="train")
For tabular analysis:
import polars as pl
pdb = pl.scan_parquet("hf://datasets/LiteFold/PDB/data/train-*.parquet")
high_res = pdb.filter(pl.col("resolution") <= 2.5).collect()
And because records carry normalized identifiers where possible, cross-dataset analysis becomes ordinary data work rather than custom plumbing:
import polars as pl
uniprot = pl.scan_parquet("hf://datasets/LiteFold/UniProtKB/data/train-*.parquet")
alphafold = pl.scan_parquet("hf://datasets/LiteFold/AlphaFoldDB/data/train-*.parquet")
joined = uniprot.join(alphafold, on="uniprot_accession", how="inner")
What you can build
AminoWeb supports several classes of biological modeling work:
- Protein language models trained on UniRef, UniProtKB, BFD, MGnify, or NCBI-scale sequence corpora.
- Structure-conditioned models that combine PDB, AlphaFoldDB, ESMAtlas, CATH, PDB-CCD, and protenix-data.
- Variant-effect predictors evaluated on ProteinGym, FLIP2, and assay-specific DMS landscapes.
- Stability models trained on MegaScale, FireProtDB, and related thermodynamic datasets.
- Binding and interface models using SKEMPI2, STRING, PDB complexes, and ligand-rich structure data.
- Function prediction models using GO, GOA, InterPro, Pfam, Human Protein Atlas, and IEDB.
- MSA-aware or MSA-free models using OpenProteinSet and LiteFold's Evolutionary dataset.
The common thread is reproducibility. A model trained on AminoWeb should be easier to audit because the data sources, filters, splits, and schema choices are visible.
Why open this
Open protein data already exists, but the usable layer is still uneven. The field needs shared curation, not just shared downloads.
If AminoWeb works, a structural biologist should be able to inspect ligand-rich PDB subsets without writing an mmCIF parser. A protein engineer should be able to compare stability and DMS measurements without guessing mutation notation. A model developer should be able to train on UniRef50 and evaluate on homology-aware splits without silently leaking near-identical proteins. A wet-lab group should be able to ask whether a model's claimed improvement survives a fair dataset boundary.
That is the standard we are trying to move toward.
AminoWeb is a foundation release with explicit gaps. Several resources that fit our scope are not yet included and are on the next-release roadmap: SAbDab for antibody structures, SCOP/SCOPe as a complement to CATH, DSSP-derived secondary structure, BioLiP for biologically relevant ligand annotations, PDBBind and BindingDB for protein-ligand affinity, BioGRID and IntAct for interaction data beyond STRING, and richer peptide bioactivity resources (DBAASP, APD3) beyond CycPepMPDB and SwissSidechain. Alongside those, we are working on precomputed embeddings, materialized cross-dataset join tables, additional high-throughput stability and variant datasets, and smaller high-quality subsets for compute-constrained training.
The data are here. The caveats are here too. Use both.
Explore AminoWeb: huggingface.co/LiteFold
Licensing
AminoWeb does not relicense upstream data. Each dataset retains its source license: UniProt and AlphaFoldDB under CC BY 4.0, PDB coordinates under CC0 with mixed terms for derived metadata, ProteinGym constituents under their respective per-paper licenses, and benchmark releases under their original terms. The processing scripts and AminoWeb-specific metadata are released under Apache 2.0. Users adopting AminoWeb for commercial training should review the per-dataset license fields exposed in each dataset card.
References
- Tsuboyama, K. et al. Mega-scale experimental analysis of protein folding stability in biology and design. Nature 620, 434-444 (2023). https://doi.org/10.1038/s41586-023-06328-6
- Notin, P. et al. ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design. NeurIPS Datasets and Benchmarks (2023). https://www.proteingym.org/
- ByteDance Protenix authors. Protenix: Advancing Structure Prediction Through a Comprehensive AlphaFold3 Reproduction. bioRxiv (2025). https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1
Citation
@misc{litefold2026aminoweb,
title={AminoWeb: A Curated Protein-Data Atlas for Biology and Protein Machine Learning},
author={LiteFold Team},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/LiteFold}
}