Overview#
This guide provides usage examples for data cleaning modules organized by dataset source:
Human Domainome Dataset: Site-saturation mutagenesis of 500 human protein domains.
ProteinGym DMS Substitutions Dataset: Large-scale benchmarks for protein design and fitness prediction.
cDNA Proteolysis Dataset: Mega-scale experimental analysis of protein folding stability in biology and design.
ddG-dTm Dataset: A collection of datasets providing single- and multiple-mutant measurements, labeled by the thermodynamic parameter (ΔΔG or ΔTm).
ArchStabMS1E10 Epistasis Dataset: High-order multi-mutant libraries (“1e10”) measuring protein stability for GRB2-SH3 and SRC.
Antitoxin ParD3 Epistasis Dataset: The antitoxin ParD3 3-position library is a combinatorially exhaustive dataset of 8,000 variants demonstrating that simple, independent per-residue mutation preferences are sufficient to almost perfectly predict combinatorial protein fitness.
TrpB Epistasis Dataset: A combinatorially complete sequence-fitness landscape comprising 160,000 variants across four active-site residues of the enzyme tryptophan synthase, capturing significant epistatic interactions to serve as a benchmark for model-guided enzyme engineering.
Human Myoglobin Epistasis Dataset: A deep mutational scanning library detailing the expression fitness scores for near-comprehensive single-codon mutations and a small fraction of double-codon mutations in yeast surface-displayed human myoglobin, which was used to train machine learning models for predicting epistatic effects and discovering stability-enhancing variants.
CTXM Epistasis Dataset: A large-scale pairwise deep mutational scanning dataset of the CTX-M-14 β-lactamase active site, covering 49,096 double mutants across 17 active-site residues. Fitness measurements were obtained from functional selection under ampicillin and cefotaxime, providing substrate-dependent fitness landscapes for studying epistasis, compensatory mutations, and antibiotic resistance prediction.
RBD-ACE2 Dataset: SARS-CoV-2 RBD sequences with ACE2 binding affinity scores, labeled by
log10Kawhere higher values indicate stronger ACE2 binding affinity.RBD-Antibody Dataset: SARS-CoV-2 RBD antibody escape data with mutation-level antibody escape scores.
Chitosanase dTm Dataset: Wet-lab generated chitosanase mutation dataset containing wild-type protein sequences, amino acid mutation annotations, and experimentally measured melting temperatures (Tm). The Tm value serves as the label for protein thermostability prediction and mutation effect analysis.
MGnify ddG Dataset:A computationally derived thermodynamic dataset constructed from MGnify absolute free energy (dG) data. By applying specific sequence-length constraints and clustering operations, variants are systematically paired to establish wild-type and mutant relationships, yielding relative stability (ΔΔG) labels for protein engineering models.