Overview#

This guide provides usage examples for data cleaning modules organized by dataset source:

  • Human Domainome Dataset: Site-saturation mutagenesis of 500 human protein domains.

  • ProteinGym DMS Substitutions Dataset: Large-scale benchmarks for protein design and fitness prediction.

  • cDNA Proteolysis Dataset: Mega-scale experimental analysis of protein folding stability in biology and design.

  • ddG-dTm Dataset: A collection of datasets providing single- and multiple-mutant measurements, labeled by the thermodynamic parameter (ΔΔG or ΔTm).

  • ArchStabMS1E10 Epistasis Dataset: High-order multi-mutant libraries (“1e10”) measuring protein stability for GRB2-SH3 and SRC.

  • Antitoxin ParD3 Epistasis Dataset: The antitoxin ParD3 3-position library is a combinatorially exhaustive dataset of 8,000 variants demonstrating that simple, independent per-residue mutation preferences are sufficient to almost perfectly predict combinatorial protein fitness.

  • TrpB Epistasis Dataset: A combinatorially complete sequence-fitness landscape comprising 160,000 variants across four active-site residues of the enzyme tryptophan synthase, capturing significant epistatic interactions to serve as a benchmark for model-guided enzyme engineering.

  • Human Myoglobin Epistasis Dataset: A deep mutational scanning library detailing the expression fitness scores for near-comprehensive single-codon mutations and a small fraction of double-codon mutations in yeast surface-displayed human myoglobin, which was used to train machine learning models for predicting epistatic effects and discovering stability-enhancing variants.

  • CTXM Epistasis Dataset: A large-scale pairwise deep mutational scanning dataset of the CTX-M-14 β-lactamase active site, covering 49,096 double mutants across 17 active-site residues. Fitness measurements were obtained from functional selection under ampicillin and cefotaxime, providing substrate-dependent fitness landscapes for studying epistasis, compensatory mutations, and antibiotic resistance prediction.

  • RBD-ACE2 Dataset: SARS-CoV-2 RBD sequences with ACE2 binding affinity scores, labeled by log10Ka where higher values indicate stronger ACE2 binding affinity.

  • RBD-Antibody Dataset: SARS-CoV-2 RBD antibody escape data with mutation-level antibody escape scores.

  • Chitosanase dTm Dataset: Wet-lab generated chitosanase mutation dataset containing wild-type protein sequences, amino acid mutation annotations, and experimentally measured melting temperatures (Tm). The Tm value serves as the label for protein thermostability prediction and mutation effect analysis.

  • MGnify ddG Dataset:A computationally derived thermodynamic dataset constructed from MGnify absolute free energy (dG) data. By applying specific sequence-length constraints and clustering operations, variants are systematically paired to establish wild-type and mutant relationships, yielding relative stability (ΔΔG) labels for protein engineering models.