MolBasis
The datalake

Every dataset, one provenance-stamped atlas

All kernels are fit against a single atlas of σ-profiles, generated with one recipe applied identically to every molecule. Here is the complete census — every experimental dataset across every property category, with honest coverage.

6,876
production σ-profiles
v0.91.1 SVP atlas
25
experimental datasets
across 11 property categories
16d3ca90
recipe SHA
pinned in every provenance
3,140
SMILES still to compute
Tier A: 942 (~$9.42)

Atlas composition

6,876
production v0.91.1 SVP
45
reference solvents
55
v0.92.0 TZVPP smoke
64
hetero v3 (legacy)
6,903
unique molecules

Recipe: wB97M-V / def2-SVP / C-PCM (ε=80, Bondi cavity, Klamt-purge) · vendor SHA 56c80cea · regenerates for ~$10–21 compute, <2 days from raw data.

Solvent atlas: 45/49 reference solvents computed (water, the 6 MNSol solvents, alcohols, alkanes, chlorinated, polar aprotic…). Missing: tetrahydrofuran, formic acid, hexafluoroisopropanol, 2-octanol. These feed the bilinear solvent-partner kernels.

Hydration & solvation

8 datasets

Free energies of transfer into water and organic solvents — the home turf of σ-profile thermodynamics.

MNSol · water
Marenich/Cramer/Truhlar — Minnesota Solvation Database
ΔG_hydration · kcal/mol
3,037 raw
91%
shipped
MNSol · octanol
MNSol (solvent = octanol)
ΔG_solvation · kcal/mol
119 mol
94%
shipped
MNSol · hexadecane
MNSol (solvent = hexadecane)
ΔG_solvation · kcal/mol
108 mol
93%
shipped
MNSol · chloroform
MNSol (solvent = chloroform)
ΔG_solvation · kcal/mol
60 mol
98%
active
MNSol · cyclohexane
MNSol (solvent = cyclohexane)
ΔG_solvation · kcal/mol
57 mol
98%
active
MNSol · carbon-tet
MNSol (solvent = carbon tetrachloride)
ΔG_solvation · kcal/mol
47 mol
98%
active
CompSol 2024
Moine/Trinh/Ramos 2024
ΔG_hydration · kcal/mol
35,624 raw
29%
shipped
FreeSolv
Mobley FreeSolv (raw .txt + .xlsx variant)
ΔG_hydration · kcal/mol
potential

Aqueous solubility (logS)

5 datasets

Log aqueous solubility — cavity desolvation cost vs H-bond hydration reward.

ESOL (Delaney 2004)
Delaney 2004, J. Chem. Inf. Model.
logS · log mol/L
1,129 raw
100%
shipped
Aqueous cosolvent
Aqueous cosolvent set
logS / logP · log
323 raw
100%
100% covered
Llinás
Llinás et al., JCIM
logS0 / logP · log
132 raw
95%
shipped
FDA drugs · LogW
FDA Orange Book + literature
logW (solubility) · log
71 raw
10%
potential
Wiki_pS0
Wikipedia pS0 solubility compilation
logS0 + melting point · log / K
pooled

Melting point (Tm)

2 datasets

Lattice cohesion as a size × polarity-spread family.

Bradley DPG
Bradley Double-Plus-Good MP
Tm (melting point) · K
3,041 raw
39%
shipped
FDA drugs · MP
FDA Orange Book + literature
Tm (melting point) · K
72 raw
10%
deferred

Hansen parameters & miscibility

2 datasets

Hansen solubility parameters (δ_d/δ_p/δ_h) and drug–polymer miscibility χ.

HSP dataset
Hansen Solubility Parameters
δ_d / δ_p / δ_h · MPa^½
1,206 raw
60%
shipped
ASD miscibility χ
MolForge ASD — 9 drugs × 10 polymers
Flory χ interaction · dimensionless
90 raw
blocked

Surface tension (γ)

1 dataset

Interfacial cooperative regime — the ¼-integer law predicts a 3/2 exponent here.

CRC / Lange 25 °C
CRC + Lange's Handbook
γ surface tension · mN/m
34 raw
94%
potential

PC-SAFT parameters

1 dataset

Equation-of-state segment number m — a bridge to process thermodynamics.

Gross–Sadowski 2001
Gross & Sadowski 2001
m (segment number) · dimensionless
24 raw
96%
active

Polymer glass transition (Tg)

1 dataset

Amorphous-polymer Tg — hits a chain-length ceiling with a pentamer atlas.

Polymer Tg
Polymer Tg compilation
Tg · K
1,440 raw
4%
fail

Protein–ligand binding (ΔG)

1 dataset

Binding free energy from pocket-fragment σ-profiles — awaiting a schema regen.

PDBBind v2020 (refined)
PDBBind v2020 refined set
ΔG_bind · kcal/mol
1,310 raw
70%
blocked

Dimer BSSE (reference)

2 datasets

Counterpoise-corrected dimer interaction energies — v0.92.0 TZVPP smoke sets.

Homodimer BSSE
MolForge v0.92.0 TZVPP smoke
interaction energy · kcal/mol
31 raw
reference
Heterodimer BSSE (CL3)
MolForge v0.92.0 TZVPP smoke
interaction energy · kcal/mol
31 raw
reference

Binding-pocket fragments

1 dataset

Fragmented protein pockets at pH 7.4 — input source for the binding kernel.

v3 pocket manifests (pH 7.4)
MolForge pocket fragmenter
pocket σ-descriptors · 9-d vector
1,778 raw
blocked

The compute gap

3,140 unique SMILES would fully cover every kernel dataset. We prioritise by leverage, not volume — and we explicitly reject fills the learning curve says won’t move the needle.

Tier A — leverage fill

EXECUTE
942 compounds · ~$9.42

Compounds appearing in 2+ datasets — each unlocks 2–6 kernels at once.

Tier B — Bradley_Tm full

REJECTED
1,831 compounds · ~$20

Learning curve is flat (+0.0007 r per +500 anchors) → data-saturated.

Tier C — HSP full

DOWNGRADED
474 compounds · ~$5

Polarizability feature test failed the +30% gate; HSP near its physics floor.

Add your molecules to the lake

Upload a set of mfsig σ-profiles with measured property values and we can fit and benchmark a kernel against it under the same rigor gate.

Upload your mfsig dataset