MolBasis
B.UNDERFIT1-termHansen Solubility Parameters · n=705

HSP · δ_d

Hansen dispersion parameter · MPa^½

closed-form kernel
ŷ = β₀ + β₁ · z(cavity^p)
0.473
nested-CV r
in-sample, winner's-curse-safe
0.710
scaffold-blind r
25% held-out OOD
-1.1%
blind MAE-vs-NULL lift
gate = +30%
MAE (MPa^½)
grouped-LOO

In-sample → out-of-distribution

How the correlation holds when moving to unseen scaffolds.

nested-CV r (in-sample)0.473
scaffold-blind r (OOD)0.710
Blind MAE-vs-NULL lift-1.1%

Vertical marker = the +30% production gate.

Blind calibration slope

Slope of measured vs predicted on the blind set. 1.0 = no magnitude compression.

0.81
slope (ideal 1.0)
Below +30% blind gate

What it is

The hardest HSP component. r=0.473 (below the 0.50 floor) and blind lift −1.1% — the kernel orders molecules but does not predict absolute dispersion values. Honestly the weakest ship.

The physics

Dispersion cohesion needs polarizability. A dedicated test added explicit isotropic α (GFN2-CPSCF via dxtb): a {donor+α} 2-term reaches r=0.586 (MAE-vs-NULL 23.7%) — real and dispersion-specific, but still below the 30% gate, so α was NOT productionised.

caveat

B.UNDERFIT. Needs crystal-lattice / polarizability physics beyond single-geometry σ-profiles.

How the methods compare

Hansen parameters are dominated by group-contribution tables (Stefanis–Panayiotou / HSPiP). This kernel derives δ_h directly from hydrogen-bond acceptor surface area — a physical quantity — instead of counting functional groups. δ_p and δ_d remain honestly under-fit on current physics.

Input cost

Per-molecule compute/data burden to predict a new molecule — shorter is cheaper.

Stefanis–Panayiotou (group contribution)2D structure (pencil / inference)
Van Krevelen / Hansen–Beerbower2D structure (pencil / inference)
ML (XGBoost / CatBoost on HSPiP)2D structure (pencil / inference)
HSP · δ_d · this kernelone QM calculation (SCF)
COSMO-RS σ-moment mappingone QM calculation (SCF)

What each method needs

External dependencies each method carries. An amber dot means the method requires it — fewer dots means fewer things to procure or that can go wrong.

Method3D geometryMD / samplingtraining corpusa measured valueproprietary paramsdeps
HSP · δ_d · this kernel1
Stefanis–Panayiotou (group contribution)1
Van Krevelen / Hansen–Beerbower0
ML (XGBoost / CatBoost on HSPiP)1
COSMO-RS σ-moment mapping2

This kernel needs only a 3D geometry for its one SCF — no MD, no training corpus, no measured value, no proprietary software.

Competing methods

How this property is predicted elsewhere — with the input each method needs (a key differentiator) and the literature reference. Numbers are each method’s own reported figure on its own benchmark, so they are indicative, not a head-to-head on an identical split.

MethodClassReported performanceInput neededReference
HSP · δ_dthis kernelclosed-formPearson r 0.473 (nested-CV) · 0.710 scaffold-blindone DFT SCF σ-profile · no training set · no MDMF-FQSL (this lab)
Stefanis–Panayiotou (group contribution)group-contributionThe standard GC route; 1st + 2nd-order groups, implemented in HSPiP2D functional-group counts (UNIFAC + conjugation groups)Stefanis & Panayiotou, Int. J. Thermophys. 2008
Van Krevelen / Hansen–Beerbowergroup-contributionClassic additive GC; component-dependent accuracy2D functional-group countsVan Krevelen; Hansen, HSP Handbook 2007
ML (XGBoost / CatBoost on HSPiP)ML / GNNRecent gradient-boosted models on the extended HSPiP corpusmolecular descriptors + the HSPiP training setrecent HSP ML studies (2023–2024)
COSMO-RS σ-moment mappingphysicsHSP from σ-profile moments; parametrisation-dependentDFT σ-profileKlamt; σ-moment → HSP correlations

Metrics are as published by each method on its own dataset (different splits, different cohorts) — treat them as an orientation of the landscape, not a controlled benchmark. The differentiator for this kernel is the input column: a single closed-form solve from one σ-profile, with no training corpus, no MD, and no measured melting point.

Descriptors used

cavitycavity_area_aa2 — solvent-accessible surface area of the COSMO cavity (Ų)

Version history

  1. 2026-05-31v0.91.1 B.UNDERFIT

    r=0.473, blind lift −1.1%. Orders but doesn't predict.

  2. 2026-06-01Program 0 (α test)

    Explicit polarizability α complementary to donor (r 0.586) but sub-gate → not productionised.