MolBasis
B.UNDERFIT1-termHansen Solubility Parameters · n=705

HSP · δ_p

Hansen polar parameter · MPa^½

closed-form kernel
ŷ = β₀ + β₁ · z(dipole^p)
0.700
nested-CV r
in-sample, winner's-curse-safe
0.849
scaffold-blind r
25% held-out OOD
+48.1%
blind MAE-vs-NULL lift
gate = +30%
MAE (MPa^½)
grouped-LOO

In-sample → out-of-distribution

How the correlation holds when moving to unseen scaffolds.

nested-CV r (in-sample)0.700
scaffold-blind r (OOD)0.849
Blind MAE-vs-NULL lift48.1%

Vertical marker = the +30% production gate.

Blind calibration slope

Slope of measured vs predicted on the blind set. 1.0 = no magnitude compression.

0.98
slope (ideal 1.0)
Form stable, magnitude noisy
blind cohort n ≈ 33

What it is

The Hansen polar component. In-sample r=0.700 at +27% MAE-vs-NULL (below the 30% gate) → B.UNDERFIT. Blind r looks high (0.849) but n_blind=33 makes it noisy.

The physics

Polar cohesion tracks dipole, but a single descriptor leaves headroom the current feature set can't fill.

caveat

B.UNDERFIT on the honest in-sample gate; the flattering blind r rests on a small (n=33) held-out set.

How the methods compare

Hansen parameters are dominated by group-contribution tables (Stefanis–Panayiotou / HSPiP). This kernel derives δ_h directly from hydrogen-bond acceptor surface area — a physical quantity — instead of counting functional groups. δ_p and δ_d remain honestly under-fit on current physics.

Input cost

Per-molecule compute/data burden to predict a new molecule — shorter is cheaper.

Stefanis–Panayiotou (group contribution)2D structure (pencil / inference)
Van Krevelen / Hansen–Beerbower2D structure (pencil / inference)
ML (XGBoost / CatBoost on HSPiP)2D structure (pencil / inference)
HSP · δ_p · this kernelone QM calculation (SCF)
COSMO-RS σ-moment mappingone QM calculation (SCF)

What each method needs

External dependencies each method carries. An amber dot means the method requires it — fewer dots means fewer things to procure or that can go wrong.

Method3D geometryMD / samplingtraining corpusa measured valueproprietary paramsdeps
HSP · δ_p · this kernel1
Stefanis–Panayiotou (group contribution)1
Van Krevelen / Hansen–Beerbower0
ML (XGBoost / CatBoost on HSPiP)1
COSMO-RS σ-moment mapping2

This kernel needs only a 3D geometry for its one SCF — no MD, no training corpus, no measured value, no proprietary software.

Competing methods

How this property is predicted elsewhere — with the input each method needs (a key differentiator) and the literature reference. Numbers are each method’s own reported figure on its own benchmark, so they are indicative, not a head-to-head on an identical split.

MethodClassReported performanceInput neededReference
HSP · δ_pthis kernelclosed-formPearson r 0.700 (nested-CV) · 0.849 scaffold-blindone DFT SCF σ-profile · no training set · no MDMF-FQSL (this lab)
Stefanis–Panayiotou (group contribution)group-contributionThe standard GC route; 1st + 2nd-order groups, implemented in HSPiP2D functional-group counts (UNIFAC + conjugation groups)Stefanis & Panayiotou, Int. J. Thermophys. 2008
Van Krevelen / Hansen–Beerbowergroup-contributionClassic additive GC; component-dependent accuracy2D functional-group countsVan Krevelen; Hansen, HSP Handbook 2007
ML (XGBoost / CatBoost on HSPiP)ML / GNNRecent gradient-boosted models on the extended HSPiP corpusmolecular descriptors + the HSPiP training setrecent HSP ML studies (2023–2024)
COSMO-RS σ-moment mappingphysicsHSP from σ-profile moments; parametrisation-dependentDFT σ-profileKlamt; σ-moment → HSP correlations

Metrics are as published by each method on its own dataset (different splits, different cohorts) — treat them as an orientation of the landscape, not a controlled benchmark. The differentiator for this kernel is the input column: a single closed-form solve from one σ-profile, with no training corpus, no MD, and no measured melting point.

Descriptors used

dipoledipole_debye — molecular dipole moment (Debye)

Version history

  1. 2026-05-31v0.91.1 B.UNDERFIT

    r=0.700, +27% MAE-vs-NULL, slope 0.982.