MolBasis
The kernel catalogue

Every kernel, with its honest number

Each kernel maps mfsig σ-profile descriptors to a physical property through a closed-form ridge equation. Cards show the in-sample nested-CV r and the 25% scaffold-blind r — the commercial out-of-distribution estimate.

Production kernels

A.IMPROVE · 8

Honest under-fits

B.UNDERFIT · 4

These clear r ≥ 0.5 but fall below the +30% MAE-vs-NULL blind-lift gate. We ship them labelled as under-fits rather than dressing them up — the physics ceiling is real, and the honest verdict is the useful one.

Pipeline & roadmap

7 targets

Properties we’ve probed that aren’t shipped — shown with their honest status so the scope is complete: what failed and why, what’s blocked on infrastructure, and what’s a data-prep away from a fit.

FAILr_cv ≈ 0.29

Polymer Tg

Glass transition temperature

The kernel fails the 5-fold CV gate. The signal is physical (donor^-0.25) but a pentamer atlas plus a 9-descriptor vector hits a chain-length ceiling.

Next: Nonamer/decamer chains + lattice-aware features — not more SVP atlas growth.

BLOCKED

ASD miscibility χ

Drug–polymer Flory χ

Needs the hetero (Boys–Bernardi counterpoise) pipeline, whose orchestrator shipped but was never validated on a GPU pod.

Next: 1–2 smoke pairs on a pod (~$2) to confirm the code path before any cohort spawn.

BLOCKEDlegacy MAE 1.72, r 0.56 (n=119)

PDBBind binding ΔG

Protein–ligand binding free energy

Blocked on a pocket schema regen — 305 v28 pocket fragments carry only 5/9 features (missing HOMO/LUMO/gap/g_polar) — plus hetero pipeline validation.

Next: Pocket v29 regen (~5×H100, ~3h, ~$50), then an apples-to-apples re-fit.

POTENTIAL

Surface tension γ

Liquid surface tension @ 25 °C

Data is 94% covered and the ¼-integer law predicts a clean 3/2 exponent here — but the raw file needs a SMILES column added (French headers) before audit.

Next: Data prep, then run through the standard rigor gate.

POTENTIAL

FreeSolv ΔG_hyd

Hydration free energy (independent baseline)

A valuable independent test set for the hydration kernels; needs an xlsx column fix before it can be wired in.

Next: Fix the input, use as an out-of-distribution check on MNSol_water / CompSol.

POTENTIALn=71 (7 in atlas)

FDA drugs LogW

Drug aqueous solubility

Small, drug-relevant cohort — underpowered on its own but useful pooled with ESOL/AqCoSol.

Next: Compute the 64 missing (~$1); treat as exploratory.

DEFERREDnested r ≈ 0.32

Tm v1 (FDA + cosolvent)

Melting point (small cohort)

An honest but weak result on a small cohort — the Bradley cohort (n=1183) gives a far stronger Tm predictor.

Next: Deferred in favour of Bradley_Tm as the primary melting-point anchor set.

Reading the cards

  • nested-CV r — in-sample Pearson r under per-fold pair selection (winner’s-curse-safe).
  • blind r — Pearson r on a 25% Bemis–Murcko scaffold split held out up front, re-fit on the 75%. The OOD estimate.
  • A.IMPROVE — clears the +30% MAE-vs-NULL gate. B.UNDERFIT — clears r ≥ 0.5 but not the +30% gate.
  • 1-term / 2-term / family — number of descriptor terms; family means many exponent pairs tie within ±0.01 r (no single ‘magic’ pair).

12 kernels total. All numbers from internal audit docs (May–June 2026); corrected values shown where an earlier figure was superseded.