Categories
Nevin Manimala Statistics

Data Compression of the D1200 Suite: Achieving the Optimal Balance Between Size and Accuracy with the Diet-D200 Subset

J Comput Chem. 2026 Sep 5;47(23):e70488. doi: 10.1002/jcc.70488.

ABSTRACT

The evaluation of non-covalent interactions remains a cornerstone of density functional theory benchmarking, yet the increasing size of reference datasets like the D1200 presents a significant computational bottleneck. This work introduces an optimized “Diet-D1200” protocol, utilizing a genetic algorithm to identify highly representative subsets of interaction energies across an expansive chemical space, including p-block elements such as Boron, Sulfur, Phosphorus, halogens, and noble gases. By employing a WTMAD-4 (Weighted Total Mean Absolute Deviation) scheme-normalized specifically to the D1200 intrinsic energy scale-we demonstrate that subsets as small as 100 to 250 systems can faithfully reproduce the statistical performance of the full dataset. To evaluate the generalization performance of the optimized subsets and mitigate the risk of overfitting to the training manifold, a rigorous external validation protocol was performed. This assessment utilized twenty density functionals approximations, from GGAs and meta-GGAs to range-separated hybrids and double-hybrids. Our results demonstrate that the Diet-D200 protocol effectively preserves the topological features of the error surface, reducing the computational overhead by 83.3% while maintaining an exceptional Pearson correlation ( r > 0.995 ) with the parent D1200.

PMID:42634015 | DOI:10.1002/jcc.70488

By Nevin Manimala

Portfolio Website for Nevin Manimala