Data-Efficient Reaction Prediction Using Quantum-Mechanically Informed Surrogate Representations
11:30 - 11:45
The efficient exploration and screening of chemical reaction space require machine-learning (ML) models capable of making accurate predictions both within and across reaction families. However, when only small, high-quality, task-specific datasets are available, conventional ML architectures often suffer from limited accuracy and poor generalizability. One strategy for addressing this challenge is to augment the models with physically motivated quantum-mechanical (QM) descriptors that encode bonding and electronic-structure information. This approach has been shown to enable robust and data-efficient predictions of chemical reactivity, even when only a few hundred training data points are available.
Unfortunately, generating QM descriptors is computationally demanding, which limits the practicality of this strategy for truly large-scale or high-throughput screening. Surrogate models that predict such descriptors directly from molecular graphs provide a promising alternative. When combined with large, precomputed QM datasets from the literature, these surrogate models can support generalizable screening pipelines that require minimal task-specific training data and incur only marginal computational cost.
In this talk, I will present our recent work in this area, focusing on results showing that downstream predictive models often perform better when they use the internal hidden representations of the surrogate model rather than the predicted descriptors themselves. An exception arises when the descriptor set is closely aligned with the mechanistic physics of the target reaction class. These latent embeddings encode rich and transferable chemical information, making them a powerful foundation for data-efficient predictive chemistry.
Finally, I will discuss how these pre-trained hidden representations can also be used to substantially accelerate Bayesian optimization campaigns.