Understanding Machine Learning Models Trained on DNA-Encoded Libraries for Virtual Screening

Artur Menzeleev , Sathya Chitturi , Geraint Davies , Tony Schroeder , Alpha Lee

Digital Discovery

DOI: 10.1039/d6dd00162a

Abstract

DNA-encoded library (DEL) screening enables the rapid generation of billion-scale structure-activity datasets and has become a powerful approach for identifying chemical starting points against challenging biological targets. Recent work has leveraged DEL data to train machine learning (ML) models that screen external databases of purchasable compounds, expanding accessible chemical space and reducing the need for off-DNA synthesis. However, DEL datasets contain substantial noise and sampling biases arising in part from synthesis inefficiencies, uneven sampling, and sequencing variability. These factors can obscure true structure-activity relationships and complicate the evaluation and benchmarking of ML models trained on these data. Here, we present DEL Simulator, an open-source framework for systematically studying how DEL experimental parameters influence data quality and downstream ML model performance. The simulator models the key stages of DEL workflows, including library synthesis, selection, and sequencing, and provides a controlled environment for analyzing how factors such as library composition, reaction efficiency, and sequencing depth affect enrichment measurements and predictive modeling. Using this framework, we evaluate preprocessing strategies such as disynthon aggregation and identify regimes in which they improve predictive accuracy. We also observe counterintuitive trends, including reduced model performance with increasing library size under certain conditions. By connecting experimental design choices with data fidelity and model behavior, DEL Simulator provides a principled framework for analyzing DEL datasets and improving ML-driven virtual screening workflows.

logo
logo