Many of the most widely used benchmarks for AI-driven drug-discovery have data issues that can alter which algorithms appear to work best, an audit of 51 datasets has found.
For researchers developing machine-learning models for drug discovery, benchmarks serve as a testing ground. Benchmark performance often determines where drug-discovery models rank on leaderboards and which are considered state of the art. However, subtle issues with the composition of benchmark datasets – for example, when the same molecule or a very similar one appears in both the training and test sets, known as train–test leakage – can also make some models seem better than they are.
Concerns about benchmark datasets are not new. As well as train–test leakage, various researchers have noticed inconsistent annotations and misparsed structures in benchmark datasets. Maximilian Schuh at the Technical University of Munich in Germany began wondering quite how widespread these issues might be after attending a machine-learning conference and seeing how heavily researchers relied on benchmark scores to compare their methods. ‘It always has been suspected that some benchmarks are not that great,’ says Schuh. ‘But you always should look into the data, and it should be quantifiable to have proof it’s actually bad.’
Schuh and colleagues designed an auditing framework and applied it to 51 benchmark configurations spanning four of the most widely used datasets: 7 Polaris datasets, 22 from Therapeutics Data Commons (TDC), 9 from MoleculeNet and 13 drug–target interaction (DTI) benchmarks, including datasets derived from BindingDB and PDBbind. Together, these resources have more than 10,000 citations.
The researchers found problems in most datasets, notably train–test leakage and contradictory labels assigned to identical molecules. DTI benchmarks were particularly prone to overlap, with an average nearest-neighbour similarity score of 0.779, compared with 0.494 for Polaris.
To test whether these issues actually matter, the team deliberately introduced them into datasets and used them to retrain a range of machine-learning models. They found that enriching test sets with molecules that were identical or highly similar to training examples boosted how a model performed, while conflicting labels generally reduced it. In a separate experiment, they recalculated benchmark leaderboards using alternative test sets. Changing the mix of familiar compounds and conflicting annotations was often enough to change which model came out on top.

‘If you have data leak, training and test that are similar or duplicated, the performance will become much better than it actually is,’ says Yang Zhang from the National University of Singapore, whose AI-driven therapeutic discovery research includes database construction and benchmark design. ‘This is a common problem. The value of this paper is that it quantitatively checks this problem across many different databases and methods.’
‘The fundamental problem is that many benchmarks are not explicitly linked to a particular application scenario,’ comments Pedro Ballester, whose work at Imperial College London, UK, developing and applying computational methods to boost biomedical discoveries includes benchmarking virtual screening . Similar molecules may be appropriate for predicting activity during lead optimisation, whereas virtual screening may require models to generalise to chemically distinct compounds, he notes.
Schuh and colleagues are now calling on their fellow researchers to report what their test sets contain and routinely audit benchmark datasets. To help with this, they have created BenchAudit, an open-source toolkit that identifies issues such as train–test contamination and label conflicts. Schuh says the goal is not to eliminate every flaw, but to ‘raise awareness and be as honest as possible’.
References
This article is open access
M G Schuh, A Daniluk and S A Sieber, Chem. Sci., 2026, DOI: 10.1039/d6sc01799a





No comments yet