Audit identifies subtle but consequential issues, such as conflicting labels and train–test leakage, in numerous benchmark datasets
Many of the most widely used benchmarks for AI-driven drug-discovery have data issues that can alter which algorithms appear to work best, an audit of 51 datasets has found.