Salient

The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI

Nelen, J., Khan, O., Adams, E., Aschenbrenner, J. C., Thompson, W., Ebrahim, A., Capkin, E., Vallee, C., OpenBind,, Shotton, E. J., et al.
10.64898/2026.08.27.747600 · was preprinted
method development benchmarked
Surfaced because: benchmarked against baselines.
relevance 0.42 openness 0.25 novelty 0.32

Abstract

High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.

Lifecycle

Discussion

No qualifying discussion yet.