Class Imbalance In Toxicity Prediction: The Effect of Performance Metric and Decision Threshold Selection on Model Ranking in the Tox21 Dataset


Creative Commons License

Yalçın Özkat G., Özkat E. C.

2. Istinye International Interdisciplinary Academic Studies Congress, İstanbul, Türkiye, 1 - 02 Eylül 2026, ss.197-206, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.197-206
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Recep Tayyip Erdoğan Üniversitesi Adresli: Evet

Özet

Machine learning models for toxicity prediction are usually reported with a single performance metric, most often the area under the receiver operating characteristic curve (AUROC). Toxicology datasets, however, are severely imbalanced, and as the positive class becomes rarer AUROC reflects the quality of the decisions a model would make in practice ever more weakly. In this study, 12 assays from the Tox21 screening programme, accessed through the Therapeutics Data Commons, were examined across an imbalance range in which the positive class accounts for 2.9% to 16.2% of the molecules. Two split protocols, five random seeds, two molecular representations and four learning algorithms were fully crossed, giving 960 independently trained models, each evaluated both with threshold-free metrics and with a decision threshold selected on a validation partition. The results show that the choice of metric changes the ranking of models: the configuration selected by AUROC is also the best configuration under the prevalence-aware metric in only 61.7% of the blocks. Leaving the decision threshold at its default value of 0.5 costs on average 0.024 units of the Matthews correlation coefficient: that threshold recovers only 52.4% of the positive molecules an assay actually contains. The size of the gain is governed not by the positive class rate but by how far a model's probabilities are displaced relative to the default cut-off. We therefore recommend that toxicity prediction studies report a prevalence-aware metric alongside AUROC, together with a tuned decision threshold and a multi-seed evaluation.