Sequence Homology Inflates Reported Accuracy in Antimicrobial Peptide Prediction: A Machine Learning and Deep Learning Benchmark


Creative Commons License

Özkat E. C., Yalçın Özkat G.

5th International Conference on Recent and Innovative Results in Engineering and Technology, Konya, Türkiye, 25 - 26 Ağustos 2026, ss.255-262, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Basıldığı Şehir: Konya
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.255-262
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Recep Tayyip Erdoğan Üniversitesi Adresli: Evet

Özet

Machine learning is now routine for screening antimicrobial peptide candidates, and reported accuracies are high. This study asks how much of that accuracy depends on sequence homology between training and test data. A benchmark of 6288 mature peptides was assembled from curated UniProtKB annotation, using one extraction rule for both classes, so negative-set bias is excluded by construction. All 19,766,328 peptide pairs were aligned with local Smith-Waterman scoring. Partitions were drawn at 0.8 identity, the only threshold that clears a shuffled-sequence noise floor without chaining the set into one component. Twenty pipelines were then trained under two partitioning regimes with five seeds each. They comprise sixteen classical pipelines, three neural architectures and a control that only looks up the most similar training peptide. The peptide set proved strongly redundant, and a random partition left 62 per cent of test peptides with a close relative in training. Every trained pipeline lost accuracy when whole identity components were held out, by 0.039 in area under the curve on average, and the leading pipeline changed. The lookup control reached 0.958 under the random partition against 0.819 under the homology-aware one. Accuracy fell steadily as identity to the nearest training peptide decreased. Holding out whole greedy clusters, which is common practice, still left 87 per cent of held-out peptides with a close relative in training. Sequence-based peptide classifiers should therefore be partitioned by identity component, and test-to-training similarity should be reported alongside the headline metric.