Comparing Global and Local Likelihood Score Thresholds in Multiclass Laplacian-Modified Naive Bayes Protein Target Prediction

Georgios      Drakakis; Alexios      Koutsoukas; Suzanne   C.   Brewerton; Michael   J.   Bodkin; David   A.   Evans; Andreas      Bender

doi:10.2174/1386207318666150305145012

Abstract

The increase of publicly available bioactivity data has led to the extensive development and usage of in silico bioactivity prediction algorithms. A particularly popular approach for such analyses is the multiclass Naïve Bayes, whose output is commonly processed by applying empirically-derived likelihood score thresholds. In this work, we describe a systematic way for deriving score cut-offs on a per-protein target basis and compare their performance with global thresholds on a large scale using both 5-fold cross-validation (ChEMBL 14, 189k ligand-protein pairs over 477 protein targets) and external validation (WOMBAT, 63k pairs, 421 targets). The individual protein target cut-offs derived were compared to global cut-offs ranging from -10 to 40 in score bouts of 2.5. The results indicate that individual thresholds had equal or better performance in all comparisons with global thresholds, ranging from 95% of protein targets to 57.96%. It is shown that local thresholds behave differently for particular families of targets (CYPs, GPCRs, Kinases and TFs). Furthermore, we demonstrate the discrepancy in performance when we move away from the training dataset chemical space, using Tanimoto similarity as a metric (from 0 to 1 in steps of 0.2). Finally, the individual protein score cut-offs derived for the in silico bioactivity application used in this work are released, as well as the reproducible and transferable KNIME workflows used to carry out the analysis.

Keywords: Cheminformatics, in silico bioactivity prediction, likelihood score thresholds.

« Previous

Rights & Permissions Print Cite

Article Metrics

22

Journal Information

For Authors

For Editors

For Reviewers

Explore Articles

Open Access

Open Access Articles

For Visitors

DOI https://dx.doi.org/10.2174/1386207318666150305145012	Print ISSN 1386-2073
Publisher Name Bentham Science Publisher	Online ISSN 1875-5402

Combinatorial Chemistry & High Throughput Screening

Comparing Global and Local Likelihood Score Thresholds in Multiclass Laplacian-Modified Naive Bayes Protein Target Prediction

Abstract Play Pause

Related Journals

Abstract