Literature DB >> 33892802

Extended similarity indices: the benefits of comparing more than two objects simultaneously. Part 1: Theory and characteristics.

Ramón Alain Miranda-Quintana1, Dávid Bajusz2, Anita Rácz3, Károly Héberger4.   

Abstract

Quantification of the similarity of objects is a key concept in many areas of computational science. This includes cheminformatics, where molecular similarity is usually quantified based on binary fingerprints. While there is a wide selection of available molecular representations and similarity metrics, there were no previous efforts to extend the computational framework of similarity calculations to the simultaneous comparison of more than two objects (molecules) at the same time. The present study bridges this gap, by introducing a straightforward computational framework for comparing multiple objects at the same time and providing extended formulas for as many similarity metrics as possible. In the binary case (i.e. when comparing two molecules pairwise) these are naturally reduced to their well-known formulas. We provide a detailed analysis on the effects of various parameters on the similarity values calculated by the extended formulas. The extended similarity indices are entirely general and do not depend on the fingerprints used. Two types of variance analysis (ANOVA) help to understand the main features of the indices: (i) ANOVA of mean similarity indices; (ii) ANOVA of sum of ranking differences (SRD). Practical aspects and applications of the extended similarity indices are detailed in the accompanying paper: Miranda-Quintana et al. J Cheminform. 2021. https://doi.org/10.1186/s13321-021-00504-4 . Python code for calculating the extended similarity metrics is freely available at: https://github.com/ramirandaq/MultipleComparisons .

Entities:  

Keywords:  ANOVA; Comparisons; Consistency; Extended similarity indices; Molecular fingerprints; Rankings; Sum of ranking differences

Year:  2021        PMID: 33892802     DOI: 10.1186/s13321-021-00505-3

Source DB:  PubMed          Journal:  J Cheminform        ISSN: 1758-2946            Impact factor:   5.514


  3 in total

1.  Extended continuous similarity indices: theory and application for QSAR descriptor selection.

Authors:  Anita Rácz; Timothy B Dunn; Dávid Bajusz; Taewon D Kim; Ramón Alain Miranda-Quintana; Károly Héberger
Journal:  J Comput Aided Mol Des       Date:  2022-03-15       Impact factor: 3.686

2.  Molecular Dynamics Simulations and Diversity Selection by Extended Continuous Similarity Indices.

Authors:  Anita Rácz; Levente M Mihalovits; Dávid Bajusz; Károly Héberger; Ramón Alain Miranda-Quintana
Journal:  J Chem Inf Model       Date:  2022-07-14       Impact factor: 6.162

3.  Extended many-item similarity indices for sets of nucleotide and protein sequences.

Authors:  Dávid Bajusz; Ramón Alain Miranda-Quintana; Anita Rácz; Károly Héberger
Journal:  Comput Struct Biotechnol J       Date:  2021-06-16       Impact factor: 7.271

  3 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.