Literature DB >> 11206366

Ties in proximity and clustering compounds.

J MacCuish1, C Nicolaou, N E MacCuish.   

Abstract

Hierarchical clustering algorithms such as Wards or complete-link are commonly used in compound selection and diversity analysis. Many such applications utilize binary representations of chemical structures, such as MACCS keys or Daylight fingerprints, and dissimilarity measures, such as the Euclidean or the Soergel measure. However, hierarchical clustering algorithms can generate ambiguous results owing to what is known in the cluster analysis literature as the ties in proximity problem, i.e., compounds or clusters of compounds that are equidistant from a compound or cluster in a given collection. Ambiguous ties can occur when clustering only a few hundred compounds, and the larger the number of compounds to be clustered, the greater the chance for significant ambiguity. Namely, as the number of "ties in proximity" increases relative to the total number of proximities, the possibility of ambiguity also increases. To ensure that there are no ambiguous ties, we show by a probabilistic argument that the number of compounds needs to be less than 2(n 1/4), where n is the total number of proximities, and the measure used to generate the proximities creates a uniform distribution without statistically preferred values. The common measures do not produce uniformly distributed proximities, but rather statistically preferred values that tend to increase the number of ties in proximity. Hence, the number of possible proximities and the distribution of statistically preferred values of a similarity measure, given a bit vector representation of a specific length, are directly related to the number of ties in proximities for a given data set. We explore the ties in proximity problem, using a number of chemical collections with varying degrees of diversity, given several common similarity measures and clustering algorithms. Our results are consistent with our probabilistic argument and show that this problem is significant for relatively small compound sets.

Year:  2001        PMID: 11206366     DOI: 10.1021/ci000069q

Source DB:  PubMed          Journal:  J Chem Inf Comput Sci        ISSN: 0095-2338


  4 in total

Review 1.  Global analysis of large-scale chemical and biological experiments.

Authors:  David E Root; Brian P Kelley; Brent R Stockwell
Journal:  Curr Opin Drug Discov Devel       Date:  2002-05

Review 2.  QSAR without borders.

Authors:  Eugene N Muratov; Jürgen Bajorath; Robert P Sheridan; Igor V Tetko; Dmitry Filimonov; Vladimir Poroikov; Tudor I Oprea; Igor I Baskin; Alexandre Varnek; Adrian Roitberg; Olexandr Isayev; Stefano Curtarolo; Denis Fourches; Yoram Cohen; Alan Aspuru-Guzik; David A Winkler; Dimitris Agrafiotis; Artem Cherkasov; Alexander Tropsha
Journal:  Chem Soc Rev       Date:  2020-05-01       Impact factor: 54.564

3.  JEDA: Joint entropy diversity analysis. An information-theoretic method for choosing diverse and representative subsets from combinatorial libraries.

Authors:  Melissa R Landon; Scott E Schaus
Journal:  Mol Divers       Date:  2006-09-21       Impact factor: 2.943

4.  How frequently do clusters occur in hierarchical clustering analysis? A graph theoretical approach to studying ties in proximity.

Authors:  Wilmer Leal; Eugenio J Llanos; Guillermo Restrepo; Carlos F Suárez; Manuel Elkin Patarroyo
Journal:  J Cheminform       Date:  2016-01-25       Impact factor: 5.514

  4 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.