Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Scalable model-based clustering for large databases based on data summarization.

Literature DB >> 16285371

Scalable model-based clustering for large databases based on data summarization.

Huidong Jin¹, Man-Leung Wong, K S Leung.

Abstract

The scalability problem in data mining involves the development of methods for handling large databases with limited computational resources such as memory and computation time. In this paper, two scalable clustering algorithms, bEMADS and gEMADS, are presented based on the Gaussian mixture model. Both summarize data into subclusters and then generate Gaussian mixtures from their data summaries. Their core algorithm, EMADS, is defined on data summaries and approximates the aggregate behavior of each subcluster of data under the Gaussian mixture model. EMADS is provably convergent. Experimental results substantiate that both algorithms can run several orders of magnitude faster than expectation-maximization with little loss of accuracy.

Mesh：

Year: 2005 PMID： 16285371 DOI： 10.1109/TPAMI.2005.226

Source DB: PubMed Journal: IEEE Trans Pattern Anal Mach Intell ISSN： 0098-5589 Impact factor: 6.226

Keyword Cloud
Cited

3 in total

Scalable model-based clustering for large databases based on data summarization.

1. A Scalable Framework For Cluster Ensembles.

2. Similarity measure and domain adaptation in multiple mixture model clustering: An application to image processing.

3. Compatibility Evaluation of Clustering Algorithms for Contemporary Extracellular Neural Spike Sorting.