Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Batch Mode Active Sampling based on Marginal Probability Distribution Matching.

Literature DB >> 25309807

Batch Mode Active Sampling based on Marginal Probability Distribution Matching.

Rita Chattopadhyay¹, Zheng Wang¹, Wei Fan², Ian Davidson³, Sethuraman Panchanathan⁴, Jieping Ye¹.

Abstract

Active Learning is a machine learning and data mining technique that selects the most informative samples for labeling and uses them as training data; it is especially useful when there are large amount of unlabeled data and labeling them is expensive. Recently, batch-mode active learning, where a set of samples are selected concurrently for labeling, based on their collective merit, has attracted a lot of attention. The objective of batch-mode active learning is to select a set of informative samples so that a classifier learned on these samples has good generalization performance on the unlabeled data. Most of the existing batch-mode active learning methodologies try to achieve this by selecting samples based on varied criteria. In this paper we propose a novel criterion which achieves good generalization performance of a classifier by specifically selecting a set of query samples that minimizes the difference in distribution between the labeled and the unlabeled data, after annotation. We explicitly measure this difference based on all candidate subsets of the unlabeled data and select the best subset. The proposed objective is an NP-hard integer programming optimization problem. We provide two optimization techniques to solve this problem. In the first one, the problem is transformed into a convex quadratic programming problem and in the second method the problem is transformed into a linear programming problem. Our empirical studies using publicly available UCI datasets and a biomedical image dataset demonstrate the effectiveness of the proposed approach in comparison with the state-of-the-art batch-mode active learning methods. We also present two extensions of the proposed approach, which incorporate uncertainty of the predicted labels of the unlabeled data and transfer learning in the proposed formulation. Our empirical studies on UCI datasets show that incorporation of uncertainty information improves performance at later iterations while our studies on 20 Newsgroups dataset show that transfer learning improves the performance of the classifier during initial iterations.

Entities: Chemical Disease Gene Species

Keywords: Active learning; Maximum Mean Discrepancy; marginal probability distribution

Year: 2012 PMID： 25309807 PMCID： PMC4191836 DOI： 10.1145/2339530.2339647

Source DB: PubMed Journal: KDD ISSN： 2154-817X

4 in total

1. Domain adaptation via transfer component analysis.

Authors: Sinno Jialin Pan; Ivor W Tsang; James T Kwok; Qiang Yang
Journal: IEEE Trans Neural Netw Date: 2010-11-18

2. Integrating structured biological data by Kernel Maximum Mean Discrepancy.

Authors: Karsten M Borgwardt; Arthur Gretton; Malte J Rasch; Hans-Peter Kriegel; Bernhard Schölkopf; Alex J Smola
Journal: Bioinformatics Date: 2006-07-15 Impact factor: 6.937

3. Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition.

Authors: Chengjun Liu; Harry Wechsler
Journal: IEEE Trans Image Process Date: 2002 Impact factor: 10.856

4. Global analysis of mRNA localization reveals a prominent role in organizing cellular architecture and function.

Authors: Eric Lécuyer; Hideki Yoshida; Neela Parthasarathy; Christina Alm; Tomas Babak; Tanja Cerovina; Timothy R Hughes; Pavel Tomancak; Henry M Krause
Journal: Cell Date: 2007-10-05 Impact factor: 41.582

4 in total

3 in total

1. Constrained Active Learning for Anchor Link Prediction Across Multiple Heterogeneous Social Networks.

Authors: Junxing Zhu; Jiawei Zhang; Quanyuan Wu; Yan Jia; Bin Zhou; Xiaokai Wei; Philip S Yu
Journal: Sensors (Basel) Date: 2017-08-03 Impact factor: 3.576

2. On the improvement of reinforcement active learning with the involvement of cross entropy to address one-shot learning problem.

Authors: Honglan Huang; Jincai Huang; Yanghe Feng; Jiarui Zhang; Zhong Liu; Qi Wang; Li Chen
Journal: PLoS One Date: 2019-06-19 Impact factor: 3.240

3. Double-Criteria Active Learning for Multiclass Brain-Computer Interfaces.

Authors: Qingshan She; Kang Chen; Zhizeng Luo; Thinh Nguyen; Thomas Potter; Yingchun Zhang
Journal: Comput Intell Neurosci Date: 2020-03-10

3 in total