Literature DB >> 28163983

DPClass: An Effective but Concise Discriminative Patterns-Based Classification Framework.

Jingbo Shang1, Wenzhu Tong1, Jian Peng1, Jiawei Han1.   

Abstract

Pattern-based classification was originally proposed to improve the accuracy using selected frequent patterns, where many efforts were paid to prune a huge number of non-discriminative frequent patterns. On the other hand, tree-based models have shown strong abilities on many classification tasks since they can easily build high-order interactions between different features and also handle both numerical and categorical features as well as high dimensional features. By taking the advantage of both modeling methodologies, we propose a natural and effective way to resolve pattern-based classification by adopting discriminative patterns which are the prefix paths from root to nodes in tree-based models (e.g., random forest). Moreover, we further compress the number of discriminative patterns by selecting the most effective pattern combinations that fit into a generalized linear model. As a result, our discriminative pattern-based classification framework (DPClass) could perform as good as previous state-of-the-art algorithms, provide great interpretability by utilizing only very limited number of discriminative patterns, and predict new data extremely fast. More specifically, in our experiments, DPClass could gain even better accuracy by only using top-20 discriminative patterns. The framework so generated is very concise and highly explanatory to human experts.

Entities:  

Year:  2016        PMID: 28163983      PMCID: PMC5287366          DOI: 10.1137/1.9781611974348.64

Source DB:  PubMed          Journal:  Proc SIAM Int Conf Data Min


  2 in total

1.  The spectrum kernel: a string kernel for SVM protein classification.

Authors:  Christina Leslie; Eleazar Eskin; William Stafford Noble
Journal:  Pac Symp Biocomput       Date:  2002

2.  DROP: an SVM domain linker predictor trained with optimal features selected by random forest.

Authors:  Teppei Ebina; Hiroyuki Toh; Yutaka Kuroda
Journal:  Bioinformatics       Date:  2010-12-17       Impact factor: 6.937

  2 in total
  2 in total

1.  Mining Discriminative Patterns to Predict Health Status for Cardiopulmonary Patients.

Authors:  Qian Cheng; Jingbo Shang; Joshua Juen; Jiawei Han; Bruce Schatz
Journal:  ACM BCB       Date:  2016-10

2.  An Interpretable Classification Framework for Information Extraction from Online Healthcare Forums.

Authors:  Jun Gao; Ninghao Liu; Mark Lawley; Xia Hu
Journal:  J Healthc Eng       Date:  2017-08-03       Impact factor: 2.682

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.