Literature DB >> 30247540

Generalized integrative principal component analysis for multi-type data with block-wise missing structure.

Huichen Zhu1, Gen Li1, Eric F Lock2.   

Abstract

High-dimensional multi-source data are encountered in many fields. Despite recent developments on the integrative dimension reduction of such data, most existing methods cannot easily accommodate data of multiple types (e.g. binary or count-valued). Moreover, multi-source data often have block-wise missing structure, i.e. data in one or more sources may be completely unobserved for a sample. The heterogeneous data types and presence of block-wise missing data pose significant challenges to the integration of multi-source data and further statistical analyses. In this article, we develop a low-rank method, called generalized integrative principal component analysis (GIPCA), for the simultaneous dimension reduction and imputation of multi-source block-wise missing data, where different sources may have different data types. We also devise an adapted Bayesian information criterion (BIC) criterion for rank estimation. Comprehensive simulation studies demonstrate the efficacy of the proposed method in terms of rank estimation, signal recovery, and missing data imputation. We apply GIPCA to a mortality study. We achieve accurate block-wise missing data imputation and identify intriguing latent mortality rate patterns with sociological relevance.
© The Author 2018. Published by Oxford University Press. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.

Keywords:  Block-wise missing imputation; Exponential family; Exponential principal component analysis; Joint and individual variation explained; Multi-view data

Mesh:

Year:  2020        PMID: 30247540     DOI: 10.1093/biostatistics/kxy052

Source DB:  PubMed          Journal:  Biostatistics        ISSN: 1465-4644            Impact factor:   5.899


  7 in total

1.  Integrative factorization of bidimensionally linked matrices.

Authors:  Jun Young Park; Eric F Lock
Journal:  Biometrics       Date:  2019-11-10       Impact factor: 2.571

2.  sJIVE: Supervised Joint and Individual Variation Explained.

Authors:  Elise F Palzer; Christine H Wendt; Russell P Bowler; Craig P Hersh; Sandra E Safo; Eric F Lock
Journal:  Comput Stat Data Anal       Date:  2022-06-14       Impact factor: 2.035

Review 3.  Heterogeneous data integration methods for patient similarity networks.

Authors:  Jessica Gliozzo; Marco Mesiti; Marco Notaro; Alessandro Petrini; Alex Patak; Antonio Puertas-Gallardo; Alberto Paccanaro; Giorgio Valentini; Elena Casiraghi
Journal:  Brief Bioinform       Date:  2022-07-18       Impact factor: 13.994

4.  BIDIMENSIONAL LINKED MATRIX FACTORIZATION FOR PAN-OMICS PAN-CANCER ANALYSIS.

Authors:  Eric F Lock; Jun Young Park; Katherine A Hoadley
Journal:  Ann Appl Stat       Date:  2022-03-28       Impact factor: 1.959

5.  Generalized Co-Clustering Analysis via Regularized Alternating Least Squares.

Authors:  Gen Li
Journal:  Comput Stat Data Anal       Date:  2020-05-04       Impact factor: 1.681

Review 6.  Metabolomics and Multi-Omics Integration: A Survey of Computational Methods and Resources.

Authors:  Tara Eicher; Garrett Kinnebrew; Andrew Patt; Kyle Spencer; Kevin Ying; Qin Ma; Raghu Machiraju; And Ewy A Mathé
Journal:  Metabolites       Date:  2020-05-15

7.  A hierarchical spike-and-slab model for pan-cancer survival using pan-omic data.

Authors:  Sarah Samorodnitsky; Katherine A Hoadley; Eric F Lock
Journal:  BMC Bioinformatics       Date:  2022-06-17       Impact factor: 3.307

  7 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.