Literature DB >> 30046354

A Novel Approach for Predicting Disease-lncRNA Associations Based on the Distance Correlation Set and Information of the miRNAs.

Haochen Zhao1,2, Linai Kuang1,2, Lei Wang1,2, Zhanwei Xuan1,2.   

Abstract

Recently, accumulating laboratorial studies have indicated that plenty of long noncoding RNAs (lncRNAs) play important roles in various biological processes and are associated with many complex human diseases. Therefore, developing powerful computational models to predict correlation between lncRNAs and diseases based on heterogeneous biological datasets will be important. However, there are few approaches to calculating and analyzing lncRNA-disease associations on the basis of information about miRNAs. In this article, a new computational method based on distance correlation set is developed to predict lncRNA-disease associations (DCSLDA). Comparing with existing state-of-the-art methods, we found that the major novelty of DCSLDA lies in the introduction of lncRNA-miRNA-disease network and distance correlation set; thus DCSLDA can be applied to predict potential lncRNA-disease associations without requiring any known disease-lncRNA associations. Simulation results show that DCSLDA can significantly improve previous existing models with reliable AUC of 0.8517 in the leave-one-out cross-validation. Furthermore, while implementing DCSLDA to prioritize candidate lncRNAs for three important cancers, in the first 0.5% of forecast results, 17 predicted associations are verified by other independent studies and biological experimental studies. Hence, it is anticipated that DCSLDA could be a great addition to the biomedical research field.

Entities:  

Mesh:

Substances:

Year:  2018        PMID: 30046354      PMCID: PMC6038663          DOI: 10.1155/2018/6747453

Source DB:  PubMed          Journal:  Comput Math Methods Med        ISSN: 1748-670X            Impact factor:   2.238


1. Introduction

For long time, RNA was just considered to be transcriptional noise and intermediary between a DNA sequence and its encoded protein [1, 2]. However, sequence analyses point out that more than 98% of the human genome does not encode protein sequences [3]. Furthermore, increasing studies based on biological experiments have indicated that ncRNAs play important roles in numerous critical biological processes such as chromosome dosage compensation, epigenetic regulation, and cell growth [4]. In particular, the lncRNAs, as a class of important ncRNAs with a length more than 200 nucleotides [5], have been found to be associated with a wide range of human diseases, such as breast cancer [6], colorectal cancer [7], lung cancer [8], and cardiovascular diseases [9]. Hence, the study of finding novel disease-lncRNA associations has captured the attention of a lot of researchers and has been considered as one of the hottest topics in the research fields of diseases and lncRNAs. The identification of disease-lncRNA association can not only accelerate the understanding of human complex disease mechanism at the lncRNA level, but also serve as a biomarker identification for human disease diagnosis, treatment, and prevention [10]. So far, a lot of studies have generated a large amount of lncRNAs related biological data about sequence, expression, function, and so on [11-13]. However, compared with the rapidly increasing number of newly discovered lncRNAs, only few known lncRNA-disease associations have been reported. Hence, it is challenging and urgently needed to develop efficient and successful computational approaches to predict potential lncRNA-disease associations. In recent years, some computational methods have been proposed to predict novel lncRNA-disease associations, which can significantly decrease the time and cost of biological experiments by calculating the association probability of lncRNA-disease pairs. For example, Chen G et al. presented the first prediction method (genomic locus based) and constructed a lncRNA-disease association database as well [14]. Liang et al. proposed a genetic mediator and key regulator model to unveil the subtle relationships between lncRNAs and lung cancer. Liu et al. developed a computational framework to accomplish this by combining human lncRNA expression profiles, gene expression profiles, and human disease-associated gene data. Applying this framework to available human long intergenic noncoding RNAs (lincRNAs) expression data, Chen et al. developed a semi-supervised learning method based on framework of Laplacian Regularized Least Squares, LRLSLDA, to infer potential lncRNA-disease associations which did not need negative samples and could obtain a reliable AUC of 0.7760 in the leave-one-out cross-validations [15]. In 2014, Sun et al. constructed a lncRNA functional similarity network and applied random walk with restart (RWR) to infer potential lncRNA-disease associations [16]. In the same year, Li et al. presented a bioinformatics method based on genomic location to predict the lncRNAs associated with vascular disease [17]. Then, Zhao et al. developed a computational method based on the naïve Bayesian classifier to identify cancer-related lncRNAs by integrating genome, regulome, and transcriptome data [18]. In 2015 Zhou et al. proposed a novel rank-based method named RWRHLDA to prioritize candidate lncRNA-disease associations by integrating miRNA-associated lncRNA-lncRNA crosstalk network, disease-disease similarity network, and known lncRNA-disease association network into a heterogeneous network and implemented a random walk with restart on the newly generated heterogeneous network [19]. Nowadays, with advent of many biological datasets, such as LncRNADisease [14], lncRNAdb [20], and NONCODE [13], the number of lncRNA-disease associations is still very limited. In 2015, Chen developed a method, named HGLDA, based on the information of miRNA [21], which predicted lncRNA-disease associations by integrating disease-miRNA associations with lncRNA-miRNA interactions and did not rely on known lncRNA-disease associations. Different from the method of HGLDA proposed by Chen et al., in this article, on the basis of experimentally reported lncRNA-disease associations collected from the HMDD database [22] and miRNA-lncRNA associations collected from the starBase database [23], a novel model based on distance correlation set is developed to predict potential lncRNA-disease associations by integrating known lncRNA-miRNA associations and known miRNA-disease associations. Compared with HGLDA, the advantage of DCSLDA lies in the introduction of the similarity of disease pairs and lncRNA pairs and distance correlation set. In addition, to optimize the prediction performance of DCSLDA, new methods to calculate the similarity of disease-disease pairs and lncRNA-lncRNA pairs are developed simultaneously. Finally, to evaluate the prediction performance of DCSLDA, LOOCV is implemented on the basis of the known lncRNA-disease associations and known lncRNA-cancer associations separately, and simulation results demonstrate that DCSLDA is superior to the state-of-the-art methods and can achieve a reliable AUC of 0.8517 in the LOOCV when the pregiven threshold parameter r is set at 6. Additionally, to further evaluate the prediction performance of DCSLDA, case studies of breast cancer, colorectal cancer, and lung cancer are implemented for DCSLDA; as a result, among the first 0.5% of predictive results, 9, 6, and 2 predicted potential associations are confirmed by recent experimental reports, respectively. Hence, considering the excellent prediction performance of DCSLDA, it is obvious that DSCLDA can become a useful and efficient computational tool for biomedical researches.

2. Materials and Methods

2.1. Disease-miRNA Associations

We downloaded known disease-miRNA associations from the Human MicroRNA Disease Database (HMDD) in July 2017 (see Supplementary file 1), which included 10381 experimentally verified disease-miRNA associations (including 572 miRNAs and 383 diseases). After merging miRNAs which produce the same mature miRNA and eliminating duplicate data, we obtained dataset1 including 5430 disease-miRNA associations (including 383 human diseases and 495 lncRNAs). Let D be the number of different diseases and M1 be the number of different miRNAs collected from the dataset1, respectively, S = {d1, d2,…, d} represent the set of these D different diseases, and S = {m1, m1,…, m1} represent the set of these M1 different miRNAs; then for any given d ∈ S and m1 ∈ S, we can define the Association Strong Correlation (ASC1) between d and m1 as follows:

2.2. miRNA-lncRNA Associations

We downloaded known miRNA-lncRNA associations dataset from starBase v2.0 dataset in July 2017, which provided the most comprehensive experimentally confirmed lncRNA-miRNA interactions based on large scale CLIP-seq data. After data preprocessing (including elimination of duplicate values, erroneous data, disorganized data, and so on), dataset2 (including 10195 lncRNA-miRNA associations, 275 miRNAs, and 1127 lncRNAs) was obtained from the starBase v2.0 (see Supplementary file 2). Let M2 be the number of different miRNAs and L be the number of different lncRNAs collected from the dataset2, S = {m21, m22,…, m2} represent the set of these M2 different miRNAs, and S = {l, l,…, l} represent the set of these L different lncRNAs; then, for any given m2 ∈ S and l ∈ S, we can define the ASC2 between m2i and l as follows:

2.3. lncRNA-Disease Associations

In order to evaluate the performance of DCSLDA, the newly lncRNA-disease associations were downloaded from LncRNADisease database, which integrated more than 1000 lncRNA-disease entries and 475 lncRNA interaction entries, including 321 lncRNAs and 221 diseases from ~500 publications. In this dataset, after duplicate associations and the lncRNA-disease associations involved in either diseases or lncRNAs which were not contained in the dataset1 or dataset2 were removed, 203 high-quality lncRNA-disease associations were obtained finally (see Supplementary file 3).

2.4. Disease Functional Similarity Based on miRNAs

For calculating the functional similarity between diseases, we introduced the concept of social network. In the social network, for any two nodes, we can calculate the similarities between them by comparing and integrating the similarities of nodes associated with these two nodes. In this section, based on the assumption that similar diseases tend to show a similar interaction and noninteraction pattern with the miRNAs, we calculated the disease similarity in the disease-miRNA interactive network. As illustrated in Figure 1, the calculation procedures of disease functional similarity based on miRNAs include 3 steps. First, we constructed miRNA-disease interactive network from known miRNA-disease associations (dataset1), whose topology can be abstracted as an undirected graph G1 = (V1, E1), where V1 = S ∪ S = {d1, d2,…, d, m1, m1,…, m1} is the set of vertices, E1 is the set of edges, and, for any two nodes a, b ∈ V1, there is an edge between a and b in E1, if and only if there are a ∈ S, b ∈ S, and ASC1(a, b) = 1. However, since different miRNA terms in the dataset1 may relate to different numbers of diseases, it is not suitable to assign the same contribution value to different miRNAs. Hence, we define the contribution value of each miRNA as follows:Finally, we defined the functional similarity between diseases di and dj by integrating the miRNAs related to di, dj, or both of them as follows:where FSD is the disease functional similarity matrix calculated based on miRNA and D(d) and D(d) are the number of di related edges and dj related edges in E1, respectively. As an example, in Figure 1, there is FSD  (d1, d2) = exp⁡(C(m1) + C(m3) + C(m4))/(4 + 5 − 3).
Figure 1

The flowchart of functional similarity calculation based on information of miRNA includes three steps: (1) constructing known disease-miRNA association and miRNA-lncRNA association network respectively; (2) obtaining contribution of each miRNA; (3) calculating functional similarity for diseases and lncRNAs, respectively.

2.5. lncRNA Functional Similarity Based on miRNAs

Based on the assumption that similar lncRNAs tend to show a similar interaction and noninteraction pattern with the miRNAs, we can calculate the lncRNA similarity in the lncRNA-miRNA interactive network. Similar to the calculation procedures of disease functional similarity, first, we constructed lncRNA-miRNA interactive network from known lncRNA-miRNA associations (dataset2), whose topology can be abstracted as an undirected graph G2 = (V2, E2), where V2 = S ∪ S = {m21, m22,…, l, l,…, l} is the set of vertices, E2 is the set of edges, and, for any two nodes a, b ∈ V2, there is an edge between a and b in E2, if and only if there are a ∈ S, b ∈ S, and ASC2(a, b) = 1. Then, considering the number of lncRNA-miRNA associations, we defined the contribution value of each miRNA as follows:Additionally, we defined the functional similarity between lncRNA l and l by integrating the miRNAs related to l, l, or both of them as follows:where FSL is the disease functional similarity matrix calculated based on miRNA and D(l) and D(l) are the number of l related edges and l related edges in E2, respectively.

2.6. Method for Predicting Potential Association between lncRNAs and Diseases

Based on the assumptions that similar diseases tend to show a similar interaction and noninteraction pattern with the miRNAs and similar miRNAs tend to show a similar interaction and noninteraction pattern with the lncRNAs, we proposed a novel model, DCSLDA, based on miRNAs and distance correlation set to predict potential disease-lncRNA associations. As illustrated in Figure 2, the procedures of DCSLDA consist of the following 6 major steps.
Figure 2

The procedures of DCSLDA.

Step 1 (construction of the disease-miRNA-lncRNA interaction network).

On the basis of the above descriptions and letting M = M1∩M2, we can construct a disease-miRNA-lncRNA interaction network based on dataset1 and dataset2, whose topology can be abstracted to an undirected graph G3 = (V3, E3), where V3 = S ∪ S ∪ S = {d1, d2,…, d, m, m,…, m, l, l,… , l} is the set of vertices, E3 is the edge set of G3, and ∀l ∈ L, m ∈ M, d ∈ D. There is an edge between l and m in E3, if and only if the lncRNA l relates to the miRNA m. Moreover, there is an edge between m and d in E3, if and only if the miRNA m is related to the disease d. Then, for any given a, b ∈ V3, we can define the ASC3 between a and b as follows:In addition, although we did not use any known disease-lncRNA associations, the diseases and lncRNAs can still be linked by integrating edges between diseases node and miRNAs node and edges between miRNAs nodes and lncRNAs nodes in the G3.

Step 2 (construction of the Adjacency Matrix based on the disease-miRNA-lncRNA interactive network).

We can construct a (D + M + L)×(D + M + L) dimensional Adjacency Matrix (AM) based on the disease-miRNA-lncRNA interactive network as follows: where i ∈ [1, D + M + L] and j ∈ [1, D + M + L].

Step 3 (construction of the shortest distance matrix based on the disease-miRNA-lncRNA interactive network).

Let r be a pregiven positive integer; then we can obtain r matrixes such as AM1, AM2, …, AM based on the Adjacency Matrix. Then, we can construct a (D + M + L)×(D + M + L) dimensional Shortest Path Matrix (SPM) as follows:where i ∈ [1, D + M + L], j ∈ [1, D + M + L], k ∈ [2, r], and k satisfies AM(i, j) ≠ 0 while AM1(i, j) = AM2(i, j) = ⋯ = AM(i, j) = 0.

Step 4 (collection of the distance correlation sets for nodes in the interactive network).

In G = (V, E), let V = {d1, d2,…, d, m, m,…, m, l, l,…, l} = {v1, v2,…, v, v, v,…, v, v , v,…, v}; then for each node v ∈ V, we can obtain its distance correlation set DCS according to the shortest distance matrix as follows:For instance, in the disease-miRNA-lncRNA interaction network illustrated in Figure 3, supposing that we hope to collect the DCSD1, then according to the above description, we can easily know that the distance correlation sets of D1 will be {M1, M2, M3, M4, L1, L2, L3, L4, L5} when r = 2.
Figure 3

Distance correlation set of D1 with r=2.

And thereafter, for any given node v ∈ DCS, where j ≠ i, we can compute the distance correlation coefficient P(i, j) between the node v and v as follows:Hence, based on (11), we can further obtain a (D + M + L)×(D + M + L) dimensional Distance Correlation Coefficient Matrix (DCCM) as follows:where i ∈ [1, D + M + L] and j ∈ [1, D + M + L].

Step 5 (estimation of association degree between a pair of nodes in the disease-miRNA-lncRNA interactive network).

Based on (12), we can obtain distance correlation coefficient of each nodes pair. For any given nodes pair (v, v) in G = (V, E), where V = {d1, d2,…, d, l, l,…, l} = {v1, v2,…, v, v, v,…, v} and {v, v}⊆V, we can obtain the association degree (AD) between them as follows:where i ∈ [1, D + M + L] and j ∈ [1, D + M + L].

Step 6 (construction of the Final Prediction Result Matrix).

Based on (13), let , where C11 is a D × D matrix, C12 is a D × M matrix, C13 is a D × L matrix, C21 is a M × D matrix, C22 is a M × M matrix, C23 is a M × L matrix, C31 is a L × D matrix, C32 is a L × M matrix, and C33 is a L × L matrix. It can be easily inferred that the matrix C13 will be our prediction results, which provided the association probability between each disease and lncRNA. Moreover, we can introduce disease functional similarity and lncRNA functional similarity for C13 as follows:where the entity FAD(i, j) in row i column j reflects the probability that the lncRNA l(j) is related to the disease d(i).

3. Results and Case Studies

To evaluate the prediction performance of DCSLDA, first of all, we implemented LOOCV (leave-one-out cross-validation) to compare DCSLDA with HGLDA [21] based on the lncRNA-disease association dataset downloaded from LncRNADisease database [14]. Next, LOOCV would be implemented to further evaluate the prediction performance of DCSLDA based on the known experimentally verified lncRNA-cancer associations. And then, the effects of the disease functional similarity and the lncRNA functional similarity to the prediction performance of DCSLDA would be analyzed also. Finally, experimental results about the prediction of associations between lncRNAs and three cancers were listed (see Table 1), and the performance comparisons between DCLSDA and HGLDA were implemented according to the rankings of these new disease-related lncRNAs in the case studies of three cancers (see Table 2).
Table 1

17 predicted lncRNA-disease pairs with high predicted value while DCSLDA was applied to three important kinds of cancer (breast cancer, colorectal cancer, and lung cancer).

CancerLncRNAPMID
Breast cancerKCNQ1OT121304052; 26323944

Breast cancerMALAT124525122; 19379481

Breast cancerXIST27248326

Breast cancerNEAT125417700; 28034643

Breast cancerLINC0065726942882

Breast cancerSNHG1628232182

Breast cancerCASP8AP228388918

Breast cancerPPP1R9B26387546

Breast cancerTUG127791993

Colorectal cancerKCNQ1OT116965397; 11340379

Colorectal cancerMALAT125025966

Colorectal cancerXIST17143621

Colorectal cancerNEAT126552600

Colorectal cancerSNHG1626823726

Colorectal cancerCASP8AP222216762

Lung cancerMALAT120937273; 24757675; 24667321

Lung cancerXIST27501756
Table 2

Performance comparisons between DCSLDA and HGLDA based on the rankings of ten lncRNA-disease associations related to three important kinds of cancer (breast cancer, colorectal cancer, and lung cancer).

CancerLncRNADCSLDAHGLDA
Breast cancerKCNQ1OT118

Breast cancerMALAT1430

Breast cancerXIST51

Breast cancerNEAT1812

Breast cancerSNHG16123

Colorectal cancerKCNQ1OT115

Colorectal cancerMALAT143

Colorectal cancerXIST51

Lung cancerMALAT149

Lung cancerXIST51

Average ranks4.97.3

3.1. Performance Evaluation of Potential Disease-lncRNA Association Prediction

According to the lncRNA-disease association datasets downloaded from LncRNADisease database, DCSLDA and HGLDA were applied in the framework of LOOCV, respectively. While the LOOCV was implemented for investigated diseases and lncRNAs, each known lncRNA-disease association would be left out in turn as test sample, and then we further evaluated how well this association ranked relatively to the candidate samples. Here, the candidate samples comprised all potential lncRNA-disease pairs without confirmed associations. Therefore, after the implementation of DCSLDA was completed, the rank of each left-out testing sample relative to the candidate samples could be further obtained. And then, the testing samples with a prediction rank higher than the given threshold were considered successfully predicted. Thus, we could further obtain the corresponding true positive rates (TPR, sensitivity) and false positive rates (FPR, 1-specificity) by setting different thresholds. Here, sensitivity refers to the percentage of test samples that were predicted with ranks higher than the given threshold, and the specificity was computed as the percentage of negative samples with ranks lower than the threshold. Therefore, the receiver-operating characteristics (ROC) curves could be drawn by plotting TPR versus FPR at different thresholds. And then, the areas under ROC curve (AUC) would be further calculated to evaluate the prediction performance of DCSLDA. An AUC value of 1 represented a perfect prediction while an AUC value of 0.5 indicated purely random performance. The results of the performance comparison between DCSLDA and HGLDA were shown in Figure 4. Since the HGLDA method predicts lncRNA-disease associations without relying on the information of known disease-lncRNA association, it was selected for performance comparison with our method DCSLDA. As a result, it is clear that our newly proposed method DCSLDA achieved the AUC of 0.8517 in the framework of LOOCV, which is much higher than the AUC of 0.7621 achieved by HGLDA [21]. Simulation results indicate that DCSLDA significantly improved the performance of HGLDA by at least 0.0896 in the term of AUC values and fully demonstrate the performance superiority of HGLDA.
Figure 4

Performance comparisons between DCSLDA and HGDLA in terms of ROC curve and AUC based on LOOCV.

3.2. Performance Evaluation of Potential lncRNA-Cancer Association Prediction

Cancer has become one of the most dangerous killers for human beings [24, 25], and there is a high incidence of cancer in both developed countries and developing countries. Therefore, to further evaluate the prediction performance of DCSLDA, LOOCV was implemented on the basis of 117 lncRNA-cancer associations collected from the LncRNADisease dataset, and the simulation results were illustrated in Figure 5.
Figure 5

Performance evaluation of potential lncRNA-cancer association prediction in terms of ROC curve and AUC based on LOOCV.

From Figure 5, it is easy to find that DCSLDA achieved the AUC of 0.9015 in the frameworks of LOOCV when r is set as 6, which indicates that our newly proposed method DCSLDA has a reliable predictive performance of cancers, and therefore it is a precise and high efficient method for the lncRNA-disease association prediction.

3.3. Effects of the Disease Functional Similarity and lncRNA Functional Similarity

In formula (14), we defined  FAD = FSD × C13 × FSL. Then, in this section, we will analyze the effects of the disease similarity matrix FSD and the lncRNA similarity matrix FSL through comparing the prediction performances of DCSLDA in the framework of LOOCV while letting FAD = C13 and FAD = FSD × C13 × FSL, respectively. The simulation results are illustrated in Figure 6. It is obvious that DCSLDA achieved the AUCs of 0.8517 while matrixes FSD and FSL were considered, but the AUC achieved by DCSLDA is 0.8352 only when letting FAD = C13. Simulation results indicated that the prediction performance of DCSLDA will be significantly improved by introducing the similarity matrixes FSD and FSC. Moreover, in Table 1, DCSLDA was applied to three important kinds of cancer (breast cancer, colorectal cancer, and lung cancer). As a result, 17 predicted lncRNA-disease pairs with high predicted value were publicly released to benefit the biological experimental validation.
Figure 6

Comparison of effects of the disease functional similarity and lncRNA functional similarity to the prediction performance of PCSLDA in the framework of LOOCV with r =6.

3.4. Case Studies

Obviously, DCSLDA can predict all potential relationships between diseases and lncRNAs in dataset1 and dataset2 simultaneously. And of course, potential associations with high predicted value can be publicly released to benefit the biological experimental validation. It is anticipated that these potential disease-lncRNA associations that significantly share common miRNAs could be validated by biological experiments and provide important complement for experimental studies. Moreover, plentiful evidence has indicated that lncRNAs played important roles in various kinds of human cancers. The predicted results were sorted from best to worse, among which the first 0.5% results are selected to be analyzed (see Supplementary file 4). Case studies about three important kinds of cancers based on top 0.5% of predicted results were implemented to show the predictive performance of DCSLDA. Prediction results were verified based on the recent updates in the LncRNADisease dataset and recently published experimental literature (ranking results have been listed in Table 1). In the world, breast cancer is the most prevalent cancer in women and a major public health problem. Several studies have focused on studying this disease, but more are needed, especially at the genetic and molecular levels [26, 27]. Therefore, it is necessary to predict breast cancer-related lncRNAs and identify lncRNA biomarkers. DCSLDA was implemented to prioritize candidate lncRNAs for breast cancer. Among the first 5% of predictive results, nine breast cancer-related lncRNAs have been confirmed based on recent experimental literature (see Table 1). For example, KCNQ1OT1, MALAT1, XIST, and NEAT1 are experimentally confirmed breast cancer-related lncRNAs, which have been ranked 2nd, 11th, 12th, and 19th in the predicted list based on the model of DCSLDA, respectively. KCNQ1OT1 had significantly higher expression levels in invasive breast carcinoma and was induced by estrogen in estrogen receptor-alpha expressing breast cancer cells [28]. 17β-Estradiol treatment affects breast tumor or nontumor cells proliferation, migration, and invasion in an ERα-independent, but a dose-dependent, way by decreasing the MALAT1 RNA level [29]. XIST expression is significantly reduced in breast cancer cell lines and breast cancer samples [30]. Breast cancer patients with high level of NEAT1 expression show low survival rate [31]. Colorectal cancer (CRC) is a leading cause of cancer deaths worldwide, one of the fundamental processes driving the initiation and progression of CRC is the accumulation of a variety of genetic and epigenetic changes in colon epithelial cells. Colorectal cancer is usually caused by the combination of various factors, such as genetic and epigenetic changes [32, 33]. Specially, lncRNAs have been demonstrated to play a critical role in the development and progression of colon cancer [34]. As a result, six colorectal cancer-related lncRNAs were listed in Table 1. For example, Tanaka K et al. proved that Loss of imprinting of KCNQ1OT1 is considered as a useful marker for diagnosis of colorectal cancer because of its frequent occurrences in colorectal cancer samples [35]. Ji Q et al. findings implied that MALAT1 might be a potential predictor for tumor metastasis and prognosis [36]. Furthermore, the interaction between MALAT1 and SFPQ could be a novel therapeutic target for CRC. Lassmann S et al. proved that expression level change of or DNA amplification of XIST is associated with colorectal cancer [37]. Over the past 30 years, the morbidity and mortality of lung cancer have been increasing and the cancer has the highest incidence and mortality across the world [38]. Due to the early diagnosis of lung cancer and the lack of effective treatment, its survival rate is around 10% within five years, which seriously endangers human health. More and more evidence has shown that lncRNAs play a critical role in treatment of lung cancers. Among the first 5% of predictive results, three predicted lncRNAs have been confirmed by published experimental literature [39]. According to this literature, MALAT1 has been shown to be highly associated with metastasis of lung cancer and promote lung cancer cell motility by regulating motility related gene expression [40, 41]. Long noncoding RNA XIST acts as an oncogene in non-small cell lung cancer by epigenetically repressing KLF2 expression [42]. In addition, performance comparisons between DCSLDA and HGLDA were implemented according to the rankings of these disease-related lncRNAs in the case studies of breast cancer, colorectal cancer, and lung cancer (see Table 2). By ranging the predicated results by HGLDA and our methods from good to bad, we selected the intersection of the underlying disease-lncRNA relationship predicated by HGLDA and the first 0.5 percent of the predicted results by our methods and listed the lncRNA items related to breast cancer, colorectal cancer, and lung cancer in this intersection in Table 2. As a result, DCSLDA significantly improved the prediction ability of HGLDA with higher ranks for these new disease-related lncRNAs.

4. Discussion and Conclusions

In recent years, plenty of studies have generated an enormous amount of biological data related to lncRNAs. Accumulating evidence shows that lncRNAs have played a very important role in the biological functions, and the study of lncRNA-disease association prediction is of great significance to human beings. However, there is a few computational models for predicting potential disease-lncRNA associations based on the information of miRNA. To utilize the wealth of disease-miRNA, miRNA-lncRNA, and disease-lncRNA associations data collected from three datasets and recently published in experimental literature, in this article, the novel model of DCSLDA was developed to predict potential disease-lncRNA associations. We calculated distance correlation set of each node based on disease-miRNA-lncRNA interactive network first and then further integrated disease functional similarity and lncRNA functional similarity for DCSLDA. The important difference from previous computational model is that DCSLDA does not rely on any known disease-lncRNA associations and it predicts disease-lncRNA associations only based on disease-miRNA-lncRNA interactive network. In order to evaluate the prediction performance of DCSLDA, the validation frameworks of LOOCV were implemented based on known disease-lncRNA and cancer-related-lncRNA associations downloaded from LncRNADisease database. And case studies were further implemented to three important cancers (breast cancer, colorectal cancer, and lung cancer) based on recently published experimental literature. The simulation results show that DCSLDA can achieve reliable and excellent prediction performance and is superior to the state-of-the-art methods. Hence, it is anticipated that DCSLDA could play an important role in the prospective biomedical researches. Disease functional similarity plays an important role in disease-related molecular function research. Functional associations between disease-related genes are often used to identify pairs of similar diseases from different perspectives. Calculating lncRNA functional similarity could benefit lncRNA function inference and disease-related lncRNA prioritization. Therefore, based on the two assumptions that (1) similar diseases tend to show a similar interaction and noninteraction pattern with the miRNAs and (2) similar lncRNAs tend to show a similar interaction and noninteraction pattern with the miRNAs, DCSLDA was developed to predict potential disease-related lncRNA by integrating lncRNA functional similarity and disease functional similarity. Simulation results indicated that the prediction performance of DCSLDA will be significantly improved by disease similarity and lncRNA similarity. However, there are also some limitations in our method. Firstly, DCSLDA measures the correlations between lncRNAs and investigated diseases by integrating walks with different lengths in a lncRNA-miRNA-disease network, which is constructed by combining the known disease-miRNA network, miRNA-lncRNA network, and disease similarity network. The value of distance threshold parameters r is an important factor in DCSLDA, and how to select this parameter is not yet solved well. Secondly, although DCSLDA does not rely on any known experimentally verified lncRNA-disease relationships, the performance of DCSLDA was not very satisfactory compared with that of several existing methods. In the future, we will further integrate data of diseases and lncRNAs that do not rely on the lncRNA-disease interactive network, disease-miRNA interactive network, or miRNA-lncRNA interactive network; then these above problems may be well solved. Finally, introducing more reliable measure of disease similarity and lncRNA similarity and developing more reliable similarity integration method would improve the performance of DCSLDA. In particular, disease similarity and lncRNA similarity in this model totally rely on known disease-miRNA and miRNA-lncRNA associations. The performance of DCSLDA would be further improved when sequence similarity of lncRNA and semantic similarity of disease are introduced.
  42 in total

Review 1.  The genetic basis of colorectal cancer: insights into critical pathways of tumorigenesis.

Authors:  D C Chung
Journal:  Gastroenterology       Date:  2000-09       Impact factor: 22.682

2.  Microarray expression profile analysis of long non-coding RNAs in human breast cancer: a study of Chinese women.

Authors:  Nan Xu; Fengliang Wang; Mingming Lv; Lu Cheng
Journal:  Biomed Pharmacother       Date:  2014-12-12       Impact factor: 6.529

3.  Local consolidative therapy may be beneficial in patients with oligometastatic non-small cell lung cancer.

Authors:  Mary Kay Barton
Journal:  CA Cancer J Clin       Date:  2017-01-17       Impact factor: 508.702

4.  Array CGH identifies distinct DNA copy number profiles of oncogenes and tumor suppressor genes in chromosomal- and microsatellite-unstable sporadic colorectal carcinomas.

Authors:  Silke Lassmann; Roland Weis; Frank Makowiec; Jasmine Roth; Mihai Danciu; Ulrich Hopt; Martin Werner
Journal:  J Mol Med (Berl)       Date:  2006-12-02       Impact factor: 4.599

5.  Loss of imprinting of long QT intronic transcript 1 in colorectal cancer.

Authors:  K Tanaka; G Shiota; M Meguro; K Mitsuya; M Oshimura; H Kawasaki
Journal:  Oncology       Date:  2001       Impact factor: 2.935

6.  Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.

Authors:  Ewan Birney; John A Stamatoyannopoulos; Anindya Dutta; Roderic Guigó; Thomas R Gingeras; Elliott H Margulies; Zhiping Weng; Michael Snyder; Emmanouil T Dermitzakis; Robert E Thurman; Michael S Kuehn; Christopher M Taylor; Shane Neph; Christoph M Koch; Saurabh Asthana; Ankit Malhotra; Ivan Adzhubei; Jason A Greenbaum; Robert M Andrews; Paul Flicek; Patrick J Boyle; Hua Cao; Nigel P Carter; Gayle K Clelland; Sean Davis; Nathan Day; Pawandeep Dhami; Shane C Dillon; Michael O Dorschner; Heike Fiegler; Paul G Giresi; Jeff Goldy; Michael Hawrylycz; Andrew Haydock; Richard Humbert; Keith D James; Brett E Johnson; Ericka M Johnson; Tristan T Frum; Elizabeth R Rosenzweig; Neerja Karnani; Kirsten Lee; Gregory C Lefebvre; Patrick A Navas; Fidencio Neri; Stephen C J Parker; Peter J Sabo; Richard Sandstrom; Anthony Shafer; David Vetrie; Molly Weaver; Sarah Wilcox; Man Yu; Francis S Collins; Job Dekker; Jason D Lieb; Thomas D Tullius; Gregory E Crawford; Shamil Sunyaev; William S Noble; Ian Dunham; France Denoeud; Alexandre Reymond; Philipp Kapranov; Joel Rozowsky; Deyou Zheng; Robert Castelo; Adam Frankish; Jennifer Harrow; Srinka Ghosh; Albin Sandelin; Ivo L Hofacker; Robert Baertsch; Damian Keefe; Sujit Dike; Jill Cheng; Heather A Hirsch; Edward A Sekinger; Julien Lagarde; Josep F Abril; Atif Shahab; Christoph Flamm; Claudia Fried; Jörg Hackermüller; Jana Hertel; Manja Lindemeyer; Kristin Missal; Andrea Tanzer; Stefan Washietl; Jan Korbel; Olof Emanuelsson; Jakob S Pedersen; Nancy Holroyd; Ruth Taylor; David Swarbreck; Nicholas Matthews; Mark C Dickson; Daryl J Thomas; Matthew T Weirauch; James Gilbert; Jorg Drenkow; Ian Bell; XiaoDong Zhao; K G Srinivasan; Wing-Kin Sung; Hong Sain Ooi; Kuo Ping Chiu; Sylvain Foissac; Tyler Alioto; Michael Brent; Lior Pachter; Michael L Tress; Alfonso Valencia; Siew Woh Choo; Chiou Yu Choo; Catherine Ucla; Caroline Manzano; Carine Wyss; Evelyn Cheung; Taane G Clark; James B Brown; Madhavan Ganesh; Sandeep Patel; Hari Tammana; Jacqueline Chrast; Charlotte N Henrichsen; Chikatoshi Kai; Jun Kawai; Ugrappa Nagalakshmi; Jiaqian Wu; Zheng Lian; Jin Lian; Peter Newburger; Xueqing Zhang; Peter Bickel; John S Mattick; Piero Carninci; Yoshihide Hayashizaki; Sherman Weissman; Tim Hubbard; Richard M Myers; Jane Rogers; Peter F Stadler; Todd M Lowe; Chia-Lin Wei; Yijun Ruan; Kevin Struhl; Mark Gerstein; Stylianos E Antonarakis; Yutao Fu; Eric D Green; Ulaş Karaöz; Adam Siepel; James Taylor; Laura A Liefer; Kris A Wetterstrand; Peter J Good; Elise A Feingold; Mark S Guyer; Gregory M Cooper; George Asimenos; Colin N Dewey; Minmei Hou; Sergey Nikolaev; Juan I Montoya-Burgos; Ari Löytynoja; Simon Whelan; Fabio Pardi; Tim Massingham; Haiyan Huang; Nancy R Zhang; Ian Holmes; James C Mullikin; Abel Ureta-Vidal; Benedict Paten; Michael Seringhaus; Deanna Church; Kate Rosenbloom; W James Kent; Eric A Stone; Serafim Batzoglou; Nick Goldman; Ross C Hardison; David Haussler; Webb Miller; Arend Sidow; Nathan D Trinklein; Zhengdong D Zhang; Leah Barrera; Rhona Stuart; David C King; Adam Ameur; Stefan Enroth; Mark C Bieda; Jonghwan Kim; Akshay A Bhinge; Nan Jiang; Jun Liu; Fei Yao; Vinsensius B Vega; Charlie W H Lee; Patrick Ng; Atif Shahab; Annie Yang; Zarmik Moqtaderi; Zhou Zhu; Xiaoqin Xu; Sharon Squazzo; Matthew J Oberley; David Inman; Michael A Singer; Todd A Richmond; Kyle J Munn; Alvaro Rada-Iglesias; Ola Wallerman; Jan Komorowski; Joanna C Fowler; Phillippe Couttet; Alexander W Bruce; Oliver M Dovey; Peter D Ellis; Cordelia F Langford; David A Nix; Ghia Euskirchen; Stephen Hartman; Alexander E Urban; Peter Kraus; Sara Van Calcar; Nate Heintzman; Tae Hoon Kim; Kun Wang; Chunxu Qu; Gary Hon; Rosa Luna; Christopher K Glass; M Geoff Rosenfeld; Shelley Force Aldred; Sara J Cooper; Anason Halees; Jane M Lin; Hennady P Shulha; Xiaoling Zhang; Mousheng Xu; Jaafar N S Haidar; Yong Yu; Yijun Ruan; Vishwanath R Iyer; Roland D Green; Claes Wadelius; Peggy J Farnham; Bing Ren; Rachel A Harte; Angie S Hinrichs; Heather Trumbower; Hiram Clawson; Jennifer Hillman-Jackson; Ann S Zweig; Kayla Smith; Archana Thakkapallayil; Galt Barber; Robert M Kuhn; Donna Karolchik; Lluis Armengol; Christine P Bird; Paul I W de Bakker; Andrew D Kern; Nuria Lopez-Bigas; Joel D Martin; Barbara E Stranger; Abigail Woodroffe; Eugene Davydov; Antigone Dimas; Eduardo Eyras; Ingileif B Hallgrímsdóttir; Julian Huppert; Michael C Zody; Gonçalo R Abecasis; Xavier Estivill; Gerard G Bouffard; Xiaobin Guan; Nancy F Hansen; Jacquelyn R Idol; Valerie V B Maduro; Baishali Maskeri; Jennifer C McDowell; Morgan Park; Pamela J Thomas; Alice C Young; Robert W Blakesley; Donna M Muzny; Erica Sodergren; David A Wheeler; Kim C Worley; Huaiyang Jiang; George M Weinstock; Richard A Gibbs; Tina Graves; Robert Fulton; Elaine R Mardis; Richard K Wilson; Michele Clamp; James Cuff; Sante Gnerre; David B Jaffe; Jean L Chang; Kerstin Lindblad-Toh; Eric S Lander; Maxim Koriabine; Mikhail Nefedov; Kazutoyo Osoegawa; Yuko Yoshinaga; Baoli Zhu; Pieter J de Jong
Journal:  Nature       Date:  2007-06-14       Impact factor: 49.962

7.  LncDisease: a sequence based bioinformatics tool for predicting lncRNA-disease associations.

Authors:  Junyi Wang; Ruixia Ma; Wei Ma; Ji Chen; Jichun Yang; Yaguang Xi; Qinghua Cui
Journal:  Nucleic Acids Res       Date:  2016-02-16       Impact factor: 16.971

Review 8.  LncRNAs: the bridge linking RNA and colorectal cancer.

Authors:  Yanfei Yang; Linjie Zhao; Lingzi Lei; Wayne Bond Lau; Bonnie Lau; Qilian Yang; Xiaobing Le; Huiliang Yang; Chenlu Wang; Zhongyue Luo; Yu Xuan; Yi Chen; Xiangbing Deng; Lian Xu; Min Feng; Tao Yi; Xia Zhao; Yuquan Wei; Shengtao Zhou
Journal:  Oncotarget       Date:  2017-02-14

9.  starBase v2.0: decoding miRNA-ceRNA, miRNA-ncRNA and protein-RNA interaction networks from large-scale CLIP-Seq data.

Authors:  Jun-Hao Li; Shun Liu; Hui Zhou; Liang-Hu Qu; Jian-Hua Yang
Journal:  Nucleic Acids Res       Date:  2013-12-01       Impact factor: 16.971

10.  A four-long non-coding RNA signature in predicting breast cancer survival.

Authors:  Jin Meng; Ping Li; Qing Zhang; Zhangru Yang; Shen Fu
Journal:  J Exp Clin Cancer Res       Date:  2014-10-06
View more
  3 in total

1.  Bioinformatics Approaches for Functional Prediction of Long Noncoding RNAs.

Authors:  Fayaz Seifuddin; Mehdi Pirooznia
Journal:  Methods Mol Biol       Date:  2021

2.  DSCMF: prediction of LncRNA-disease associations based on dual sparse collaborative matrix factorization.

Authors:  Jin-Xing Liu; Ming-Ming Gao; Zhen Cui; Ying-Lian Gao; Feng Li
Journal:  BMC Bioinformatics       Date:  2021-05-12       Impact factor: 3.169

3.  Computational Methods and Applications for Identifying Disease-Associated lncRNAs as Potential Biomarkers and Therapeutic Targets.

Authors:  Congcong Yan; Zicheng Zhang; Siqi Bao; Ping Hou; Meng Zhou; Chongyong Xu; Jie Sun
Journal:  Mol Ther Nucleic Acids       Date:  2020-05-21       Impact factor: 8.886

  3 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.