| Literature DB >> 29297335 |
Morteza Zihayat1, Heidar Davoudi2, Aijun An2.
Abstract
BACKGROUND: Mining frequent gene regulation sequential patterns in time course microarray datasets is an important mining task in bioinformatics. Although finding such patterns are of paramount important for studying a disease, most existing work do not consider gene-disease association during gene regulation sequential pattern discovery. Moreover, they consider more absent/existence effects of genes during the mining process than taking the degrees of genes expression into account. Consequently, such techniques discover too many patterns which may not represent important information to biologists to investigate the relationships between the disease and underlying reasons hidden in gene regulation sequences.Entities:
Keywords: Gene regulation sequential patterns; High utility pattern mining; Time-course microarray datasets
Mesh:
Year: 2017 PMID: 29297335 PMCID: PMC5751562 DOI: 10.1186/s12918-017-0475-4
Source DB: PubMed Journal: BMC Syst Biol ISSN: 1752-0509
An example of a time course microarray dataset
| Patient IDs | Genes |
|
|
|
|
|---|---|---|---|---|---|
|
|
| 2420 | 546 | 100 | 50 |
|
| 321 | 98 | 454 | 974 | |
|
| 410 | 350 | 251 | 243 | |
|
|
| 128 | 786 | 135 | 344 |
|
| 253 | 820 | 482 | 90 | |
|
| 290 | 150 | 256 | 864 | |
|
|
| 600 | 188 | 99 | 40 |
|
| 500 | 555 | 510 | 80 | |
|
| 200 | 400 | 350 | 450 |
Fold changes of gene/probe values
| Patient IDs | Genes |
|
|
|
|
|---|---|---|---|---|---|
|
|
| 1 | 2.2 | -2.4 | -4.8 |
|
| 1 | -3.2 | 1.4 | 3.0 | |
|
| 1 | -1.1 | -1.6 | -1.6 | |
|
|
| 1 | 6.1 | 1.0 | 2.6 |
|
| 1 | 3.2 | 1.9 | -2.8 | |
|
| 1 | -1.9 | -1.1 | 2.9 | |
|
|
| 1 | -3.1 | -6.6 | -15 |
|
| 1 | 1.1 | 1.0 | -6.2 | |
|
| 1 | 2 | 1.7 | 2.2 |
A time course sequential dataset from time course microarray dataset in Table 1
| Patiend IDs | Sequence |
|---|---|
|
|
|
|
|
|
|
|
|
Importance of genes
| Gene |
|
|
|
|---|---|---|---|
| Score | 0.8 | 0.6 | 0.1 |
Fig. 1ItemUtilLists of G 1+, G 2− and G 3− in Tables 3 and 4
Fig. 2An example of HUSP-Tree for the dataset in Tables 3 and 4
Top-20 genes related to pneumonia
| Rank | Gene | Rank | Gene | Rank | Gene | Rank | Gene |
|---|---|---|---|---|---|---|---|
| 1 |
| 6 |
| 11 |
| 16 |
|
| 2 |
| 7 |
| 12 |
| 17 |
|
| 3 |
| 8 |
| 13 |
| 18 |
|
| 4 |
| 9 |
| 14 |
| 19 |
|
| 5 |
| 10 |
| 15 |
| 20 |
|
Top-10 diseases that share genes with pneumonia
| Disease name | Shared genes |
|---|---|
|
| 295 |
|
| 285 |
|
| 274 |
|
| 267 |
|
| 258 |
|
| 258 |
|
| 249 |
|
| 237 |
Top-20 genes related to Rheumatoid Arthritis
| Rank | Gene | Rank | Gene | Rank | Gene | Rank | Gene |
|---|---|---|---|---|---|---|---|
| 1 |
| 6 |
| 11 |
| 16 |
|
| 2 |
| 7 |
| 12 |
| 17 |
|
| 3 |
| 8 |
| 13 |
| 18 |
|
| 4 |
| 9 |
| 14 |
| 19 |
|
| 5 |
| 10 |
| 15 |
| 20 |
|
Top-20 genes related to Asthma
| Rank | Gene | Rank | Gene | Rank | Gene | Rank | Gene |
|---|---|---|---|---|---|---|---|
| 1 |
| 6 |
| 11 |
| 16 |
|
| 2 |
| 7 |
| 12 |
| 17 |
|
| 3 |
| 8 |
| 13 |
| 18 |
|
| 4 |
| 9 |
| 14 |
| 19 |
|
| 5 |
| 10 |
| 15 |
| 20 |
|
Top-4 HUGSs versus Top-4 FGSs with respect to Pneumonia
| Algorithm | ID. | Sequence of genes (e.g., |
|
|
|---|---|---|---|---|
| Top-HUGS |
| (CAT) (CAT MBL2) (CAT) | 9 | 250600 |
|
| (GOLPH3 PDPN) (CAT) (PDPN) (CAT) (PDPN) | 9 | 250325 | |
|
| (CAT MBL2) (CAT) (CAT) | 9 | 249741 | |
|
| (PDPN)(CAT) (PDPN) (CAT) (PDPN) | 9 | 243037 | |
| CTGR-Span |
| (LCN2 S100A12) (LCDN2) | 11 | 59981 |
|
| (LCN2) (S100A12) (CSF3 LCN2) | 11 | 59962 | |
|
| (CSF3 S100A12) | 11 | 59931 | |
|
| (LCN2 S100A12) (LCN2 S100A12) | 11 | 58514 |
Top-4 HUGSs versus Top-4 FGSs with respect to Rheumatoid Arthritis
| Algorithm | ID. | Sequence of genes (e.g., |
|
|
|---|---|---|---|---|
| Top-HUGS |
| (TRAF1 CTLA4 IL1B) (IL2RA CD40) (TRAF1 PADI4 CTLA4) (STAT4 IL2RA) | 6 | 5857 |
|
| (TRAF1 CTLA4 IL1B) (PTPN2 CD40) (TRAF1 PADI4 CTLA4) (STAT4 IL2RA) | 6 | 5856 | |
|
| (TRAF1 CTLA4 IL1B) (CD40) (TRAF1 PADI4 CTLA4 IL1B) (STAT4 IL2RA) | 6 | 5843 | |
|
| (TRAF1 CTLA4 IL1B) (IL2RA PTPN2 CD40) (TRAF1 PADI4 CTLA4) (STAT4) | 6 | 5834 | |
| CTGR-Span |
| (CTLA4) (IL1B) | 11 | 2981 |
|
| (PTPN2 PADI4) (TRAF1) (PTPN2) | 11 | 2964 | |
|
| (TRAF1) (TRAF1) (PTPN2) | 11 | 2961 | |
|
| (ANXA3) (PTPN2 PADI4) | 11 | 2947 |
Top-4 HUGSs versus Top-4 FGSs with respect to Asthma
| Algorithm | ID. | Sequence of genes (e.g., |
|
|
|---|---|---|---|---|
| Top-HUGS |
| (TAF9 KCMF1) (TBCE) (TAF9) (TAF9) | 10 | 218301 |
|
| (TAF9) (TBCE) (TAF9) (VAMP4) | 10 | 215541 | |
|
| (TAF9) (TBCE) (TAF9) (TBCE) | 10 | 207292 | |
|
| (TAF9) (TBCE) (TBCE) | 10 | 201802 | |
| CTGR-Span |
| (CAMP) (CAMP) (CAMP) | 11 | 39216 |
|
| (CAMP) | 11 | 86800 | |
|
| (CAMP) (CAMP) (CAMP) (CAMP) | 11 | 55486 | |
|
| (CAMP) (CAMP) (FCN2 CAMP) (CAMP) | 11 | 15136 |
The average value of Sup, GU, Pop, GU-Pop and Sup-Pop for top-1000 sequences returned by the method
| Method | Sup | GU | Pop | GU-Pop | Sup-Pop |
|---|---|---|---|---|---|
| Top-HUGS | 5 | 198939 | 12.5 | 24.96 | 7.32 |
| CTGR-Span | 10 | 44691 | 1.02 | 2.05 | 1.86 |
Different versions of Top-HUGS
| Method | Baseline | PES | RSO |
|---|---|---|---|
|
| ✓ | × | × |
|
| ✓ | ✓ | × |
|
| ✓ | × | ✓ |
|
| ✓ | ✓ | ✓ |
Fig. 3Run time on the GSE3677 Dataset
Fig. 4Memory usage on the GSE3677 Dataset
Fig. 5First page of the system with parameters
Fig. 6Second page of the system for discovered patterns