Literature DB >> 32393195

Knowledge-based analyses reveal new candidate genes associated with risk of hepatitis B virus related hepatocellular carcinoma.

Deke Jiang1, Jiaen Deng2, Changzheng Dong3, Xiaopin Ma4, Qianyi Xiao5, Bin Zhou1, Chou Yang1, Lin Wei6,7, Carly Conran8, S Lilly Zheng6, Irene Oi-Lin Ng9,10, Long Yu4, Jianfeng Xu6, Pak C Sham11, Xiaolong Qi1, Jinlin Hou1, Yuan Ji7, Guangwen Cao12, Miaoxin Li13,14,15,16,17.   

Abstract

BACKGROUND: Recent genome-wide association studies (GWASs) have suggested several susceptibility loci of hepatitis B virus (HBV)-related hepatocellular carcinoma (HCC) by statistical analysis at individual single-nucleotide polymorphisms (SNPs). However, these loci only explain a small fraction of HBV-related HCC heritability. In the present study, we aimed to identify additional susceptibility loci of HBV-related HCC using advanced knowledge-based analysis.
METHODS: We performed knowledge-based analysis (including gene- and gene-set-based association tests) on variant-level association p-values from two existing GWASs of HBV-related HCC. Five different types of gene-sets were collected for the association analysis. A number of SNPs within the gene prioritized by the knowledge-based association tests were selected to replicate genetic associations in an independent sample of 965 cases and 923 controls.
RESULTS: The gene-based association analysis detected four genes significantly or suggestively associated with HBV-related HCC risk: SLC39A8, GOLGA8M, SMIM31, and WHAMMP2. The gene-set-based association analysis prioritized two promising gene sets for HCC, cell cycle G1/S transition and NOTCH1 intracellular domain regulates transcription. Within the gene sets, three promising candidate genes (CDC45, NCOR1 and KAT2A) were further prioritized for HCC. Among genes of liver-specific expression, multiple genes previously implicated in HCC were also highlighted. However, probably due to small sample size, none of the genes prioritized by the knowledge-based association analyses were successfully replicated by variant-level association test in the independent sample.
CONCLUSIONS: This comprehensive knowledge-based association mining study suggested several promising genes and gene-sets associated with HBV-related HCC risks, which would facilitate follow-up functional studies on the pathogenic mechanism of HCC.

Entities:  

Keywords:  Hepatitis B virus; Hepatocellular carcinoma; Knowledge-based genetic association; Susceptibility

Mesh:

Substances:

Year:  2020        PMID: 32393195      PMCID: PMC7216662          DOI: 10.1186/s12885-020-06842-0

Source DB:  PubMed          Journal:  BMC Cancer        ISSN: 1471-2407            Impact factor:   4.430


Background

Hepatocellular carcinoma (HCC) is one of the most common cancers worldwide. With 750,000 new HCC cases diagnosed each year, it is the third leading cause of cancer mortality [1]. As many as 30% of patients diagnosed with hepatitis, fibrosis or cirrhosis ultimately develop HCC. In high endemic areas such as Africa and Asia, at least 60% of HCC is associated with hepatitis B virus (HBV) [2]. However, only a minority of HBV carriers develops HCC. HBV carriers with a family history of HCC were estimated to have over two-fold risk for HCC compared with those without a family history of HCC [3]. Furthermore, genetic complex segregation analysis suggested that major genes may be involved in the genetic predisposition to develop HCC at an earlier age [4]. Genome-wide association study (GWAS) is a widely used strategy for identifying risk loci of complex diseases. Recently, several GWASs on risk of HBV-related HCC were conducted using single-nucleotide polymorphisms (SNPs)-based statistical association tests. Multiple susceptibility loci were identified, including rs17401966 in intron 24 of KIF1B at 1p36.22, rs7574865 in intron 3 of STAT4 at 2q32.2–32.3, rs9275319 between HLA-DQB1 and HLA-DQA2 at 6p21.3, rs9272105 between HLA-DQA1 and HLA-DRB1 at 6p21.3, and rs455804 in intron 1 of GRIK1 at 21q21.3 [5-7]. However, these susceptibility loci account for only a small fraction of the contribution of genetics to HBV-related HCC. Identifying additional genetic alterations associated with HBV-related HCC may be difficult due to the relatively weak effects of many individual risk SNPs, which may be unidentifiable with the currently available but relatively small sample sizes [8]. SNP-based statistical association tests alone in GWAS do not have enough power to discover most risk loci for human complex diseases. Gene- and biological pathway-based association analysis has been proposed to enhance statistical power compared with conventional statistical tests, as the former can relieve multiple testing and enrich signals [9]. Moreover, gene- and biological pathway-based analysis also lends itself to introducing more disease-specific knowledge into the analysis. In the present study, we performed a series of knowledge-based analyses (including gene- and gene-set-based association tests) on variant-level association p-values from two in-house GWASs of HBV-related HCC. SNPs within genes prioritized by the knowledge-based analyses were selected for replication in two independent HBV-related HCC case/control samples.

Methods

Two existing GWASs on HBV-related HCC

The association p-values were obtained from two previous GWASs on HBV-related HCC in Chinese populations for meta-analysis and knowledge-based association analysis. One study [7] contained 2689 chronic HBV carriers (1212 HBV-related HCC cases and 1477 controls) recruited from May 2006 to December 2012 by the Qidong Liver Cancer Institute in Jiangsu Province of Mainland China. The other study [10] consisted of 95 HBV-infected HCC patients (cases) and 97 HBV-infected patients without HCC (controls) recruited at Queen Mary Hospital, Hong Kong. The sample inclusion and exclusion criteria were described in the original papers [7, 10].

Subjects in replication studies

The subjects in replication, including 965 chronic HBV carriers with HCC as cases and 923 chronic HBV carriers without HCC as controls, were recruited from the affiliated hospitals of the Second Military Medical University, Shanghai, China. All the samples are of Han Chinese descent and have participated in previously published studies [7, 11]. The inclusion and exclusion criteria for all the subjects have been previously described [7, 11]. Briefly, all the subjects were negative for antibodies to hepatitis C virus, or human immunodeficiency virus; and had no other types of liver disease, such as autoimmune hepatitis, toxic hepatitis, and primary biliary cirrhosis. All the controls were chronic HBV carriers and had, by self-report, no history of HCC or other cancers. Chronic HBV carriers were defined as positive for both hepatitis B surface antigen and antibody immunoglobulin G to hepatitis B core antigen for at least 6 months. All the cases were chronic HBV carriers and diagnosed as HCC patients. The diagnosis of HCC was based on a) positive findings on cytological or pathological examination and/or b) positive images on angiogram, ultrasonography, computed tomography and/or magnetic resonance imaging, combined with an Alpha-fetoprotein level ≥ 400 ng/ml. All the cases were confirmed to not have other cancers by an initial screening. The mean (standard deviation) ages of the cases and controls were 50.8 (±12.2) years and 52.9 (±11.2) years, respectively. The male to female ratio were 5.3 in cases and 1.6 in controls, respectively. The study was performed in accordance with guidelines approved by the local ethical committees from all participating centers involved in both the GWAS stage and the replication stage. A written informed consent to participate in the study was obtained from each subject in accordance with the declaration of Helsinki principles. All study participants approved the storage of their frozen DNA specimens, for research purposes, in our laboratory.

Genotyping and quality control in replication

Genomic DNA from the peripheral blood of all participants in replication was extracted using the QIAamp DNA Blood Mini Kit (QIAGEN GmbH, Hilden, Germany). Genotyping analyses for replication samples were conducted using the Sequenom MassArray system (Sequenom) according to the manufacturer’s instructions. Genotyping quality was examined by a detailed QC procedure consisting of a 95% successful call rate, duplicate calling of genotypes, and internal positive control samples and two water samples (PCR negative controls) included in each 96-well plate. Genotype analysis was performed by technicians in a blind fashion.

Meta-analysis of variants

The association p-values of untyped SNPs were imputed directly by the tool FAPI (http://grass.cgs.hku.hk/limx/fapi/) [12] with default settings. The p-values of the two GWASs were then combined by Stouffer’s Z-score method for meta-analysis on FAPI as well: where in which N is the number of GWASs, zi is the individual z-score of the ith GWAS study, and ni is the sample size of the ith study.

Gene-based and gene-set-based analysis

The knowledge-based secondary analysis platform KGG Version 4.0 (http://grass.cgs.hku.hk/limx/kgg/) was used to map the SNPs onto reference genes (UCSC RefGene hg19), and to perform gene-based and gene-set-based association analysis with default settings. Two types of gene-based association tests, GATES [13] and ECS [14], were employed for the analysis which combined SNP-level association signal according to the best significance and accumulated significance respectively. In addition, LDRT [15] was adopted for gene-set-based association analysis. The phased genotypes of Eastern Asian samples in the 1000 Genomes Project [16] were used to account for linkage disequilibrium of SNPs through KGG. The Benjamini-Hochberg approach was used to control false discovery rate (FDR) of genome-wide genes or genes within gene-sets, which is a more powerful multiple testing approach than Bonferroni correction when there are multiple susceptibility genes.

Variants functional annotation

The genomic annotation tools, HaploReg v4.1 (http://www.broadinstitute.org/mammals/haploreg/haploreg.php) [17] and RegulomeDB Version 1.1 (http://regulomedb.org/) [18], were used to annotate SNPs with epigenomic markers and potential regulatory elements, including regions of DNase I hypersensitivity, binding sites for transcription factors (TFs), promoter regions that have been biochemically characterized to regulate transcription, chromatin states as well as DNase foot printing, PWMs, and DNA Methylation. KGGSeq (Version 1.0) [19, 20] was used to annotate selected SNP with four regulatory or functional prediction scores (including CADD.CScore [21], SuRFR [22], FunSeq2 [23] and cepip [24]).

Results

We first combined the association p-values of variants by meta-analysis from two independent GWASs. Association analyses at genes and multiple functional gene-sets were carried to prioritize potential HBV-related HCC susceptibility genes. A series of prioritized variants were selected from the knowledge-based association analyses to replicate their genetic associations in a group of independent case-control samples. The overall workflow is shown in Fig. 1.
Fig. 1

Knowledge-based prioritization framework of SNPs’ statistical p-values for association with HCC

Knowledge-based prioritization framework of SNPs’ statistical p-values for association with HCC

Genome-wide meta-analysis of two HBV-related HCC GWASs in Chinese populations

Association p-values were imputed based on the linkage disequilibrium (LD) pattern in the Eastern Asian Panel from the 1000 Genomes Project. A genome-wide meta-analysis was then performed with SNP p-values from two existing Chinese HCC GWASs using the tool FAPI [12]. After quality control (QC), 5,375,073 meta-analysis p-values of SNPs were obtained. The Manhattan plot and QQ plots of p-values are shown in Supplementary Figure 1 and Supplementary Figure 2, respectively. At the upper tail of the QQ plot, there is a deviation from the 95% confidence level of the non-hypothesis line, suggesting the existence of association signals at some SNPs. The small proportion of significant signals was consistent with the estimated low heritability in the samples by GCTA, 0.063 (±0.028) on the underlying liability scale [25].

Gene-based association analysis

We then used the meta-analysis p-values for gene-based association analysis by GATES [13] and ECS [14] on KGG (version 4.0) [26]. In addition to SNPs within the untranslated regions, introns and exons, the meta-analysis p-values of SNPs within 5 kb upstream and downstream of a gene were also included in the gene-based association test by GATES and ECS. SNPs in overlapping regions of multiple genes were assigned to all involved genes. The QQ plots of gene-based p-values are shown in Fig. 2.
Fig. 2

Quantile-quantile plot of gene-based p-values and SNP-based p-values a) the p-values produced by GATES b) the p-values produced by ECS

Quantile-quantile plot of gene-based p-values and SNP-based p-values a) the p-values produced by GATES b) the p-values produced by ECS According to the gene-based p-values by GATES, two genes, SLC39A8 and GOLGA8M passed the multiple-testing correction by FDR, 0.05 (Table 1). In addition, two genes, SMIM31 and WHAMMP2, had nearly significant q-values (< 0.06 by GATES) on the genome (Table 1). Interestingly, SMIM31, encoding small integral membrane protein 31, was annotated as a long noncoding RNA gene (LINC01207) previously. We further annotated the pseudogene, WHAMMP2, with known regulatory elements and epigenomic markers by the UCSC genome browser (http://genome.ucsc.edu). Although it is annotated as a pseudogene, there are multiple regulatory factors binding sites and epigenomic markers in WHAMMP2 (See Supplementary Figure 3). These annotations imply that this gene is also functionally active despite its non-protein-coding function. The other gene-based test, ECS, detected no significant gene. The gene with smallest p-value (7.5E-06) is RNF157-AS1.
Table 1

The top 5 genes according to gene-based p-values by GATES and ECS, respectively

GeneCHRType#SNPGATESECS
NominalpCorrectedpaNominalpCorrectedpa
SLC39A84protein-coding gene2221.63E-060.041380.224950.83656
GOLGA8M15protein-coding gene113.19E-060.041381.57E-050.20422
SMIM314protein-coding gene3006.43E-060.055600.004240.52238
WHAMMP215pseudogene149.03E-060.058580.000310.32182
CLDN522protein-coding gene222.76E-050.125960.016650.62074
RNF157-AS117non-coding RNA203.84E-050.126557.50E-060.19448
LRRC914other3480.002490.627654.46E-050.32182
LINC020625non-coding RNA1500.024020.699270.000070.32182
TTL2protein-coding gene1010.012540.656107.64E-050.32182

Note. CHR: chromosome

a The p-values are corrected by the Benjamini-Hochberg FDR approach. SLC39A8, GOLGA8M, SMIM31, WHAMMP2 and CLDN5 are the top five genes according to GATES. RNF157-AS1, GOLGA8M, LRRC9, LINC02062 and TTL are the top five genes according to ECS

The top 5 genes according to gene-based p-values by GATES and ECS, respectively Note. CHR: chromosome a The p-values are corrected by the Benjamini-Hochberg FDR approach. SLC39A8, GOLGA8M, SMIM31, WHAMMP2 and CLDN5 are the top five genes according to GATES. RNF157-AS1, GOLGA8M, LRRC9, LINC02062 and TTL are the top five genes according to ECS

Prioritization of genes in different gene-sets

To select more promising candidate genes for replication in independent samples, we resorted to a series of gene-set resources to prioritize genes with suggestive association p-values. We first examined the association with HCC in 1057 canonical pathways curated in the Molecular Signatures Database (MSigDB V 4.0), after removing the pathways containing too few (< 5) or too many (> 300) genes. The gene-set-based association p-value was performed by LDRT [15] on KGG. Although no gene-sets passed multiple testing (FDR q < 0.05), several promising functional gene sets are prioritized. The top two gene sets according to the p-value are the cell cycle G1/S transition (p = 5.5E-4) and the NOTCH1 intracellular domain regulates transcription (p = 7.1E-4). In the G1/S transition gene set, 12 out 99 genes had gene-based association (p < 0.05, See details in Supplementary Excel Table 1). The gene with the smallest p-value is CDC45 (p = 1.1E-4) in this gene set. In the gene set of NOTCH1 intracellular domain regulates transcription, 10 out 40 genes had gene-based association (p < 0.05, See details in Supplementary Excel Table 1). In the set, NCOR1 had the smallest p-value (p = 5.8E-3). The second gene, KAT2A, had similar p-value (6.6E-3). Then, we investigated whether the genes highly and specifically expressed in human liver were associated with HCC. In the database, Tissue-specific Gene Expression and Regulation (TiGER, http://bioinfo.wilmer.jhu.edu/tiger/), 309 genes preferentially expressed in liver were retrieved. In the human proteome atlas (http://www.proteinatlas.org/humanproteome), 433 genes showing elevated expression of proteins in liver compared to other tissue types were retrieved as well. To reduce potential false positives, we only used overlapping genes in the two sets. As a result, a total of 189 genes were obtained. Three genes (PAH, UGT2B10 and UROC1) had the FDR q values < 0.1 by ECS while GATES did not detect any significant gene (See the genes and p-values in Table 2 and Supplementary Table 1).
Table 2

Genetic association p-values of genes preferentially expressed in liver

Gene SymbolaGATESpECSpECSqCHRStart PositionLength (BP)Number of SNPs
PAH> 0.050.000350.06412103,230,66680,356266
UGT2B100.015040.000790.073469,870,294172,553122
UROC10.027280.001380.0853126,200,00836,60892
TF0.002930.013880.3863133,465,23650,249288
C4A> 0.050.014720.386631,949,83320,62420
SLCO1B1> 0.050.015280.3861221,284,127108,603298
C5> 0.050.016050.3869123,761,95050,603150
GSTA20.040130.018640.386652,614,88413,38959
C4B> 0.050.020120.386631,982,57112,11338
HAO10.032760.021950.386207,863,63157,474118
NAT20.010230.023080.386818,248,791993470
GSTA10.01104> 0.050.697652,656,17012,44450
AQP90.03199> 0.050.7331558,430,57947,531182
SAA20.04169> 0.050.8041118,266,786342955
APOA20.04277> 0.050.7331161,192,081133722

Note. CHR: chromosome; BP base pairs

a Only the genes with a p-value less than 0.05 are listed in this table. The whole gene list is shown in Supplementary Table 1

Genetic association p-values of genes preferentially expressed in liver Note. CHR: chromosome; BP base pairs a Only the genes with a p-value less than 0.05 are listed in this table. The whole gene list is shown in Supplementary Table 1 We also examined the association of recurrent integrated genes by HBV reported in previous studies [27-30], the genes reported to be genetically associated with HBV-related HCC risk in previous studies, and HCC risk genes defined by COSMIC database (http://cancer.sanger.ac.uk/cosmic). However, none of the genes had a promising association p-value with HCC in our samples (see the genes and p-values in Supplementary Tables 2, 3 and 4).

Replication study in independent samples

We replicated genetic association at genes prioritized by the above gene-based and gene-set-based associations in a group of independent HBV-related HCC case-control samples. Due to budget limit, only 21 SNPs were selected for the replication. The SNPs were at prioritized genes according to consistency of their allele frequencies in ancestry matched reference panel in the 1000 Genomes Project and HapMap Project, and/or their predicted functional importance by RegulomeDB (http://regulomedb.org/) with regulatory elements (See examples in Supplementary Figures 3 and 4). After the genotype quality assessment, two SNPs were excluded because they failed to pass the Hardy-Weinberg equilibrium test (p < 0.001). Three genetic models (additive, dominant and recessive) were considered under a logistic regression framework in which the HCC status was adjusted for sex and age. None of the 19 SNPs survived the multiple Bonferroni correction for family-wise error rate 0.05. Only two SNPs, rs17343667 and rs389883, had a nominal p-value below 0.05. The rs17343667, which is located in the first intron of EIF2AK1, had an association p-value equal to 0.02 under the dominant model with an odds ratio of 1.27 for the minor allele, which was found to have a risk effect in both original Qidong and Hong Kong GWAS samples (Table 3). However, its p-value was only 0.15 under the additive model. The rs389883, which is in intron region of STK19, had p-values of 0.026 and 0.032 for HCC association under additive and recessive models, respectively, with a protective effect at the minor allele G. However, in the original Qidong GWAS sample and Hong Kong GWAS sample, G was estimated to have a risk effect. Therefore, the SNP-level replication was generally negative.
Table 3

Summary of genetic association results in the replication

CHRSNPBPCADD.CScoreSuRFRFunSeq2HCCCell_ProbRegulomeDBA1A2AdditiveaDominantaRecessivea
OR (95% CI)pOR (95% CI)pOR (95% CI)p
1rs3813948207,269,858−0.03914.3560.76350.7965CT1.04 (0.87–1.24)0.6681.02 (0.83–1.25)0.8451.27 (0.72–2.23)0.412
2rs6032540216,077,8730.14417.30.18520.3705AT0.85 (0.59–1.21)0.3600.83 (0.58–1.19)0.319- b0.999
3rs7612684178,984,575−0.16319.3340.81090.3704GA0.79 (0.59–1.05)0.1050.76 (0.56–1.03)0.0801.53 (0.23–10.01)0.657
3rs76863563178,987,536−0.49815.4930.18810.3705CT0.91 (0.66–1.26)0.5670.89 (0.64–1.24)0.5062.00 (0.17–24.06)0.587
5rs11696623557,794,613−0.636.0.18520.3703aGA1.07 (0.79–1.46)0.6701.06 (0.78–1.45)0.705- b0.999
5rs125146191,783,6551.7417.5562.7050.3702bCT1.10 (0.94–1.28)0.2521.06 (0.87–1.28)0.5631.45 (0.96–2.20)0.078
6rs38988331,947,4600.14214.2131.6230.3701fGT0.86 (0.75–0.98)0.0260.86 (0.71–1.03)0.1080.73 (0.55–0.97)0.032
6rs61567232,574,171−0.1624.6270.79720.3706GC0.93 (0.81–1.07)0.2930.98 (0.81–1.17)0.7950.74 (0.54–1.01)0.056
7rs173436676,065,1940.39215.5430.88980.3701fAG1.11 (0.96–1.27)0.1511.27 (1.04–1.55)0.0200.97 (0.76–1.24)0.792
7rs5574417518,332,3962.27517.1950.69090.3705AG1.05 (0.90–1.24)0.5241.07 (0.89–1.30)0.4741.02 (0.65–1.62)0.924
8rs16898013124,138,8910.78017.31400.3703aAG0.85 (0.63–1.16)0.3060.85 (0.61–1.16)0.3040.82 (0.11–6.23)0.847
8rs227595937,455,0590.2456.3770.31140.8634AG0.98 (0.86–1.12)0.7911.02 (0.83–1.25)0.8540.93 (0.74–1.16)0.503
8rs273602015,714,529−0.0023.9779.418E-1610.3707CT1.09 (0.94–1.25)0.2551.13 (0.93–1.36)0.2091.06 (0.79–1.44)0.687
10rs300171910,409,365−0.1133.2770.18520.3705GT1.08 (0.94–1.25)0.2881.11 (0.92–1.34)0.2611.08 (0.76–1.52)0.674
11rs1089724362,043,174−0.49715.5114.535E-330.3706GC0.92 (0.79–1.08)0.3110.93 (0.77–1.13)0.4680.81 (0.55–1.20)0.296
12rs7947504539,083,557−0.26415.8220.18810.3705TG0.88 (0.73–1.06)0.1890.91 (0.74–1.12)0.3770.55 (0.29–1.06)0.072
12rs979722118,217,3040.01415.8990.43650.3707CT1.05 (0.91–1.20)0.5121.05 (0.87–1.27)0.5971.09 (0.81–1.46)0.577
16rs1291837656,558,181−0.02512.0434.562E-740.3706TG1.11 (0.96–1.27)0.1531.11 (0.91–1.35)0.3031.19 (0.92–1.54)0.182
20rs242504633,871,6610.09017.7871.780.9182bCT0.98 (0.77–1.24)0.8480.92 (0.72–1.19)0.5402.20 (0.79–6.14)0.134

Note. CHR chromosome, BP base pairs, OR odd ratio, CI confidence interval, A1 minor allele, A2 major allele, CADD.CScore, SuRFR and FunSeq2 scores are annotated by KGGSeq (V1.0). HCCCell_Prob: Probability of cell type-specific regulation in GENCODE liver cancer cells (HepG2)

a This model was tested under Logistic regression model with adjustment for age and sex

b The value is not available

Summary of genetic association results in the replication Note. CHR chromosome, BP base pairs, OR odd ratio, CI confidence interval, A1 minor allele, A2 major allele, CADD.CScore, SuRFR and FunSeq2 scores are annotated by KGGSeq (V1.0). HCCCell_Prob: Probability of cell type-specific regulation in GENCODE liver cancer cells (HepG2) a This model was tested under Logistic regression model with adjustment for age and sex b The value is not available

Discussion

This study utilized knowledge-based approaches to mine new susceptibility loci of HBV-related HCC in existing HBV-related HCC GWAS data sets. The gene-based association analysis suggested four suggestively significant genes including SLC39A8, GOLGA8M, SMIM31and WHAMMP2. The gene-set-based association analysis prioritized three top genes (CDC45, NCOR1 and KAT2A), which have been implicated with HCC previously, mainly through regulated expression. In addition, three genes, PAH, UGT2B10 and UROC1 were also highlighted when multiple-testing correction (FDR q < 0.1) was performed among genes highly and specifically expressed in human liver. However, probably due to small small sample size, no associations prioritized by the knowledge-based association analysis were successfully replicated in an independent sample. The rs17343667 of EIF2AK1 is the only one with suggestive significance. Furthermore, our analysis also suggested that the germline susceptibility loci of HBV-related HCC are unlikely to be enriched in recurrent targeted genes of HBV infection, or HCC risk genes with many somatic mutations. According to our estimation, HCC has relatively low heritability (6.3%). It is unlikely that there are susceptibility genes or loci of large effect size. The association test enriched the association signals of multiple loci in multiple genes with low effect size so that the susceptibility pathways and gene sets can be prioritized. Moreover, it is easier to prioritize potential susceptibility genes given the prioritized gene sets. In our analysis, a non-trivial fraction of genes within the gene sets achieved moderately significant p-values. It is likely that some of the genes may achieve genome-wide significance when sample sizes are increased. However, almost all of the genes would be ignored by the widely-adopted genome-wide p-value threshold (5E-8) in the present samples (1307 cases vs.1574 controls). Our study is the first to show that genetic variations of two genes (SLC39A8 and GOLGA8M) are significantly associated with the development of HBV-related HCC. SLC39A8 encodes a member of the SLC39 family of solute-carrier genes (Zrt/Irt-like protein 8, ZIP8), which may play an important role in autophagy during ethanol exposure in human hepatoma cells [31]. Liu et al. suggested that hepatic ZIP8 deficiency was associated with tumor formation [32]. Moreover, SLC39A8 has been reported to regulate IFN-γ level in T cells [33] and influence trace element homeostasis in liver [34, 35], which may be relevant to the development of HCC. GOLGA8M encodes golgin A8 family member M. Although it has not been linked to cancer, a study suggested that palindromic GOLGA8 core duplicons promoted chromosome microdeletion and evolutionary instability [36]. In addition, two other genes (SMIM31 and WHAMMP2) also achieved suggestively significant p-values. SMIM31 has been implicated as a biomarker for survival of colorectal adenocarcinoma [37] and promoting proliferation of lung adenocarcinoma [38]. RNF157-AS1, which was implicated by ECS, is an antisense RNA gene. Differential expression between tumor and non-tumor tissue at this gene has been founded in lung cancer [39] and ovarian cancer [40]. Anyhow, functional validation studies are needed to explore the mechanisms of the potential roles of these genes in risk of HBV-related HCC. The successful prioritization of two gene sets that are highly relevant to cancer development also implies the power of the knowledge-based analysis. The top two functional gene-sets are cell cycle G1/S transition and NOTCH1 intracellular domain regulates transcription. There have been numerous studies linking these functional gene sets to HCC [41-44]. For example, Wang et al. recently showed that lnc-UCID promotes G1/S Transition and hepatoma growth by preventing DHX9-Mediated CDK6 down-regulation [41]. As the gene with the smallest p-value in the cell cycle G1/S transition gene set, CDC45 encodes cell division control protein 45 and has been linked to many cancers according to its expression, including HCC [45] and colorectal cancer [46]. NCOR1, the gene with the smallest p-value in the gene set of NOTCH1 intracellular domain regulates transcription, encodes a protein that mediates ligand-independent transcription repression of thyroid-hormone and retinoic-acid receptors, which may regulate de novo fatty acids synthesis in liver regeneration and hepatocarcinogenesis in mice [47]. For another gene with similar p-value as NCOR1 in the gene set of NOTCH1 intracellular domain regulates transcription, KAT2A encodes lysine acetyltransferase 2A and was linked to HCC. For instance, Majaz et al. suggested that KAT2A may promote human HCC progression by enhancing AIB1 expression [48]. The highly and specifically expression in human liver is also an effective stratum for prioritization of HCC susceptibility genes. When multiple testing correction is carried out in this gene set, three genes PAH, UGT2B10 and UROC1 achieved suggestive significance level (FDR q < 0.1). All of the three genes have been implicated with HCC by multiple studies. The most significant gene PAH (p = 3.5E-4 and q = 0.064) has the largest number of literature supports, that is, many studies have implicated this gene in development of HCC. For example, Miller et al. showed p-Chlorphenylalanine effect on phenylalanine hydroxylase in hepatoma cells in culture [49]. Gopalakrishnan and Anderson showed the epigenetic activation of phenylalanine hydroxylase in mouse erythroleukemia cells by the cytoplast of rat hepatoma cells [50]. UGT2B10 (p = 7.9E-4) encodes UDP-Glucuronosyltransferase 2B10. Hanioka et al. showed that expression of UGT2B isoforms (including UGT2B10) was significantly increased by AFB1 in HepG2 cells [51]. UROC1 (p = 1.4E-3) encodes enzyme involved in histidine catabolism, metabolizing urocanic acid to formiminoglutamic acid. Zhang et al. showed that UROC1 may play important roles in HCC development, especially alcohol-related HCC development and progression [52]. The negative findings in the curated gene sets of recurrent targeted genes of HBV infection and HCC risk genes with many somatic mutations are unexpected to some extent. Both gene sets appeared to be biologically relevant to the development of HCC. In the analyses, there were no trends that genes with smaller HCC association p-values were enriched in the gene sets. These results suggest that the biological context or connection of underlying susceptibility genes is elusive, and that it is difficult to simply use our current knowledge to identify the unknown susceptibility genes of HCC. Using larger sample for hypothesis-free GWASs is likely the only effective way for identification of HCC risk genes at present. The issue of negative association at variants in replication sample is consistent with that in the discovery sample. Due to small effect size, no variants in the discovery GWAS sample of 1307 HBV-related HCC cases and 1574 controls had a p-value less than the widely-adopted genome-wide cutoff (5E-8). It was the gene-based association analysis combing the p-values of multiple SNPs that achieved genome-wide significant p-values at some genes. Because of budget limit, however, most genes only had one selected SNPs to maximize the total number of genes for replication. Therefore, we were unable to carry out the gene-based association in the replication study as we did in the GWAS sample. Unfortunately, probably due to low effect size, no variants achieved significant p-value in the replication sample of 965 HBV-related HCC cases and 923 controls. The SNP-level negative replication implies either more powerful knowledge-based association study or larger sample is needed for identifying HCC susceptibility genes.

Conclusion

We performed the first systematic gene- and gene-set-based association study of HCC. Our study suggested several promising genes significantly associated with HCC risk, which may shed insights into pathogenic mechanisms of this fatal disorder. However, the failure in replication study also implies small effect size of the susceptibility genes. More hypothesis-free genetic studies with larger sample sizes are needed to elucidate the susceptibility genes and mechanisms of HCC. Additional file 1: Supplementary Figure 1. Manhattan plots of p-values of 5,375,073 SNPs obtained by a meta-analysis of two HBV-related HCC GWASs in Chinese populations. Supplementary Figure 2. The quantile-quantile plots of p-values of 5,375,073 SNPs obtained by a meta-analysis of two HBV-related HCC GWASs in Chinese populations. Supplementary Figure 3. Regulatory annotations at gene WHAMMP2. Supplementary Figure 4. Visualization of regulatory annotation at rs17343667 by RegulomeDB. Supplementary Table 1. Genetic association p-values of genes preferentially expressed in liver. Supplementary Table 2. Association p-values of genes frequently intergraded by HBV. Supplementary Table 3. Association p-values of genes with significant association in previous studies. Supplementary Table 4. Association p-values of COSMIC HCC risk genes. Additional file 2: Excel table of canonical pathways of MSigDB with p < 0.05 by gene-set based association tests.
  52 in total

1.  Role of zinc in the regulation of autophagy during ethanol exposure in human hepatoma cells.

Authors:  J P Liuzzi; Chanwong Yoo
Journal:  Biol Trace Elem Res       Date:  2013-09-24       Impact factor: 3.738

2.  The effects of hepatitis B virus integration into the genomes of hepatocellular carcinoma patients.

Authors:  Zhaoshi Jiang; Suchit Jhunjhunwala; Jinfeng Liu; Peter M Haverty; Michael I Kennemer; Yinghui Guan; William Lee; Paolo Carnevali; Jeremy Stinson; Stephanie Johnson; Jingyu Diao; Stacy Yeung; Adrian Jubb; Weilan Ye; Thomas D Wu; Sharookh B Kapadia; Frederic J de Sauvage; Robert C Gentleman; Howard M Stern; Somasekar Seshagiri; Krishna P Pant; Zora Modrusan; Dennis G Ballinger; Zemin Zhang
Journal:  Genome Res       Date:  2012-01-20       Impact factor: 9.043

3.  Effect of aflatoxin B1 on UDP-glucuronosyltransferase mRNA expression in HepG2 cells.

Authors:  Nobumitsu Hanioka; Yuko Nonaka; Keita Saito; Tomoe Negishi; Keinosuke Okamoto; Hiroyuki Kataoka; Shizuo Narimatsu
Journal:  Chemosphere       Date:  2012-06-29       Impact factor: 7.086

4.  Hepatic ZIP8 deficiency is associated with disrupted selenium homeostasis, liver pathology, and tumor formation.

Authors:  Liu Liu; Xiangrong Geng; Yihong Cai; Bryan Copple; Masafumi Yoshinaga; Jian Shen; Daniel W Nebert; Hua Wang; Zijuan Liu
Journal:  Am J Physiol Gastrointest Liver Physiol       Date:  2018-06-21       Impact factor: 4.052

5.  Genetic variants in five novel loci including CFB and CD40 predispose to chronic hepatitis B.

Authors:  De-Ke Jiang; Xiao-Pin Ma; Hongjie Yu; Guangwen Cao; Dong-Lin Ding; Haitao Chen; Hui-Xing Huang; Yu-Zhen Gao; Xiao-Pan Wu; Xi-Dai Long; Hongxing Zhang; Youjie Zhang; Yong Gao; Tao-Yang Chen; Wei-Hua Ren; Pengyin Zhang; Zhuqing Shi; Wei Jiang; Bo Wan; Hexige Saiyin; Jianhua Yin; Yuan-Feng Zhou; Yun Zhai; Pei-Xin Lu; Hongwei Zhang; Xiaoli Gu; Aihua Tan; Jin-Bing Wang; Xian-Bo Zuo; Liang-Dan Sun; Jun O Liu; Qing Yi; Zengnan Mo; Gangqiao Zhou; Ying Liu; Jielin Sun; Yin Yao Shugart; S Lilly Zheng; Xue-Jun Zhang; Jianfeng Xu; Long Yu
Journal:  Hepatology       Date:  2015-04-28       Impact factor: 17.425

6.  A powerful conditional gene-based association approach implicated functionally important genes for schizophrenia.

Authors:  Miaoxin Li; Lin Jiang; Timothy Shin Heng Mak; Johnny Sheung Him Kwan; Chao Xue; Peikai Chen; Henry Chi-Ming Leung; Liqian Cui; Tao Li; Pak Chung Sham
Journal:  Bioinformatics       Date:  2019-02-15       Impact factor: 6.937

7.  A comprehensive framework for prioritizing variants in exome sequencing studies of Mendelian diseases.

Authors:  Miao-Xin Li; Hong-Sheng Gui; Johnny S H Kwan; Su-Ying Bao; Pak C Sham
Journal:  Nucleic Acids Res       Date:  2012-01-12       Impact factor: 16.971

8.  Histone acetyl transferase GCN5 promotes human hepatocellular carcinoma progression by enhancing AIB1 expression.

Authors:  Sidra Majaz; Zhangwei Tong; Kesong Peng; Wei Wang; Wenjing Ren; Ming Li; Kun Liu; Pingli Mo; Wengang Li; Chundong Yu
Journal:  Cell Biosci       Date:  2016-08-02       Impact factor: 7.133

9.  An essential role of RNF187 in Notch1 mediated metastasis of hepatocellular carcinoma.

Authors:  Lei Zhang; Jiewei Chen; Juanjuan Yong; Liang Qiao; Leibo Xu; Chao Liu
Journal:  J Exp Clin Cancer Res       Date:  2019-09-02

10.  Signatures of Evolutionary Adaptation in Quantitative Trait Loci Influencing Trace Element Homeostasis in Liver.

Authors:  Johannes Engelken; Guadalupe Espadas; Francesco M Mancuso; Nuria Bonet; Anna-Lena Scherr; Victoria Jímenez-Álvarez; Marta Codina-Solà; Daniel Medina-Stacey; Nino Spataro; Mark Stoneking; Francesc Calafell; Eduard Sabidó; Elena Bosch
Journal:  Mol Biol Evol       Date:  2015-11-17       Impact factor: 16.240

View more
  5 in total

1.  Characterization of Hepatitis B Virus Integrations Identified in Hepatocellular Carcinoma Genomes.

Authors:  Pranav P Mathkar; Xun Chen; Arvis Sulovari; Dawei Li
Journal:  Viruses       Date:  2021-02-04       Impact factor: 5.048

Review 2.  Zinc transporters and their functional integration in mammalian cells.

Authors:  Taiho Kambe; Kathryn M Taylor; Dax Fu
Journal:  J Biol Chem       Date:  2021-01-22       Impact factor: 5.157

3.  An enhancer variant at 16q22.1 predisposes to hepatocellular carcinoma via regulating PRMT7 expression.

Authors:  Ting Shen; Ting Ni; Jiaxuan Chen; Haitao Chen; Xiaopin Ma; Guangwen Cao; Tianzhi Wu; Haisheng Xie; Bin Zhou; Gang Wei; Hexige Saiyin; Suqin Shen; Peng Yu; Qianyi Xiao; Hui Liu; Yuzheng Gao; Xidai Long; Jianhua Yin; Yanfang Guo; Jiaxue Wu; Gong-Hong Wei; Jinlin Hou; De-Ke Jiang
Journal:  Nat Commun       Date:  2022-03-09       Impact factor: 14.919

4.  Identifying novel host-based diagnostic biomarker panels for COVID-19: a whole-blood/nasopharyngeal transcriptome meta-analysis.

Authors:  Samaneh Maleknia; Mohammad Javad Tavassolifar; Faezeh Mottaghitalab; Mohammad Reza Zali; Anna Meyfour
Journal:  Mol Med       Date:  2022-08-03       Impact factor: 6.376

Review 5.  Lineage tracing: technology tool for exploring the development, regeneration, and disease of the digestive system.

Authors:  Yue Zhang; Fanhong Zeng; Xu Han; Jun Weng; Yi Gao
Journal:  Stem Cell Res Ther       Date:  2020-10-15       Impact factor: 6.832

  5 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.