Literature DB >> 30010717

The small peptide world in long noncoding RNAs.

Seo-Won Choi1, Hyun-Woo Kim1, Jin-Wu Nam1.   

Abstract

Long noncoding RNAs (lncRNAs) are a group of transcripts that are longer than 200 nucleotides (nt) without coding potential. Over the past decade, tens of thousands of novel lncRNAs have been annotated in animal and plant genomes because of advanced high-throughput RNA sequencing technologies and with the aid of coding transcript classifiers. Further, a considerable number of reports have revealed the existence of stable, functional small peptides (also known as micropeptides), translated from lncRNAs. In this review, we discuss the methods of lncRNA classification, the investigations regarding their coding potential and the functional significance of the peptides they encode.
© The Author(s) 2018. Published by Oxford University Press.

Entities:  

Keywords:  coding-potential prediction; long noncoding RNA (lncRNA); small ORF; small peptide

Year:  2019        PMID: 30010717      PMCID: PMC6917221          DOI: 10.1093/bib/bby055

Source DB:  PubMed          Journal:  Brief Bioinform        ISSN: 1467-5463            Impact factor:   11.622


Introduction

Long noncoding RNAs (lncRNAs) are a heterogeneous group of RNAs >200 nucleotides (nt) in length that lack coding potential, but their gene structures resemble those of RNA polymerase II products, such as mRNAs [1-3]. For a decade, since their discovery in the 1990s, lncRNAs were arguably considered to be junk or by-products of transcription [4]. In 2007, with the aid of high-throughput sequencing technologies, the ENCODE project unveiled an extensive set of noncoding elements with biochemical functions, which largely overlapped with the lncRNA gene loci from mammalian genomes [5]. Since then, researchers have been exploring these cryptic yet possibly functional noncoding transcripts from genomes. As lncRNAs are known to be expressed in specific cell types and developmental stages, early studies aimed at the computational identification of novel transcribed regions, using complementary DNA (cDNA); RNA sequencing (RNA-seq); chromatin immunoprecipitation followed by sequencing (ChIP-seq), 3P-seq and many other types of transcriptome data; and the transcriptome assembly of high-throughput short reads from different cell types and stages [6-13]. As only sequence and locus information were available for candidate noncoding transcripts, lncRNA classifications were initially based on the features that could be derived from sequences, such as predicted open reading frame (ORF) length, sequence conservation and sequence similarity to known coding genes [10-23]. However, the introduction of high-throughput sequencing of ribosome-protected fragments (Ribo-seq) helped us to examine the ribosome association of candidate transcripts in vivo [24]. Surprisingly, many studies repeatedly reported that some lncRNAs showed a strong association with ribosomes, although the association does not always imply that they are actively translated [9, 25–29]. To address whether the ribosomes associated with lncRNAs actively translate them, several studies attempted to detect either movement of the translating ribosome along the lncRNA transcripts, using Ribo-seq [26, 30–36] or peptides coded by lncRNAs, using mass spectrometry (MS), which is an analytical tool that ionizes peptides and measures their mass-to-charge ratio to identify their amino acid (aa) sequences [37-39]. Meanwhile, functional studies of a few well-conserved lncRNAs, such as XIST [40-42], OIP5-AS1 [7], NEAT1 [43, 44] and MALAT1 [45-47], and of cancer-related lncRNAs, such as GAS5, LUCAT1, HOTAIR and ANRIL [48-51], have shed light on the various regulatory roles of lncRNAs in cells. On investigating the functions of lncRNAs, a few studies confirmed that some lncRNAs indeed had small open reading frames (sORF, length <300 nt) that could code for a short peptide with key biological functions [52-63]. The presence of functional small peptides coded by the lncRNAs suggests that these lncRNAs could play dual roles, with both RNA and peptides, and therefore should be reclassified as bifunctional RNAs [64-66]. This review provides a brief overview of computational and combinatorial approaches for the classification of coding/noncoding RNAs and for the systematic identification of small peptides coded by these transcripts and summarizes functional small peptides encoded by invertebrate and vertebrate lncRNAs. Finally, this review discusses the clinical implications of these small peptides and their host lncRNAs.

Classification and annotation of coding and noncoding RNAs

Computational approaches for lncRNA classification

The advancement in RNA-seq and bioinformatics technologies led to the genome-wide identification of novel transcripts from plant and animal genomes. Although the assessment of coding potential was originally devised to detect novel protein-coding genes, the large number of novel transcripts sequenced with RNA-seq motivated researchers to apply it to distinguish protein-coding and noncoding RNAs. Early methods of estimating coding potential were intended to predict characteristics of translated RNAs from the sequence and locus information of novel transcripts. Intrinsic sequence features, including ORF length, sequence homology to known protein sequences, sequence conservation, nucleotide composition, substitution ratio and secondary structure, were invented and used for the calculation of coding potential (Table 1). ORF length is one of the most commonly used features. The use of ORF length is based on the premise that genuine protein-coding genes would include ORFs of sufficient lengths [10, 13–16, 19–22]. Other features include ORF integrity (whether the ORF includes start and stop codons to define its range) [22]. Protein homology is used to search for conserved segments among protein families and is often assessed by alignment with a protein database [14, 15, 20, 21, 67]. Conservation is considered to be a powerful feature because it is known that lncRNAs are less conserved than mRNAs [14, 18, 20, 23, 67]. Nucleotide composition refers to the frequency of certain k-mers or codon usage in coding or noncoding sequences [10–14, 16, 17, 19, 20, 22, 23]. The substitution ratio is the ratio between synonymous and nonsynonymous mutations in a given sequence and is used to assess whether the mutation profile of a given sequence is better explained by those of protein-coding sequence or those of noncoding sequence [17, 18, 67]. Secondary structures were applied to lncRNAs, as it was hypothesized that functional noncoding RNAs would have different secondary structures from mRNAs [12, 14, 23]. However, prediction algorithms available at the time did not consider the biological characteristics of lncRNAs, and, therefore, some features could be less accurate [12]. To train a computational coding-potential model without any biases, many tools adopted machine learning techniques using known coding and noncoding transcripts as training/test data sets. The most popular method, the coding potential calculator (CPC), takes advantage of BLAST-related features (sequence homology) and ORF-related features to train their model using a support vector machine (SVM) [15]. CPC is favored by many researchers, because of its robust performance, despite relatively long running times. Recently, an updated version of CPC, coding potential calculator 2 (CPC2), was introduced [22]. Unlike the original version, CPC2 excludes the time-consuming sequence alignment step and examines four intrinsic features: Fickett TESTCODE score [68], ORF length, ORF integrity and isoelectric point (the pH at which the peptide carries zero net charge), as implemented in other computational tools (Table 1). In addition to CPC and CPC2, other SVM-powered computational approaches, such as CONC, PORTRAIT, CNCI, iSeeRNA and PLEK, have also been developed with different intrinsic features (Table 1). Moreover, Linc-SF combines genetic algorithm and SVM (GA-SVM) techniques to optimize the classification model [12]. Other machine learning approaches, such as logistic regression [coding potential assessment tool (CPAT)], random forest (LncRNA-ID, COME), deep stacking network (lncRNA-MFDL), and the expectation–maximization (EM) algorithm (PhyloCSF), have also been applied to classify the coding/noncoding transcripts (Table 1). Although there have been a few other approaches, such as RNAcode [67], that avoid the training process to enable a more generic use of the algorithm, this trend toward using machine learning approaches continued after the ribosome profiling data were used in the field of classifying coding/noncoding transcripts.
Table 1.

Computational lncRNA classification

MethodMachine learning techniqueFeature
Result
Reference
ORF lengthProtein homologyConservationNucleotide compositionSubstitution ratio (dN/dS)Secondary structuresORF detectionCoding/ noncoding prediction P value
CONCSVMOOOOOO[14]
CPCSVMOOO[15]
PORTRAITSVMOO[16]
sORF finderOOOOO[17]
PhyloCSFEMOOO[18]
RNAcodeOOOOO[67]
CNCISVMOOO[10]
CPATLogistic regressionOO[19]
iSeeRNASVMOOOOO[20]
PLEKSVMOO[11]
Linc-SFGA-SVMOO??[12]
LncRNA-IDBalanced random forestOO?[21]
lncRNA-MFDLDeep stacking networkOO??[13]
CPC2SVMOOOO[22]
COMEBalanced random forestOOOO[23]

Note: ‘?’ mark indicates that the corresponding information could not be found.

Computational lncRNA classification Note: ‘?’ mark indicates that the corresponding information could not be found.

Classification of lncRNA using experimental data

Ribosome profiling, also known as ribosome footprinting, was introduced in 2009 by Ingolia and his colleagues [24]. Ribosome profiling is a technique that reads ribosome-protected RNA fragments (RPFs), which are obtained by stalling ribosomes on RNA with translation-inhibiting chemicals, applying RNase to eliminate unprotected RNAs and sequencing the remaining RNA molecules [24]. Ribosome profiling has enabled the observation of the global translation status and the computational analysis of in vivo translation. A short time after its introduction, Ribo-seq was applied to examine not only ribosome association but also ribosome dynamics during translation to classify coding/noncoding transcripts (Table 2). Ingolia and his colleagues first devised ribosome association as a measure of translation and later adjusted the value with the expression level of each genes, which was termed the translation efficiency (TE) [24, 69]. As an initial metric, TE was based on the amount of RPFs associated with a transcript, and it could not distinguish translating ribosomes from either nonspecific or nontranslating ribosome interactions. To address this issue, diverse derivatives of RPF coverage have been developed as extensions, and some were fed into machine learning algorithms in combinatorial methods. Shortly after the introduction of TE, a method that compares the RPF depth within ORFs to those in untranslated regions (UTRs) was introduced by two groups [26, 70]. Although the metrics they used to measure the features were similar, their conclusions were different with respect to the translation activity of lncRNAs. One study claimed that most lncRNAs are not actively translated [26], while the other claimed that some lncRNAs contain actively translating regions [70]. Many others also deduced conclusions similar to the latter study on implementing certain forms of RPF coverage or the coverage of RPFs with specific lengths [27, 30–33, 39, 71] (Table 2). To further emphasize the characteristics of active translation using RPFs, features representing ribosome dynamics were described in following studies. Bazzini et al. [30] suggested a new metric, ORFscore, which tests the presence of three-nucleotide periodicity. The periodicity originates from the codon-base translocation of ribosomes during translation along mRNAs, which is often represented by the coverage of ribosome reads mapped to the first, second and third nucleotide positions of a codon (also called sub-codon position), with the fraction of the mapped reads being skewed toward the first position [72]. To assign the mapped reads to a certain nucleotide position, it is essential to predict the position of the ribosome P-site on the reads. Normally, mammalian ribosome covers approximately 30 nt of RNA, and, therefore, the P-site is considered to be located at the 15th nucleotide from the 5′-end of the protected read. The three-nucleotide periodicity was adopted by most succeeding methods, such as RibORF classifier, riboHMM, SPECtre, RiboTaper, Rp-Bp and TERIUS (Table 2) [31-36]. As the P-site in Ribo-seq reads can vary according to the RPF read length, these methods should include the information for the P-site to calculate the three-nucleotide periodicity. For instance, RiboTaper requires users to manually detect the P-sites in reads with certain lengths [34], whereas Rp-Bp automatically infers P-sites by Bayesian inference [35].
Table 2.

Combinatorial lncRNA classification

MethodExperimental dataFeature
Result
Reference
Three-nucleotide periodicityRPF coverageRPF length distributionsORF detectionCoding/ noncoding prediction P value
RRSRibo-seqO[26]
TOCRibo-seqO???[70]
FLOSSRibo-seqOO[27]
ORFscoreRibo-seqOO[30]
PROTEOFORMERRibo-seq, MSOOO[39]
ORF-RATERRibo-seqOOO[71]
RibORFRibo-seqOOOOO[31]
riboHMMRibo-seq, RNA-seqOOO[32]
SPECtreRibo-seqOOOO[33]
RiboTaperRibo-seq, RNA-seqOOOO[34]
Rp-BpRibo-seqOO[35]
TERIUSRibo-seqOO[36]

Note: ‘?’ mark indicates that the corresponding information could not be found.

Combinatorial lncRNA classification Note: ‘?’ mark indicates that the corresponding information could not be found. However, the type of experimental data used to detect translated ORFs is not limited to Ribo-seq. MS spectra and global translation initiation sequencing (GTI-seq) have also been incorporated into classification tools to add additional layers of evidence. GTI-seq is a technique that uses two translation inhibitors, lactimidomycin (LTM) and cycloheximide (CHX), to differentiate ribosome initiation from elongation. GTI-seq has the potential to identify translation initiation sites because CHX binds to all translating ribosomes, while LTM preferentially binds to initiating ribosomes with free E-sites. PROTEOFORMER uses both harringtonine- and LTM-treated Ribo-seqs to identify translated regions and translation initiation sites and integrates MS data for peptide identification [39]. Lee and his colleagues [73] developed a GTI-seq technique by combining Ribo-seq data, generated from samples treated with two different translation-inhibiting chemicals, to generate two types of Ribo-seq signal landscapes and to enhance the accuracy of annotating the translation initiation site.

Methods for detecting sORFs

The extent to which lncRNAs can produce small peptides is still debatable; however, it is now widely accepted that some lncRNAs can be translated [64, 66]. Although computational approaches that do not use Ribo-seq and/or MS data have been successful in detecting coding potential in RNAs, the majority of them lack the capability to detect sORFs that could encode small peptides. Most tools, including CONC, CPC, PORTRAIT, CNCI, CPAT, iseeRNA, LncRNA-ID, lncRNA-MFDL and CPC2, consider sORFs to be a feature of noncoding transcripts [10, 13–16, 19–22]. For instance, a group specified that the length distributions of ORFs originated from noncoding transcripts and those originating from coding transcripts were distinct and most clearly separated at approximately 300 nt [19]. Only sORF finder is capable of detecting sORFs in transcripts [17] (Table 1). In contrast, the combinatorial approaches that use Ribo-seq and MS data successfully detected sORFs in the UTRs of mRNAs and in noncoding RNAs. These approaches that rely on experimental data are free from length restrictions on ORFs, thereby enabling the detection of sORFs. Some groups aimed to design a tolerant classifier by implementing length normalization or additional translation signals independent of ORF length. For instance, translated ORF classifier (TOC) used several ribosome-protected read count per kilobase of ORF exons-based features, which were applied to a random forest classifier [70]. ORF-RATER is based on a nonnegative logistic regression of RPF coverage with harringtonine- and LTM-indicated translation start site data [71]. Although PROTEOFORMER takes similar steps as ORF-RATER, which analyzes translation start signals, it requires an RPF coverage of >85% of exons, hindering sORF detection [39]. In contrast, periodicity-based classification methods tend to be less susceptible to scarce ribosome coverage, although they could still suffer from a lack of a statistical significance arising from the scarce RPF coverage. For instance, ORFscore used a chi-square test to imply the significance of three-nucleotide periodicity and detected 190 sORFs coding for peptides of 20–100 aa in length when analyzing Ribo-seqs from zebrafish embryos [30]. RibORF calculates the maximum entropy value, which considers the fraction of RPF reads at the first and second nucleotides of codons to be a feature for building an SVM model [31]. Using RibORF, an 80-aa-ORF in the CEBPZOS lncRNA gene, along with other sORFs in upstream and downstream ORFs, were identified in a breast epithelial cell line [31]. riboHMM uses a hidden Markov model that considers nucleotide triplets associating with RPFs in all three possible frames to be emission probabilities and translated or untranslated states to be hidden states, resulting in a robust detection of sORFs in transcripts, even with a low RPF coverage [32]. In fact, more than half of the novel ORFs identified by riboHMM were shorter than 30 aa in length [32]. In addition, RiboTaper uses a Fourier transformation technique following a multitaper spectral density estimation of RPF signals at P-sites to detect the periodicity, allowing the discovery of multiple upstream ORFs in HEK293 cells [34]. Rp-Bp calculates the marginal likelihood ratios of the coding and noncoding profiles to determine which profile better describes observed data, leading to the detection of approximately 2500 sORFs in HEK293 cells [35]. In summary, computational methods can identify all possible ORFs, including those with low expression levels and without experimental data, but their results may include ORFs that are not translated. In contrast, combinatorial methods can identify ORFs that are actively translated, are non-canonical or are species specific. However, experimental data are needed to run combinatorial methods and often additional data are needed, such as matched RNA-seq; therefore, transcripts with low expression levels are likely to be neglected.

LncRNAs that encode small peptides

Numerous studies have identified translated ORFs from animal and plant lncRNAs using the previously mentioned approaches (Table 3), among which, RPFs, along with other sequence-related features, were most commonly used to detect sORFs in lncRNAs. For instance, a group profiled RPFs from breast epithelial cell and BJ fibroblast cells and found 1204 translated ORFs in 510 lncRNAs using the RibORF classifier [31]. Of 510 lncRNAs with translated ORFs, 412 encoded peptides <100 aa long, and 19 produced peptides <10-aa-long. Analyzing 93 human translated lncRNAs with orthologs in mice, they found that 41 encoded peptides are conserved in mice, presumably implicating their functional importance. Moreover, the translated lncRNAs were preferentially localized in the cytoplasm compared to other lncRNAs [31]. Crappé et al. [76] analyzed public RPFs, to find sORFs embedded in noncoding RNA (ncRNA) genes, and cryptic intergenic loci, using sORF finder, and found 528 and 226 sORFs, respectively, with supporting ribosome association with both ncRNAs and intergenic regions. Of the 528 sORFs found in ncRNAs, 514 were from lncRNAs (Table 3).
Table 3.

Studies that identified small ORFs and short peptides in lncRNA

SpeciesApproachaMethodExperimental dataTranslated ORFs detected in lncRNAsTranslated sORFs detected in lncRNAsMS evidenceReference
HumanEMS8 peptides[74]
C+EORFscoreRibo-seq261 from lncRNAs261[30]
C+ERibORFRibo-seq1204 from 510 lncRNAs[31]
C+EPhyloCSFRibo-seq, MS354 from lncRNAs35422 peptides[75]
C+EHexamer-based coding scoreRibo-seq143 from 390 lncRNAs99[28]
C+ERibORFRibo-seq, MS925 from 233 lncRNAs68618 lncRNAs[29]
MouseC+EsORF finderRibo-seq514 from lncRNAs514[76]
C+EHexamer-based coding scoreRibo-seq137 from 403 lncRNAs107s[28]
C+EPhyloCSFMS98 from lncRNAs9811 peptides[75]
ZebrafishC+EORFscoreRibo-seq, MS535 from lncRNAs 5356 peptides[30]
C+EPhyloCSFMS99 from lncRNAs99[75]
C+EHexamer-based coding scoreRibo-seq379 from 726 lncRNAs155[28]
Fruit flyC+EPhyloCSFMS53 from lncRNAs532 peptides[75]
C+EHexamer-based coding scoreRibo-seq7 from 22 lncRNAs7[28]
YeastERibo-seq, Polysome-seq47 from 331 lncRNAs47[77]
C+EHexamer-based coding scoreRibo-seq5 from 6 lncRNAs5[28]
WormC+EPhyloCSFMS81 from lncRNAs811 peptide[75]
Arabidopsis thaliana C+EHexamer-based coding scoreRibo-seq43 from 93 lncRNAs43[28]

Note: Approacha is denoted as E if the method is purely experimental, C if computational and C + E if combinatorial.

Studies that identified small ORFs and short peptides in lncRNA Note: Approacha is denoted as E if the method is purely experimental, C if computational and C + E if combinatorial. Although ribosome profiling successfully identified sORFs, ribosome occupancy does not guarantee an active translation signal that produces peptides. Therefore, several studies have adopted peptidomics that integrate RNA-seq, Ribo-seq and MS data to explore the peptide product from lncRNA sORFs (Table 3). For instance, Wang and colleagues identified actively translated sORFs using RibORF and detected 1332 ribosome-associated lncRNAs in eight human cell lines. Among those, 233 lncRNAs included 686 sORFs with RPF evidence, 18 of which were confirmed to express small peptides by MS data [29]. Conversely, Bazzini and colleagues [30] first analyzed Ribo-seq data using ORFscore in lncRNAs expressed in zebrafish, identifying 535 sORFs with ribosome association from lncRNAs. To verify the presence of peptides translated from sORFs, they used MS data from zebrafish embryos and confirmed the presence of peptides translated from six sORFs. Many translated sORFs were also detected in invertebrate and metazoan lncRNAs, using Ribo-seq and/or MS data (Table 3). Smith et al. [77] identified 47 sORFs from 331 unannotated RNAs with ribosome occupancy, 20 of which were evolutionarily conserved in other yeasts. Mackowiak et al. [75] developed a computational pipeline that uses MS data to identify conserved sORFs in five vertebrate and invertebrate species. They predicted 2002 conserved sORFs in the UTRs of mRNAs or ncRNAs and validated them using MS spectra data from human cell lines, mouse cells and tissues, and whole animal zebrafish, fly and worm samples. As a result, a number of novel peptides were discovered in each species, including novel peptides from 36 lncRNAs (Table 3).

Functional small peptides

Muscle-related small peptides

Despite the discovery of many small peptides coded by sORFs, the biological functions of only a handful of them have been described (Table 4). These peptides are usually conserved and are involved in a wide range of biological processes. Recent studies reported lncRNA-encoded small peptides related to specific muscle developmental processes in human and mouse, which participate in muscle regeneration and development (Table 4). Matsumoto and colleagues [59] used a peptidomics approach in human and mouse cell lines and tissues and identified an lncRNA encoding a peptide that is conserved in human and mouse. This small peptide, SPAR, is 90-aa-long in human and 75-aa-long in mouse, regulates mTORC1 activation and inhibits muscle regeneration. Zhang et al. [60] identified an 84-aa-long conserved peptide, Minion, which is involved in the regulation of muscle cell fusion in mouse. The human homologue of Minion was also encoded by a transcript previously annotated as an lncRNA and showed a similar function to its mouse counterpart. Similar functional small peptides related to muscle tissues were also discovered in other model organisms. One study identified a 46-aa-long evolutionarily conserved peptide from an lncRNA. This peptide, named myoregulin (MLN), interacts with the sarcoendoplasmic reticulum calcium transport ATPase (SERCA) calcium-ATPase and inhibits calcium reuptake into the sarcoplasmic reticulum [57]. The other peptide related to this function, DWORF, a 34-aa-long peptide, was also identified in mouse and was shown to regulate calcium reuptake [58]. DWORF enhanced the SERCA calcium-ATPase activity and calcium reuptake into the sarcoplasmic reticulum by displacing SERCA inhibitors, including phospholamban, sarcolipin and MLN. Invertebrates also have similar functional lncRNA-encoded peptides. In Drosophila, a member of the SERCA regulating family was initially identified as an ncRNA gene and encodes a transmembrane peptide, sarcolamban (Scl), of 28- or 29-aa [55]. Sarcolamban also inhibits the SERCA calcium-ATPase and regulates heart contractions.
Table 4.

Known functions of small peptides coded by lncRNAs

SpeciesPeptide nameLncRNAPeptide length (aa)FunctionDetailed functionReference
HumanSPARENSG0000023538790Muscle and cancer-related (oncogenic)Negatively regulates mTORC1 activation and inhibits muscle regeneration[59]
Minion/myomixerENSG0000026217984Muscle-relatedRegulates muscle development and muscle cell fusion[60]
HOXB-AS3ENSG0000023310153Cancer-related (tumor-suppressive)Suppresses colon cancer aerobic glycolysis by inhibiting hnRNP A1-dependent PKM splicing[62]
NOBODYENSG0000020427271Cancer-related and othersInvolved in mRNA processing and negatively regulates P-body association[63]
MouseMLNENSMUSG0000001993346Muscle-relatedInteracts with SERCA (calcium-ATPase) and inhibits calcium reuptake into the sarcoplasmic reticulum[57]
DWORFENSMUSG0000010347634Muscle-relatedEnhances SERCA activity and calcium reuptake into the sarcoplasmic reticulum[58]
SPARENSMUSG000000284775Muscle and cancer-related (oncogenic)Negatively regulates mTORC1 activation and inhibits muscle regeneration[59]
Minion/myomixerENSMUSG0000007947184Muscle-relatedRegulates muscle development and muscle cell fusion[60, 61]
ZebrafishToddlerENSDARG0000009472958OthersActivates G protein-coupled apelin receptor (APJ)/APJ signaling and promotes cell movement during gastrulation[56]
Fruit FlyTarsal-less/talFBgn008700311 and 32OthersActivates the transcription factor responsible for cuticle formation[53]
SclFBgn026649228 and 29Muscle-relatedRegulates calcium transport and muscle contraction[55]
PgcFBgn001605371OthersRepresses CTD2 serine phosphorylation in germline progenitor cells[54]
Soy beanENOD40GmENOD4012 and 24OthersInteracts with sucrose synthase and is required for plant–bacteria symbiotic interactions[52]
Known functions of small peptides coded by lncRNAs

Cancer-related small peptides

Recent studies found that small peptides are expressed or regulated during cancer progression, suggesting their roles in cancer development (Figure 1A and B). Matsumoto and his group [59] confirmed that the downregulation of the SPAR peptide resulted in the upregulation of mTORC1, even though the expression of SPAR RNA was unperturbed. Later, Jiang et al. [78] reported that the expression levels of the lncRNA encoding SPAR was inversely correlated with clinical outcomes in non-small cell lung cancer (NSCLC) (Figure 1A). The NOBODY peptide, a 71-aa-long peptide encoded by the lncRNA LINC01420, was also discovered before the reporting of a connection between its lncRNA and nasopharyngeal carcinoma (NPC) [79]. The lncRNA LINC01420 was negatively correlated with overall survival, and its knockdown by small interfering RNA reduced the migration and invasion of NPC cells. The authors observed that the expression of LINC01420 was elevated in both NPC cell lines and tissue samples and that NPC patients with high LINC01420 expression tended to show poor overall survival rates (Figure 1B). The authors, however, did not verify whether the molecule affecting cancer progression was the peptide or the lncRNA, leaving the function of the corresponding gene inconclusive. Huang et al. [62] reported the function of a 53-aa-long conserved peptide encoded in HOXB-AS3, which appeared to be downregulated in cancers. The expression of HOXB-AS3, previously annotated as an lncRNA, was downregulated in acute myeloid leukemia (AML) [80]. Ribo-seq implied that HOXB-AS3 could produce the encrypted peptide, which was also shown to be downregulated in cancer cells. In fact, HOXB-AS3 peptide, but not the RNA itself, suppressed cancer cell growth, colony formation, migration, invasion and tumorigenesis by inhibiting hnRNP A1-dependent PKM splicing [62] (Figure 1C).
Figure 1.

Cancer-related lncRNAs with functional peptides. Left side (gray box) of each figure shows the RNA function. (A)LINC00961 related to NSCLC. (B)LINC01420 related to NPC. (C)HOXB-AS3 transcript related to AML in OCI-AML3 cells. The right side (blue box) shows the functions for the peptides. (A) SPAR inhibiting mTORC1 activation. (B) NOBODY promoting NMD in K562 and HEK293T cells. (C) HOXB-AS3 peptide regulating PKM splicing and suppressing cancer growth.

Cancer-related lncRNAs with functional peptides. Left side (gray box) of each figure shows the RNA function. (A)LINC00961 related to NSCLC. (B)LINC01420 related to NPC. (C)HOXB-AS3 transcript related to AML in OCI-AML3 cells. The right side (blue box) shows the functions for the peptides. (A) SPAR inhibiting mTORC1 activation. (B) NOBODY promoting NMD in K562 and HEK293T cells. (C) HOXB-AS3 peptide regulating PKM splicing and suppressing cancer growth.

Other functional small peptides

Other studies showed that small peptides from lncRNAs also participate in other biological processes (Table 4). The peptide NOBODY was identified through a proteomics approach in both K562 and HEK293T cell lines [63]. This peptide is thought to regulate mRNA processing by interacting with the mRNA decapping complex and to act as a negative regulator of P-body association. Pauli and colleagues [56] identified a zebrafish peptide of 58 aa, Toddler, from a transcript annotated as an lncRNA in zebrafish, mouse and human. Toddler binds to the apelin receptor (APJ) and induces G protein-coupled receptor signaling to promote cell movement during gastrulation in zebrafish. In soybean, using peptide mass fingerprinting, two small peptides (12- and 24-aa-long) translated from the ENOD40 transcript were identified to interact with sucrose synthase, which is required for plant symbiosis [52].

Discussion

The discovery of functional small peptides translated from lncRNAs has encouraged researchers to reexamine the roles of lncRNAs. The repertoire of biological processes that lncRNAs are involved in has grown rapidly, and lncRNAs serve as biomarkers and potential drug targets in many types of diseases [81]. However, the working mechanisms of disease-associated lncRNAs are largely unknown, and their coding potential under diseased conditions are rarely discussed, even for those that are known to harbor ORFs. For example, MALAT1 and KCNQ1OT1 have been associated with cardiovascular disease and are even used as biomarkers, but whether their functions are dependent on the lncRNA or the peptide is not clear [82]. The steroid receptor RNA activator (SRA) lncRNA produces SRA protein in breast cancer cells, and recently, the lncRNA showed a strong oncogenic property in cervical cancer; yet, the contribution of the SRA protein was not explored [83, 84]. The study of the translation status of these lncRNAs might shed light on their significance and clinical implications. The existence of functional peptide products of lncRNAs emphasizes the need to thoroughly separate the RNA functions and peptide functions of lncRNAs. As in the case of HOXB-AS3 and SPAR, researchers that aim to investigate function of lncRNAs that encode a small peptide should clearly discriminate whether the acting molecule is the peptide, the lncRNA or both. In addition, when studying lncRNA function, the isoform that is responsible for the phenotype in question should be examined. This process is crucial, especially in cancer studies, where most isoform candidates are selected by differential expression. Several reports have shown that the major form of an effector lncRNA gene changes between alternative isoforms in cancer, without prominent differences in the gene expression level. It is also widely accepted that different isoforms may have different coding potentials and that changes in the expression levels of isoforms can affect the proteome of cells. Therefore, researchers should first define the exact effector and further inspect the behavior of the molecule to clarify its role. Although researchers have focused on elucidating the functions of lncRNAs, studying the regulation of lncRNA expression is undoubtedly an equally crucial research focus. Given that most known functional small peptides are conserved in other species, the short, highly conserved regions in lncRNAs may be necessary for both producing peptides and the quality control of produced RNAs. Nonsense-mediated mRNA decay (NMD) is a surveillance mechanism that degrades erroneous mRNAs and reacts to spurious translation followed by premature stop codons. As lncRNAs that harbor sORFs are likely to be targeted by NMD, the extent of NMD targeting against lncRNAs should also be explored. LncRNAs, or the peptides coded by them, are expected to be the missing pieces of many molecular mechanisms. Encouraged by the discovery of many novel sORFs residing in genomic locations that were previously thought to be noncoding, repositories of sORFs or small peptides identified by ribosome profiling and/or MS data, such as sORFs.org [85] and SmProt [86], have been developed. In the case of sORF.org, which stores all published sORFs from five species (human, mouse, rat fly, zebrafish and Caenorhabditiselegans), this database provides coding-potential evidence, such as PhyloCSF, FLOSS and ORFscore. SmProt incorporates small peptides from eight species (human, mouse, rat, zebrafish, fly, yeast, C. elegans and Escherichiacoli). However, the coding nature of lncRNAs is still largely unknown, and the group remains heterogeneous, thus far. More rigorous investigation of lncRNAs and the small peptides hidden within them would lead to a profound understanding and give new insights into numerous unsolved conundrums in the fields of biological and medical sciences. Key Points We presented computational and combinatorial approaches that use various experimental data to classify coding and noncoding RNAs and identify sORFs that encode small peptides. Based on supporting experimental data, we demonstrated the challenges in identifying small peptides and the methods to improve their detection. We summarized a list of functional peptides translated from lncRNAs, many of which are evolutionarily conserved in other species and some of which are known to be involved in muscle-related functions and cancer development.
  86 in total

Review 1.  When one is better than two: RNA with dual functions.

Authors:  Damien Ulveling; Claire Francastel; Florent Hubé
Journal:  Biochimie       Date:  2010-11-24       Impact factor: 4.079

2.  Soybean ENOD40 encodes two peptides that bind to sucrose synthase.

Authors:  Horst Rohrig; Jurgen Schmidt; Edvins Miklashevichs; Jeff Schell; Michael John
Journal:  Proc Natl Acad Sci U S A       Date:  2002-02-12       Impact factor: 11.205

3.  Identification of small ORFs in vertebrates using ribosome footprinting and evolutionary conservation.

Authors:  Ariel A Bazzini; Timothy G Johnstone; Romain Christiano; Sebastian D Mackowiak; Benedikt Obermayer; Elizabeth S Fleming; Charles E Vejnar; Miler T Lee; Nikolaus Rajewsky; Tobias C Walther; Antonio J Giraldez
Journal:  EMBO J       Date:  2014-04-04       Impact factor: 11.598

4.  The Xist lncRNA exploits three-dimensional genome architecture to spread across the X chromosome.

Authors:  Jesse M Engreitz; Amy Pandya-Jones; Patrick McDonel; Alexander Shishkin; Klara Sirokman; Christine Surka; Sabah Kadri; Jeffrey Xing; Alon Goren; Eric S Lander; Kathrin Plath; Mitchell Guttman
Journal:  Science       Date:  2013-07-04       Impact factor: 47.728

5.  Control of muscle formation by the fusogenic micropeptide myomixer.

Authors:  Pengpeng Bi; Andres Ramirez-Martinez; Hui Li; Jessica Cannavino; John R McAnally; John M Shelton; Efrain Sánchez-Ortiz; Rhonda Bassel-Duby; Eric N Olson
Journal:  Science       Date:  2017-04-06       Impact factor: 47.728

6.  An architectural role for a nuclear noncoding RNA: NEAT1 RNA is essential for the structure of paraspeckles.

Authors:  Christine M Clemson; John N Hutchinson; Sergio A Sara; Alexander W Ensminger; Archa H Fox; Andrew Chess; Jeanne B Lawrence
Journal:  Mol Cell       Date:  2009-02-12       Impact factor: 17.970

7.  A Peptide Encoded by a Putative lncRNA HOXB-AS3 Suppresses Colon Cancer Growth.

Authors:  Jin-Zhou Huang; Min Chen; Xing-Cheng Gao; Song Zhu; Hongyang Huang; Min Hu; Huifang Zhu; Guang-Rong Yan
Journal:  Mol Cell       Date:  2017-10-05       Impact factor: 17.970

8.  A human microprotein that interacts with the mRNA decapping complex.

Authors:  Nadia G D'Lima; Jiao Ma; Lauren Winkler; Qian Chu; Ken H Loh; Elizabeth O Corpuz; Bogdan A Budnik; Jens Lykke-Andersen; Alan Saghatelian; Sarah A Slavoff
Journal:  Nat Chem Biol       Date:  2016-12-05       Impact factor: 15.040

9.  COME: a robust coding potential calculation tool for lncRNA identification and characterization based on multiple features.

Authors:  Long Hu; Zhiyu Xu; Boqin Hu; Zhi John Lu
Journal:  Nucleic Acids Res       Date:  2016-09-07       Impact factor: 16.971

10.  TERIUS: accurate prediction of lncRNA via high-throughput sequencing data representing RNA-binding protein association.

Authors:  Seo-Won Choi; Jin-Wu Nam
Journal:  BMC Bioinformatics       Date:  2018-02-19       Impact factor: 3.169

View more
  69 in total

1.  Role of LINC00152 in non-small cell lung cancer.

Authors:  Hong Yu; Shu-Bin Li
Journal:  J Zhejiang Univ Sci B       Date:  2020 Mar.       Impact factor: 3.066

2.  An integrative proteogenomics approach reveals peptides encoded by annotated lincRNA in the mouse kidney inner medulla.

Authors:  Cameron T Flower; Lihe Chen; Hyun Jun Jung; Viswanathan Raghuram; Mark A Knepper; Chin-Rang Yang
Journal:  Physiol Genomics       Date:  2020-08-31       Impact factor: 3.107

Review 3.  Peptide-Liganded G Protein-Coupled Receptors as Neurotherapeutics.

Authors:  Lee E Eiden; Ki Ann Goosens; Kenneth A Jacobson; Lorenzo Leggio; Limei Zhang
Journal:  ACS Pharmacol Transl Sci       Date:  2020-03-18

4.  Multiple information carried by RNAs: total eclipse or a light at the end of the tunnel?

Authors:  Baptiste Bogard; Claire Francastel; Florent Hubé
Journal:  RNA Biol       Date:  2020-06-26       Impact factor: 4.652

5.  LINC01420 RNA structure and influence on cell physiology.

Authors:  Daria O Konina; Alexandra Yu Filatova; Mikhail Yu Skoblov
Journal:  BMC Genomics       Date:  2019-05-08       Impact factor: 3.969

Review 6.  Long Noncoding RNAs in Host-Pathogen Interactions.

Authors:  Federica Agliano; Vijay A Rathinam; Andrei E Medvedev; Sivapriya Kailasan Vanaja; Anthony T Vella
Journal:  Trends Immunol       Date:  2019-04-30       Impact factor: 16.687

7.  Analysis of Soybean Long Non-Coding RNAs Reveals a Subset of Small Peptide-Coding Transcripts.

Authors:  Xiao Lin; Wengui Lin; Yee-Shan Ku; Fuk-Ling Wong; Man-Wah Li; Hon-Ming Lam; Sai-Ming Ngai; Ting-Fung Chan
Journal:  Plant Physiol       Date:  2019-12-27       Impact factor: 8.340

Review 8.  Long Noncoding RNAs and Their Therapeutic Promise in Diabetic Nephropathy.

Authors:  Juan D Coellar; Jianyin Long; Farhad R Danesh
Journal:  Nephron       Date:  2021-04-14       Impact factor: 2.847

9.  lnc-Rps4l-encoded peptide RPS4XL regulates RPS6 phosphorylation and inhibits the proliferation of PASMCs caused by hypoxia.

Authors:  Yiying Li; Junting Zhang; Hanliang Sun; Yujie Chen; Wendi Li; Xiufeng Yu; Xijuan Zhao; Lixin Zhang; Jianfeng Yang; Wei Xin; Yuan Jiang; Guilin Wang; Wenbin Shi; Daling Zhu
Journal:  Mol Ther       Date:  2021-01-09       Impact factor: 11.454

10.  LncRBase V.2: an updated resource for multispecies lncRNAs and ClinicLSNP hosting genetic variants in lncRNAs for cancer patients.

Authors:  Troyee Das; Aritra Deb; Sibun Parida; Sudip Mondal; Sunirmal Khatua; Zhumur Ghosh
Journal:  RNA Biol       Date:  2020-10-28       Impact factor: 4.652

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.