已阅读5页,还剩12页未读, 继续免费阅读
版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
Prediction of Complete Gene Structures in Human Genomic DNA Chris Burge* and Samuel Karlin Department of Mathematics Stanford University, Stanford CA, 94305, USA We introduce a general probabilistic model of the gene structure of human genomic sequences which incorporates descriptions of the basic transcriptional, translational and splicing signals, as well as length distri- butions and compositional features of exons, introns and intergenic regions. Distinct sets of model parameters are derived to account for the many substantial differences in gene density and structure observed in distinct C ? G compositional regions of the human genome. In addition, new models of the donor and acceptor splice signals are described which capture potentially important dependencies between signal positions. The model is applied to the problem of gene identifi cation in a computer pro- gram, GENSCAN, which identifi es complete exon/intron structures of genes in genomic DNA. Novel features of the program include the ca- pacity to predict multiple genes in a sequence, to deal with partial as well as complete genes, and to predict consistent sets of genes occurring on either or both DNA strands. GENSCAN is shown to have substan- tially higher accuracy than existing methods when tested on standardized sets of human and vertebrate genes, with 75 to 80% of exons identifi ed exactly. The program is also capable of indicating fairly accurately the re- liability of each predicted exon. Consistently high levels of accuracy are observed for sequences of differing C ? G content and for distinct groups of vertebrates. # 1997 Academic Press Limited Keywords: exon prediction; gene identifi cation; coding sequence; probabilistic model; splice signal*Corresponding author Introduction The problem of identifying genes in genomic DNA sequences by computational methods has at- tracted considerable research attention in recent years. From one point of view, the problem is clo- sely related to the fundamental biochemical issues of specifying the precise sequence determinants of transcription, translation and RNA splicing. On the other hand, with the recent shift in the emphasis of the Human Genome Project from physical map- ping to intensive sequencing, the problem has taken on signifi cant practical importance, and com- puter software for exon prediction is routinely used by genome sequencing laboratories (in con- junction with other methods) to help identify genes in newly sequenced regions. Many early approaches to the problem focused on prediction of individual functional elements, e.g. promoters, splice sites, coding regions, in iso- lation (reviewed by Gelfand, 1995). More recently, a number of approaches have been developed which integrate multiple types of information in- cluding splice signal sensors, compositional prop- erties of coding and non-coding DNA and in some cases database homology searching in order to pre- dict entire gene structures (sets of spliceable exons) in genomic sequences. Some examples of such pro- grams include: FGENEH (Solovyev et al., 1994), GENMARK (Borodovsky Sp, specifi city; CC, correlation coeffi cient; AC, approximate correlation; ME, missed exons; WE, wrong exons; snRNP, small nuclear ribonucleoprotein particle; snRNA, small nuclear RNA; WMM, weight matrix model; WAM, weight array model; MDD, maximal dependence decomposition. J. Mol. Biol. (1997) 268, 7894 00222836/97/16007817 $25.00/0/mb970951# 1997 Academic Press Limited exactly one complete gene (so that, when presented with a sequence containing a partial gene or mul- tiple genes, the results generally do not make sense); and that accuracy measured by indepen- dent control sets may be considerably lower than was originally thought. The issue of the predictive accuracy of such methods has recently been ad- dressedthroughanexhaustivecomparisonof available methods using a large set of vertebrate genesequences(Burset Duret et al., 1995). Our model is similar in its overall architecture to the Generalized Hidden Markov Model approach adopted in the program Genie (Kulp et al., 1996), but differs from most existing programs in several important respects. First, we use an explicitly double-strandedgenomicsequencemodelin which potential genes occuring on both DNA strands are analyzed in simultaneous and inte- grated fashion. Second, while most existing inte- grated gene fi nding programs assume that in each input sequence there is exactly one complete gene, our model treats the general case in which the se- quence may contain a partial gene, a complete gene, multiple complete (or partial) genes, or no geneatall.Thecombinationofthedouble- stranded nature of the model and the capacity to deal with variable numbers of genes should prove particularly useful for analysis of long human genomic contigs, e.g. those of a hundred kilobases or more, which will often contain multiple genes on one or both DNA strands. Third, we introduce a novel method, Maximal Dependence Decompo- sition, to model functional signals in DNA (or pro- tein) sequences which allows for dependencies between signal positions in a fairly natural and statistically justifi able way. This method is applied to generate a model of the donor splice signal whichcapturesseveraltypesofdependencies which may relate to the mechanism of donor splice site recognition in pre-mRNA sequences by U1 smallnuclearribonucleoproteinparticle(U1 snRNP) and possible other factors. Finally, we de- monstrate that the predictive accuracy of GEN- SCAN is substantially better than other methods when tested on standardized sets of human and vertebrate genes, and show that the method can be used effectively to predict novel genes in long genomic contigs. Results GENSCAN was tested on the Burset/Guigo set of 570 vertebrate multi-exon gene sequences (Bur- set only in the category of Missed Exons did GeneID? do better (0.07 versus 0.09 for GEN- Gene Structure Prediction79 SCAN). Use of protein sequence homology infor- mation in conjunction with GENSCAN predictions is addressed in the Discussion. Going beyond exons to the level of whole gene structures, we may defi ne the gene-level accu- racy (GA) for a set of sequences as the proportion of actual genes which are predicted exactly, i.e. all coding exons predicted exactly with no additional predicted exons in the transcription unit (in prac- tice, the annotated GenBank sequence). Gene-level accuracy was 0.43 (243/570) for GENSCAN in the Burset/Guigo set, demonstrating that it is indeed possible to predict complete multi-exon gene struc- tures with a reasonable degree of success by com- puter. It should be noted that this proportion almost certainly overstates the true gene-level ac- Table 1. Performance comparison for Burset/Guigo set of 570 vertebrate genes A Comparison of GENSCAN with other gene prediction programs Accuracy per nucleotideAccuracy per exon ProgramSequencesSnSpACCCSnSpAvg.MEWE GENSCAN570 (8)0.930.930.910.920.780.810.800.090.05 FGENEH569 (22)0.770.880.780.800.610.640.640.150.12 GeneID570 (2)0.630.810.670.650.440.460.450.280.24 Genie570 (0)0.760.770.72n/a0.550.480.510.170.33 GenLang570 (30)0.720.790.690.710.510.520.520.210.22 GeneParser2562 (0)0.660.790.670.650.350.400.370.340.17 GRAIL2570 (23)0.720.870.750.760.360.430.400.250.11 SORFIND561 (0)0.710.850.730.720.420.470.450.240.14 Xpound570 (28)0.610.870.680.690.150.180.170.330.13 GeneID?478 (1)0.910.910.880.880.730.700.710.070.13 GeneParser3478 (1)0.860.910.860.850.560.580.570.140.09 B GENSCAN accuracy for sequences grouped by C ? G content and by organism Accuracy per nucleotideAccuracy per exon SubsetSequencesSnSpACCCSnSpAvg.MEWE C ? G 6056 (0)0.970.890.900.900.760.770.760.070.08 Primates237 (1)0.960.940.930.940.810.820.820.070.05 Rodents191 (4)0.900.930.890.910.750.800.780.110.05 Non-mam. Vert.72 (2)0.930.930.900.930.810.850.840.110.06 A, For each sequence in the test set of 570 vertebrate sequences constructed by Burset Genie results are from Kulp et al. (1996). Recent versions of Genie have demonstrated substantial improvements in accuracy over that given here (M. G. Reese, personal communication). To calculate accuracy statistics, each nucleotide of a test sequence is classifi ed as predicted positive (PP) if it is in a predicted coding region or predicted negative (PN) otherwise, and also as actual posi- tive (AP) if it is a coding nucleotide according to the annotation, or actual negative (AN) otherwise. These assignments are then com- pared to calculate the number of true positives, TP ? PPAP (i.e. the number of nucleotides which are both predicted positives and actual positive); false positives, FP ? PPAN; true negatives, TN ? PNAN; and false negatives, FN ? PNAP. The following mea- sures of accuracy are then calculated: Sensitivity, Sn=TP/AP; Specifi city, Sp=TP/PP; Correlation Coeffi cient, CC ? ?TP?TN? ? ?FP?FN? ? ?PP?PN?AP?AN? p; and the Approximate Correlation, AC ? 1 2 ?TP AP ? TP PP ? TN AN ? TN PN ? ?1: The rationale for each of these defi nitions is discussed by Burset true positives (TP) is the number of predicted exons which exactly match an actual exon (i.e. both endpoints exactly correct). Exon-level sensitivity (Sn) and specifi city (Sp) are then defi ned using the same for- mulas as at the nucleotide level, and the average of Sn and Sp is calculated as an overall measure of accuracy in lieu of a correlation measure. Two additional statistics are calculated at the exon level: Missed Exons (ME) is the proportion of true exons not overlapped by any predicted exon, and Wrong Exons (WE) is the proportion of predicted exons not overlapped by any real exon. Under the heading Sequences, the number of sequences (out of 570) effectively analyzed by each program is given, followed by the number of sequences for which no gene was predicted, in parentheses. Performance of the programs which make use of amino acid similarity searches, GeneID? and GeneParser3, are shown separately at the bottom of the Table: these programs were run only on sequences less than 8 kb in length. B, Results of GENSCAN for different subsets of the Burset/Guigo test set, divided either according to the C ? G% composition of the GenBank sequence or by the organism of origin. Classifi cation by organism was based on the GenBank ORGANISM key. Primate sequences are mostly of human origin; rodent sequences are mostly from mouse and rat; the non-mam- malian vertebrate set contains 22 fi sh, 17 amphibian, 5 reptilian and 28 avian sequences. 80Gene Structure Prediction curacy of GENSCAN because of the substantial bias in the Burset/Guigo set towards small genes (mean: 5.1 kb) with relatively simple intron-exon structure (mean: 4.6 exons per gene). Nevertheless, GENSCAN was able to correctly reconstruct some highly complex genes, the most dramatic example being the human gastric (H ? ? K?)-ATPase gene (accession no. J05451), containing 22 coding exons. The performance of GENSCAN was found to be relatively insensitive to C ? G content (Table 1B), with CC values of 0.93, 0.91, 0.92 and 0.90 ob- served for sequences of 60% C ? G, respectively, and similarly homo- geneous values for the AC statistic. Nor did accu- racy vary substantially for different subgroups of vertebrate species (Table 1B); CC was 0.91 for the rodent subset, 0.94 for primates and 0.93 for a di- verse collection of non-mammalian vertebrate se- quences. A feature which may prove extremely useful in practical applications of GENSCAN is the for- ward-backward probability, p, which is calcu- lated for each predicted exon as described in Methods. Specifi cally, of the 2678 exons predicted in the Burset/Guigo set: 917 had p 0.99 and, of these, 98% were exactly correct; 551 had p 2 0.95, 0.99 (92% correct); 263 had p 2 0.90, 0.95 (88% correct);337 had p 2 0.75, 0.90 (75% correct); 362 had p 2 0.50, 0.75 (54% correct); and 248 had p 2 0.00, 0.50, of which 30% were correct. Thus, the forward-backward probability provides a use- ful guide to the likelihood that a predicted exon is correct and can be used to pinpoint regions of a prediction which are more certain or less certain. From the data above, about one half of predicted exons have p 0.95, with the practical consequence that any (predicted) gene with four or more exons will likely have two or more predicted exons with p 0.95, from which PCR primers could be de- signed to screen a cDNA library with very high likelihood of success. Since for GENSCAN, as for most of the other programs tested, there was a certain degree of overlap between the learning set and the Bur- set/Guigo test set, it was important also to test the method on a truly independent test set. For this purpose, in the construction of the learning set l, we removed all genes more than 25% identical at the amino acid level to the genes of the previously published GeneParser test sets (Snyder thesomewhatlarger fl uctuationsobservedin Table 2 undoubtedly relate to the much smaller size of the GeneParser test sets. Of course, it might be argued that none of the accuracy results described above are truly indica- tive of the programs likely performance on long genomic contigs, since all three of the test sets used consist primarily of relatively short sequences con- taining single genes, whereas contigs currently being generated by genome sequencing labora- tories are often tens to hundreds of kilobases in length and may contain several genes on either or both DNA strands. To our knowledge, only one systematictestofagenepredictionprogram (GRAIL) on long human contigs has so far been re- ported in the literature (Lopez et al., 1994), and the authors encountered a number of diffi culties in car- rying out this test, e.g. it was not always clear whether predicted exons not matching the annota- tion were false positives or might indeed represent real exons which had not been found by the orig- inal submitters of the sequence. As a test of the performance of gene prediction programs on a largehumancontig,weranGENSCANand GRAIL II on the recently sequenced CD4 gene re- gion of human chromosome 12p13 (Ansari-Lari et al., 1996), a contig of 117 kb in length in which six genes have been detected and characterized ex- perimentally. Annotated genes, GENSCAN predicted genes, and GRAIL predicted exons in this sequence are displayed in Figure 1: both programs fi nd most of the known exons in this region, but signifi cant differences between the predictions are observed. Comparison of the GENSCAN predicted genes (GS1 through GS8) with the annotated (known) genes showed that: GS1 corresponds closely to the CD4 gene (the predicted exon at about 1.5 kb is ac- tually a non-coding exon of CD4); GS2 is identical to one of the alternatively spliced forms of Gene A; GS3 contains several exons from both Gene B and GNB3; GS5 is identical to ISOT, except for the ad- dition of one exon at around 74 kb; and GS6 is identical to TPI, except with a different translation start site. This leaves GS4, GS7 and GS8 as poten- tial false positives, which do not correspond to any annotated gene, of which GS7 and GS8 are over- lapped by GRAIL predicted exons. A BLASTP (Altschul et al., 1990) search of the predicted peptides corresponding to GS4, GS7 and GS8 against the non-redundant protein sequence databases revealed that: GS8 is substantially identi- cal (BLAST score 419, P ? 2.6 E-57) to mouse 60 S ribosomalprotein(SwissProtaccessionno. P47963); GS7 is highly similar (BLAST score 150, P ? 2.8 E-32) to Caenorhabditis elegans predicted protein C26E6.5 (GenBank accession no. 532806); and GS4 is not similar to any known protein (no. BLASTP hit with P 0.01). Examination of the se- quence around GS8 suggests that this is probably a 60 S ribosomal protein pseudogene. Predicted gene GS7 might be an expressed gene, but we did not detect any hits against the database of expressed sequence tags (dbEST) to confi rm this. However, we did fi nd several ESTs substantially identical to the predicted 30UTR and exons of GS4 (GenBank accessionno.AA070439,W92850,AA055898, R82668, AA070534, W93300 and others), strongly implying that this is indeed an expressed human gene which was missed by the submitters of this sequence (probably because GRAIL did not detect it). Aside from the prediction of this novel gene, this example also illustrates the potential of GEN- SCAN to predict the number of genes in a se- quence fairly well: of the eight genes predicted, seven correspond closely to known or putative genes and only one (GS3) corresponds to a fusion of exons from two known genes. Discussion As the focus of the human genome project shifts from mapping to large-scale sequencing, the need for effi cient methods for identifying genes in anon- ymous genomic DNA sequences will increase. Ex- perimental approaches will always be required to prove the exact locations, transcriptional activity and splicing patterns of novel genes, but if compu- tational methods can give accurate and reliable in- dicationsofexonlocationsbeforehand,the experimental work involved may often be signifi - cantly reduced. We have developed a probabilistic model of human genomic sequences which ap- proximates many of the important structural and compositional features of human genes, and have described the implementation of this model in the GENSCAN program to predict exon/gene locations in genomic sequences. Novel features of the method include: (1) use of distinct, explicit, em- pirically derived sets of model parameters to cap- ture differences in gene structure and composition between distinct C ? G compositional regions (iso- chores) of the human genome; (2) the capacity to predict multiple genes in a sequence, to deal with partial as well as complete genes, and to predict consistent sets of genes occuring on either or both DNA strands; and (3) new statistical models of donor and acceptor splice sites which
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2026江苏省卫生系统招聘考试(检验学)历年参考题库含答案详解
- 2026教师职称考试(英语)历年参考题库含答案详解
- 基于NLP的博客文章情感判断课程设计
- 彩色的梦微课程设计
- FPGA构建UART通信模块方法课程设计
- 基于Snort的入侵检测系统在的应用分析课程设计
- CAN总线车载通信模拟课程项目课程设计
- 泵与泵站送水课程设计
- 母婴护理高级技师考试试卷及答案
- OCR身份证识别与采集方案课程设计
- 2026年天津市辅警招聘考试试题带答案(精练)
- 2026秋教科版(新教材)小学科学六年级上册(全册)教学设计(附目录p276)
- 旅行社计调绩效考核制度
- GB/T 6188-2017螺栓和螺钉用内六角花形
- GB/T 451.2-2002纸和纸板定量的测定
- GB/T 1690-1992硫化橡胶耐液体试验方法
- 南京农业大学农业设施工程学第一章 设施农业建筑材料2013课件
- 江西南昌大学跨境教育中心国际教育学院管理人员招考聘用(全考点)模拟卷含答案
- 秸秆综合利用技术(与工艺)课件
- 部编教材九年级历史(上)全册教案
- 水知道答案(完整版)
评论
0/150
提交评论