版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1、 76 Computer-Aided Translation TechnologyIt is very important to note that corpus analysis tools do not interpret the data - it is still the responsibility of the translator to analyze the information found in the corpus.需要重点注意很重要的一点的是,语料库分析工具不解释数据分析语料库中的信息仍是译者的责任。FURTHER READINGEngwall (1994), Bowk
2、er (1996), Meyer and Mackintosh (1996),Pearson (1998). Austermuhl (2001), and Bowker and Pearson (2002)discuss issues relating to corpus design and compilation.Barnbrook (1996), Kennedy (1998), McEnery and Wilson (1996), and Bowker and Pearson (2002) provide good introductions to corpus linguistics
3、tools and techniques.Bowker (1998, 2000), Lindquist (1999), and Bowker and Pearson (2002) investigate how corpora can be exploited as translation resources.LHomme (1999a, chapter 6) and Bowker and Pearson (2002) explain how monolingual and bilingual concordancers work and explore how they can be use
4、ful to translators.Pearson (1996) and Zanettin (1998) explore how corpus analysis tools can be integrated into the translation classroom.Garside, LeechHYPERLINK /view/53311.htm/view/53311.htm著名语言学家,有固定的译名, and McEnery (1997) provide information on various types of corpus annotation.扩展阅读英格沃尔(1994),鲍克
5、(1996),梅耶和麦金托什(1996),皮尔森(1998),奥斯特穆勒(2001),,和鲍克和皮尔森(2002)讨论了有关与语料库设计和编译相关的问题。巴恩布克(1996),肯尼迪(1998),麦克恩瑞和威尔逊(1996),鲍克和皮尔森(2002)较好地介绍了提供了对于语料库语言工具和技术完善的介绍。鲍克(1998,2000),林奎斯特(1999),鲍克和皮尔森(2002)调查了如何将语料库作为翻译资源进行开发。洛姆(1999 a,,第6章)和鲍克和皮尔森(2002)解释了单语和双语词语语词检索索引的工作机制如何实现并探索译者如何将它们作为有用工具进行使用的这些语词检索工具对译员来说如何利用
6、。皮尔森(1996)和扎内廷(1998)探讨如何将语料库分析工具应用到翻译课堂中去。加赛德,利里奇,麦克恩瑞(1997)提供有关为各种不同类型语料库注释提供了的信息。4. Terminology-Management Systems_ users who try to use standard spreadsheet, database, or word-processing programs to manage terminological data almost inevitably run into problems involving compromised data integrit
7、y due to inadequate modeling features, in addition to difficulties manipulating large volumes of data as resources grow over time. Schmitz (2001, 539)4 术语管理系统那些尝试使用标准表格、,数据库、,或文字处理项目来管理术语数据的用户几乎不可避免地会遇到一些问题,除了难以操作由随着时间的流逝而不断增多的资源产生的大量数据,还。这些问题包括由于不完整建模特征功能不足而导致的破坏数据泄露完整性的问题,除了那些由于操纵迅速增长的大量数据而引起的困难。施
8、密茨(2001, 539)A major part of any translation project is identifying equivalents for specialized terms. Subject fields such as computing, manufacturing, law, and medicine all have significant amounts of field-specific terminology. In addition, many clients will have preferred in-house terminology. Re
9、searching the specific terms needed to complete any given translation is a time-consuming task, and translators do not want to have to repeat all this work each time they begin a new translation. A terminology-management system (TMS) can help with various aspects of the translators terminology-relat
10、ed tasks, including the storage, retrieval, and updating of term records. A TMS can help to ensure greater consistency in the use of terminology. which not only makes documentation easier to read and understand, but also prevents mis-communications. Effective terminology management can help to cut c
11、osts, improve linguistic quality, and reduce turnaround times for translation, which is very important in this age of intense time-to-market pressures.任何翻译项目的主要部分都是识别专业术语的等价项。诸如如计算、制造、法律和医学之类的等学科领域都拥有大量的领域专业独特术语。此外,,很多客户会优先选择内部术语。研究需要完成所有给定翻译的专业的术语需要完成所有给定翻译,这是一项非常耗时的任务,译员们并不想每次开始新的翻译工作时都要重复这项工作。术语管
12、理系统((TMS))注意中英文标点切换可以帮助在译员进行相关术语的各方面翻译工作时给予其各方面的帮助,包括存储、检索和更新术语记录。术语管理系统(TMS)能够确保在术语的使用术语时更加的一致,这不仅会使文档更易于容易阅读和理解,而且可以防止出现错误交流。有效的术语管理有助于可以降低成本,提高语言质量,,减少翻译周转时间,这在这个市场竞争激烈的时代中这些优势十分重要发挥着重要作用。 TMSs have been in existence for some time. Early efforts to use computers for terminology management began i
13、n the 1960s and eventually led to the development of several large-scale term banks, such as Eurodicautom. Termium, and the Banque de terminologie du Quebec (now known as the Grand dictionnaire terminologique), which were maintained on mainframe computers by large organizations. In the 1980s, when d
14、esktop computers became available, personal TMSs were among the first CAT tools commercially available to translators. Although they were very welcome at the time, these early TMSs had some limitations. They were designed to run on a single computer and could not easily be shared. They typically all
15、owed only simple management of bilingual terminology and imposed considerable restrictions on the type and number of data fields as well as on the maximum amount of data that could be stored in these fields. Recently, however, this type of software has become more powerful and flexible, particularly
16、 in terms of storage and retrieval options.术语管理系统已经存在了一段时间。早期利用电脑进行术语管理的前期努力始于20世纪60年代,,最终开发了促进了诸如Eurodicautom、Termium、the Banque de terminologie du Quebec((现在被称为巨型词典术语))几个大型术语存储库的发展,它们这些都是大型组织在主机上保存的存储库。在20世纪80年代,当台式机进入人们的生活,个人的术语管理系统便成为了译者可从市场上买到的最早的首批CAT工具商市售工具中的一种。虽然在当时很受欢大家迎,但这些早期的术语库管理系统仍具备存在一
17、定的局限性。这所设计的这些数据库管理系统被设计为只能在一台计算机上运行并不容易被便于进行共享。他们通常只允许进行简单对的双语术语进行简单管理并且极大限制了对于数据域的类型和数量以及和可以在这些域中存储在这些数据域中的最大数据信息量最大值。进行限制。然而,,最近这种类型的软件功能变得更加强大和灵活,,特别是在存储和检索选项方面。4.1 StorageThe most fundamental function of a TMS is that it acts as a repository for consolidating and storing terminological informati
18、on for use in future translation projects. Previously, many TMSs stored information in structured text files, mapping source-to-target terminology using a unidirectional one-to-one correspondence. This caused difficulties, for example, if a French-English term base needed to be used for an English-F
19、rench translation. The newer, more sophisticated software stores the information using a relational model. This means that the information is stored in a more onomasiological or concept-based way, which permits mapping in multiple language directions.4.1存储术语库最基本的功能是作为一个存储库来巩固并和存储术语信息,以备将来翻译项目之用。之前,许
20、多术语管理系统将信息存储在结构化文本中的术语库存储的信息应用间接,使用单向一一对应的方法进行源语和目标语之间的一一对应的方法进行从源到目标的转化。这样,就产生就引出了一些难题,比方说法英翻译需要用到的,这就像一个法语-英语术语库需要用于英语法语翻译。较更新的、更较复杂的软件采用关系模型来存储信息,也就。这意味着要,信息以通过一种更符合更加偏向于以专名学或概念为基础语义学、更基于概念的方法进行储存信息,这允许进行了多种语言方向之间的转化。 There is also increased flexibility in the type and amount of information that
21、can be stored on a term record. Formerly, users were required to choose from a predefined set of fields (e.g., subject field, definition, context, source), which had to be filled in on each term record. 此外,这也使增加了可那些储存在术语记录中的信息的类型和信息数量的灵活性更加灵活。以前,用户需要从一组预定义字段(比如主字段、定义、上下文、来源)中进行选择(这些预定义字段包括:主字段,定义,上下
22、文,资源),并且这些字段必须来自每一条术语记录。 The number of fields was often fixed, as was the number of characters that could be stored in each field. For instance, if a TMS allowed for only one context, the user was forced to record only one context, even though it may have been useful to provide several. An example o
23、f a typical conventional record template is provided in figure 4.1. Term(En):Term(Fr)Subject field:Definition:Context:Synonyms:Source:Comment:Administrative info(date,author,quality code,etc,):Figure 4.1 TMS term record with a fixed set of predefined fields 通常,字段的数目以及每个字段能够存储的字符数通常是固定的,每个字段能够存储的字符数也
24、都是固定的。例如,如果一个术语管理系统只允许记录一个文本的话,即使对用户来说它有可能记录多个文本是有好处的,但他用户还是只能记录一个而已。图4.1所示的是一个典型的传统记录模板的典型例子。术语(英文):术语(法文)主字段:定义:上下文:同义词:来源:注释:管理信息( 日期,、 作者、,质量、 编码等):图4.1 含有一套固定的预定义领域字段的术语管理系统的术语记录 图4.1表格名称漏掉了In contrast, as illustrated in figure 4.2, most contemporary TMSs have adopted a free entry structure, wh
25、ich allows users to define their own fields of information, including repeatable fields (e.g., for multiple contexts) and some even permit the inclusion of graphics. Not only can users choose their own information fields, they can also arrange and format them, choosing different layouts, fonts, or c
26、olors for easy identification of important information. This means that the software can be adapted to suit a specific users needs and can grow as future requirements change. The amount of information that can be stored in any given field or record has also increased dramatically. Different term bas
27、es can be created and maintained desired. Term(En): selected(v)Subject field: computingContext1: The item you selected does not exist Source: Computer magazine ABC, 1999Context2: When you are finished the selecting the text, click on the Format menu Source: User manual XYZ,1998Client: Company AFr: S
28、lectionner Date: June 2000Client: Company BFr: choisirDate: January 2001 Figure 4.2 TMS term record with free entry structure 如图4.2所示,相比之下,当下大多数的术语管理系统当前都采用的是自由条目结构,它可以让用户自行定义他们自己的信息字段,其中包括可重复字段(例如,处理多个文本时),还有一些甚至允许录入图表。用户不仅可以选择他们自己的信息字段,还可以将这些信息字段对它们进行排序和格式化,为它们选择不同的布局和字体,对容易识别的重要信息进行标色等。这意味着可以对这款软
29、件进行调整以可以满足适应特定用户的需求,并且它还可以随着未来用户需求的变化而发展。很明显,任何给定的字段或记录所能存储的信息量也在明显的增加。可以创建不同的术语库以满足不同的需求。 术语(英文): 已选择(v)主字段: 计算文本1: 你所选择的项目不存在 来源: 计算机基础杂志, 1999文本2: 完成文本的选择后,单击“格式”菜单 来源:用户手册指南,1998客户: A公司法文: 竞选人日期: 2000年6月客户: B公司法文: 选择日期: 2001年1月 图4.2 含有自由条目结构的术语管理系统的术语记录4.2 Retrieval4.2 检索 Once the terminology ha
30、s been stored, translators need to be able to retrieve this information. A range of search and retrieval mechanisms is available. The simplest search technique consists of a look-up to retrieve an exact match. Some TMSs permit the use of wildcards for truncated searches. A wildcard is a character. s
31、uch as an asterisk, that can be used to represent any other character or string of characters. For instance, a wildcard search using the search string comput* could be used to retrieve the term records for computer, computing, and so on. More sophisticated TMSs also employ fuzzy matching techniques.
32、 A fuzzy match will retrieve term records that are similar to the requested search pattern, but that do not match it exactly. Fuzzy matching allows translators to retrieve records for morphological variants (e.g., different forms of verbs, words with suffixes or prefixes), spelling variants (or even
33、 spelling errors), and multi-word terms, even if the translator does not know precisely how the elements of the multi-word term are ordered. Table 4.1 provides some examples of the term records that could be retrieved using fuzzy matching techniques. Table 4.1 Sample term records retrieved using fuz
34、zy matchingSearch pattern entered by user Term record retrieved using fuzzy matching “anovulatory” ovulation“discus” disk“department for dangerous goods dangerous Goods Emergency Centre emergencies”一旦术语被存储起来,就要求译员具备就需要有能够就这些信息进行检索出这些信息的能力。有一系列的搜索和检索机制可供使用。最简单的搜索技术就是由通过一个查询来检索出精确匹配项。一些术语管理系统允许使用通配符来进
35、行截断搜索。一个通配符就是一个字符,比如一个星号可以用来代表任何字符或者字符串中的字符。例如,一个通配符搜索使用的搜索字符串“comput *”可以用来检索“ computer,”、 “computing”,等术语的记录。较更复杂的术语管理系统还也可以使用模糊匹配技术。模糊匹配可以检索出与所要求搜索模式相似的术语记录,但并不是精确地匹配。即使译员不能准确理解多词术语中各种成分的组织形式,模糊匹配可以让他们译员检索出形态变体(例如,动词的不同形式的动词,带有前缀或者后缀的单词词语), 拼写变体(或者甚至是拼写错误)和多词术语的记录,即使译员不能准确理解多词术语中各种成分是如何组织的。表 4.1所
36、示是一些通过模糊匹配技术检索出的一些术语记录的例子。表 4.1 通过模糊匹配检索出的术语记录示例 用户使用的搜索模式 通过模糊匹配检索出的术语记录“anovulatory” ovulation“discus” disk“department for dangerous goods dangerous Goods Emergency Centre emergencies” When wildcard searching or fuzzy matching is used, it is possible that more than one record will be retrieved as
37、a potential match. When this happens, users are presented with a hit list of all the records in the term base that may be of interest, and they can select the record(s) that they wish to view. Sample hit lists are shown in table 4.2.Table 4.2 Sample hit lists retrieved for different Search patternsH
38、it list containing records that Hit list containing records that fuzzy match the wildcard search pattern search pattern“cake” “skate-boarding champion” “cake” champion”cheesecake championcupcake skateboard (n)fruitcake skateboard (v)pancake skateboarding International Skateboarding Championships使用通配
39、符搜索或者模糊匹配时,可能会检索出不止一条充当作为潜在的匹配角色的可能会检索出不止一条记录。出现这种情况时,用户在术语库中会看到他们会感兴趣的一个可能比较有意思的“命中列表”,在这个列表中,这样用户就可以选择他们想要查看的记录。表4.2所示的是命中列样表示例表4.2 不同检索模式下的命中列样表示例命中列表:包含匹配通配符搜索 命中列表:包含匹配模糊搜索搜索 模式的记录 模式的记录 下的“cake”记录 包含模糊匹配模式下的记录 “skate-boarding champion” 的命中列表 “skate-boarding” 的命中列表 cheesecake champion cupcake s
40、kateboard (n)fruitcake skateboard (v)pancake skateboarding International Skateboarding Championships4.3 Active terminology recognition and pre-translation4.3 主动术语识别和预翻译Another feature offered by some TMSs, particularly those that operate as part of an integrated package with word processors and tran
41、slation-memory systems (see section ) is known as active terminology recognition. This feature is essentially a type of automatic dictionary look-up. As the translator moves through the text, the terminology- recognition component compares items in the source text against the contents of the term ba
42、se, and if a match is found, the term record in question is displayed for the user to consult. 一些术语管理系统的另一功能,特别那些是作为带有文字处理器和翻译记忆系统的完整的软件包的一部分进行运作的术语管理系统(见节)功能,称为以主动术语识别著称。从本质上来说,这个功能本质上这是一种自动字典查询功能。随着译员逐渐深入研究文本,术语识别组件会将源文本中的项目与术语库中的内容进行对比,如果找到匹配项,就会把这个选中的术语记录会展现给用户,以备供其参考查询。 Some TMSs also permit a
43、more automated extension of this feature in which a translator can ask the system to do a sort of pre-translation or batch processing of the text. 还有一些术语管理系统具有自动扩展功能。,在这个功能之下这样,译者就可以利用系统完成文本的预翻译和批处理。82 Computer-Aided Translation Technology 计算机辅助翻译技术Table 4.3 Automatic replacement of source-text term
44、s with translation equivalents found in a term baseSource text sentence Term base entries for items Sentence produced following contained in the source text pre-translation The file operation disk disque The opration de fichiercannot be completed file operation- opration de fichier cannot be complet
45、ed becausebecause the disk is full full-sature the disque is sature表4.3 用术语库中的翻译等值项目自动替换源文本中的术语原文本句子 术语库条目中包含的源文本术语 预翻译后产生的句子The file operation disk disque The opration de fichiercannot be completed file operation- opration de fichier cannot be completed becausebecause the disk is full full-sature t
46、he disque is satureIn this case, the TMS will identify terms for which an entry exists in the term base, and it will then automatically insert the corresponding equivalents into the target text. The result of this pre-translation phase is a sort of hybrid text, as shown in table 4.3. In a post-editi
47、ng phase, it is up to the translator to verify the correctness of the proposed terms and to translate the remainder of the text for which no equivalents were found in the term base.在这种情况下,,术语管理系统将会在已有的术语库中识别这些术语,然后自动在目标文本中插入相应的翻译等值项。这个预翻译阶段的结果就是生成一种混合文本,如表4.3所示。在文章编辑阶段,将由译者来验证所替换术语的正确性并翻译未能在术语库中找到翻译
48、等值项的剩余部分。4.4 Term extractionAnother feature that may be included in some TMSs is a term-extraction tool, which is sometimes referred to as a term-recognition or term-identification tool. Most term-extraction tools are monolingual, and they attempt to analyze source texts in order to identify candida
49、te terms. However, some bilingual tools are being developed that analyze existing source texts along with their translations in an attempt to identify potential terms and their equivalents. This process can help a translator build a term base more quickly; however, although the initial extraction at
50、tempt is performed by a computer, the resulting list of candidates must be verified by a human, and therefore the process is best described as being computer-aided or semi-automatic rather than fully automatic. Unlike the word-frequency lists described in section 3.2.1, term-extraction tools attempt
51、 to identify multi-word units. There are two main approaches to term extraction: linguistic and statistical. For clarity, these approaches will be explained in separate sections; however, aspects of both approaches can be combined in a single term-extraction tool.4.4术语抽取一些术语管理系统可能还有另一个特点,就是包含了术语抽取工具
52、,有时也被称为术语识别或术语鉴别工具。大多数术语抽取工具是单语的,它们试图分析源文本以确定候选术语。然而,也正在开发,一些双语工具正在开发,这些工具可分析现有的源文本以及他们的翻译在以期识别出它们的潜在术语及等值项。这一过程可以帮助译者更迅速的建立一个术语库;尽管最初的提取尝试是由有计算机执行的,但是必须由人来验证最终产生的候选列表,因此对它的最佳描述应该是计算机辅助或半自动翻译过程而非全自动。与3.2.1节中所描述的词频列表不同,术语抽取工具试图识别多词单位。术语抽取主要有两种方法:语言学方法和统计学方法。为了清楚起见,将在不同的章节对这两种方法进行分别解释;然而,这两种方法的某些方面也可以
53、结合成一个单一术语提取工具。Terminology -Management Systems 83术语管理系统 83Antivirus programs now include a number of options. Integrity checking performs checks of the status of the files against the information that is stored in a database. Behavior blocking performs before-the-fact detection. Heuristic analysis is
54、 a form of after-the-fact detection. Figure 4.3 A short text that has been processed using a linguistic approach to term extraction. 图4.3 一个使用语言学方法进行术语抽取加工的简短文本Antivirus programs now include more options. Integrity checking performs periodic checks of the current status of the files against the info
55、rmation that is stored information. Behavior blocking performs before-the-fact detection. Heuristic analysis is a form of after-the-fact detection.Figure 4.4 A slightly modified version of the text that has been processed using a linguistic approach to term extraction.图4.4 使用语言学方法进行术语抽取加工并轻微修正过的文本4.
56、4.1 Linguistic approachTerm-extraction tools that use a linguistic approach typically attempt to identify word combinations that match particular part-of-speech patterns. For example, in English, many terms consist of NOUN+NOUN or ADJECTIVE+NOUN combinations. In order to implement such an approach,
57、each word in the text must first be tagged with its appropriate part of speech, as described in section 3.3. Once the text has been correctly tagged, the term-extraction tool simply identifies all the occurrences that match the specified part-of-speech patterns. For instance, a tool that has been pr
58、ogrammed to identify NOUN+NOUN and ADJECTIVE+NOUN combinations as potential terms would identify all lexical combinations matching those patterns from a given text, as illustrated in figure 4.3.Unfortunately, not all texts can be processed this neatly. If the text is modified slightly, as illustrate
59、d in figure 4.4, problems such as noise and silence become apparent.First, not all of the combinations that follow the will qualify specified patterns as terms. Of the NOUN+NOUN and ADJECTIVE+NOUN candidates that were identified in figure 4.4, some qualify as terms4.4.1 语言学方法使用语言学方法的术语抽取工具的典型特点是:试图通
60、过匹配特定的词性模式来识别单词组合。例如,许多英语术语的构成模式是:名词+名词 或者 形容词+名词。为了适应这种方法,首先必须适当标记出文本中每个单词的词性,如3.3节所述。一旦文本被正确标记,术语提取工具将很容易识别出与特定词性模式相匹配的所有术语。例如,一个术语抽取工具编程的潜在条件是识别名词+名词组合和形容词+名词组合,那么该工具可以从给定文本中识别出与这一模式相匹配的所有词汇组合,如图4.3所示。不幸的是,并不是所有的文本都可以被加工的这么整齐。如果对文本稍作修改,如图4.4所示,“噪声”和“无声”之类的问题将变得很显而易见。首先,并非所有的词汇组合都按照指定的术语模式以合格特定术语模
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2025年安康南宫山技师学院高职单招职业适应性测试考试模拟试卷附完整答案详解(名师系列)
- 2025年陕西服装工程学院单招职业技能考试题库(综合题)附答案详解
- 2024年张家口技师学院高职部高职单招职业技能考试模拟试卷【培优B卷】附答案详解
- 2026年山东鲁南职业学院单招职业技能考试题库含答案详解【新】
- 2027年河南省三门峡市单招综合素质考试模拟试卷(典型题)附答案详解
- 2026年山东清泉职业学院单招综合素质考试题库含答案详解AB卷
- 2025年桥山职业学院单招综合素质考试题库及答案详解(全优)
- 2027年甘肃省白银市高职单招职业技能考试题库【突破训练】附答案详解
- 2024年吉林交通职院高职单招职业技能考试题库附参考答案详解(轻巧夺冠)
- 2024年东营现代化工职业学院高职单招职业适应性测试考试模拟试卷(黄金题型)附答案详解
- 2026年安徽合肥经开区社区工作者招聘考试试卷-含答案解析
- 2026年广东省中考化学试卷(含答案及解析)
- 2026广西来宾市机关事务管理中心招聘事业单位后勤服务控制数人员1人笔试备考题库及答案详解
- 2026-2030泡沫金属行业市场发展分析及发展趋势与投资前景研究报告
- 2026海南万宁市总工会招聘工会社会工作者11人(第1号)笔试参考题库及答案详解
- 功能性消化不良诊疗指南(2026版)
- 2026年东风汽车校招人才测评题库
- 2026年湖北省荆门市东宝区3年级数学期中人教版考试及答案
- (必刷)山东省大数据专业高级职称(大数据应用分析专业)考点精粹必做500题-含答案
- 颅脑手术的麻醉课件
- 写字楼验收移交工作方案
评论
0/150
提交评论