Tem-4Talk讲座Test7-9参考答案(week8)_第1页
Tem-4Talk讲座Test7-9参考答案(week8)_第2页
Tem-4Talk讲座Test7-9参考答案(week8)_第3页
Tem-4Talk讲座Test7-9参考答案(week8)_第4页
Tem-4Talk讲座Test7-9参考答案(week8)_第5页
已阅读5页,还剩80页未读, 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

1、约500词的篇章10道填空题,每空不超过3个单词;独白或演讲词口语特征句式简单,解构明晰(信息衔接词)Pre-reading 预读考题 (30秒) 了解主题和大致内容,明晰文章的结构和各部分的内容 Focus on Key information 专注关键信息 关注细节实词,训练用简单符号做笔记 Relate with the context 联系语篇 注意所填词语的词性和语义和语篇是否搭配。所填内容不超过3词 Top-down skimming (自上而下跳读) Distinguish the genre of the passage(区分文章的体裁和题材) Find out the hig

2、h-frequent words (找出高频词汇明确主题词) Whats the procedures of Pre-reading一、训练快速记笔记的能力 速记法举例:Heart disease: heart dis. Similarity: simty.Cholesterol: chol. Difference: diffr.Cigarette: cigat. Financial: finan.Exercise: ex. Responsibility: respty.Especially: esp. Hostility: hosty.二、训练语句间意义的推测和预测能力三、训练区分重点(论点

3、与主题句)与论据的能力. Key points in Note-takingHow to make use of context informationTest 6 1 communicate ; 2. bilingual 3. dialects 4. Sign language; 5. values; 6. real; 7. fables; 8. honor; 9. picnic10. religious Test 7. 1. seminar paper; 2. preparation 3.circulating; 4. discussion; 5. lack; 6. boring; 7.

4、time limit 8. main points; 9. outline note 10. ending Test 8. 1. educational; 2. benefit; 3. summary 4. experience; 5. sales instruments; 6. attractive; 7. noticeable; 8. tailor 9. shorthand; 10. cheat Test 9. 1. understood and appreciated; 2. credibility 3. casual; 4. physical contact; 5. a must/ n

5、ecessary/ important; 6. titles and positions; 7. chopsticks 8. Taboos: The term taboo comes from the Tongan tapu or Fijian tabu (prohibited, disallowed, forbidden),3 related among others to the Maori tapu, Hawaiian kapu, Malagasy fady. 9. cutting; 10. death Abstraction: Abstraction consists of the t

6、ranslation (mapping) of terms in the scheme to terms in a theoretically motivated model or dataset. Abstraction typically includes linguist-directed search but may include e.g., rule-learning for parsers. Analysis: Analysis consists of statistically probing, manipulating and generalising from the da

7、taset. Analysis might include statistical evaluations, optimisation of rule-bases or knowledge discovery methods. Most lexical corpora today are part-of speech-tagged (POS-tagged). However even corpus linguists who work with unannotated plain text inevitably apply some method to isolate terms that t

8、hey are interested in from surrounding words. In such situations annotation and abstraction are combined in a lexical search. What is the advantage of publishing an annotated corpusThe advantage of publishing an annotated corpus is that other users can then perform experiments on the corpus. Linguis

9、ts with other interests and differing perspectives than the originators can exploit this work. By sharing data, corpus linguists are able to treat the corpus as a locus of linguistic debate, rather than as an exhaustive fount of Knowledge. How many types of corpora are there? There are many differen

10、t kinds of corpora. They can contain written or spoken (transcribed) language, modern or old texts, texts from one language or several languages. The texts can be whole books, newspapers,journals, speeches etc, or consist of extracts of varying length. The kind of texts included and the combination

11、of different texts vary between different corpora and corpus types. General corpora: General corpora consist of general texts, texts that do not belong to a single text type, subject field, or register. An example of a general corpus is the British National Corpus. Contemporary American English Corp

12、us is another example.Some corpora contain texts that are sampled (chosen from) a particular variety of a language, for example, from a particular dialect or from a particular subject area. These corpora are sometimes called Sublanguage Corpora. Specialized corpora Historical corpora The use of coll

13、ections of text in the study of language is, as we have seen, not a new invention. Among those involved in historical linguistics were some that soon saw the potential usefulness of computerised historical corpora. A diachronic corpus with English texts from different periods was compiled at the Uni

14、versity of Helsinki. The Helsinki Corpus of English Texts contains texts from the Old, Middle and Early Modern English periods, 1,5 million words in total. nxxnxxnxxnxxnxxnxxnxxnxxnxxnxxnxxnxxAnother historical corpus is the recently released Lampeter Corpus of Early Modern English Tracts. This coll

15、ection consists of pamphlets and tracts published in the century between 1640 and 1740 from six different domains. The Lampeter Corpus can be seen as one example of a corpus covering a more specialized area. The corpora described above are general collections of text, collected to be used for resear

16、ch in various fields. There is a large, and growing, amount of highly specialized corpora that are created for a special purpose. Many of these are used for work on spoken language systems. Examples of such are, for example, the Air Traffic Control Corpus, ATC0 , created to be used in the area of ro

17、bust speech recognition in domains of air traffic control , and the TRAINS Spoken Dialogue Corpus collected as part of a project set up to createa conversationally proficient planning assistant (railroad freight system). A number of highly specialized corpora are held at the Centre for Spoken Langua

18、ge Understanding, CSLU, in Oregon. These corpora are specialized in a different way to the ones mentioned above. They are not restricted to be used within a particular subject field, but are called specialized because their content. Many of the corpora/databases consist of recordings of people asked

19、 to perform a particular task over the telephone, such as saying and spelling their name or repeating certain words/ phrases/ numbers/ or letters. As we have seen above, there is a great variety of corpora in English. So far much corpus work has indeed concerned the English language, for various rea

20、sons. There are, however, a growing number of corpora available in other languages as well. Some of them are monolingual corpora - collections of text from one language. Here the Oslo Corpus of Bosnian text and the Contemporary Portuguese Corpus can be mentioned as two examples. A number of multilin

21、gual corpora also exist. Many of these are parallel corpora; corpora with the same text in several languages. These corpora are often used in the field of Translation. The English-Norwegian Parallel Corpus is one example, the English Turkish Aligned Parallel Corpora another. The Linguistic Data Cons

22、ortium (LDC) holds a collection of telephone conversations in various languages: CALLFRIEND and CALLHOME. The Use of the Internet The increased availability and use of the Internet have made it possible to find great amounts of texts readily available in electronic format. Apart from all the web-pag

23、es containing information of different kinds, it is also possible to find whole collections of text. Among these collections can be mentioned all the on-line newspapers and journals (example), and sites where whole books can be found on-line (example). Other examples yet include dictionaries and wor

24、d- lists of various kinds. Although these collections may not be considered corpora for one reason or another, they can be analysed with corpus linguistic tools and methods. This is an area which has not yet been explored in detail, although some attempts have been made at using the Internet as one

25、big corpus. Some Large Corpora Projects ICE: the International Corpus of English In twenty centres around the world, compilers are busy collecting material for the ICE corpora. Each ICE corpus will consist of 1 million words (written and spoken) of a national variety of English. The first ICE corpus

26、 to be completed is the British component, ICE-GB. On their own, the ICE corpora will be a valuable resource to exploit in order to learn about different varieties of English. As a whole, the 20 corpora will be useful for variational studies of various kinds. Learn more about the ICE project at the

27、ICE-GB site. On their own, the ICE corpora will be a valuable resource to exploit in order to learn about different varieties of English. As a whole, the 20 corpora will be useful for variational studies of various kinds. Learn more about the ICE project at the ICE-GB site. ICLE: the International C

28、orpus of Learner English Like ICE, ICLE is an international project involving several countries. Unlike ICE, however, the ICLE corpora do not consist of native speaker language. Instead they are corpora of English language produced by learners in the different countries. This will constitute a valua

29、ble resource for research on second language acquisition. The interest for computerised corpora and corpus linguistics is growing. More and more universities offer courses in corpus linguistics and/or use corpora in their teaching and research. The number and diversity of corpora being compiled are

30、great and corpora as used in many projects. It is not possible to go into detail and present all the corpora, all the courses, all the projects here. This has been meant as a brief introduction. More information can be found by browsing the net and reading journals and books. Corpus-Related Research

31、 This is a short introduction to some of the research areas where corpora can be and have been used. 1. Computational Linguistics Computational Linguistics is an interdisciplinary field which centers around the use of computers to process or produce human language.” In some ways, computational lingu

32、istics and corpus linguistics can be seen as overlapping disciplines. Computational linguists are dependent on computer-readable linguistic data to use in their research, while corpus linguists often use computational methods when analysing their data. One main difference can be said to be that in c

33、orpus linguistics it is the data in the corpus that is the main object of study. In computational linguistics, corpora are not studied as such but used as a resource to solve various problems. 2.Cultural Studies The existence of comparable corpora makes it possible to compare the language use in, fo

34、r example, different countries. The result of such comparisons can point to differences in culture. It has been suggested, for example, that the lower proportion of expressions of future in the Kolhapur Corpus of Indian English, as compared to Brown Corpora, can be explained with cultural difference

35、s. Maybe the Indian mind is not given to thinking much in terms of the future. So far, the use of corpora in cultural studies is not a particular well developed field. Perhaps the ongoing work of compiling 20 corpora of different varieties of English within the ICE project (International Corpora of

36、English) will help make this a more fruitful research area in the future. 3.Discourse Analysis and Pragmatics Pragmatics is the study of the way language is used in particular situations, and is therefore concerned with the functions of words as opposed to their forms. It deals with the intentions o

37、f the speaker, and the way in which the hearer interprets what is said (from The Collins Cobuild English Language Dictionary (1987). Corpora have not been extensively used much in discourse analysis or pragmatic studies. One explanation to that is that it has been difficult to find material suitable

38、 for this kind of research. As more corpora are being compiled and annotated with the relevant information, more corpus based research is also being performed in this area. Examples of such studies are can be found among the work done by scholars in Bergen (Norway), on their corpus of London teenage

39、 language, COLT. See, for example, They like wanna see how we talk and all that.The use of like as a discourse marker in London teenage speech. More examples about using corpora in discourse analysis and pragmatics: _An Introduction to Spoken Interaction by Anna-Brita Stenstrom. Questions and Respon

40、ses in English Conversation by Anna-Brita Stenstrom. The Discourse Resource Initiative project COLT-based research with abstracts of papers/articles based on COLT material (The Bergen Corpus of London Teenage 4.Grammar/Syntax Much research on grammar and syntax has been based on the researchers intu

41、ition about the language, on his/her competence. The existence of large corpora has made it easier to study the language as it is produced, to study the performance of many people. Every (formal) grammar is initially written on the basis of intuitive data; by confronting the grammar with unrestricte

42、d corpus data it can be tested on its correctness and its completeness. Corpus data are being used to a larger or smaller extent for the production of grammar books. One example of a book, based completely on corpus evidence is An Empirical Grammar of the English Verb: Modal Verbs by Dieter Mindt (2

43、005). An other example of how corpora can be used to corpus-based research on grammar and syntax is, for example, Clause patterns in Modern British English: A corpus-based (quantitative) study by N. Oostdijk and P. de Haan (2004). 5.Historical Linguistics The possibility of having representative sam

44、ples of the language at different points in history in machine-readable form allows historical linguists to conduct their research faster and more efficiently. The Helsinki Corpus is a well-known and much used corpus of texts from different periods. The Lampeter Corpus of Early Modern English Tracts

45、 contains a collection of pamphlets published between 1640 and 1740. 6.Language Acquisition The ICLE corpus (International Corpus of Learner English) contains data produced by learners of English as a foreign language from different countries. It is being use for a variety of research purposes, some

46、 of which were presented at the AILA96 conference (abstracts). Learn more about this research in the book Learner English on Computer by Sylviane Granger. The CHILDES database contains transcripts of language spoken by children. This material can be used for research in a number of fields, language

47、acquisition being one. An annotated bibliography of research in child language and language disorders can be found by using some links. 7.Language Teaching There are many articles and books of how corpora can be, and have been, used in language teaching. See, for example: Classroom Concordancing / D

48、ata-driven learning Bibliography can be identified on the internet which offers very good references to the direct use of data from linguistic corpora for language teaching and language learning. 8.Language Variation Much work with corpora concerns language variation. Corpora are used to study how l

49、aguage varies between different text types, domains, times, regions, speakers, writers, etc. In these kinds of studies, one variant of the language is compared to another. These variants can be different parts of one and the same corpus or similar parts of different corpora. An example of the former

50、 would be, for example, the Science Fiction texts in the LOB corpus compared to the Romantic Fiction texts in the same corpus. An example of a study of variation between two corpora would be, for example, an examination of the Science fiction texts in the LOB corpus as compared to the Science fictio

51、n texts in the Brown corpus. Language variation can also concern how speakers vary their production depending on the situation, how the language has changed over time, or how the language varies within an area (dialect). 9.Lexicography Corpora are increasingly used in lexicography today. The first e

52、xtensive use of large corpora in dictionary and grammar book production was the Collins Cobuild. Longman has consulted the British National Corpus (BNC) and the Longman Corpus Network for their latest edition of the Longman Dictinary of Contemporary English. You can read more about the use of corpor

53、a in dictionary making in Computer Corpus Lexicography by Vincent B. Y. Ooi. For examples of corpus based studies in the field of lexicography, see, for example, the Bibliography of papers by Cobuild staff members. 10. Psycholinguistics Corpora are important sources of data for almost all the areas

54、within the wide scope of Linguistics. Analyses of language use provide an important complementary perspective to traditional linguistic descriptions. Douglas Biber Observing the language found in a corpus can contribute to the creation of hypothesis about the way language is processed by the mind. T

55、he use of corpora can also contribute to research in language pathologies. In order to analyse a particular language impairment, it is important to have a very clear picture of the structural and formal differences between the impairment and its correct form. 11.Semantics There are various ways in w

56、hich you can study the meaning of words/utterances. One way is to look at the context in which the word/phrase occurs. Concordances and collocations are often used for this. Attempts have been made at annotating corpora with semantic Example of how such information can be given can be found by looki

57、ng at the Word Net, a semantically annotated lexical database for English. 12.Sociolinguistics With the existence of corpora provided with sociolinguistic information about the speakers and/or authors of a text has come the possibility of using corpora in sociolinguistic research. The British Nation

58、al Corpus BNC has been extensively annotated for various sociolinguistic parameters, such as speakers age, sex, and social class, writers age, sex, domicile, etc. This information is used in a number of studies, for example by Paul Rayson et al in Social Differentiation in the Use of English Vocabul

59、ary: Some Analyses of the Conversational Component of the British National Corpus (IJCL 1997:2:1). Another corpus with sociolinguistic annotation is the COLT Corpus of London teenager language. See, for example, Girls conflict talk: a sociolinguistic investigation of variation in the verbal disputes

60、 of adolescent females by A-B Stenstrom and I.K. Hasund. Historical corpora are also being used for sociolinguistic research. See, for example, the Sociolinguistics and Language History Project. 13.Speech The first computer-readable corpus of spoken discourse was the London-Lund Corpus (LLC). It con

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论