版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1、约500词的篇章10道填空题,每空不超过3个单词;独白或演讲词口语特征句式简单,解构明晰(信息衔接词)Pre-reading 预读考题 (30秒) 了解主题和大致内容,明晰文章的结构和各部分的内容 Focus on Key information 专注关键信息 关注细节实词,训练用简单符号做笔记 Relate with the context 联系语篇 注意所填词语的词性和语义和语篇是否搭配。所填内容不超过3词 Top-down skimming (自上而下跳读) Distinguish the genre of the passage(区分文章的体裁和题材) Find out the hig
2、h-frequent words (找出高频词汇明确主题词) Whats the procedures of Pre-reading一、训练快速记笔记的能力 速记法举例:Heart disease: heart dis. Similarity: simty.Cholesterol: chol. Difference: diffr.Cigarette: cigat. Financial: finan.Exercise: ex. Responsibility: respty.Especially: esp. Hostility: hosty.二、训练语句间意义的推测和预测能力三、训练区分重点(论点
3、与主题句)与论据的能力. Key points in Note-takingHow to make use of context informationTest 6 1 communicate ; 2. bilingual 3. dialects 4. Sign language; 5. values; 6. real; 7. fables; 8. honor; 9. picnic10. religious Test 7. 1. seminar paper; 2. preparation 3.circulating; 4. discussion; 5. lack; 6. boring; 7.
4、time limit 8. main points; 9. outline note 10. ending Test 8. 1. educational; 2. benefit; 3. summary 4. experience; 5. sales instruments; 6. attractive; 7. noticeable; 8. tailor 9. shorthand; 10. cheat Test 9. 1. understood and appreciated; 2. credibility 3. casual; 4. physical contact; 5. a must/ n
5、ecessary/ important; 6. titles and positions; 7. chopsticks 8. Taboos: The term taboo comes from the Tongan tapu or Fijian tabu (prohibited, disallowed, forbidden),3 related among others to the Maori tapu, Hawaiian kapu, Malagasy fady. 9. cutting; 10. death Abstraction: Abstraction consists of the t
6、ranslation (mapping) of terms in the scheme to terms in a theoretically motivated model or dataset. Abstraction typically includes linguist-directed search but may include e.g., rule-learning for parsers. Analysis: Analysis consists of statistically probing, manipulating and generalising from the da
7、taset. Analysis might include statistical evaluations, optimisation of rule-bases or knowledge discovery methods. Most lexical corpora today are part-of speech-tagged (POS-tagged). However even corpus linguists who work with unannotated plain text inevitably apply some method to isolate terms that t
8、hey are interested in from surrounding words. In such situations annotation and abstraction are combined in a lexical search. What is the advantage of publishing an annotated corpusThe advantage of publishing an annotated corpus is that other users can then perform experiments on the corpus. Linguis
9、ts with other interests and differing perspectives than the originators can exploit this work. By sharing data, corpus linguists are able to treat the corpus as a locus of linguistic debate, rather than as an exhaustive fount of Knowledge. How many types of corpora are there? There are many differen
10、t kinds of corpora. They can contain written or spoken (transcribed) language, modern or old texts, texts from one language or several languages. The texts can be whole books, newspapers,journals, speeches etc, or consist of extracts of varying length. The kind of texts included and the combination
11、of different texts vary between different corpora and corpus types. General corpora: General corpora consist of general texts, texts that do not belong to a single text type, subject field, or register. An example of a general corpus is the British National Corpus. Contemporary American English Corp
12、us is another example.Some corpora contain texts that are sampled (chosen from) a particular variety of a language, for example, from a particular dialect or from a particular subject area. These corpora are sometimes called Sublanguage Corpora. Specialized corpora Historical corpora The use of coll
13、ections of text in the study of language is, as we have seen, not a new invention. Among those involved in historical linguistics were some that soon saw the potential usefulness of computerised historical corpora. A diachronic corpus with English texts from different periods was compiled at the Uni
14、versity of Helsinki. The Helsinki Corpus of English Texts contains texts from the Old, Middle and Early Modern English periods, 1,5 million words in total. nxxnxxnxxnxxnxxnxxnxxnxxnxxnxxnxxnxxAnother historical corpus is the recently released Lampeter Corpus of Early Modern English Tracts. This coll
15、ection consists of pamphlets and tracts published in the century between 1640 and 1740 from six different domains. The Lampeter Corpus can be seen as one example of a corpus covering a more specialized area. The corpora described above are general collections of text, collected to be used for resear
16、ch in various fields. There is a large, and growing, amount of highly specialized corpora that are created for a special purpose. Many of these are used for work on spoken language systems. Examples of such are, for example, the Air Traffic Control Corpus, ATC0 , created to be used in the area of ro
17、bust speech recognition in domains of air traffic control , and the TRAINS Spoken Dialogue Corpus collected as part of a project set up to createa conversationally proficient planning assistant (railroad freight system). A number of highly specialized corpora are held at the Centre for Spoken Langua
18、ge Understanding, CSLU, in Oregon. These corpora are specialized in a different way to the ones mentioned above. They are not restricted to be used within a particular subject field, but are called specialized because their content. Many of the corpora/databases consist of recordings of people asked
19、 to perform a particular task over the telephone, such as saying and spelling their name or repeating certain words/ phrases/ numbers/ or letters. As we have seen above, there is a great variety of corpora in English. So far much corpus work has indeed concerned the English language, for various rea
20、sons. There are, however, a growing number of corpora available in other languages as well. Some of them are monolingual corpora - collections of text from one language. Here the Oslo Corpus of Bosnian text and the Contemporary Portuguese Corpus can be mentioned as two examples. A number of multilin
21、gual corpora also exist. Many of these are parallel corpora; corpora with the same text in several languages. These corpora are often used in the field of Translation. The English-Norwegian Parallel Corpus is one example, the English Turkish Aligned Parallel Corpora another. The Linguistic Data Cons
22、ortium (LDC) holds a collection of telephone conversations in various languages: CALLFRIEND and CALLHOME. The Use of the Internet The increased availability and use of the Internet have made it possible to find great amounts of texts readily available in electronic format. Apart from all the web-pag
23、es containing information of different kinds, it is also possible to find whole collections of text. Among these collections can be mentioned all the on-line newspapers and journals (example), and sites where whole books can be found on-line (example). Other examples yet include dictionaries and wor
24、d- lists of various kinds. Although these collections may not be considered corpora for one reason or another, they can be analysed with corpus linguistic tools and methods. This is an area which has not yet been explored in detail, although some attempts have been made at using the Internet as one
25、big corpus. Some Large Corpora Projects ICE: the International Corpus of English In twenty centres around the world, compilers are busy collecting material for the ICE corpora. Each ICE corpus will consist of 1 million words (written and spoken) of a national variety of English. The first ICE corpus
26、 to be completed is the British component, ICE-GB. On their own, the ICE corpora will be a valuable resource to exploit in order to learn about different varieties of English. As a whole, the 20 corpora will be useful for variational studies of various kinds. Learn more about the ICE project at the
27、ICE-GB site. On their own, the ICE corpora will be a valuable resource to exploit in order to learn about different varieties of English. As a whole, the 20 corpora will be useful for variational studies of various kinds. Learn more about the ICE project at the ICE-GB site. ICLE: the International C
28、orpus of Learner English Like ICE, ICLE is an international project involving several countries. Unlike ICE, however, the ICLE corpora do not consist of native speaker language. Instead they are corpora of English language produced by learners in the different countries. This will constitute a valua
29、ble resource for research on second language acquisition. The interest for computerised corpora and corpus linguistics is growing. More and more universities offer courses in corpus linguistics and/or use corpora in their teaching and research. The number and diversity of corpora being compiled are
30、great and corpora as used in many projects. It is not possible to go into detail and present all the corpora, all the courses, all the projects here. This has been meant as a brief introduction. More information can be found by browsing the net and reading journals and books. Corpus-Related Research
31、 This is a short introduction to some of the research areas where corpora can be and have been used. 1. Computational Linguistics Computational Linguistics is an interdisciplinary field which centers around the use of computers to process or produce human language.” In some ways, computational lingu
32、istics and corpus linguistics can be seen as overlapping disciplines. Computational linguists are dependent on computer-readable linguistic data to use in their research, while corpus linguists often use computational methods when analysing their data. One main difference can be said to be that in c
33、orpus linguistics it is the data in the corpus that is the main object of study. In computational linguistics, corpora are not studied as such but used as a resource to solve various problems. 2.Cultural Studies The existence of comparable corpora makes it possible to compare the language use in, fo
34、r example, different countries. The result of such comparisons can point to differences in culture. It has been suggested, for example, that the lower proportion of expressions of future in the Kolhapur Corpus of Indian English, as compared to Brown Corpora, can be explained with cultural difference
35、s. Maybe the Indian mind is not given to thinking much in terms of the future. So far, the use of corpora in cultural studies is not a particular well developed field. Perhaps the ongoing work of compiling 20 corpora of different varieties of English within the ICE project (International Corpora of
36、English) will help make this a more fruitful research area in the future. 3.Discourse Analysis and Pragmatics Pragmatics is the study of the way language is used in particular situations, and is therefore concerned with the functions of words as opposed to their forms. It deals with the intentions o
37、f the speaker, and the way in which the hearer interprets what is said (from The Collins Cobuild English Language Dictionary (1987). Corpora have not been extensively used much in discourse analysis or pragmatic studies. One explanation to that is that it has been difficult to find material suitable
38、 for this kind of research. As more corpora are being compiled and annotated with the relevant information, more corpus based research is also being performed in this area. Examples of such studies are can be found among the work done by scholars in Bergen (Norway), on their corpus of London teenage
39、 language, COLT. See, for example, They like wanna see how we talk and all that.The use of like as a discourse marker in London teenage speech. More examples about using corpora in discourse analysis and pragmatics: _An Introduction to Spoken Interaction by Anna-Brita Stenstrom. Questions and Respon
40、ses in English Conversation by Anna-Brita Stenstrom. The Discourse Resource Initiative project COLT-based research with abstracts of papers/articles based on COLT material (The Bergen Corpus of London Teenage 4.Grammar/Syntax Much research on grammar and syntax has been based on the researchers intu
41、ition about the language, on his/her competence. The existence of large corpora has made it easier to study the language as it is produced, to study the performance of many people. Every (formal) grammar is initially written on the basis of intuitive data; by confronting the grammar with unrestricte
42、d corpus data it can be tested on its correctness and its completeness. Corpus data are being used to a larger or smaller extent for the production of grammar books. One example of a book, based completely on corpus evidence is An Empirical Grammar of the English Verb: Modal Verbs by Dieter Mindt (2
43、005). An other example of how corpora can be used to corpus-based research on grammar and syntax is, for example, Clause patterns in Modern British English: A corpus-based (quantitative) study by N. Oostdijk and P. de Haan (2004). 5.Historical Linguistics The possibility of having representative sam
44、ples of the language at different points in history in machine-readable form allows historical linguists to conduct their research faster and more efficiently. The Helsinki Corpus is a well-known and much used corpus of texts from different periods. The Lampeter Corpus of Early Modern English Tracts
45、 contains a collection of pamphlets published between 1640 and 1740. 6.Language Acquisition The ICLE corpus (International Corpus of Learner English) contains data produced by learners of English as a foreign language from different countries. It is being use for a variety of research purposes, some
46、 of which were presented at the AILA96 conference (abstracts). Learn more about this research in the book Learner English on Computer by Sylviane Granger. The CHILDES database contains transcripts of language spoken by children. This material can be used for research in a number of fields, language
47、acquisition being one. An annotated bibliography of research in child language and language disorders can be found by using some links. 7.Language Teaching There are many articles and books of how corpora can be, and have been, used in language teaching. See, for example: Classroom Concordancing / D
48、ata-driven learning Bibliography can be identified on the internet which offers very good references to the direct use of data from linguistic corpora for language teaching and language learning. 8.Language Variation Much work with corpora concerns language variation. Corpora are used to study how l
49、aguage varies between different text types, domains, times, regions, speakers, writers, etc. In these kinds of studies, one variant of the language is compared to another. These variants can be different parts of one and the same corpus or similar parts of different corpora. An example of the former
50、 would be, for example, the Science Fiction texts in the LOB corpus compared to the Romantic Fiction texts in the same corpus. An example of a study of variation between two corpora would be, for example, an examination of the Science fiction texts in the LOB corpus as compared to the Science fictio
51、n texts in the Brown corpus. Language variation can also concern how speakers vary their production depending on the situation, how the language has changed over time, or how the language varies within an area (dialect). 9.Lexicography Corpora are increasingly used in lexicography today. The first e
52、xtensive use of large corpora in dictionary and grammar book production was the Collins Cobuild. Longman has consulted the British National Corpus (BNC) and the Longman Corpus Network for their latest edition of the Longman Dictinary of Contemporary English. You can read more about the use of corpor
53、a in dictionary making in Computer Corpus Lexicography by Vincent B. Y. Ooi. For examples of corpus based studies in the field of lexicography, see, for example, the Bibliography of papers by Cobuild staff members. 10. Psycholinguistics Corpora are important sources of data for almost all the areas
54、within the wide scope of Linguistics. Analyses of language use provide an important complementary perspective to traditional linguistic descriptions. Douglas Biber Observing the language found in a corpus can contribute to the creation of hypothesis about the way language is processed by the mind. T
55、he use of corpora can also contribute to research in language pathologies. In order to analyse a particular language impairment, it is important to have a very clear picture of the structural and formal differences between the impairment and its correct form. 11.Semantics There are various ways in w
56、hich you can study the meaning of words/utterances. One way is to look at the context in which the word/phrase occurs. Concordances and collocations are often used for this. Attempts have been made at annotating corpora with semantic Example of how such information can be given can be found by looki
57、ng at the Word Net, a semantically annotated lexical database for English. 12.Sociolinguistics With the existence of corpora provided with sociolinguistic information about the speakers and/or authors of a text has come the possibility of using corpora in sociolinguistic research. The British Nation
58、al Corpus BNC has been extensively annotated for various sociolinguistic parameters, such as speakers age, sex, and social class, writers age, sex, domicile, etc. This information is used in a number of studies, for example by Paul Rayson et al in Social Differentiation in the Use of English Vocabul
59、ary: Some Analyses of the Conversational Component of the British National Corpus (IJCL 1997:2:1). Another corpus with sociolinguistic annotation is the COLT Corpus of London teenager language. See, for example, Girls conflict talk: a sociolinguistic investigation of variation in the verbal disputes
60、 of adolescent females by A-B Stenstrom and I.K. Hasund. Historical corpora are also being used for sociolinguistic research. See, for example, the Sociolinguistics and Language History Project. 13.Speech The first computer-readable corpus of spoken discourse was the London-Lund Corpus (LLC). It con
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2025-2026学年《天鹅的故事》说课稿
- 2025-2026学年中国梦班会说课稿
- 2025-2026学年《飞天凌空》说课稿设计
- 2025-2026学年大树立体画美术说课稿
- 2025-2026学年大班非洲鼓说课稿
- 化学合成制药工7S考核试卷含答案
- 维纶热处理操作工安全宣贯水平考核试卷含答案
- 2025-2026学年动物的纹理说课稿
- 木雕工竞争评优考核试卷含答案
- 2025-2026学年八分钟说课稿多少字合适
- 第一月考过关测试卷(试卷)2026-2027学年五年级语文上册统编版(含答案)
- 3.1《买文具》课件 -2026-2027学年五年级上册数学北师大版
- 2026年湖南工业职业技术学院高职单招笔试职业技能测验试题库含答案解析3套试卷
- 医院职工绩效考核制度(2026版)
- 临床科学防治颈椎病守护颈部健康关键策略
- 油田分层注水技术
- 电气控制技术说课
- 灌装工专业技能培训课件
- 中药黄芪课件
- 2025部编版三年级道德与法治上册全册教案
- 可溃式制动踏板
评论
0/150
提交评论