Corpus linguistics
'''Corpus linguistics''' is the [[study of language]] as expressed in [[sample]]s ''([[Text corpus|corpora]])'' or "real world" text. This method represents a [[digest]]ive approach to deriving a set of abstract rules by which a [[natural language]] is governed or else relates to another language. Originally done by hand, corpora are largely derived by an automated process, which is corrected. 

Computational methods had once been viewed as a [[holy grail]] of [[linguistics|linguistic]] research, which would ultimately manifest a [[ruleset]] for 
[[natural language processing]] and [[machine translation]] at a high level. Such has not been the case, and since the [[cognitive revolution]], cognitive linguistics has been largely critical of many claimed practical uses for corpora. However, as [[computation]] capacity and speed have increased, the use of corpora to study language and term relationships en masse has gained some respectability.

The corpus approach runs counter to [[Noam Chomsky]]'s view that real language is riddled with performance-related errors, thus requiring careful analysis of small speech samples obtained in a highly controlled laboratory setting. 
Corpus linguistics does away with Chomsky's ''competence/performance'' split; adherents believe that reliable language analysis best occurs on field-collected samples, in natural contexts and with minimal experimental interference.{{Fact|date=February 2007}}

{{Linguistics}}

== History ==
A landmark in modern corpus linguistics was the publication by [[Henry Kucera]] and [[Nelson Francis]] of ''Computational Analysis of Present-Day American English'' in 1967, a work based on the analysis of the [[Brown Corpus]], a carefully compiled selection of current American English, totalling about a million words drawn from a wide variety of sources. Kucera and Francis subjected it to a variety of computational analyses, from which they compiled a rich and variegated opus, combining elements of linguistics, language teaching, [[psychology]], [[statistics]], and [[sociology]]. A further key publication was [[Randolph Quirk]]'s 'Towards a description of English Usage' (1960, Transactions of the Philological Society, 40-61) in which he introduced ''The Survey of English Usage''.

Shortly thereafter, Boston publisher [[Houghton-Mifflin]] approached Kucera to supply a million word, three-line citation base for its new ''[[The American Heritage Dictionary of the English Language|American Heritage Dictionary]]'', the first [[dictionary]] to be compiled using corpus linguistics. The AHD made the innovative step of combining prescriptive elements (how language ''should'' be used) with descriptive information (how it actually ''is'' used).

Other publishers followed suit. The British publisher Collins' [[COBUILD]] [[monolingual learner's dictionary]], designed for users learning [[English language learning and teaching|English as a foreign language]], was compiled using the [[Bank of English]].

The [[Brown Corpus]] has also spawned a number of similarly structured corpora: the [[LOB Corpus]] (1960s [[British English]]), Kolhapur ([[Indian English]]), Wellington ([[New Zealand English]]), Australian Corpus of English ([[Australian English]]), the Frown Corpus ([[early 1990s]] [[American English]]),  and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the [[International Corpus of English]], and the [[British National Corpus]], a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities ([[Oxford University|Oxford]] and [[Lancaster University|Lancaster]]) and the [[British Library]]. For contemporary American English, work has stalled on the [[American National Corpus]], but the 360 million word [[Corpus of Contemporary American English (COCA)]] (1990-present) is now available. 

== Methods ==
This means dealing with real input data, where descriptions based on a linguist's intuition are not usually helpful.

==References==
===Journals===
There are several international peer-reviewed journals dedicated to corpus linguistics, for example, 
[[Corpora (journal)|Corpora]], 
[[Corpus Linguistics and Linguistic Theory (Journal)|Corpus Linguistics and Linguistic Theory]],
[http://icame.uib.no/journal.html ICAME Journal] and the 
[[International Journal of Corpus Linguistics]].

===Book series===
Book series in this field include
[[Language and Computers]],
[http://www.benjamins.com/cgi-bin/t_seriesview.cgi?series=SCL Studies in Corpus Linguistics] and [http://www.peterlang.com/Index.cfm?vSiteName=SearchSeriesResult.cfm&vSeriesID=ECL  English Corpus Linguistics]

===Other===
* Biber, Douglas, Susan Conrad, Randi Reppen ''Corpus Linguistics, Investigating Language Structure and Use'', Cambridge: Cambridge UP, 1998. ISBN 0-521-49957-7
* Diana McCarthy, Geoffrey Sampson ''Corpus Linguistics: Readings in a Widening Discipline'', Continuum, 2005. ISBN 0-826-48803-X

==See also==

* [[Concordance (publishing)|Concordance]] ([[Key Word in Context|KWIC]])
* [[Collocation]]
* [[Collostructional analysis]]
* [[Keyword (linguistics)]]
* [[Linguistic Data Consortium]]
* [[Machine translation]]
* [[Natural Language Toolkit]]
* [[Search engines]]: they access the "web corpus".
* [[Semantic prosody]]
* [[Text corpus]]
* [[Translation memory]]

==External links==

* [http://personal.cityu.edu.hk/~davidlee/devotedtocorpora/CBLLinks.htm Bookmarks for Corpus-based Linguists -- very comprehensive site with categorized and annotated links to language corpora, software, references, etc.]
* [http://torvald.aksis.uib.no/corpora/ Corpora discussion list]
* [http://corpus.byu.edu/ Freely-available, web-based corpora (100 million - 360 million words each): American, British (BNC), TIME, Spanish, Portuguese]
* [http://www.bmanuel.org/index.html Manuel Barbera's overview site]
* [http://ifa.amu.edu.pl/~kprzemek/biblios/corpling.zip Przemek Kaszubski's list of references]
* [http://www.corpus4u.org/ Corpus4u Community] a Chinese online forum for corpus linguistics
* [http://www.lancs.ac.uk/fss/courses/ling/corpus McEnery and Wilson's Corpus Linguistics Page]
* [http://groups.google.com/group/corpling-with-r Corpus Linguistics with R mailing list]
* [http://rdues.bcu.ac.uk/ Research and Development Unit for English Studies]
* [http://www.corpus.bham.ac.uk/ The Centre for Corpus Linguistics at Birmingham University]
* [http://www.corpus-linguistics.com Gateway to Corpus Linguistics on the Internet]: an annotated guide to corpus resources on the web
* [http://compbio.uchsc.edu/corpora Biomedical corpora]
* [http://ldc.upenn.edu Linguistic Data Consortium], a major distributor of corpora
* [http://www.xaira.org XAIRA]: a general purpose XML aware [[open-source]] corpus analysis tool 
* [http://corsis.sourceforge.net Corsis]: (formerly Tenka Text) an [[open-source]] ([[GPL]]ed) corpus analysis tool
* [http://www.arts-humanities.net/text_mining Discussion group text mining]

[[Category:Discourse analysis]]
[[Category:Corpus linguistics|*]]

[[bg:Корпусна лингвистика]]
[[cs:Korpusová lingvistika]]
[[de:Korpuslinguistik]]
[[et:Korpuslingvistika]]
[[ja:コーパス言語学]]
[[pt:Lingüística de Corpus]]
[[ru:Корпусная лингвистика]]