Changes

Corpora (view source)

Revision as of 02:04, 14 December 2007

24 bytes removed , 02:04, 14 December 2007

no edit summary

Line 2: Line 2:

[[Image:Copora.jpg|right|]]

== Text Corpus ==

−

+

In [[linguistics]], a ''corpus'' (plural '''corpora''') or textcorpora) or text corpus is a large and structured set of texts (now usually electronically stored and processed). They are used to do statistical analysis, checking occurrences or validating linguistic rules on a specific universe.

−

In [[linguistics]], a corpus (plural corpora) or textcorpora) or text corpus is a large and structured set of texts (now usually electronically stored and processed). They are used to do statistical analysis, checking occurrences or validating linguistic rules on a specific universe.

A corpus may contain texts in a single language (monolingual corpus) or text data in multiple languages (multilingual corpus). Multilingual corpora that have been specially formatted for side-by-side comparison are called aligned parallel corpora.

Line 10: Line 9:

Corpora are the main knowledge base in corpus linguistics. The analysis and processing of various types of corpora are also the subject of much work in [[computational linguistics]], [[speech recognition]] and [[machine translation]], where they are often used to create hidden [[Markov]] models for POS-tagging and other purposes. Corpora and frequency lists derived from them are useful for language teaching.

−

~~[[Category: General Reference]]~~

−

~~[[Category: Linguistics]]~~

== Archaeological corpora ==

Line 40: Line 36:

[[Category: General Reference]]

+

[[Category: Linguistics]]

Rdavis

Bureaucrats, Administrators

102,807

edits

Changes

Corpora (view source)

Revision as of 02:04, 14 December 2007

Navigation menu

Page actions

Page actions

Personal tools

Navigation

Tools

Search