UB Paderborn / Katalog / Suche / Details

Ergebnis 2 von 457

Language Resources and Evaluation, 2017-09, Vol.51 (3), p.643-662

2017

Autor(en) / Beteiligte

Titel

Building and evaluating web corpora representing national varieties of English

Ist Teil von

Ort / Verlag

Dordrecht: Springer

Erscheinungsjahr

2017

Link zum Volltext

Quelle

SpringerLink (Online service)

Beschreibungen/Notizen

Corpora are essential resources for language studies, as well as for training statistical natural language processing systems. Although very large English corpora have been built, only relatively small corpora are available for many varieties of English. National top-level domains (e.g.,. au,. ca) could be exploited to automatically build web corpora, but it is unclear whether such corpora would reflect the corresponding national varieties of English; i.e., would a web corpus built from the. ca domain correspond to Canadian English? In this article we build web corpora from national top-level domains corresponding to countries in which English is widely spoken. We then carry out statistical analyses of these corpora in terms of keywords, measures of corpus comparison based on the Chi-square test and spelling variants, and the frequencies of words known to be marked in particular varieties of English. We find evidence that the web corpora indeed reflect the corresponding national varieties of English. We then demonstrate, through a case study on the analysis of Canadianisms, that these corpora could be valuable lexicographical resources.

Sprache: Englisch
Identifikatoren: ISSN: 1574-020X
eISSN: 1572-8412, 1574-0218
DOI: 10.1007/s10579-016-9378-z
Titel-ID: cdi_proquest_journals_1925971512

Format: –
Schlagworte: Canadian English, Case studies, Computational Linguistics, Computer Science, Computerized corpora, Corpus analysis, Corpus linguistics, English language, Keywords, Language and Literature, Language varieties, Linguistics, Natural language processing, Original Paper, Social Sciences, Spelling, Statistical tests, Word frequency

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX