Sie befinden Sich nicht im Netzwerk der Universität Paderborn. Der Zugriff auf elektronische Ressourcen ist gegebenenfalls nur via VPN oder Shibboleth (DFN-AAI) möglich. mehr Informationen...
Ergebnis 2 von 457
Language Resources and Evaluation, 2017-09, Vol.51 (3), p.643-662
2017

Details

Autor(en) / Beteiligte
Titel
Building and evaluating web corpora representing national varieties of English
Ist Teil von
  • Language Resources and Evaluation, 2017-09, Vol.51 (3), p.643-662
Ort / Verlag
Dordrecht: Springer
Erscheinungsjahr
2017
Link zum Volltext
Quelle
SpringerLink (Online service)
Beschreibungen/Notizen
  • Corpora are essential resources for language studies, as well as for training statistical natural language processing systems. Although very large English corpora have been built, only relatively small corpora are available for many varieties of English. National top-level domains (e.g.,. au,. ca) could be exploited to automatically build web corpora, but it is unclear whether such corpora would reflect the corresponding national varieties of English; i.e., would a web corpus built from the. ca domain correspond to Canadian English? In this article we build web corpora from national top-level domains corresponding to countries in which English is widely spoken. We then carry out statistical analyses of these corpora in terms of keywords, measures of corpus comparison based on the Chi-square test and spelling variants, and the frequencies of words known to be marked in particular varieties of English. We find evidence that the web corpora indeed reflect the corresponding national varieties of English. We then demonstrate, through a case study on the analysis of Canadianisms, that these corpora could be valuable lexicographical resources.
Sprache
Englisch
Identifikatoren
ISSN: 1574-020X
eISSN: 1572-8412, 1574-0218
DOI: 10.1007/s10579-016-9378-z
Titel-ID: cdi_proquest_journals_1925971512

Weiterführende Literatur

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX