UB Paderborn / Katalog / Suche / Details

Ergebnis 15 von 171

PeerJ. Computer science, 2022-04, Vol.8, p.e896-e896, Article e896

2022

Volltextzugriff (PDF)

Autor(en) / Beteiligte

Titel

Multi-label emotion classification of Urdu tweets

Ist Teil von

Ort / Verlag

United States: PeerJ. Ltd

Erscheinungsjahr

2022

Quelle

EZB Free E-Journals

Beschreibungen/Notizen

Urdu is a widely used language in South Asia and worldwide. While there are similar datasets available in English, we created the first multi-label emotion dataset consisting of 6,043 tweets and six basic emotions in the Urdu Nastalíq script. A multi-label (ML) classification approach was adopted to detect emotions from Urdu. The morphological and syntactic structure of Urdu makes it a challenging problem for multi-label emotion detection. In this paper, we build a set of baseline classifiers such as machine learning algorithms (Random forest (RF), Decision tree (J48), Sequential minimal optimization (SMO), AdaBoostM1, and Bagging), deep-learning algorithms (Convolutional Neural Networks (1D-CNN), Long short-term memory (LSTM), and LSTM with CNN features) and transformer-based baseline (BERT). We used a combination of text representations: stylometric-based features, pre-trained word embedding, word-based n-grams, and character-based n-grams. The paper highlights the annotation guidelines, dataset characteristics and insights into different methodologies used for Urdu based emotion classification. We present our best results using micro-averaged F1, macro-averaged F1, accuracy, Hamming loss (HL) and exact match (EM) for all tested methods.

Sprache: Englisch
Identifikatoren: ISSN: 2376-5992
eISSN: 2376-5992
DOI: 10.7717/peerj-cs.896
Titel-ID: cdi_doaj_primary_oai_doaj_org_article_ccf19b56fa80484fa5fceed97be5f77d

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX